Baibaichuchu at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text?
The BAIBAICHUCHU team participated in the NTCIR-19 FinArg-3 Social Media Subtask, ranking Chinese investor posts by Maximum Possible Profit (MPP). Their ensemble model—combining lexical features, a MacBERT ranker fine-tuned on FinArg-2, and an LLM judge—achieved a development evaluation score of 0.734 but settled on an official run scoring 0.517. All submitted runs ranged between 0.4598 and…
Key points
- BAIBAICHUCHU’s ensemble model scored 0.517 on NTCIR-19 FinArg-3’s official task, with development scores up to 0.734
- Post-hoc audit found a long-only rule bias in bearish posts, lowering development accuracy to 0.622 after correction
- Pre-posting volatility correlated with MPP outcomes (Spearman ρ=0.320), while transferred text showed minimal predictive value
A post-hoc audit revealed the judge applied a long-only rule to bearish posts, despite MPP being stance-aware. Adjusting this rule altered 28 of 87 predictions without changing accuracy but reduced development accuracy from 0.680 to 0.622. The team also analyzed ranking factors—directional text, volatility, horizon, and pairwise margin—and found a running-extremum model predicted $\sigma\sqrt{T}$ scaling. Pre-posting volatility correlated with later MPP outcomes (Spearman $\rho=0.320$), while transferred text showed near-zero correlation with July 2026 outcomes.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- GitHub releases ReviewBench to evaluate AI code reviews across 219 pull requests · 2 src
- Anthropic study: task understanding beats job title for AI success · 1 src
- Researchers propose FEM-ASM to separate storage, execution, and coordination in language models · 1 src
- Researchers propose surrogate log-probabilities to audit LLM agents · 3 src
- Behavioral history outperforms descriptions for LLM synthetic personas · 1 src
Comments
via GitHub Discussions