Competence-Gated Pooling Improves Hybrid Event Forecasts Over External Signals
A new study introduces a "competence gate" that decides when a language model should be combined with existing market, crowd, or statistical forecasts for binary event prediction. By estimating domain‑level source weights from resolved outcomes and shrinking uncertain estimates toward a global baseline, the gate recalibrates pooled forecasts and boosts performance.
Key points
- Competence gate reduces Brier score from 0.0771 to 0.0732 across 2,357 binary questions
- Four Qwen models and other LMs tested; gate outperforms global pooling and leakage‑safe priors
- Verbal confidence fails to predict model usefulness; outcome‑based competence guides abstention
Evaluated on 2,357 resolved questions using five language models—including four Qwen variants—the method lowered the Brier score from 0.0771 to 0.0732, outperforming simple global pooling. The improvement held under strict leakage controls and against a time‑series prior, though it offered no gain on the ForecastBench market subset, where the gate largely deferred to the market. The authors also show that verbal confidence from the models is a poor indicator of competence, while outcome‑based estimates enable better abstention decisions.
The work provides a practical framework for selectively deploying language models only when they add measurable marginal value, helping forecasters avoid over‑reliance on AI predictions that may not improve existing signals.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- R2VC Boosts Fact‑Checking Accuracy by 13.74% on FEVER Using Modular Retrieval and Calibration · 1 src
- Study Finds Deictic Ambiguity Can Undermine Draft‑Verify‑Revise LLM Pipelines · 1 src
- New GLARE model improves meeting continuation forecasting on MDFB benchmark · 1 src
- Chopthin-Consensus Power Sampling Boosts LLM Reasoning Accuracy Without Retraining · 1 src
- Context-Augmented KG Training Boosts Multi-Hop QA Accuracy on Disease Graphs · 1 src
Comments
via GitHub Discussions