benchmarks evals
The Forecasting Gate Learned When to Ignore the Model
Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.

Summary
Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.
A language model is not always the best signal in a forecast, especially when a market, crowd or statistical prior already exists. This study estimates each source’s marginal value by domain, shrinks uncertain weights toward a global value and recalibrates the pooled result. Across 2,357 resolved binary questions and five language models, the gate improved the main external baseline’s Brier score from 0.0771 to 0.0732 and beat global combinations under the reported leakage controls. On ForecastBench’s official market subset it found no significant gain and mostly deferred to the market. Verbal confidence across four Qwen models did not reliably identify when the model added value.
Why it matters
Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.
Limits and context
- A language model is not always the best signal in a forecast, especially when a market, crowd or statistical prior already exists.
- Verbal confidence across four Qwen models did not reliably identify when the model added value.
Key claims
Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.
Qualification: A language model is not always the best signal in a forecast, especially when a market, crowd or statistical prior already exists.
Evidence: source-2026-09-14-003
Sources
- arXiv preprint 2609.12101arXiv · primary research
Corrections
No corrections have been recorded for this story.