TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Forecasting Gate Learned When to Ignore the Model

Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.

Published Updated Story ID: mp-2026-09-14-003
Read the complete editionStory JSON

Summary

Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.

A language model is not always the best signal in a forecast, especially when a market, crowd or statistical prior already exists. This study estimates each source’s marginal value by domain, shrinks uncertain weights toward a global value and recalibrates the pooled result. Across 2,357 resolved binary questions and five language models, the gate improved the main external baseline’s Brier score from 0.0771 to 0.0732 and beat global combinations under the reported leakage controls. On ForecastBench’s official market subset it found no significant gain and mostly deferred to the market. Verbal confidence across four Qwen models did not reliably identify when the model added value.

Why it matters

Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.

Limits and context

  • A language model is not always the best signal in a forecast, especially when a market, crowd or statistical prior already exists.
  • Verbal confidence across four Qwen models did not reliably identify when the model added value.

Key claims

  1. Across 2,357 resolved questions, domain-level competence weights improved the main external baseline’s Brier score from 0.0771 to 0.0732.

    Qualification: A language model is not always the best signal in a forecast, especially when a market, crowd or statistical prior already exists.

    Evidence: source-2026-09-14-003

Sources

  1. arXiv preprint 2609.12101arXiv · primary research

Corrections

No corrections have been recorded for this story.