safety
The Dashboard Made the Unknowable Feel Actionable
Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted.

Summary
Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted.
The preregistered study reports that commitment on provably unpredictable questions rose from 6.5 percent with a bare question to 54 percent as evidence packaging intensified. Fully fabricated displays lifted commitment from 24.5 to 36.8 percent, statistically indistinguishable from the 37.6 percent produced by genuine market data. Models often recognized that a question was unknowable when asked first, suggesting a failure at the act-or-abstain gate rather than simple incapacity. Fine-tuning one 3B model on 540 synthetic examples eliminated commitment on the original cases, but the behavior returned under rigid response formats. These are author-reported experimental results, not proof that every model or deployment fails this way.
Why it matters
Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted.
Limits and context
- These are author-reported experimental results, not proof that every model or deployment fails this way.
Key claims
Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted.
Qualification: These are author-reported experimental results, not proof that every model or deployment fails this way.
Evidence: source-2026-08-28-001
Sources
- arXiv preprint 2608.27167arXiv · primary research
Corrections
No corrections have been recorded for this story.