benchmarks evals
The Broad Finance Score Hid Decision-Level Regret
FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.
Summary
FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.
The Chinese-language suite contains 9,742 static instances across 53 task families plus 680 replayed pre-action states from 104 de-identified trajectories. Across 33 model configurations, operation rankings correlated only 0.42 on average, and knowledge-based shortlisting incurred up to 18.01 points of regret on individual operations.
Why it matters
FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.
Limits and context
- Across 33 model configurations, operation rankings correlated only 0.42 on average, and knowledge-based shortlisting incurred up to 18.01 points of regret on individual operations.
Key claims
FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.
Qualification: Across 33 model configurations, operation rankings correlated only 0.42 on average, and knowledge-based shortlisting incurred up to 18.01 points of regret on individual operations.
Evidence: source-2026-08-27-014
Sources
- arXiv preprint 2608.25325arXiv · primary research
Corrections
No corrections have been recorded for this story.