TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

The Broad Finance Score Hid Decision-Level Regret

FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.

Published Updated Story ID: mp-2026-08-27-014
Read the complete editionStory JSON

Summary

FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.

The Chinese-language suite contains 9,742 static instances across 53 task families plus 680 replayed pre-action states from 104 de-identified trajectories. Across 33 model configurations, operation rankings correlated only 0.42 on average, and knowledge-based shortlisting incurred up to 18.01 points of regret on individual operations.

Why it matters

FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.

Limits and context

  • Across 33 model configurations, operation rankings correlated only 0.42 on average, and knowledge-based shortlisting incurred up to 18.01 points of regret on individual operations.

Key claims

  1. FinRiskAtlas evaluates specific review operations and whether the available evidence supports a defensible next step.

    Qualification: Across 33 model configurations, operation rankings correlated only 0.42 on average, and knowledge-based shortlisting incurred up to 18.01 points of regret on individual operations.

    Evidence: source-2026-08-27-014

Sources

  1. arXiv preprint 2608.25325arXiv · primary research

Corrections

No corrections have been recorded for this story.