TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

benchmarks evals

Safety Evidence Lost to Retrieval Friction

Four frontier models reacted strongly to severity and cost, while a stated risk jump from 10% to 70% moved inspection by at most 21 points.

Published Updated Story ID: mp-2026-09-17-006
Read the complete editionStory JSON

Summary

Four frontier models reacted strongly to severity and cost, while a stated risk jump from 10% to 70% moved inspection by at most 21 points.

SAFE tests whether a model asks for safety-relevant evidence before making a deployment decision. Across GPT-5.5, o3, Claude Opus 4.8 and Claude Sonnet 4.6, inspection increased with severity and fell with retrieval cost. Stated probability mattered less: moving a problem from 10% to 70% likelihood changed inspection by no more than 21 percentage points. Opus inspected almost by default; o3 skipped most and reacted most sharply to thresholds. Counterfactual framing changed some decisions without appearing in the explanations, exposing a gap between rationales and acquisition policy.

Why it matters

Four frontier models reacted strongly to severity and cost, while a stated risk jump from 10% to 70% moved inspection by at most 21 points.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Four frontier models reacted strongly to severity and cost, while a stated risk jump from 10% to 70% moved inspection by at most 21 points.

    Evidence: source-2026-09-17-006

Sources

  1. arXiv preprint 2609.17865arXiv · primary research

Corrections

No corrections have been recorded for this story.