safety
Robustness Failed When the Model Ignored Good Advice
MIST tests clean, misleading, correct and irrelevant context together so resistance cannot masquerade as selective judgment.

Summary
MIST tests clean, misleading, correct and irrelevant context together so resistance cannot masquerade as selective judgment.
The MIST benchmark renders each reasoning item under four matched context conditions and measures how often misleading context flips an otherwise correct answer. The authors say susceptibility appeared across the open models they tested; their SCOPE training method reduced those flips while preserving accuracy when context was correct, clean or irrelevant. The preprint argues for evaluating selective trust rather than blanket resistance, but does not establish immunity to adversarial context in deployed systems.
Why it matters
MIST tests clean, misleading, correct and irrelevant context together so resistance cannot masquerade as selective judgment.
Limits and context
- The preprint argues for evaluating selective trust rather than blanket resistance, but does not establish immunity to adversarial context in deployed systems.
Key claims
MIST tests clean, misleading, correct and irrelevant context together so resistance cannot masquerade as selective judgment.
Qualification: The preprint argues for evaluating selective trust rather than blanket resistance, but does not establish immunity to adversarial context in deployed systems.
Evidence: source-2026-08-08-003
Sources
- arXiv preprint 2608.06377arXiv · primary research
Corrections
No corrections have been recorded for this story.