safety security
The Auditor Trained Against Hidden Behaviors
Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.

Summary
Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.
The training environment planted hidden behaviors through target system prompts and rewarded investigations by pairwise comparison with references. The authors report stronger investigations, more concerning behaviors surfaced in unmodified production models, improved realism and cross-scaffold generalization, with false positives below one percent in tested settings.
Why it matters
Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.
Limits and context
No additional limitation was separately recorded.
Key claims
Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.
Evidence: source-2026-08-27-008
Sources
- arXiv preprint 2608.25460arXiv · primary research
Corrections
No corrections have been recorded for this story.