TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

safety security

The Auditor Trained Against Hidden Behaviors

Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.

Published Updated Story ID: mp-2026-08-27-008
Read the complete editionStory JSON

Summary

Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.

The training environment planted hidden behaviors through target system prompts and rewarded investigations by pairwise comparison with references. The authors report stronger investigations, more concerning behaviors surfaced in unmodified production models, improved realism and cross-scaffold generalization, with false positives below one percent in tested settings.

Why it matters

Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Reinforcement learning improved model investigations while negative examples helped keep false positives below one percent.

    Evidence: source-2026-08-27-008

Sources

  1. arXiv preprint 2608.25460arXiv · primary research

Corrections

No corrections have been recorded for this story.