safety security
A Benign Reply Carried Four Digits of the Secret
Researchers reconstructed in-context secrets from ordinary outputs even when models refused direct extraction.
Summary
Researchers reconstructed in-context secrets from ordinary outputs even when models refused direct extraction.
Across eight proprietary models in controlled experiments, the authors report near-perfect recovery of two-digit secrets and 82 percent exact recovery for four digits from benign responses. They also trained attacks that infer predicates about user memories and extract longer identifiers in a production-style agent, framing context sensitivity itself as a covert leakage channel that capability may amplify.
Why it matters
Researchers reconstructed in-context secrets from ordinary outputs even when models refused direct extraction.
Limits and context
No additional limitation was separately recorded.
Key claims
Researchers reconstructed in-context secrets from ordinary outputs even when models refused direct extraction.
Evidence: source-2026-08-23-012
Sources
- arXiv preprint 2608.19857arXiv · primary research
Corrections
No corrections have been recorded for this story.