safety security
The Poison Looked Like an Update
A retrieval attack avoided direct contradiction by presenting false material as newer information, increasing success across many tested settings.
Summary
A retrieval attack avoided direct contradiction by presenting false material as newer information, increasing success across many tested settings.
PURSUE tests a black-box retrieval-augmented-generation poisoning strategy that frames injected misinformation as a non-contradictory update rather than an explicit conflict. Across three question-answering benchmarks, five generators and three conflict-resolution methods, the authors report the highest attack success in 35 of 45 settings and a mean gain of 9.7 percentage points over their strongest prior baseline. This is defensive preprint research on benchmark systems; it does not establish compromise of a named deployed service, and this report omits operational injection instructions.
Why it matters
A retrieval attack avoided direct contradiction by presenting false material as newer information, increasing success across many tested settings.
Limits and context
- This is defensive preprint research on benchmark systems; it does not establish compromise of a named deployed service, and this report omits operational injection instructions.
Key claims
A retrieval attack avoided direct contradiction by presenting false material as newer information, increasing success across many tested settings.
Qualification: This is defensive preprint research on benchmark systems; it does not establish compromise of a named deployed service, and this report omits operational injection instructions.
Evidence: source-2026-08-06-005
Sources
- arXiv preprint 2608.04756arXiv · primary research
Corrections
No corrections have been recorded for this story.