safety security
A Learned Prefix Cut Model Utility by 26.8 Points
ENDOPROMPT optimized harmful prefixes from unlabeled instructions and produced losses across almost every tested split.
Summary
ENDOPROMPT optimized harmful prefixes from unlabeled instructions and produced losses across almost every tested split.
ENDOPROMPT uses white-box optimization to derive input prefixes without requiring labeled attack examples. Across four models, seven benign benchmarks and 28 model-benchmark combinations, the authors report a mean utility loss of 26.8 percentage points and negative effects in 27 of 28 cases. Their controls did not establish a compensating request-matching benefit; the preprint promises code upon acceptance, so independent reproduction remains a next step.
Why it matters
ENDOPROMPT optimized harmful prefixes from unlabeled instructions and produced losses across almost every tested split.
Limits and context
- Their controls did not establish a compensating request-matching benefit; the preprint promises code upon acceptance, so independent reproduction remains a next step.
Key claims
ENDOPROMPT optimized harmful prefixes from unlabeled instructions and produced losses across almost every tested split.
Qualification: Their controls did not establish a compensating request-matching benefit; the preprint promises code upon acceptance, so independent reproduction remains a next step.
Evidence: source-2026-09-27-007
Sources
- arXiv preprint 2609.29948arXiv · primary research
Corrections
No corrections have been recorded for this story.