safety security
The Safe Score Hid the Privacy Collapse
A black-box audit across more than 120 language models found that safety, security, privacy and benign usefulness can move in different directions.

Summary
A black-box audit across more than 120 language models found that safety, security, privacy and benign usefulness can move in different directions.
The aiXamine preprint combines 46 tests across nine services so that alignment, adversarial robustness, privacy and over-refusal are measured as related properties rather than separate leaderboards. Across more than 5,000 runs, the authors report three recurring trade-offs: stronger safety enforcement often rejected more benign requests; privacy was nearly orthogonal to the other trust dimensions; and one off-policy distillation setting collapsed robustness from 56.9 to 2.6 on the same base architecture. These are benchmark findings, not a universal ranking of deployed models, but they make a practical point: a high score on one trust axis cannot certify the others.
Why it matters
A black-box audit across more than 120 language models found that safety, security, privacy and benign usefulness can move in different directions.
Limits and context
- These are benchmark findings, not a universal ranking of deployed models, but they make a practical point: a high score on one trust axis cannot certify the others.
Key claims
A black-box audit across more than 120 language models found that safety, security, privacy and benign usefulness can move in different directions.
Qualification: These are benchmark findings, not a universal ranking of deployed models, but they make a practical point: a high score on one trust axis cannot certify the others.
Evidence: source-2026-08-24-001
Sources
- arXiv preprint 2608.20554arXiv · primary research
Corrections
No corrections have been recorded for this story.