benchmarks evals
Code Attribution Mostly Found Length, Not Identity
Across 15 model-benchmark combinations, balanced accuracy stayed near chance and tracked solution length.

Summary
Across 15 model-benchmark combinations, balanced accuracy stayed near chance and tracked solution length.
A study of model self-attribution found balanced accuracy between 49% and 58% across 15 model-benchmark combinations. Pairwise attribution scores correlated at r=0.93 with longer solutions, suggesting that style and verbosity were doing much of the work. After normalization, 10 of 12 tested conditions fell to chance and the remaining two continued to follow length, erasing the apparent self-preference in this setup.
Why it matters
Across 15 model-benchmark combinations, balanced accuracy stayed near chance and tracked solution length.
Limits and context
No additional limitation was separately recorded.
Key claims
Across 15 model-benchmark combinations, balanced accuracy stayed near chance and tracked solution length.
Evidence: source-2026-09-27-008
Sources
- arXiv preprint 2609.30048arXiv · primary research
Corrections
No corrections have been recorded for this story.