frontier models
Steering Changed Less of the Language Model
MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.

Summary
MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.
Pre-logit steering can raise a test-time reward but also distort the rest of a frozen model's output distribution. MISVO penalizes interventions using the local Fisher geometry of token probabilities and optimizes position-specific vectors without updating model weights. Across preference and code-generation tasks on roughly one- to fourteen-billion-parameter models, the authors report higher reward with restrained distributional change. The comparison is limited to the selected models, tasks and reward functions.
Why it matters
MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.
Limits and context
No additional limitation was separately recorded.
Key claims
MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.
Evidence: source-2026-09-26-008
Sources
- arXiv preprint 2609.30218arXiv · primary research
Corrections
No corrections have been recorded for this story.