TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

Steering Changed Less of the Language Model

MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.

Published Updated Story ID: mp-2026-09-26-008
Read the complete editionStory JSON

Summary

MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.

Pre-logit steering can raise a test-time reward but also distort the rest of a frozen model's output distribution. MISVO penalizes interventions using the local Fisher geometry of token probabilities and optimizes position-specific vectors without updating model weights. Across preference and code-generation tasks on roughly one- to fourteen-billion-parameter models, the authors report higher reward with restrained distributional change. The comparison is limited to the selected models, tasks and reward functions.

Why it matters

MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. MISVO led mean reward in six of seven model-task settings while keeping diversity and coherence near Best-of-N.

    Evidence: source-2026-09-26-008

Sources

  1. arXiv preprint 2609.30218arXiv · primary research

Corrections

No corrections have been recorded for this story.