frontier models
The Diffusion Decoder Earned More Space Before It Grew
CARVE inserts new masked positions only when the expanded canvas leaves unresolved predictions stable.
Summary
CARVE inserts new masked positions only when the expanded canvas leaves unresolved predictions stable.
Masked diffusion language models normally commit to an answer length before denoising begins, risking truncation or wasted computation. CARVE starts short, proposes more masked space and retains it only when aligned unresolved positions show low Jensen-Shannon divergence under the counterfactual expansion. Across code and mathematical reasoning benchmarks, the training-free method improved average performance across the evaluated model families and in some settings used half the FLOPs of fixed-length decoding. Those gains remain benchmark- and model-specific.
Why it matters
CARVE inserts new masked positions only when the expanded canvas leaves unresolved predictions stable.
Limits and context
- CARVE starts short, proposes more masked space and retains it only when aligned unresolved positions show low Jensen-Shannon divergence under the counterfactual expansion.
- Those gains remain benchmark- and model-specific.
Key claims
CARVE inserts new masked positions only when the expanded canvas leaves unresolved predictions stable.
Qualification: CARVE starts short, proposes more masked space and retains it only when aligned unresolved positions show low Jensen-Shannon divergence under the counterfactual expansion.
Evidence: source-2026-09-01-010
Sources
- arXiv preprint 2608.30922arXiv · primary research
Corrections
No corrections have been recorded for this story.