infrastructure
Each Attention Head Got Its Own Memory Window
HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.
Summary
HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.
On Qwen3.6-27B, the fixed-model study reduced sampled peak device memory 8.59% at 112K context and extended the largest verified context from 114K to 161K while retaining near-full-cache quality on the tested suites.
Why it matters
HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.
Limits and context
No additional limitation was separately recorded.
Key claims
HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.
Evidence: source-2026-09-03-021
Sources
- arXiv preprint 2609.02029arXiv · primary research
Corrections
No corrections have been recorded for this story.