TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

infrastructure

Each Attention Head Got Its Own Memory Window

HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.

Published Updated Story ID: mp-2026-09-03-019
Read the complete editionStory JSON

Summary

HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.

On Qwen3.6-27B, the fixed-model study reduced sampled peak device memory 8.59% at 112K context and extended the largest verified context from 114K to 161K while retaining near-full-cache quality on the tested suites.

Why it matters

HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. HeadWiseKV assigns static per-head histories under a total cache budget without retraining the model.

    Evidence: source-2026-09-03-021

Sources

  1. arXiv preprint 2609.02029arXiv · primary research

Corrections

No corrections have been recorded for this story.