research
Each Cache Page Found Its Own Low-Rank Basis
PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.
Summary
PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.
At roughly 60 percent of original KV storage, PuzzleKV retained more than 96 percent of Full KV performance across both evaluated models and all reported settings. Combined with quantization, it retained more than 93 percent using 18.7 percent of storage, with attention computed directly over dense and factorized pages.
Why it matters
PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.
Limits and context
No additional limitation was separately recorded.
Key claims
PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.
Evidence: source-2026-08-26-012
Sources
- arXiv preprint 2608.23843arXiv · primary research
Corrections
No corrections have been recorded for this story.