TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

Each Cache Page Found Its Own Low-Rank Basis

PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.

Published Updated Story ID: mp-2026-08-26-012
Read the complete editionStory JSON

Summary

PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.

At roughly 60 percent of original KV storage, PuzzleKV retained more than 96 percent of Full KV performance across both evaluated models and all reported settings. Combined with quantization, it retained more than 93 percent using 18.7 percent of storage, with attention computed directly over dense and factorized pages.

Why it matters

PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. PuzzleKV compresses completed per-head pages independently instead of sharing one projection across a broad cache region.

    Evidence: source-2026-08-26-012

Sources

  1. arXiv preprint 2608.23843arXiv · primary research

Corrections

No corrections have been recorded for this story.