chips infrastructure
Cache Compression Won the Cost Axis Until Weights Became the Wall
A profiled simulator places tensor parallelism and KV compression on one cost-latency chart.
Summary
A profiled simulator places tensor parallelism and KV compression on one cost-latency chart.
Across two model sizes and three GPU types, compression was 1.20 to 2.00 times cheaper in every comparable memory-relief configuration. Above roughly 36 billion parameters on an 80 GB device, weights—not cache—became the binding constraint, making tensor parallelism an entry requirement despite higher cost.
Why it matters
A profiled simulator places tensor parallelism and KV compression on one cost-latency chart.
Limits and context
- Above roughly 36 billion parameters on an 80 GB device, weights—not cache—became the binding constraint, making tensor parallelism an entry requirement despite higher cost.
Key claims
A profiled simulator places tensor parallelism and KV compression on one cost-latency chart.
Qualification: Above roughly 36 billion parameters on an 80 GB device, weights—not cache—became the binding constraint, making tensor parallelism an entry requirement despite higher cost.
Evidence: source-2026-08-26-021
Sources
- arXiv preprint 2608.23962arXiv · primary research
Corrections
No corrections have been recorded for this story.