infrastructure
Tier Capacity Beat Clever Cache Prediction
GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.
Summary
GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.
Long-lived chats and agent loops fill expensive GPU memory with key-value cache blocks. A calibrated simulator compared recency, reuse frequency, predicted reuse and look-ahead prefetch across GPU HBM, CPU DRAM and SSD. The three-tier capacity model supported 73.02 times more concurrent sessions per GPU and cut cost per session 62.04 times, but those gains came from capacity rather than sophisticated placement. Recency minimized chat migration while reuse frequency led for agents and document work. Even an oracle prefetch policy failed to beat no prefetch on migration traffic, challenging recommendations that prediction alone improves this hierarchy.
Why it matters
GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.
Limits and context
No additional limitation was separately recorded.
Key claims
GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.
Evidence: source-2026-09-16-006
Sources
- arXiv preprint 2609.16215arXiv · primary research
Corrections
No corrections have been recorded for this story.