TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

infrastructure

Tier Capacity Beat Clever Cache Prediction

GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.

Published Updated Story ID: mp-2026-09-16-006
Read the complete editionStory JSON

Summary

GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.

Long-lived chats and agent loops fill expensive GPU memory with key-value cache blocks. A calibrated simulator compared recency, reuse frequency, predicted reuse and look-ahead prefetch across GPU HBM, CPU DRAM and SSD. The three-tier capacity model supported 73.02 times more concurrent sessions per GPU and cut cost per session 62.04 times, but those gains came from capacity rather than sophisticated placement. Recency minimized chat migration while reuse frequency led for agents and document work. Even an oracle prefetch policy failed to beat no prefetch on migration traffic, challenging recommendations that prediction alone improves this hierarchy.

Why it matters

GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.

    Evidence: source-2026-09-16-006

Sources

  1. arXiv preprint 2609.16215arXiv · primary research

Corrections

No corrections have been recorded for this story.