chips infrastructure
The Architecture Learned the Order the CPU Could Stream
Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.
Summary
Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.
One tested architecture halved critical-path weight bandwidth from 9.00 to 4.50 MB per token while staying within 0.24 perplexity of the best candidate. On a 30.9-billion-parameter mixture-of-experts model, the cflow runtime reported 5.94 tokens per second on 32 Ice Lake vCPUs; the authors also report one design claim refuted and another inconclusive.
Why it matters
Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.
Limits and context
No additional limitation was separately recorded.
Key claims
Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.
Evidence: source-2026-08-26-010
Sources
- arXiv preprint 2608.23841arXiv · primary research
Corrections
No corrections have been recorded for this story.