TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

chips infrastructure

The Architecture Learned the Order the CPU Could Stream

Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.

Published Updated Story ID: mp-2026-08-26-010
Read the complete editionStory JSON

Summary

Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.

One tested architecture halved critical-path weight bandwidth from 9.00 to 4.50 MB per token while staying within 0.24 perplexity of the best candidate. On a 30.9-billion-parameter mixture-of-experts model, the cflow runtime reported 5.94 tokens per second on 32 Ice Lake vCPUs; the authors also report one design claim refuted and another inconclusive.

Why it matters

Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Pipeline-native transformers co-design dependency graphs with a stage-major runtime for bandwidth-bound decoding.

    Evidence: source-2026-08-26-010

Sources

  1. arXiv preprint 2608.23841arXiv · primary research

Corrections

No corrections have been recorded for this story.