TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

Two Billion Active Parameters Carried a Twenty-Billion Model

Turing-20B-A2B combines dynamic expert routing and hybrid attention for long-context, latency-sensitive physical-AI workloads.

Published Updated Story ID: mp-2026-09-01-026
Read the complete editionStory JSON

Summary

Turing-20B-A2B combines dynamic expert routing and hybrid attention for long-context, latency-sensitive physical-AI workloads.

The mixture-of-experts model activates about two billion of its twenty billion parameters per token, using quantile routing to vary expert allocation while controlling average compute. Lightning Attention is mixed with a small number of full-attention layers; continued pretraining extends native context to 128K and YaRN is used for 512K inference-time extension. The authors report base-model capability above Qwen3-8B Base and near Qwen3.5-9B Base with favorable prefill scaling. Those are self-reported benchmark comparisons, not independent deployment results.

Why it matters

Turing-20B-A2B combines dynamic expert routing and hybrid attention for long-context, latency-sensitive physical-AI workloads.

Limits and context

  • Those are self-reported benchmark comparisons, not independent deployment results.

Key claims

  1. Turing-20B-A2B combines dynamic expert routing and hybrid attention for long-context, latency-sensitive physical-AI workloads.

    Qualification: Those are self-reported benchmark comparisons, not independent deployment results.

    Evidence: source-2026-09-01-015

Sources

  1. arXiv preprint 2608.30567arXiv · primary research

Corrections

No corrections have been recorded for this story.