TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

infrastructure

Attention Clustered Once, Then Ran Sparse

ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.

Published Updated Story ID: mp-2026-08-29-015
Read the complete editionStory JSON

Summary

ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.

The authors report two-to-six-times speedups on a tabular foundation model while retaining at least 99 percent of dense accuracy, plus a 1.8-times video-generation speedup with outputs closer to dense attention than a tested specialist method.

Why it matters

ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. ClusterAttention uses fast per-head recursive clustering and centroid compensation without retraining or offline calibration.

    Evidence: source-2026-08-29-017

Sources

  1. arXiv preprint 2608.26965arXiv · primary research

Corrections

No corrections have been recorded for this story.