TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

frontier models

A Seven-Billion-Parameter Model Opened Its Entire Training Trail

ZGCM-1 reports a 4.2-fold improvement in 16K pretraining time-to-loss and releases weights, checkpoints, code, data recipes and logs.

Published Updated Story ID: mp-2026-09-15-003
Read the complete editionStory JSON

Summary

ZGCM-1 reports a 4.2-fold improvement in 16K pretraining time-to-loss and releases weights, checkpoints, code, data recipes and logs.

ZGCM-1 is a dense seven-billion-parameter foundation model built around the claim that compact models should combine internal reasoning with external tool use instead of trying to memorize the open web. Its recipe mixes sliding-window and full attention, FP8 Muon optimization, progressive context scaling to 256K and interaction traces recast as Markov decision processes. The team reports roughly 4.2× better 16K pretraining time-to-loss and competitive results against larger models on selected math and agentic-search suites. It also releases stage-by-stage weights, checkpoints, code, data recipes and experiment logs; independent replication remains necessary.

Why it matters

ZGCM-1 reports a 4.2-fold improvement in 16K pretraining time-to-loss and releases weights, checkpoints, code, data recipes and logs.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. ZGCM-1 reports a 4.2-fold improvement in 16K pretraining time-to-loss and releases weights, checkpoints, code, data recipes and logs.

    Evidence: source-2026-09-15-003

Sources

  1. arXiv preprint 2609.13356arXiv · primary research

Corrections

No corrections have been recorded for this story.