TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

The Model Learned a Fact Better When the Corpus Showed Another View

With a fixed token budget, reformulations beat spending the same tokens on simple document repetition—even for factual recall.

Published Updated Story ID: mp-2026-09-06-009
Read the complete editionStory JSON

Summary

With a fixed token budget, reformulations beat spending the same tokens on simple document repetition—even for factual recall.

Controlled pretraining experiments found that repetition remained necessary, but reallocating some repeated tokens to auxiliary representations improved knowledge acquisition. Paraphrases helped under smaller batches, while contextual and foundational views aided learning when prior knowledge was missing. The effect did not depend on a stronger teacher generating the reformulation. The experiments isolate learning mechanisms under controlled conditions and do not prove that every synthetic rewrite improves a production corpus.

Why it matters

With a fixed token budget, reformulations beat spending the same tokens on simple document repetition—even for factual recall.

Limits and context

  • The effect did not depend on a stronger teacher generating the reformulation.
  • The experiments isolate learning mechanisms under controlled conditions and do not prove that every synthetic rewrite improves a production corpus.

Key claims

  1. With a fixed token budget, reformulations beat spending the same tokens on simple document repetition—even for factual recall.

    Qualification: The effect did not depend on a stronger teacher generating the reformulation.

    Evidence: source-2026-09-06-009

Sources

  1. arXiv preprint 2609.04180arXiv · primary research

Corrections

No corrections have been recorded for this story.