TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

business enterprise

Search Rewards Improved Query Understanding

Roblox experiments raised NDCG@20 by 8.9 points over supervised fine-tuning and 3.5 over one end-to-end reward.

Published Updated Story ID: mp-2026-09-26-026
Read the complete editionStory JSON

Summary

Roblox experiments raised NDCG@20 by 8.9 points over supervised fine-tuning and 3.5 over one end-to-end reward.

The framework first distills a teacher into a schema-compliant query-understanding policy, then optimizes intent classification, query expansion and other components with rewards drawn from their actual interaction with the search engine. Giving each component an operational reward improved both component utility and downstream retrieval in the reported Roblox experiments. The result is specific to that game-search pipeline and does not show that reinforcement learning will improve every production search system.

Why it matters

Roblox experiments raised NDCG@20 by 8.9 points over supervised fine-tuning and 3.5 over one end-to-end reward.

Limits and context

  • The result is specific to that game-search pipeline and does not show that reinforcement learning will improve every production search system.

Key claims

  1. Roblox experiments raised NDCG@20 by 8.9 points over supervised fine-tuning and 3.5 over one end-to-end reward.

    Qualification: The result is specific to that game-search pipeline and does not show that reinforcement learning will improve every production search system.

    Evidence: source-2026-09-26-015

Sources

  1. arXiv preprint 2609.30177arXiv · primary research

Corrections

No corrections have been recorded for this story.