TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

The Same Forty Slots Received Different Questions

A two-study simulation traces how embeddings, structural screens and selection policy determine what psychometricians ever review.

Published Updated Story ID: mp-2026-08-26-007
Read the complete editionStory JSON

Summary

A two-study simulation traces how embeddings, structural screens and selection policy determine what psychometricians ever review.

Across 32,000 selected Big Five items, broad semantic agreement concealed large local changes in evidence and survival. Every evaluable form filled all content cells, yet inclusive primary forms from different embedding configurations shared a median of only six of 40 items, making the computational evaluator part of measurement design rather than neutral plumbing.

Why it matters

A two-study simulation traces how embeddings, structural screens and selection policy determine what psychometricians ever review.

Limits and context

  • Every evaluable form filled all content cells, yet inclusive primary forms from different embedding configurations shared a median of only six of 40 items, making the computational evaluator part of measurement design rather than neutral plumbing.

Key claims

  1. A two-study simulation traces how embeddings, structural screens and selection policy determine what psychometricians ever review.

    Qualification: Every evaluable form filled all content cells, yet inclusive primary forms from different embedding configurations shared a median of only six of 40 items, making the computational evaluator part of measurement design rather than neutral plumbing.

    Evidence: source-2026-08-26-007

Sources

  1. arXiv preprint 2608.23766arXiv · primary research

Corrections

No corrections have been recorded for this story.