TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

research

Fifty-Seven Thousand Prompts Became a Dataset

Researchers collected transactional prompts embedded in software and mapped their formal, semantic and usage patterns.

Published Updated Story ID: mp-2026-08-16-007
Read the complete editionStory JSON

Summary

Researchers collected transactional prompts embedded in software and mapped their formal, semantic and usage patterns.

The corpus contains 57,500 unique prompts gathered from GitHub, focused on reproducible instructions integrated into code rather than one-off chat messages. A structured ontology records properties, formal components and semantics, revealing a Zipf-like mix of dominant patterns and a long tail across languages, domains, tasks and modalities. The team released both the dataset and an exploration interface, along with an error analysis of its automated annotations.

Why it matters

Researchers collected transactional prompts embedded in software and mapped their formal, semantic and usage patterns.

Limits and context

No additional limitation was separately recorded.

Key claims

  1. Researchers collected transactional prompts embedded in software and mapped their formal, semantic and usage patterns.

    Evidence: source-2026-08-16-007

Sources

  1. arXiv preprint 2608.12905arXiv · primary research

Corrections

No corrections have been recorded for this story.