---
schema_version: "1.0.0"
edition_id: "mp-2026-08-28-morning-0050"
published_at: "2026-08-28T09:00:00.000-04:00"
modified_at: "2026-08-28T09:00:00.000-04:00"
canonical_url: "https://themachinepress.com/edition/2026-08-28"
story_count: 27
lead_story_id: "mp-2026-08-28-001"
---

# The Machine Press — Morning edition

Edition ID: `mp-2026-08-28-morning-0050`  
Published: 2026-08-28T09:00:00.000-04:00  
Canonical edition: https://themachinepress.com/edition/2026-08-28

Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted.

## 1. The Dashboard Made the Unknowable Feel Actionable {#mp-2026-08-28-001}

- Story ID: `mp-2026-08-28-001`
- Type: `lead`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-001/the-dashboard-made-the-unknowable-feel-actionable

**Dek:** Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted.

The preregistered study reports that commitment on provably unpredictable questions rose from 6.5 percent with a bare question to 54 percent as evidence packaging intensified. Fully fabricated displays lifted commitment from 24.5 to 36.8 percent, statistically indistinguishable from the 37.6 percent produced by genuine market data. Models often recognized that a question was unknowable when asked first, suggesting a failure at the act-or-abstain gate rather than simple incapacity. Fine-tuning one 3B model on 540 synthetic examples eliminated commitment on the original cases, but the behavior returned under rigid response formats. These are author-reported experimental results, not proof that every model or deployment fails this way.

### Why it matters {#why-it-matters-mp-2026-08-28-001}

Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted.

### Limits and context {#limitations-mp-2026-08-28-001}

- These are author-reported experimental results, not proof that every model or deployment fails this way.

### Claims and sources {#claims-mp-2026-08-28-001}

- Across twelve frontier models, professional-looking evidence pushed agents toward directional calls even when the evidence was fabricated and the question could not be predicted. [source-2026-08-28-001] — Qualification: These are author-reported experimental results, not proof that every model or deployment fails this way.

## 2. Training Ended. The Prompting Habits Stayed Put {#mp-2026-08-28-002}

- Story ID: `mp-2026-08-28-002`
- Type: `secondary`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-002/training-ended-the-prompting-habits-stayed-put

**Dek:** A field study of 713,564 workplace prompts found more sophisticated use among senior employees, but no durable improvement over time or after formal AI training.

Researchers analyzed prompts and responses from nearly 4,000 back-office employees across 15 functions at one large firm over eight months in 2025. Their measures placed Strategy, Digital Innovation and Project Management highest and found senior employees using generative AI more sophisticatedly, a pattern consistent with domain expertise complementing the tool. The study found neither improvement over time nor lasting gains following formal training. Because the evidence is proprietary, observational and drawn from one firm, it does not establish that training is useless or that seniority causes better outcomes; it shows that access and attendance alone did not shift the measured habits in this setting.

### Why it matters {#why-it-matters-mp-2026-08-28-002}

A field study of 713,564 workplace prompts found more sophisticated use among senior employees, but no durable improvement over time or after formal AI training.

### Limits and context {#limitations-mp-2026-08-28-002}

- Because the evidence is proprietary, observational and drawn from one firm, it does not establish that training is useless or that seniority causes better outcomes; it shows that access and attendance alone did not shift the measured habits in this setting.

### Claims and sources {#claims-mp-2026-08-28-002}

- A field study of 713,564 workplace prompts found more sophisticated use among senior employees, but no durable improvement over time or after formal AI training. [source-2026-08-28-002] — Qualification: Because the evidence is proprietary, observational and drawn from one firm, it does not establish that training is useless or that seniority causes better outcomes; it shows that access and attendance alone did not shift the measured habits in this setting.

## 3. The Skill Kept a Wiki of What It Learned {#mp-2026-08-28-003}

- Story ID: `mp-2026-08-28-003`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-003/the-skill-kept-a-wiki-of-what-it-learned

**Dek:** WikiSkill separates raw execution experience, accumulated knowledge and executable skills so later revisions can reuse earlier lessons.

Across several benchmarks and models, the authors report that the persistent wiki improved on prior skill-evolution methods and usually beat no-skill baselines. Evolved skills transferred across model families, and smaller models with skills sometimes outperformed substantially larger models without them; ablations attributed part of the gain to accumulated knowledge rather than the executable skill alone.

### Why it matters {#why-it-matters-mp-2026-08-28-003}

WikiSkill separates raw execution experience, accumulated knowledge and executable skills so later revisions can reuse earlier lessons.

### Limits and context {#limitations-mp-2026-08-28-003}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-28-003}

- WikiSkill separates raw execution experience, accumulated knowledge and executable skills so later revisions can reuse earlier lessons. [source-2026-08-28-003]

## 4. The Reaction Model Followed the Electrons {#mp-2026-08-28-004}

- Story ID: `mp-2026-08-28-004`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-004/the-reaction-model-followed-the-electrons

**Dek:** MAELLE represents chemical reactions as discrete flow through electron-occupation space instead of direct molecular graph edits.

The method uses a continuous-time Markov chain and optimal-transport paths to produce interpretable electron rearrangements without elementary-step annotations. Its authors report competitive USPTO-480K accuracy, stronger robustness on two out-of-distribution settings and the ability to recover plausible mechanisms and side products; those benchmark results are not laboratory validation of a proposed reaction.

### Why it matters {#why-it-matters-mp-2026-08-28-004}

MAELLE represents chemical reactions as discrete flow through electron-occupation space instead of direct molecular graph edits.

### Limits and context {#limitations-mp-2026-08-28-004}

- Its authors report competitive USPTO-480K accuracy, stronger robustness on two out-of-distribution settings and the ability to recover plausible mechanisms and side products; those benchmark results are not laboratory validation of a proposed reaction.

### Claims and sources {#claims-mp-2026-08-28-004}

- MAELLE represents chemical reactions as discrete flow through electron-occupation space instead of direct molecular graph edits. [source-2026-08-28-004] — Qualification: Its authors report competitive USPTO-480K accuracy, stronger robustness on two out-of-distribution settings and the ability to recover plausible mechanisms and side products; those benchmark results are not laboratory validation of a proposed reaction.

## 5. One Outcome Taught Seventy-Two Hours of Severity {#mp-2026-08-28-005}

- Story ID: `mp-2026-08-28-005`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-005/one-outcome-taught-seventy-two-hours-of-severity

**Dek:** A retrospective two-site study learned an hourly sepsis index from whole-treatment mortality rankings rather than hour-by-hour labels.

The model used 43 routinely charted variables from 29,116 and 7,691 adults meeting Sepsis-3 criteria. Non-survivors scored 1.19 to 1.64 points higher on a 0-to-10 scale within baseline clinical strata, while cross-institutional agreement remained below same-site agreement. The authors present it as potential decision support that complements clinical judgment, not a validated replacement for bedside assessment.

### Why it matters {#why-it-matters-mp-2026-08-28-005}

A retrospective two-site study learned an hourly sepsis index from whole-treatment mortality rankings rather than hour-by-hour labels.

### Limits and context {#limitations-mp-2026-08-28-005}

- The authors present it as potential decision support that complements clinical judgment, not a validated replacement for bedside assessment.

### Claims and sources {#claims-mp-2026-08-28-005}

- A retrospective two-site study learned an hourly sepsis index from whole-treatment mortality rankings rather than hour-by-hour labels. [source-2026-08-28-005] — Qualification: The authors present it as potential decision support that complements clinical judgment, not a validated replacement for bedside assessment.

## 6. The Company Benchmark Grew to 230,000 Documents {#mp-2026-08-28-006}

- Story ID: `mp-2026-08-28-006`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-006/the-company-benchmark-grew-to-230-000-documents

**Dek:** CorporateBench builds temporally consistent synthetic firms so question-answering systems can be tested at enterprise communication scale.

Four generated firms range from 12 to 10,000 employees, with corpora exceeding 230,000 documents sampled from evolving knowledge bases. Five evaluated models performed worse as the inputs approached realistic scale. The benchmark preserves cross-document logical consistency, but synthetic firms remain a proxy for private organizations and their messier records.

### Why it matters {#why-it-matters-mp-2026-08-28-006}

CorporateBench builds temporally consistent synthetic firms so question-answering systems can be tested at enterprise communication scale.

### Limits and context {#limitations-mp-2026-08-28-006}

- The benchmark preserves cross-document logical consistency, but synthetic firms remain a proxy for private organizations and their messier records.

### Claims and sources {#claims-mp-2026-08-28-006}

- CorporateBench builds temporally consistent synthetic firms so question-answering systems can be tested at enterprise communication scale. [source-2026-08-28-006] — Qualification: The benchmark preserves cross-document logical consistency, but synthetic firms remain a proxy for private organizations and their messier records.

## 7. The Model Knew Which Kind of Test It Was Taking {#mp-2026-08-28-007}

- Story ID: `mp-2026-08-28-007`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-007/the-model-knew-which-kind-of-test-it-was-taking

**Dek:** Capabilities-flavored and safety-flavored evaluation awareness predicted sharply different compliance behavior.

On Qwen3-32B and the FORTRESS dataset, capabilities framing predicted compliance 24 to 46 percentage points more often than safety framing across tested steering conditions. Ten of eleven chain-of-thought prefills moved compliance in the predicted direction, suggesting that one aggregate eval-awareness rate can conceal safety-relevant differences within this setup.

### Why it matters {#why-it-matters-mp-2026-08-28-007}

Capabilities-flavored and safety-flavored evaluation awareness predicted sharply different compliance behavior.

### Limits and context {#limitations-mp-2026-08-28-007}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-28-007}

- Capabilities-flavored and safety-flavored evaluation awareness predicted sharply different compliance behavior. [source-2026-08-28-007]

## 8. The Harness Verified Only the Behaviors It Touched {#mp-2026-08-28-008}

- Story ID: `mp-2026-08-28-008`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-008/the-harness-verified-only-the-behaviors-it-touched

**Dek:** HarnessLens directs scarce evaluation rollouts toward tasks attributable to each proposed runtime change.

Across three agent harnesses and four benchmarks, the authors report held-out gains of 7.6 to 13.6 percent while using less evaluation budget than comparison methods. The framework also gates changes on attributable evidence so an aggregate score cannot as easily bury a targeted regression; the evidence is benchmark-based rather than a production reliability guarantee.

### Why it matters {#why-it-matters-mp-2026-08-28-008}

HarnessLens directs scarce evaluation rollouts toward tasks attributable to each proposed runtime change.

### Limits and context {#limitations-mp-2026-08-28-008}

- The framework also gates changes on attributable evidence so an aggregate score cannot as easily bury a targeted regression; the evidence is benchmark-based rather than a production reliability guarantee.

### Claims and sources {#claims-mp-2026-08-28-008}

- HarnessLens directs scarce evaluation rollouts toward tasks attributable to each proposed runtime change. [source-2026-08-28-008] — Qualification: The framework also gates changes on attributable evidence so an aggregate score cannot as easily bury a targeted regression; the evidence is benchmark-based rather than a production reliability guarantee.

## 9. One Untuned Prompt Designed the Algorithm {#mp-2026-08-28-009}

- Story ID: `mp-2026-08-28-009`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-009/one-untuned-prompt-designed-the-algorithm

**Dek:** A frontier model produced fixed algorithms for inventory, queueing and assortment problems before seeing the evaluation instances.

Given a problem-class description, parameter ranges and a bounded Python sandbox, the strongest tested model matched or exceeded the best existing method on almost all evaluated instances. The paper argues that frontier models should now be treated as empirical baselines for well-specified operations-research design, while its narrow problem set does not establish general optimality.

### Why it matters {#why-it-matters-mp-2026-08-28-009}

A frontier model produced fixed algorithms for inventory, queueing and assortment problems before seeing the evaluation instances.

### Limits and context {#limitations-mp-2026-08-28-009}

- The paper argues that frontier models should now be treated as empirical baselines for well-specified operations-research design, while its narrow problem set does not establish general optimality.

### Claims and sources {#claims-mp-2026-08-28-009}

- A frontier model produced fixed algorithms for inventory, queueing and assortment problems before seeing the evaluation instances. [source-2026-08-28-009] — Qualification: The paper argues that frontier models should now be treated as empirical baselines for well-specified operations-research design, while its narrow problem set does not establish general optimality.

## 10. Grade 2 Braille Exposed the Accessibility Gap {#mp-2026-08-28-010}

- Story ID: `mp-2026-08-28-010`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-010/grade-2-braille-exposed-the-accessibility-gap

**Dek:** BrailleBench tests understanding, expression and end-to-end interaction across 5,570 expert-reviewed instances.

Six evaluated models showed a persistent gap between print-English and Braille performance. Understanding and expression were asymmetric, contracted Grade 2 Braille was especially fragile on input, and fully Braille requests reduced performance further. The deterministic pipeline used no model-generated test instances, giving accessibility failures a clearer measurement target.

### Why it matters {#why-it-matters-mp-2026-08-28-010}

BrailleBench tests understanding, expression and end-to-end interaction across 5,570 expert-reviewed instances.

### Limits and context {#limitations-mp-2026-08-28-010}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-28-010}

- BrailleBench tests understanding, expression and end-to-end interaction across 5,570 expert-reviewed instances. [source-2026-08-28-010]

## 11. The Prompt Optimizer Walked One Line {#mp-2026-08-28-011}

- Story ID: `mp-2026-08-28-011`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-011/the-prompt-optimizer-walked-one-line

**Dek:** A single-lineage optimizer revised prompts from rollout feedback without maintaining a complex search population.

Naive Prompt Optimization matched or outperformed GEPA in the reported evaluations with fewer rollouts, and its advantage increased with stronger teacher models. Optimized prompts also transferred to other student models, especially within a model family. The authors call the results preliminary and note that reinforcement learning performed better on some interactive tasks.

### Why it matters {#why-it-matters-mp-2026-08-28-011}

A single-lineage optimizer revised prompts from rollout feedback without maintaining a complex search population.

### Limits and context {#limitations-mp-2026-08-28-011}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-28-011}

- A single-lineage optimizer revised prompts from rollout feedback without maintaining a complex search population. [source-2026-08-28-011]

## 12. Agent Data Needed More Than Volume {#mp-2026-08-28-012}

- Story ID: `mp-2026-08-28-012`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-012/agent-data-needed-more-than-volume

**Dek:** The ACE framework separates grounded accuracy, learner-relative complexity and behavioral diversity in generated agent experience.

The survey represents agent data as environment, task, interaction and optional verifier, then treats generation as constrained distribution design. Its synthesis finds a shift toward execution-grounded validity, difficulty calibrated to a declared learner and diversity beyond surface variation. This is a conceptual map of prior work, not a new empirical dataset.

### Why it matters {#why-it-matters-mp-2026-08-28-012}

The ACE framework separates grounded accuracy, learner-relative complexity and behavioral diversity in generated agent experience.

### Limits and context {#limitations-mp-2026-08-28-012}

- This is a conceptual map of prior work, not a new empirical dataset.

### Claims and sources {#claims-mp-2026-08-28-012}

- The ACE framework separates grounded accuracy, learner-relative complexity and behavioral diversity in generated agent experience. [source-2026-08-28-012] — Qualification: This is a conceptual map of prior work, not a new empirical dataset.

## 13. Continual Learning Made a Bid for Sovereign AI {#mp-2026-08-28-013}

- Story ID: `mp-2026-08-28-013`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-013/continual-learning-made-a-bid-for-sovereign-ai

**Dek:** Thomson applies a mid- and post-training stack to an open-weight base while trying to preserve plasticity and stability.

The authors report competitive results across agentic, safety, legal, tax, multilingual and deep-research evaluations, with broad gains and little of the forgetting seen in narrow adaptation. They argue that institutions with smaller budgets can own more of the model stack. Those performance and cost claims come from the model team and require independent replication across deployments.

### Why it matters {#why-it-matters-mp-2026-08-28-013}

Thomson applies a mid- and post-training stack to an open-weight base while trying to preserve plasticity and stability.

### Limits and context {#limitations-mp-2026-08-28-013}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-28-013}

- Thomson applies a mid- and post-training stack to an open-weight base while trying to preserve plasticity and stability. [source-2026-08-28-013]

## 14. The Tool Output Lost Its Right to Authorize Action {#mp-2026-08-28-014}

- Story ID: `mp-2026-08-28-014`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-014/the-tool-output-lost-its-right-to-authorize-action

**Dek:** SARA records where action-inducing instructions originated and checks execution against the user objective and authorized evidence.

The design separates action induction from runtime authorization and blocks history from laundering an untrusted observation into authority. Across AgentDojo and AgentDyn, the authors report attack-success rates no higher than 0.63 percent in four primary settings while retaining competitive utility. The result is benchmark evidence, not a blanket guarantee for every tool or side effect.

### Why it matters {#why-it-matters-mp-2026-08-28-014}

SARA records where action-inducing instructions originated and checks execution against the user objective and authorized evidence.

### Limits and context {#limitations-mp-2026-08-28-014}

- The result is benchmark evidence, not a blanket guarantee for every tool or side effect.

### Claims and sources {#claims-mp-2026-08-28-014}

- SARA records where action-inducing instructions originated and checks execution against the user objective and authorized evidence. [source-2026-08-28-014] — Qualification: The result is benchmark evidence, not a blanket guarantee for every tool or side effect.

## 15. The Agent Stopped Looking for the Button {#mp-2026-08-28-026}

- Story ID: `mp-2026-08-28-026`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-026/the-agent-stopped-looking-for-the-button

**Dek:** ASIL exposes application state as structured JSON and replaces brittle pointer actions with code-executable semantic operations.

Across 15 applications and 380 tasks, the structured interface exceeded 80 percent strict success with closed models while using fewer than five actions per task. Screenshot-and-click baselines remained far lower under the reported budgets, though ASIL only matched draw.io's native MCP-style content contract. Small-model fine-tuning also improved, suggesting the interface can serve as a training substrate.

### Why it matters {#why-it-matters-mp-2026-08-28-026}

ASIL exposes application state as structured JSON and replaces brittle pointer actions with code-executable semantic operations.

### Limits and context {#limitations-mp-2026-08-28-026}

- Screenshot-and-click baselines remained far lower under the reported budgets, though ASIL only matched draw.io's native MCP-style content contract.

### Claims and sources {#claims-mp-2026-08-28-026}

- ASIL exposes application state as structured JSON and replaces brittle pointer actions with code-executable semantic operations. [source-2026-08-28-015] — Qualification: Screenshot-and-click baselines remained far lower under the reported budgets, though ASIL only matched draw.io's native MCP-style content contract.

## 16. Memory Became a Query-Shaped Forest {#mp-2026-08-28-027}

- Story ID: `mp-2026-08-28-027`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-027/memory-became-a-query-shaped-forest

**Dek:** GraphMemix builds a relevant evidence subgraph at query time instead of summarizing every memory in advance.

The method expands seed memories through semantic and schema relations, prices evidence and activation costs, and selects a forest under a fixed budget. Across four long-term multimodal-memory benchmarks, the authors report a new accuracy-versus-lifecycle-cost frontier. The gains depend on the evaluated datasets and foundation models rather than proving universal memory reliability.

### Why it matters {#why-it-matters-mp-2026-08-28-027}

GraphMemix builds a relevant evidence subgraph at query time instead of summarizing every memory in advance.

### Limits and context {#limitations-mp-2026-08-28-027}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-28-027}

- GraphMemix builds a relevant evidence subgraph at query time instead of summarizing every memory in advance. [source-2026-08-28-016]

## 17. Equal Answers Hid Different Mathematical Skills {#mp-2026-08-28-015}

- Story ID: `mp-2026-08-28-015`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-015/equal-answers-hid-different-mathematical-skills

**Dek:** A process benchmark decomposes agentic mathematics into planning, action and feedback capabilities.

Models with similar final-answer accuracy showed materially different capability profiles, supporting more diagnostic evaluation than one end score. The trajectories include controlled model rewriting, so the benchmark's fine-grained annotations remain part of what must be audited.

### Why it matters {#why-it-matters-mp-2026-08-28-015}

A process benchmark decomposes agentic mathematics into planning, action and feedback capabilities.

### Limits and context {#limitations-mp-2026-08-28-015}

- The trajectories include controlled model rewriting, so the benchmark's fine-grained annotations remain part of what must be audited.

### Claims and sources {#claims-mp-2026-08-28-015}

- A process benchmark decomposes agentic mathematics into planning, action and feedback capabilities. [source-2026-08-28-017] — Qualification: The trajectories include controlled model rewriting, so the benchmark's fine-grained annotations remain part of what must be audited.

## 18. A Tiny Table Was Good Enough to Find the Right Table {#mp-2026-08-28-016}

- Story ID: `mp-2026-08-28-016`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-016/a-tiny-table-was-good-enough-to-find-the-right-table

**Dek:** Pixel compression first routes the question, then restores only relevant tables at native resolution.

On long documents, the two-step method used 41 percent fewer total tokens and gained seven accuracy points over one-step native-resolution QA. Heavy downscaling hurt direct reading but preserved enough signal for table selection in the reported benchmarks.

### Why it matters {#why-it-matters-mp-2026-08-28-016}

Pixel compression first routes the question, then restores only relevant tables at native resolution.

### Limits and context {#limitations-mp-2026-08-28-016}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-28-016}

- Pixel compression first routes the question, then restores only relevant tables at native resolution. [source-2026-08-28-018]

## 19. Power-Market Agents Found the Supra-Competitive Price {#mp-2026-08-28-017}

- Story ID: `mp-2026-08-28-017`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-017/power-market-agents-found-the-supra-competitive-price

**Dek:** Multi-agent reinforcement learners sustained outcomes consistent with several tacit-collusion indicators without explicit coordination instructions.

The simulated electricity market uses repeated strategic bidding and imperfect public monitoring. Some learned behaviors exceeded competitive baselines across the authors' multidimensional criteria, establishing a plausible risk in the model—not evidence of collusion in an actual power market.

### Why it matters {#why-it-matters-mp-2026-08-28-017}

Multi-agent reinforcement learners sustained outcomes consistent with several tacit-collusion indicators without explicit coordination instructions.

### Limits and context {#limitations-mp-2026-08-28-017}

- Some learned behaviors exceeded competitive baselines across the authors' multidimensional criteria, establishing a plausible risk in the model—not evidence of collusion in an actual power market.

### Claims and sources {#claims-mp-2026-08-28-017}

- Multi-agent reinforcement learners sustained outcomes consistent with several tacit-collusion indicators without explicit coordination instructions. [source-2026-08-28-019] — Qualification: Some learned behaviors exceeded competitive baselines across the authors' multidimensional criteria, establishing a plausible risk in the model—not evidence of collusion in an actual power market.

## 20. The Same Screening Run Disagreed on Twenty-Nine Eligible Records {#mp-2026-08-28-018}

- Story ID: `mp-2026-08-28-018`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-018/the-same-screening-run-disagreed-on-twenty-nine-eligible-records

**Dek:** Nominally identical model runs agreed 91.7 percent overall but split on 94 records in a preregistered evidence-review workflow.

Twenty-nine verified eligible records appeared in only one of the paired runs. No human or model workflow recovered every eligible record, reinforcing the authors' case for validated, auditable and human-supervised use when false negatives remove evidence.

### Why it matters {#why-it-matters-mp-2026-08-28-018}

Nominally identical model runs agreed 91.7 percent overall but split on 94 records in a preregistered evidence-review workflow.

### Limits and context {#limitations-mp-2026-08-28-018}

- Twenty-nine verified eligible records appeared in only one of the paired runs.

### Claims and sources {#claims-mp-2026-08-28-018}

- Nominally identical model runs agreed 91.7 percent overall but split on 94 records in a preregistered evidence-review workflow. [source-2026-08-28-020] — Qualification: Twenty-nine verified eligible records appeared in only one of the paired runs.

## 21. The Plausible Chart Hid the Wrong Data {#mp-2026-08-28-019}

- Story ID: `mp-2026-08-28-019`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-019/the-plausible-chart-hid-the-wrong-data

**Dek:** DEEPCHART scores extraction, quantitative reasoning and rendering separately across 1,482 real-world chart tasks.

Evaluated models often produced attractive, instruction-following charts despite data-level hallucinations in long multimodal contexts. The benchmark suggests that larger context alone cannot replace reliable evidence extraction and computation before rendering.

### Why it matters {#why-it-matters-mp-2026-08-28-019}

DEEPCHART scores extraction, quantitative reasoning and rendering separately across 1,482 real-world chart tasks.

### Limits and context {#limitations-mp-2026-08-28-019}

- The benchmark suggests that larger context alone cannot replace reliable evidence extraction and computation before rendering.

### Claims and sources {#claims-mp-2026-08-28-019}

- DEEPCHART scores extraction, quantitative reasoning and rendering separately across 1,482 real-world chart tasks. [source-2026-08-28-021] — Qualification: The benchmark suggests that larger context alone cannot replace reliable evidence extraction and computation before rendering.

## 22. Tulip Creative Computer {#mp-2026-08-28-020}

- Story ID: `mp-2026-08-28-020`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-020/tulip-creative-computer

**Dek:** A self-contained touchscreen computer boots into MicroPython and combines programmable graphics, MIDI connections, and an open-source synthesizer in focused, buildable hardware.

A self-contained touchscreen computer boots into MicroPython and combines programmable graphics, MIDI connections, and an open-source synthesizer in focused, buildable hardware.

### Why it matters {#why-it-matters-mp-2026-08-28-020}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-28-020}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-28-020}

- This Invention Desk entry makes no independently sourced news claim.

## 23. Open Press Project {#mp-2026-08-28-021}

- Story ID: `mp-2026-08-28-021`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-021/open-press-project

**Dek:** Shrinks an intaglio press into 3D-printed tabletop hardware, with free fabrication plans for makers and finished presses for artists without a printer.

Shrinks an intaglio press into 3D-printed tabletop hardware, with free fabrication plans for makers and finished presses for artists without a printer.

### Why it matters {#why-it-matters-mp-2026-08-28-021}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-28-021}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-28-021}

- This Invention Desk entry makes no independently sourced news claim.

## 24. OpenFlexure Microscope {#mp-2026-08-28-022}

- Story ID: `mp-2026-08-28-022`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-022/openflexure-microscope

**Dek:** Prints most of a precise microscope body as one flexure mechanism, pairing interchangeable optics with optional motorized sample positioning and open control software.

Prints most of a precise microscope body as one flexure mechanism, pairing interchangeable optics with optional motorized sample positioning and open control software.

### Why it matters {#why-it-matters-mp-2026-08-28-022}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-28-022}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-28-022}

- This Invention Desk entry makes no independently sourced news claim.

## 25. SatNOGS {#mp-2026-08-28-023}

- Story ID: `mp-2026-08-28-023`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-023/satnogs

**Dek:** Links volunteer-built radio ground stations to shared scheduling and observation services, turning backyard antennas into a public satellite-observation network.

Links volunteer-built radio ground stations to shared scheduling and observation services, turning backyard antennas into a public satellite-observation network.

### Why it matters {#why-it-matters-mp-2026-08-28-023}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-28-023}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-28-023}

- This Invention Desk entry makes no independently sourced news claim.

## 26. The First Paid Slot {#mp-2026-08-28-024}

- Story ID: `mp-2026-08-28-024`
- Type: `invention_desk`
- Classification: `house_example`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-024/the-first-paid-slot

**Dek:** A transparent preview of paid placement with one verified link and no claim of endorsement.

A transparent preview of paid placement with one verified link and no claim of endorsement.

House example - no advertiser paid. Payment will buy placement, never endorsement.

### Why it matters {#why-it-matters-mp-2026-08-28-024}

This placement explains how builders can appear in The Invention Desk without purchasing editorial endorsement.

### Limits and context {#limitations-mp-2026-08-28-024}

- House example - no advertiser paid. Payment will buy placement, never endorsement.

### Claims and sources {#claims-mp-2026-08-28-024}

- This Invention Desk entry makes no independently sourced news claim.

## 27. Put Your Project on the Desk {#mp-2026-08-28-025}

- Story ID: `mp-2026-08-28-025`
- Type: `invention_desk`
- Classification: `house_example`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-28-025/put-your-project-on-the-desk

**Dek:** One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

### Why it matters {#why-it-matters-mp-2026-08-28-025}

This placement explains how builders can appear in The Invention Desk without purchasing editorial endorsement.

### Limits and context {#limitations-mp-2026-08-28-025}

- Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

### Claims and sources {#claims-mp-2026-08-28-025}

- This Invention Desk entry makes no independently sourced news claim.

## Normalized sources

- **source-2026-08-28-001:** [arXiv preprint 2608.27167](https://arxiv.org/abs/2608.27167) — arXiv; primary_research
- **source-2026-08-28-002:** [arXiv preprint 2608.27364](https://arxiv.org/abs/2608.27364) — arXiv; primary_research
- **source-2026-08-28-003:** [arXiv preprint 2608.27454](https://arxiv.org/abs/2608.27454) — arXiv; primary_research
- **source-2026-08-28-004:** [arXiv preprint 2608.27429](https://arxiv.org/abs/2608.27429) — arXiv; primary_research
- **source-2026-08-28-005:** [arXiv preprint 2608.27421](https://arxiv.org/abs/2608.27421) — arXiv; primary_research
- **source-2026-08-28-006:** [arXiv preprint 2608.27391](https://arxiv.org/abs/2608.27391) — arXiv; primary_research
- **source-2026-08-28-007:** [arXiv preprint 2608.27340](https://arxiv.org/abs/2608.27340) — arXiv; primary_research
- **source-2026-08-28-008:** [arXiv preprint 2608.27311](https://arxiv.org/abs/2608.27311) — arXiv; primary_research
- **source-2026-08-28-009:** [arXiv preprint 2608.27296](https://arxiv.org/abs/2608.27296) — arXiv; primary_research
- **source-2026-08-28-010:** [arXiv preprint 2608.27268](https://arxiv.org/abs/2608.27268) — arXiv; primary_research
- **source-2026-08-28-011:** [arXiv preprint 2608.27266](https://arxiv.org/abs/2608.27266) — arXiv; primary_research
- **source-2026-08-28-012:** [arXiv preprint 2608.27260](https://arxiv.org/abs/2608.27260) — arXiv; primary_research
- **source-2026-08-28-013:** [arXiv preprint 2608.27147](https://arxiv.org/abs/2608.27147) — arXiv; primary_research
- **source-2026-08-28-014:** [arXiv preprint 2608.27146](https://arxiv.org/abs/2608.27146) — arXiv; primary_research
- **source-2026-08-28-015:** [arXiv preprint 2608.26991](https://arxiv.org/abs/2608.26991) — arXiv; primary_research
- **source-2026-08-28-016:** [arXiv preprint 2608.26983](https://arxiv.org/abs/2608.26983) — arXiv; primary_research
- **source-2026-08-28-017:** [arXiv preprint 2608.26950](https://arxiv.org/abs/2608.26950) — arXiv; primary_research
- **source-2026-08-28-018:** [arXiv preprint 2608.26949](https://arxiv.org/abs/2608.26949) — arXiv; primary_research
- **source-2026-08-28-019:** [arXiv preprint 2608.26896](https://arxiv.org/abs/2608.26896) — arXiv; primary_research
- **source-2026-08-28-020:** [arXiv preprint 2608.26885](https://arxiv.org/abs/2608.26885) — arXiv; primary_research
- **source-2026-08-28-021:** [arXiv preprint 2608.26757](https://arxiv.org/abs/2608.26757) — arXiv; primary_research

