---
schema_version: "1.0.0"
edition_id: "mp-2026-08-25-morning-0047"
published_at: "2026-08-25T09:00:00.000-04:00"
modified_at: "2026-08-25T09:00:00.000-04:00"
canonical_url: "https://themachinepress.com/edition/2026-08-25"
story_count: 27
lead_story_id: "mp-2026-08-25-001"
---

# The Machine Press — Morning edition

Edition ID: `mp-2026-08-25-morning-0047`  
Published: 2026-08-25T09:00:00.000-04:00  
Canonical edition: https://themachinepress.com/edition/2026-08-25

A controlled logic-puzzle experiment found that cheaper AI help increased use—and that assisted performance overstated what participants could later do alone.

## 1. The Shortcut Improved the Score and Weakened the Skill {#mp-2026-08-25-001}

- Story ID: `mp-2026-08-25-001`
- Type: `lead`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-001/the-shortcut-improved-the-score-and-weakened-the-skill

**Dek:** A controlled logic-puzzle experiment found that cheaper AI help increased use—and that assisted performance overstated what participants could later do alone.

Participants worked through logic puzzles before, during and after access to on-demand AI assistance. Randomly lowering the request cost increased how often people called for help. Those who requested assistance during the access phase performed worse after it was removed, while a Bayesian latent-ability model associated more independent reasoning with larger skill gains. The experiment does not establish that every use of AI impairs learning, but it shows why an assisted score can be a poor proxy for the skill a person retains.

### Why it matters {#why-it-matters-mp-2026-08-25-001}

A controlled logic-puzzle experiment found that cheaper AI help increased use—and that assisted performance overstated what participants could later do alone.

### Limits and context {#limitations-mp-2026-08-25-001}

- The experiment does not establish that every use of AI impairs learning, but it shows why an assisted score can be a poor proxy for the skill a person retains.

### Claims and sources {#claims-mp-2026-08-25-001}

- A controlled logic-puzzle experiment found that cheaper AI help increased use—and that assisted performance overstated what participants could later do alone. [source-2026-08-25-001] — Qualification: The experiment does not establish that every use of AI impairs learning, but it shows why an assisted score can be a poor proxy for the skill a person retains.

## 2. The Design Loop Ended at an External Gate {#mp-2026-08-25-002}

- Story ID: `mp-2026-08-25-002`
- Type: `secondary`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-002/the-design-loop-ended-at-an-external-gate

**Dek:** A closed-loop agent connected language models to deterministic engineering solvers, then sent its floating-wind design to a classification society.

The AI Engineer turns natural-language requirements into geometry and a mesh, then couples topology and member-size searches to structural, aero-hydro-servo-elastic and cost checks. Its automated reviewer stops only when capacity, steel intensity, unit cost, constructability and fatigue scores clear stated floors. The top floating-wind support design passed China Classification Society Approval in Principle and used 8.1 percent less steel and 8.1 percent lower unit capital cost than the human-optimized TuQiang baseline. Approval in Principle is an external design-stage check, not final construction certification, and the authors list detailed design and fabrication constraints as remaining work.

### Why it matters {#why-it-matters-mp-2026-08-25-002}

A closed-loop agent connected language models to deterministic engineering solvers, then sent its floating-wind design to a classification society.

### Limits and context {#limitations-mp-2026-08-25-002}

- Its automated reviewer stops only when capacity, steel intensity, unit cost, constructability and fatigue scores clear stated floors.
- Approval in Principle is an external design-stage check, not final construction certification, and the authors list detailed design and fabrication constraints as remaining work.

### Claims and sources {#claims-mp-2026-08-25-002}

- A closed-loop agent connected language models to deterministic engineering solvers, then sent its floating-wind design to a classification society. [source-2026-08-25-002] — Qualification: Its automated reviewer stops only when capacity, steel intensity, unit cost, constructability and fatigue scores clear stated floors.

## 3. The Best Agent Finished the Pieces and Lost the Chain {#mp-2026-08-25-003}

- Story ID: `mp-2026-08-25-003`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-003/the-best-agent-finished-the-pieces-and-lost-the-chain

**Dek:** EarthVerse tests 405 reproducible investigations across 199 documented events and 19 natural-hazard families.

The strongest system reached 84.65 percent mean answer-unit accuracy, but only 34.81 percent Strict@95. The gap shows that agents can complete many local steps without preserving a consistent chain across evidence, units, calculations and physical interpretation.

### Why it matters {#why-it-matters-mp-2026-08-25-003}

EarthVerse tests 405 reproducible investigations across 199 documented events and 19 natural-hazard families.

### Limits and context {#limitations-mp-2026-08-25-003}

- The strongest system reached 84.65 percent mean answer-unit accuracy, but only 34.81 percent Strict@95.

### Claims and sources {#claims-mp-2026-08-25-003}

- EarthVerse tests 405 reproducible investigations across 199 documented events and 19 natural-hazard families. [source-2026-08-25-003] — Qualification: The strongest system reached 84.65 percent mean answer-unit accuracy, but only 34.81 percent Strict@95.

## 4. The Skill Learned What the Brief Forgot {#mp-2026-08-25-004}

- Story ID: `mp-2026-08-25-004`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-004/the-skill-learned-what-the-brief-forgot

**Dek:** SkillAlchemy searches open-world sources for requirements omitted by an underspecified skill brief.

Across 87 SkillsBench tasks, its source-grounded skill packages improved pass rate 19.9 points over no-skill execution and 8.6 points over the strongest automated baseline, while performing comparably to human-curated skills. The system admits procedures only at the scope supported by evidence.

### Why it matters {#why-it-matters-mp-2026-08-25-004}

SkillAlchemy searches open-world sources for requirements omitted by an underspecified skill brief.

### Limits and context {#limitations-mp-2026-08-25-004}

- The system admits procedures only at the scope supported by evidence.

### Claims and sources {#claims-mp-2026-08-25-004}

- SkillAlchemy searches open-world sources for requirements omitted by an underspecified skill brief. [source-2026-08-25-004] — Qualification: The system admits procedures only at the scope supported by evidence.

## 5. The World Model Learned Energy and Let It Drift {#mp-2026-08-25-005}

- Story ID: `mp-2026-08-25-005`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-005/the-world-model-learned-energy-and-let-it-drift

**Dek:** A frozen video world model encoded an energy-like invariant that its own imagined rollouts failed to preserve.

A label-free search recovered the same scalar across independently trained conservative pendulum models and found no comparable quantity in damped controls. Projecting latent state back toward the initial level set reduced rollout error in all three conservative models; matched random constraints usually made it worse.

### Why it matters {#why-it-matters-mp-2026-08-25-005}

A frozen video world model encoded an energy-like invariant that its own imagined rollouts failed to preserve.

### Limits and context {#limitations-mp-2026-08-25-005}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-005}

- A frozen video world model encoded an energy-like invariant that its own imagined rollouts failed to preserve. [source-2026-08-25-005]

## 6. Reasoning Training Moved Along the Safety Direction {#mp-2026-08-25-006}

- Story ID: `mp-2026-08-25-006`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-006/reasoning-training-moved-along-the-safety-direction

**Dek:** A representation-space audit links some reasoning fine-tunes to safety shifts and proposes a penalty at the implicated layers.

The authors stress that reasoning-induced misalignment does not appear across every architecture, scale or dataset. On Qwen2.5 3B and 7B experiments, a learned safety-direction penalty restored measured safety while preserving benchmark reasoning performance, with diagnostics guiding which layers to include.

### Why it matters {#why-it-matters-mp-2026-08-25-006}

A representation-space audit links some reasoning fine-tunes to safety shifts and proposes a penalty at the implicated layers.

### Limits and context {#limitations-mp-2026-08-25-006}

- The authors stress that reasoning-induced misalignment does not appear across every architecture, scale or dataset.

### Claims and sources {#claims-mp-2026-08-25-006}

- A representation-space audit links some reasoning fine-tunes to safety shifts and proposes a penalty at the implicated layers. [source-2026-08-25-006] — Qualification: The authors stress that reasoning-induced misalignment does not appear across every architecture, scale or dataset.

## 7. The Clinical Agent Had to Earn the Next Step {#mp-2026-08-25-007}

- Story ID: `mp-2026-08-25-007`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-007/the-clinical-agent-had-to-earn-the-next-step

**Dek:** MediSkill-Evo separates clinical skills, process rules, schemas and measurements behind publication and safety gates.

On 300 held-out simulated Qwen encounters, the complete system raised diagnosis accuracy from 61.33 to 69 percent and treatment-intent coverage from 33.62 to 66.44 percent while halving automatically scored critical failures relative to AgentClinic. The authors explicitly describe this as fixed-suite system evidence, not clinical validation.

### Why it matters {#why-it-matters-mp-2026-08-25-007}

MediSkill-Evo separates clinical skills, process rules, schemas and measurements behind publication and safety gates.

### Limits and context {#limitations-mp-2026-08-25-007}

- The authors explicitly describe this as fixed-suite system evidence, not clinical validation.

### Claims and sources {#claims-mp-2026-08-25-007}

- MediSkill-Evo separates clinical skills, process rules, schemas and measurements behind publication and safety gates. [source-2026-08-25-007] — Qualification: The authors explicitly describe this as fixed-suite system evidence, not clinical validation.

## 8. The Harness Stopped Turning Recovery Into a Model Failure {#mp-2026-08-25-008}

- Story ID: `mp-2026-08-25-008`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-008/the-harness-stopped-turning-recovery-into-a-model-failure

**Dek:** Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories.

The open-source harness pairs a persistent IPython environment with standardized execution, recovery, verification and resource accounting. Its reported evaluations include an ARC-AGI-3 RHAE Best@1 increase from 30 to 95.5 percent and competitive results across coding, kernel and emulator tasks; those are harness-specific benchmarks, not general capability proof.

### Why it matters {#why-it-matters-mp-2026-08-25-008}

Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories.

### Limits and context {#limitations-mp-2026-08-25-008}

- Its reported evaluations include an ARC-AGI-3 RHAE Best@1 increase from 30 to 95.5 percent and competitive results across coding, kernel and emulator tasks; those are harness-specific benchmarks, not general capability proof.

### Claims and sources {#claims-mp-2026-08-25-008}

- Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories. [source-2026-08-25-008] — Qualification: Its reported evaluations include an ARC-AGI-3 RHAE Best@1 increase from 30 to 95.5 percent and competitive results across coding, kernel and emulator tasks; those are harness-specific benchmarks, not general capability proof.

## 9. The Model Turned Its Own Failure Into the Training Signal {#mp-2026-08-25-009}

- Story ID: `mp-2026-08-25-009`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-009/the-model-turned-its-own-failure-into-the-training-signal

**Dek:** SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.

With a Qwen3-8B base, the authors report 73.3 percent on AIME 2024 using 8 percent of the training FLOPs of scaled supervised fine-tuning, alongside gains on WebShop, ALFWorld and SWE-Bench-Lite. The method uses reflection-conditioned teacher scores without a separate critic or larger teacher model.

### Why it matters {#why-it-matters-mp-2026-08-25-009}

SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.

### Limits and context {#limitations-mp-2026-08-25-009}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-009}

- SRPO converts completed trajectories into concise reflection patches and dense token-level supervision. [source-2026-08-25-009]

## 10. The Rubric Became a Program {#mp-2026-08-25-010}

- Story ID: `mp-2026-08-25-010`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-010/the-rubric-became-a-program

**Dek:** ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge.

Across HealthBench, HelpSteer and ArgQuality, executable rubrics matched or improved natural-language rubric baselines at best preference accuracies of 53, 78 and 92 percent while cutting latency by as much as 320 times. The approach makes dependencies, penalties and override conditions explicit and editable.

### Why it matters {#why-it-matters-mp-2026-08-25-010}

ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge.

### Limits and context {#limitations-mp-2026-08-25-010}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-010}

- ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge. [source-2026-08-25-010]

## 11. The Old Screenshot Returned Only When It Mattered {#mp-2026-08-25-011}

- Story ID: `mp-2026-08-25-011`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-011/the-old-screenshot-returned-only-when-it-mattered

**Dek:** CausalCache spends a fixed visual-memory budget on the past events with the highest conditional utility.

On OSWorld-Verified, restoring selected history images was worth about 13 success points over summary-only memory, although same-budget allocation methods were indistinguishable on desktop. A zero-shot mobile test produced a 3.7-point overall gain and an 8.6-point gain on the predeclared memory-critical split.

### Why it matters {#why-it-matters-mp-2026-08-25-011}

CausalCache spends a fixed visual-memory budget on the past events with the highest conditional utility.

### Limits and context {#limitations-mp-2026-08-25-011}

- On OSWorld-Verified, restoring selected history images was worth about 13 success points over summary-only memory, although same-budget allocation methods were indistinguishable on desktop.

### Claims and sources {#claims-mp-2026-08-25-011}

- CausalCache spends a fixed visual-memory budget on the past events with the highest conditional utility. [source-2026-08-25-011] — Qualification: On OSWorld-Verified, restoring selected history images was worth about 13 success points over summary-only memory, although same-budget allocation methods were indistinguishable on desktop.

## 12. A Useful Skill Harmed the Coalition {#mp-2026-08-25-012}

- Story ID: `mp-2026-08-25-012`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-012/a-useful-skill-harmed-the-coalition

**Dek:** Skill-bank audits found coalition pollution and utility reversals after cross-domain transfer.

The study uses sampled Shapley marginals to select skills and an unlabeled target-domain mask to suppress harmful transfers. Across LoCoMo, LongMemEval, HotpotQA and ALFWorld, the interventions improved performance and generalization while showing why isolation tests can miss a skill that damages the surrounding bank.

### Why it matters {#why-it-matters-mp-2026-08-25-012}

Skill-bank audits found coalition pollution and utility reversals after cross-domain transfer.

### Limits and context {#limitations-mp-2026-08-25-012}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-012}

- Skill-bank audits found coalition pollution and utility reversals after cross-domain transfer. [source-2026-08-25-012]

## 13. Stable Tokens Stopped Paying for More Denoising {#mp-2026-08-25-013}

- Story ID: `mp-2026-08-25-013`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-013/stable-tokens-stopped-paying-for-more-denoising

**Dek:** CAI-DLLM uses first-step confidence to commit easy tokens and reserve later denoising for harder ones.

The training-free method reports up to 18.2-times wall-clock speedup on LLaDA GSM8K and 13.1-times on Dream HumanEval with slightly higher measured accuracy in those settings. On harder tasks speedups reached 44.8 times with a largest 4.4-point accuracy drop, exposing the quality boundary.

### Why it matters {#why-it-matters-mp-2026-08-25-013}

CAI-DLLM uses first-step confidence to commit easy tokens and reserve later denoising for harder ones.

### Limits and context {#limitations-mp-2026-08-25-013}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-013}

- CAI-DLLM uses first-step confidence to commit easy tokens and reserve later denoising for harder ones. [source-2026-08-25-013]

## 14. The Middle of the Results Page Became the Blind Spot {#mp-2026-08-25-014}

- Story ID: `mp-2026-08-25-014`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-014/the-middle-of-the-results-page-became-the-blind-spot

**Dek:** Randomized hotel listings show that shopping agents inspect deeper than people but retain a weak, non-monotonic position bias.

Across 5,000 sessions and four language models, the middle of a 100-result page was least likely to be inspected—not the bottom. Position reached the choice stage for some models but not others, while all selected the same undominated listing; displayed attributes mattered more than rank.

### Why it matters {#why-it-matters-mp-2026-08-25-014}

Randomized hotel listings show that shopping agents inspect deeper than people but retain a weak, non-monotonic position bias.

### Limits and context {#limitations-mp-2026-08-25-014}

- Across 5,000 sessions and four language models, the middle of a 100-result page was least likely to be inspected—not the bottom.
- Position reached the choice stage for some models but not others, while all selected the same undominated listing; displayed attributes mattered more than rank.

### Claims and sources {#claims-mp-2026-08-25-014}

- Randomized hotel listings show that shopping agents inspect deeper than people but retain a weak, non-monotonic position bias. [source-2026-08-25-014] — Qualification: Across 5,000 sessions and four language models, the middle of a 100-result page was least likely to be inspected—not the bottom.

## 15. The World Came Back to the Place It Had Shown {#mp-2026-08-25-026}

- Story ID: `mp-2026-08-25-026`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-026/the-world-came-back-to-the-place-it-had-shown

**Dek:** ReWorld combines local attention with a pose-indexed landmark bank to revisit scenes under a fixed memory budget.

The interactive world model streams 704-by-1280 video in a four-step real-time mode and retrieves landmarks nearest the current pose. In 64-second out-and-back tests, a 12-chunk cache regenerated the starting view after a sliding window had evicted the evidence and full-KV attention exhausted memory.

### Why it matters {#why-it-matters-mp-2026-08-25-026}

ReWorld combines local attention with a pose-indexed landmark bank to revisit scenes under a fixed memory budget.

### Limits and context {#limitations-mp-2026-08-25-026}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-026}

- ReWorld combines local attention with a pose-indexed landmark bank to revisit scenes under a fixed memory budget. [source-2026-08-25-015]

## 16. The Repository Changed and the Skill Stayed Quiet {#mp-2026-08-25-027}

- Story ID: `mp-2026-08-25-027`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-027/the-repository-changed-and-the-skill-stayed-quiet

**Dek:** Every one of 105 selected release transitions invalidated part of the prior repository skill set.

Across 57 repositories, six frontier agents reached only 29.9 to 69.7 percent avg@3 macro F1 when asked to update V1 skills from official V1-to-V2 patches. Missed files left stale instructions intact, while broader edits raised recall at the cost of precision. This final text-only report fills the front rail beneath the feature.

### Why it matters {#why-it-matters-mp-2026-08-25-027}

Every one of 105 selected release transitions invalidated part of the prior repository skill set.

### Limits and context {#limitations-mp-2026-08-25-027}

- Across 57 repositories, six frontier agents reached only 29.9 to 69.7 percent avg@3 macro F1 when asked to update V1 skills from official V1-to-V2 patches.
- This final text-only report fills the front rail beneath the feature.

### Claims and sources {#claims-mp-2026-08-25-027}

- Every one of 105 selected release transitions invalidated part of the prior repository skill set. [source-2026-08-25-016] — Qualification: Across 57 repositories, six frontier agents reached only 29.9 to 69.7 percent avg@3 macro F1 when asked to update V1 skills from official V1-to-V2 patches.

## 17. The Blame Test Started With Tamper-Evident Logs {#mp-2026-08-25-015}

- Story ID: `mp-2026-08-25-015`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-015/the-blame-test-started-with-tamper-evident-logs

**Dek:** AUDITA pairs authenticated command records with graded causal attribution for multi-agent failures.

The preprint reports roughly threefold lower responsibility error than a judge baseline and formal limits on what evidence can certify, including overdetermination and omissions.

### Why it matters {#why-it-matters-mp-2026-08-25-015}

AUDITA pairs authenticated command records with graded causal attribution for multi-agent failures.

### Limits and context {#limitations-mp-2026-08-25-015}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-015}

- AUDITA pairs authenticated command records with graded causal attribution for multi-agent failures. [source-2026-08-25-017]

## 18. Privacy Had to Survive the Whole Writing Bundle {#mp-2026-08-25-016}

- Story ID: `mp-2026-08-25-016`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-016/privacy-had-to-survive-the-whole-writing-bundle

**Dek:** AAST jointly selects synthetic texts against account-level authorship re-identification.

Across same-genre, cross-genre, neural and non-neural attacks, the method reduced bundle-level linkability as the number of released texts grew while preserving measured meaning, acceptability and sentiment.

### Why it matters {#why-it-matters-mp-2026-08-25-016}

AAST jointly selects synthetic texts against account-level authorship re-identification.

### Limits and context {#limitations-mp-2026-08-25-016}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-016}

- AAST jointly selects synthetic texts against account-level authorship re-identification. [source-2026-08-25-018]

## 19. Tool Waiting Moved Off the GPU's Critical Path {#mp-2026-08-25-017}

- Story ID: `mp-2026-08-25-017`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-017/tool-waiting-moved-off-the-gpu-s-critical-path

**Dek:** MCP-Universe RL provisions isolated tool environments and overlaps long trajectories stalled on external calls.

One configuration trained software-engineering, deep-research and general tool-use agents on gpt-oss-20b, improving task reward in all three domains without domain-specific RL integration code.

### Why it matters {#why-it-matters-mp-2026-08-25-017}

MCP-Universe RL provisions isolated tool environments and overlaps long trajectories stalled on external calls.

### Limits and context {#limitations-mp-2026-08-25-017}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-017}

- MCP-Universe RL provisions isolated tool environments and overlaps long trajectories stalled on external calls. [source-2026-08-25-019]

## 20. Router Agreement Picked the Patch {#mp-2026-08-25-018}

- Story ID: `mp-2026-08-25-018`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-018/router-agreement-picked-the-patch

**Dek:** Risa uses native mixture-of-experts routing traces to diversify attempts and select a software patch.

On SWE-bench Verified, routing arbitration raised the gpt-oss macro resolved rate from 44.9 to 48.2 percent and matched text consensus without comparing answer strings.

### Why it matters {#why-it-matters-mp-2026-08-25-018}

Risa uses native mixture-of-experts routing traces to diversify attempts and select a software patch.

### Limits and context {#limitations-mp-2026-08-25-018}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-018}

- Risa uses native mixture-of-experts routing traces to diversify attempts and select a software patch. [source-2026-08-25-020]

## 21. A Strategy Helped Only When Its Executor Could Use It {#mp-2026-08-25-019}

- Story ID: `mp-2026-08-25-019`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-019/a-strategy-helped-only-when-its-executor-could-use-it

**Dek:** StrategyBench separates the quality of an induced rule from its downstream usefulness.

Across strategy-inducible BIG-Bench tasks, explicit strategy utility varied by category and depended on both generator and executor, demonstration design and adaptation setting.

### Why it matters {#why-it-matters-mp-2026-08-25-019}

StrategyBench separates the quality of an induced rule from its downstream usefulness.

### Limits and context {#limitations-mp-2026-08-25-019}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-25-019}

- StrategyBench separates the quality of an induced rule from its downstream usefulness. [source-2026-08-25-021]

## 22. Tulip Creative Computer {#mp-2026-08-25-020}

- Story ID: `mp-2026-08-25-020`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-020/tulip-creative-computer

**Dek:** A self-contained touchscreen computer boots into MicroPython and combines programmable graphics, MIDI connections, and an open-source synthesizer in focused, buildable hardware.

A self-contained touchscreen computer boots into MicroPython and combines programmable graphics, MIDI connections, and an open-source synthesizer in focused, buildable hardware.

### Why it matters {#why-it-matters-mp-2026-08-25-020}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-25-020}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-25-020}

- This Invention Desk entry makes no independently sourced news claim.

## 23. Open Press Project {#mp-2026-08-25-021}

- Story ID: `mp-2026-08-25-021`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-021/open-press-project

**Dek:** Shrinks an intaglio press into 3D-printed tabletop hardware, with free fabrication plans for makers and finished presses for artists without a printer.

Shrinks an intaglio press into 3D-printed tabletop hardware, with free fabrication plans for makers and finished presses for artists without a printer.

### Why it matters {#why-it-matters-mp-2026-08-25-021}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-25-021}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-25-021}

- This Invention Desk entry makes no independently sourced news claim.

## 24. OpenFlexure Microscope {#mp-2026-08-25-022}

- Story ID: `mp-2026-08-25-022`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-022/openflexure-microscope

**Dek:** Prints most of a precise microscope body as one flexure mechanism, pairing interchangeable optics with optional motorized sample positioning and open control software.

Prints most of a precise microscope body as one flexure mechanism, pairing interchangeable optics with optional motorized sample positioning and open control software.

### Why it matters {#why-it-matters-mp-2026-08-25-022}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-25-022}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-25-022}

- This Invention Desk entry makes no independently sourced news claim.

## 25. SatNOGS {#mp-2026-08-25-023}

- Story ID: `mp-2026-08-25-023`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-023/satnogs

**Dek:** Links volunteer-built radio ground stations to shared scheduling and observation services, turning backyard antennas into a public satellite-observation network.

Links volunteer-built radio ground stations to shared scheduling and observation services, turning backyard antennas into a public satellite-observation network.

### Why it matters {#why-it-matters-mp-2026-08-25-023}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-25-023}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-25-023}

- This Invention Desk entry makes no independently sourced news claim.

## 26. The First Paid Slot {#mp-2026-08-25-024}

- Story ID: `mp-2026-08-25-024`
- Type: `invention_desk`
- Classification: `house_example`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-024/the-first-paid-slot

**Dek:** A transparent preview of paid placement with one verified link and no claim of endorsement.

A transparent preview of paid placement with one verified link and no claim of endorsement.

House example - no advertiser paid. Payment will buy placement, never endorsement.

### Why it matters {#why-it-matters-mp-2026-08-25-024}

This placement explains how builders can appear in The Invention Desk without purchasing editorial endorsement.

### Limits and context {#limitations-mp-2026-08-25-024}

- House example - no advertiser paid. Payment will buy placement, never endorsement.

### Claims and sources {#claims-mp-2026-08-25-024}

- This Invention Desk entry makes no independently sourced news claim.

## 27. Put Your Project on the Desk {#mp-2026-08-25-025}

- Story ID: `mp-2026-08-25-025`
- Type: `invention_desk`
- Classification: `house_example`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-25-025/put-your-project-on-the-desk

**Dek:** One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

### Why it matters {#why-it-matters-mp-2026-08-25-025}

This placement explains how builders can appear in The Invention Desk without purchasing editorial endorsement.

### Limits and context {#limitations-mp-2026-08-25-025}

- Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

### Claims and sources {#claims-mp-2026-08-25-025}

- This Invention Desk entry makes no independently sourced news claim.

## Normalized sources

- **source-2026-08-25-001:** [arXiv preprint 2608.23543](https://arxiv.org/abs/2608.23543) — arXiv; primary_research
- **source-2026-08-25-002:** [arXiv preprint 2608.21976](https://arxiv.org/abs/2608.21976) — arXiv; primary_research
- **source-2026-08-25-003:** [arXiv preprint 2608.23525](https://arxiv.org/abs/2608.23525) — arXiv; primary_research
- **source-2026-08-25-004:** [arXiv preprint 2608.23417](https://arxiv.org/abs/2608.23417) — arXiv; primary_research
- **source-2026-08-25-005:** [arXiv preprint 2608.23526](https://arxiv.org/abs/2608.23526) — arXiv; primary_research
- **source-2026-08-25-006:** [arXiv preprint 2608.23497](https://arxiv.org/abs/2608.23497) — arXiv; primary_research
- **source-2026-08-25-007:** [arXiv preprint 2608.23397](https://arxiv.org/abs/2608.23397) — arXiv; primary_research
- **source-2026-08-25-008:** [arXiv preprint 2608.23552](https://arxiv.org/abs/2608.23552) — arXiv; primary_research
- **source-2026-08-25-009:** [arXiv preprint 2608.23493](https://arxiv.org/abs/2608.23493) — arXiv; primary_research
- **source-2026-08-25-010:** [arXiv preprint 2608.22559](https://arxiv.org/abs/2608.22559) — arXiv; primary_research
- **source-2026-08-25-011:** [arXiv preprint 2608.22577](https://arxiv.org/abs/2608.22577) — arXiv; primary_research
- **source-2026-08-25-012:** [arXiv preprint 2608.22610](https://arxiv.org/abs/2608.22610) — arXiv; primary_research
- **source-2026-08-25-013:** [arXiv preprint 2608.22646](https://arxiv.org/abs/2608.22646) — arXiv; primary_research
- **source-2026-08-25-014:** [arXiv preprint 2608.22697](https://arxiv.org/abs/2608.22697) — arXiv; primary_research
- **source-2026-08-25-015:** [arXiv preprint 2608.23565](https://arxiv.org/abs/2608.23565) — arXiv; primary_research
- **source-2026-08-25-016:** [arXiv preprint 2608.21964](https://arxiv.org/abs/2608.21964) — arXiv; primary_research
- **source-2026-08-25-017:** [arXiv preprint 2608.22160](https://arxiv.org/abs/2608.22160) — arXiv; primary_research
- **source-2026-08-25-018:** [arXiv preprint 2608.22161](https://arxiv.org/abs/2608.22161) — arXiv; primary_research
- **source-2026-08-25-019:** [arXiv preprint 2608.22167](https://arxiv.org/abs/2608.22167) — arXiv; primary_research
- **source-2026-08-25-020:** [arXiv preprint 2608.22191](https://arxiv.org/abs/2608.22191) — arXiv; primary_research
- **source-2026-08-25-021:** [arXiv preprint 2608.23475](https://arxiv.org/abs/2608.23475) — arXiv; primary_research

