---
schema_version: "1.0.0"
edition_id: "mp-2026-08-30-morning-0052"
published_at: "2026-08-30T09:00:00.000-04:00"
modified_at: "2026-08-30T09:00:00.000-04:00"
canonical_url: "https://themachinepress.com/edition/2026-08-30"
story_count: 27
lead_story_id: "mp-2026-08-30-001"
---

# The Machine Press — Morning edition

Edition ID: `mp-2026-08-30-morning-0052`  
Published: 2026-08-30T09:00:00.000-04:00  
Canonical edition: https://themachinepress.com/edition/2026-08-30

A six-stage audit fully or partially reproduced 6.52 percent of 1,304 eligible neuro-symbolic AI studies from their published artifacts.

## 1. Only Eighty-Five Studies Made It Through the Artifact Door {#mp-2026-08-30-001}

- Story ID: `mp-2026-08-30-001`
- Type: `lead`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-001/only-eighty-five-studies-made-it-through-the-artifact-door

**Dek:** A six-stage audit fully or partially reproduced 6.52 percent of 1,304 eligible neuro-symbolic AI studies from their published artifacts.

The authors began with 5,497 records, deduplicated and screened them, then sought verifiable public code for 1,304 eligible studies. They report that 849 had no qualifying code artifact; 455 entered the artifact inventory and bounded rerun, where 85 studies were fully or partially reproduced. Missing non-code artifacts blocked 321 attempts, while missing or unusable repositories blocked 42. The audit measures one subfield through its stated protocol, not reproducibility across all AI research, but it shows why a code-available label alone does not complete an experimental record.

### Why it matters {#why-it-matters-mp-2026-08-30-001}

A six-stage audit fully or partially reproduced 6.52 percent of 1,304 eligible neuro-symbolic AI studies from their published artifacts.

### Limits and context {#limitations-mp-2026-08-30-001}

- The audit measures one subfield through its stated protocol, not reproducibility across all AI research, but it shows why a code-available label alone does not complete an experimental record.

### Claims and sources {#claims-mp-2026-08-30-001}

- A six-stage audit fully or partially reproduced 6.52 percent of 1,304 eligible neuro-symbolic AI studies from their published artifacts. [source-2026-08-30-001] — Qualification: The audit measures one subfield through its stated protocol, not reproducibility across all AI research, but it shows why a code-available label alone does not complete an experimental record.

## 2. The Hypothesis Reached the Lab Bench {#mp-2026-08-30-002}

- Story ID: `mp-2026-08-30-002`
- Type: `secondary`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-002/the-hypothesis-reached-the-lab-bench

**Dek:** A Google-affiliated preprint reports a Gemini-based multi-agent system moving from proposals into constrained materials, biology and computer-science experiments.

The authors describe Co-Scientist connecting to a semi-automated chemical-vapor-deposition reactor, adapting crystal-growth recipes to laboratory constraints, predicting engineered E. coli swarming from sparse images and searching for a medical-reasoning inference architecture. They report single-attempt growth of three monolayer semiconductor materials and a blinded review of generated papers by 30 experts across 450 reviews. One MXene-like material still needs atomic-structure confirmation, and the biological comparison uses unpublished measurements. These are author-reported results from a system developed by the paper's contributors, not independent validation of a general autonomous scientist.

### Why it matters {#why-it-matters-mp-2026-08-30-002}

A Google-affiliated preprint reports a Gemini-based multi-agent system moving from proposals into constrained materials, biology and computer-science experiments.

### Limits and context {#limitations-mp-2026-08-30-002}

- The authors describe Co-Scientist connecting to a semi-automated chemical-vapor-deposition reactor, adapting crystal-growth recipes to laboratory constraints, predicting engineered E.
- These are author-reported results from a system developed by the paper's contributors, not independent validation of a general autonomous scientist.

### Claims and sources {#claims-mp-2026-08-30-002}

- A Google-affiliated preprint reports a Gemini-based multi-agent system moving from proposals into constrained materials, biology and computer-science experiments. [source-2026-08-30-002] — Qualification: The authors describe Co-Scientist connecting to a semi-automated chemical-vapor-deposition reactor, adapting crystal-growth recipes to laboratory constraints, predicting engineered E.

## 3. Five Thousand GPU-Hours Searched the Folding Model {#mp-2026-08-30-003}

- Story ID: `mp-2026-08-30-003`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-003/five-thousand-gpu-hours-searched-the-folding-model

**Dek:** AgentFold changed, ran and remembered executable protein-folding model variants inside a closed search loop.

Starting from ESMFold, the system proposed and debugged code changes, evaluated variants and stored both successful and failed interventions. The authors report roughly 80 variants, 170 million language-model tokens and about 5,000 GPU-hours. Under their matched budget, the best lDDT improved 7.5 percent over independent Codex proposals and beat random search; the result is a costly benchmarked search, not a claim of a universally better folding model.

### Why it matters {#why-it-matters-mp-2026-08-30-003}

AgentFold changed, ran and remembered executable protein-folding model variants inside a closed search loop.

### Limits and context {#limitations-mp-2026-08-30-003}

- Under their matched budget, the best lDDT improved 7.5 percent over independent Codex proposals and beat random search; the result is a costly benchmarked search, not a claim of a universally better folding model.

### Claims and sources {#claims-mp-2026-08-30-003}

- AgentFold changed, ran and remembered executable protein-folding model variants inside a closed search loop. [source-2026-08-30-003] — Qualification: Under their matched budget, the best lDDT improved 7.5 percent over independent Codex proposals and beat random search; the result is a costly benchmarked search, not a claim of a universally better folding model.

## 4. The Paraphrase Kept What the Seed Forgot {#mp-2026-08-30-004}

- Story ID: `mp-2026-08-30-004`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-004/the-paraphrase-kept-what-the-seed-forgot

**Dek:** GRAPHSU expands deletion pressure from named forget examples into neighboring support routes.

The method builds a weighted graph around aliases, paraphrases and connected training samples, then applies graded forgetting to high-risk neighbors. On TOFU and PISTOL with GPT-2 Medium and Llama-3.2-3B-Instruct, the authors report up to a 49.5-percentage-point leakage reduction over a matched seed-only baseline while retaining feasible utility. Both benchmarks are controlled evaluation settings, so enterprise deletion claims still need domain-specific testing.

### Why it matters {#why-it-matters-mp-2026-08-30-004}

GRAPHSU expands deletion pressure from named forget examples into neighboring support routes.

### Limits and context {#limitations-mp-2026-08-30-004}

- On TOFU and PISTOL with GPT-2 Medium and Llama-3.2-3B-Instruct, the authors report up to a 49.5-percentage-point leakage reduction over a matched seed-only baseline while retaining feasible utility.

### Claims and sources {#claims-mp-2026-08-30-004}

- GRAPHSU expands deletion pressure from named forget examples into neighboring support routes. [source-2026-08-30-004] — Qualification: On TOFU and PISTOL with GPT-2 Medium and Llama-3.2-3B-Instruct, the authors report up to a 49.5-percentage-point leakage reduction over a matched seed-only baseline while retaining feasible utility.

## 5. Old Training Evidence Needed a New Permit {#mp-2026-08-30-005}

- Story ID: `mp-2026-08-30-005`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-005/old-training-evidence-needed-a-new-permit

**Dek:** BCIT checks whether a previously useful update still applies after later training changed the parent model.

The controller binds each observed update effect to its source model, data and training stage, vetoes hard conflicts and can demand a bounded current-state trial before reuse. In experiments adapting one 4B model across finance reasoning, text-to-SQL and function calling, the authors report fewer harmful authorizations and higher equal-budget final quality than their alternatives. The evidence covers one model and three adaptation contexts.

### Why it matters {#why-it-matters-mp-2026-08-30-005}

BCIT checks whether a previously useful update still applies after later training changed the parent model.

### Limits and context {#limitations-mp-2026-08-30-005}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-005}

- BCIT checks whether a previously useful update still applies after later training changed the parent model. [source-2026-08-30-005]

## 6. Editing Changed the Detector's Verdict {#mp-2026-08-30-006}

- Story ID: `mp-2026-08-30-006`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-006/editing-changed-the-detector-s-verdict

**Dek:** A 135,389-pair study isolates professional English editing as a confound in AI-text detection.

The authors compared non-native academic manuscripts with native-edited versions while holding authorship and content together. Across 13 detectors, false-positive rates on human writing ranged from zero to 100 percent, and the same edits pushed scores upward for some systems and downward for others. Score movement also tracked editing extent. The study identifies style as a major confound; it does not prove that text origin can never be detected.

### Why it matters {#why-it-matters-mp-2026-08-30-006}

A 135,389-pair study isolates professional English editing as a confound in AI-text detection.

### Limits and context {#limitations-mp-2026-08-30-006}

- The study identifies style as a major confound; it does not prove that text origin can never be detected.

### Claims and sources {#claims-mp-2026-08-30-006}

- A 135,389-pair study isolates professional English editing as a confound in AI-text detection. [source-2026-08-30-006] — Qualification: The study identifies style as a major confound; it does not prove that text origin can never be detected.

## 7. The Agent's Permission Moved Into the Runtime {#mp-2026-08-30-007}

- Story ID: `mp-2026-08-30-007`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-007/the-agent-s-permission-moved-into-the-runtime

**Dek:** A governance paper derives five controls for ephemeral, model-directed agents and reports four running in private pilots.

The proposed primitives are discovery, identity, governance, attestation and supply chain. The implementation mediates actions before execution, authorizes them against a tenant vocabulary and records them in a signed hash-linked ledger. Its authors also name the costs: enforcement sits on the critical path, identity needs a workload sidecar and fail-closed mediation converts outages into denial. Four primitives are in private pilots; the fifth remains separate tooling.

### Why it matters {#why-it-matters-mp-2026-08-30-007}

A governance paper derives five controls for ephemeral, model-directed agents and reports four running in private pilots.

### Limits and context {#limitations-mp-2026-08-30-007}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-007}

- A governance paper derives five controls for ephemeral, model-directed agents and reports four running in private pilots. [source-2026-08-30-007]

## 8. Sentence Transitions Became the Detection Signal {#mp-2026-08-30-008}

- Story ID: `mp-2026-08-30-008`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-008/sentence-transitions-became-the-detection-signal

**Dek:** A graph-based detector looks for deviations between adjacent sentences instead of treating style features independently.

The paper calls the signal relational over-regularization: recurring similarity bursts and transition patterns create sentence-pair variance that differs from human text in the tested data. Its CSFG implementation reports 97.14 percent binary accuracy, a 1.57 percent false-positive rate and an 11.14-point gain over the strongest graph baseline. The authors also show the boundary: performance falls when a generator's transition variance reaches or drops below the human baseline.

### Why it matters {#why-it-matters-mp-2026-08-30-008}

A graph-based detector looks for deviations between adjacent sentences instead of treating style features independently.

### Limits and context {#limitations-mp-2026-08-30-008}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-008}

- A graph-based detector looks for deviations between adjacent sentences instead of treating style features independently. [source-2026-08-30-008]

## 9. Noisy Teammates Learned in Local Groups {#mp-2026-08-30-009}

- Story ID: `mp-2026-08-30-009`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-009/noisy-teammates-learned-in-local-groups

**Dek:** SIGMA clusters cooperating agents before combining information across the full team.

The framework starts from the claim that independent observation noise acquires local structure through task dependencies. It uses density-based grouping and within-group consensus to smooth agent-specific deviations, then integrates groups with attention. StarCraft II experiments reported stronger robustness under noisy observations while remaining competitive without noise. The evidence is simulation-based and does not establish robustness in open physical teams.

### Why it matters {#why-it-matters-mp-2026-08-30-009}

SIGMA clusters cooperating agents before combining information across the full team.

### Limits and context {#limitations-mp-2026-08-30-009}

- The evidence is simulation-based and does not establish robustness in open physical teams.

### Claims and sources {#claims-mp-2026-08-30-009}

- SIGMA clusters cooperating agents before combining information across the full team. [source-2026-08-30-009] — Qualification: The evidence is simulation-based and does not establish robustness in open physical teams.

## 10. Harder Workflows Flattened Every Judge {#mp-2026-08-30-010}

- Story ID: `mp-2026-08-30-010`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-010/harder-workflows-flattened-every-judge

**Dek:** AgentJudgeBench tests language-model judges on 3,808 dependency-ordered tool-calling workflows.

Across six workflow graph shapes and three difficulty levels, alignment with the programmatic reference declined as tasks grew harder and fell faster when ground truth was hidden. On hard no-ground-truth cases, six judges converged in a narrow 77-to-82-percent band despite scale differences. Structured rubrics improved alignment by as much as 6.5 points, while reasoning traces and temperature changes had little effect. Ground truth sometimes reduced alignment through apparent over-anchoring.

### Why it matters {#why-it-matters-mp-2026-08-30-010}

AgentJudgeBench tests language-model judges on 3,808 dependency-ordered tool-calling workflows.

### Limits and context {#limitations-mp-2026-08-30-010}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-010}

- AgentJudgeBench tests language-model judges on 3,808 dependency-ordered tool-calling workflows. [source-2026-08-30-010]

## 11. The Real Session Brought Its Mess With It {#mp-2026-08-30-011}

- Story ID: `mp-2026-08-30-011`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-011/the-real-session-brought-its-mess-with-it

**Dek:** DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.

The benchmark spans eight broad scenarios and 17 capability categories, with most tasks requiring several capabilities at once. Isolated containers add insufficient, unstable and noisy environmental conditions. Five agent frameworks paired with four models showed substantial gaps in strict completion, and both the model and harness shaped robustness. The source platform supplies the sessions and evaluation design, so the benchmark still needs broader replication across organizations.

### Why it matters {#why-it-matters-mp-2026-08-30-011}

DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact.

### Limits and context {#limitations-mp-2026-08-30-011}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-011}

- DuMateBench reconstructs 200 privacy-screened production-agent tasks with histories, configuration and workspace state intact. [source-2026-08-30-011]

## 12. Benign Inputs Combined Into Harm {#mp-2026-08-30-012}

- Story ID: `mp-2026-08-30-012`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-012/benign-inputs-combined-into-harm

**Dek:** Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.

The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context. Representative safety guards missed both kinds of compositional evidence across time and modality in the authors' evaluation. The dataset is scheduled for release in October 2026, so current claims rest on the paper's reported protocol rather than an independently inspectable public benchmark.

### Why it matters {#why-it-matters-mp-2026-08-30-012}

Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning.

### Limits and context {#limitations-mp-2026-08-30-012}

- The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context.

### Claims and sources {#claims-mp-2026-08-30-012}

- Multi2AV-Safety tests all 11 multi-input combinations of text, image, audio and video conditioning. [source-2026-08-30-012] — Qualification: The 11,024-instance benchmark is designed around harm that appears only when modalities interact, as well as explicit harmful cues diluted by benign context.

## 13. The Supervisor Could Interrupt the Run {#mp-2026-08-30-013}

- Story ID: `mp-2026-08-30-013`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-013/the-supervisor-could-interrupt-the-run

**Dek:** PILOT separates a long-running worker from a supervisor that can steer or abort it and retain lessons as skills.

The harness couples live steering with live self-evolution, distilling procedures and failure modes while a task is still active. Across two frozen model backbones and three benchmarks, the authors report first place in five of six configurations, gains up to 9.8 points on Terminal-Bench 2.0 and lower output-token use. In self-improvement settings they report 12.4- and 14.6-point gains. These are benchmarked harness effects, not evidence that intervention always improves open-ended work.

### Why it matters {#why-it-matters-mp-2026-08-30-013}

PILOT separates a long-running worker from a supervisor that can steer or abort it and retain lessons as skills.

### Limits and context {#limitations-mp-2026-08-30-013}

- These are benchmarked harness effects, not evidence that intervention always improves open-ended work.

### Claims and sources {#claims-mp-2026-08-30-013}

- PILOT separates a long-running worker from a supervisor that can steer or abort it and retain lessons as skills. [source-2026-08-30-013] — Qualification: These are benchmarked harness effects, not evidence that intervention always improves open-ended work.

## 14. Incomplete Proofs Kept Their Verified Pieces {#mp-2026-08-30-014}

- Story ID: `mp-2026-08-30-014`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-014/incomplete-proofs-kept-their-verified-pieces

**Dek:** ProofEvolve stores kernel-checked partial proof graphs so unfinished work can contribute to later theorems.

Neural models propose decompositions, repairs and schema combinations while Lean verifies each transition. Within a problem, the system evolves partial AND-OR proof graphs; across problems, verified subgraphs enter a persistent schema library with unresolved premises exposed as new goals. The authors report the highest average solve rate among evaluated systems on three competition-level Lean benchmarks. Formal verification preserves soundness of accepted steps, not the usefulness of every proposal.

### Why it matters {#why-it-matters-mp-2026-08-30-014}

ProofEvolve stores kernel-checked partial proof graphs so unfinished work can contribute to later theorems.

### Limits and context {#limitations-mp-2026-08-30-014}

- Formal verification preserves soundness of accepted steps, not the usefulness of every proposal.

### Claims and sources {#claims-mp-2026-08-30-014}

- ProofEvolve stores kernel-checked partial proof graphs so unfinished work can contribute to later theorems. [source-2026-08-30-014] — Qualification: Formal verification preserves soundness of accepted steps, not the usefulness of every proposal.

## 15. Lean Checked the First Broken Step {#mp-2026-08-30-026}

- Story ID: `mp-2026-08-30-026`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-026/lean-checked-the-first-broken-step

**Dek:** FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.

The system decomposes natural-language proofs into local units, extracts typed obligations and uses a semantic alignment gate before accepting Lean validation. With a GPT-5.4 backbone, it reports 81.43 percent exact first-error accuracy on 350 Olympiad problems versus 72.29 percent for direct judging, and 84.5 percent on 200 university-level problems versus 75 percent. The benchmarks are expert-verified but finite and model-specific.

### Why it matters {#why-it-matters-mp-2026-08-30-026}

FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning.

### Limits and context {#limitations-mp-2026-08-30-026}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-026}

- FaithSieve admits formal evidence only when an auto-formalized obligation still matches the informal proof's meaning. [source-2026-08-30-015]

## 16. Approval Expired Before the Action {#mp-2026-08-30-027}

- Story ID: `mp-2026-08-30-027`
- Type: `dispatch`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-027/approval-expired-before-the-action

**Dek:** A guardrail verdict can be correct when checked and unsafe by the time a self-adaptive system acts.

Across five reproducible environments, fixed-action replay produced verdict-change rates from 5.3 to 48.4 percent after eight simulator steps. The proposed Freshness-Bounded Shield estimates an approval horizon from safety margin and recent feature volatility; under fixed settings, the authors report reducing oracle-labeled expiry from a 3.4-to-24.7-percent range to zero-to-1.8 percent. Four language-model judges still showed nonzero use-time invalidity, motivating a check-time-and-use-time freshness contract.

### Why it matters {#why-it-matters-mp-2026-08-30-027}

A guardrail verdict can be correct when checked and unsafe by the time a self-adaptive system acts.

### Limits and context {#limitations-mp-2026-08-30-027}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-027}

- A guardrail verdict can be correct when checked and unsafe by the time a self-adaptive system acts. [source-2026-08-30-016]

## 17. Thinking Tokens Had Diminishing Returns {#mp-2026-08-30-015}

- Story ID: `mp-2026-08-30-015`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-015/thinking-tokens-had-diminishing-returns

**Dek:** A 151-run analysis finds task structure predicts marginal reasoning efficiency better than nominal difficulty.

Sequential inference tasks showed stronger token-normalized gains than knowledge recall, while higher reasoning effort sometimes lowered accuracy. The paper argues for selecting reasoning mode by task, effort and deployment context rather than enabling it universally.

### Why it matters {#why-it-matters-mp-2026-08-30-015}

A 151-run analysis finds task structure predicts marginal reasoning efficiency better than nominal difficulty.

### Limits and context {#limitations-mp-2026-08-30-015}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-015}

- A 151-run analysis finds task structure predicts marginal reasoning efficiency better than nominal difficulty. [source-2026-08-30-017]

## 18. Retries Needed Delegation Identity {#mp-2026-08-30-016}

- Story ID: `mp-2026-08-30-016`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-016/retries-needed-delegation-identity

**Dek:** A 147-incident study says service-mesh retry and breaker assumptions fail on mutating agent work.

The authors trace failures to inadequate identity and evidence, including accumulated effects across repeated invocations and enforcement that blocked correct work. They derive seven delegation-level reliability primitives but explicitly stop short of a controlled evaluation.

### Why it matters {#why-it-matters-mp-2026-08-30-016}

A 147-incident study says service-mesh retry and breaker assumptions fail on mutating agent work.

### Limits and context {#limitations-mp-2026-08-30-016}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-016}

- A 147-incident study says service-mesh retry and breaker assumptions fail on mutating agent work. [source-2026-08-30-018]

## 19. The Harness Changed the Solver {#mp-2026-08-30-017}

- Story ID: `mp-2026-08-30-017`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-017/the-harness-changed-the-solver

**Dek:** Keeping model weights fixed, context compaction and stall handling changed coding-agent results.

On a 169-task tight-context SWE-bench Verified cohort, the treatment raised complete solutions from 43 to 72 and improved mean fail-to-pass fraction. Wider-context results were closer, reinforcing that evaluations measure a model-harness pair rather than model weights alone.

### Why it matters {#why-it-matters-mp-2026-08-30-017}

Keeping model weights fixed, context compaction and stall handling changed coding-agent results.

### Limits and context {#limitations-mp-2026-08-30-017}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-017}

- Keeping model weights fixed, context compaction and stall handling changed coding-agent results. [source-2026-08-30-019]

## 20. Four Contracts Drew an Enterprise Boundary {#mp-2026-08-30-018}

- Story ID: `mp-2026-08-30-018`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-018/four-contracts-drew-an-enterprise-boundary

**Dek:** A proposed runtime divides responsibility among Skill, Harness, Scaffold and an external data substrate.

The architecture turns six design conditions into measurable obligations and proposes a crossover experiment for a capability-capacity separation hypothesis. The authors report no implementation, dataset or completed experiment, so this is a falsifiable design proposal rather than measured performance.

### Why it matters {#why-it-matters-mp-2026-08-30-018}

A proposed runtime divides responsibility among Skill, Harness, Scaffold and an external data substrate.

### Limits and context {#limitations-mp-2026-08-30-018}

- No additional limitation was separately recorded.

### Claims and sources {#claims-mp-2026-08-30-018}

- A proposed runtime divides responsibility among Skill, Harness, Scaffold and an external data substrate. [source-2026-08-30-020]

## 21. One Fairness Score Missed the Rank Shift {#mp-2026-08-30-019}

- Story ID: `mp-2026-08-30-019`
- Type: `ticker`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-019/one-fairness-score-missed-the-rank-shift

**Dek:** A synthetic correspondence-audit workflow checks candidate matching across five protected-characteristic axes.

In an example with five jobs, 100 base candidates and 10 treatments, score and retention measures stayed within tolerance while rank-stability and nDCG surfaced borderline findings. The small synthetic example supports multi-metric auditing, not a finding about hiring systems generally.

### Why it matters {#why-it-matters-mp-2026-08-30-019}

A synthetic correspondence-audit workflow checks candidate matching across five protected-characteristic axes.

### Limits and context {#limitations-mp-2026-08-30-019}

- The small synthetic example supports multi-metric auditing, not a finding about hiring systems generally.

### Claims and sources {#claims-mp-2026-08-30-019}

- A synthetic correspondence-audit workflow checks candidate matching across five protected-characteristic axes. [source-2026-08-30-021] — Qualification: The small synthetic example supports multi-metric auditing, not a finding about hiring systems generally.

## 22. MNT Reform Next {#mp-2026-08-30-020}

- Story ID: `mp-2026-08-30-020`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-020/mnt-reform-next

**Dek:** Reworks a laptop into public, swappable modules: processor, port boards, keyboard, trackpad, and user-serviceable battery packs can evolve without sealing the whole machine.

Reworks a laptop into public, swappable modules: processor, port boards, keyboard, trackpad, and user-serviceable battery packs can evolve without sealing the whole machine.

### Why it matters {#why-it-matters-mp-2026-08-30-020}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-30-020}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-30-020}

- This Invention Desk entry makes no independently sourced news claim.

## 23. Maslow 4 {#mp-2026-08-30-021}

- Story ID: `mp-2026-08-30-021`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-021/maslow-4

**Dek:** Pulls a compact router sled across full sheets with four measured belts, trading a bulky gantry for corner anchors and community-developed control software.

Pulls a compact router sled across full sheets with four measured belts, trading a bulky gantry for corner anchors and community-developed control software.

### Why it matters {#why-it-matters-mp-2026-08-30-021}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-30-021}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-30-021}

- This Invention Desk entry makes no independently sourced news claim.

## 24. FarmBot Genesis {#mp-2026-08-30-022}

- Story ID: `mp-2026-08-30-022`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-022/farmbot-genesis

**Dek:** Moves an interchangeable tool head across a raised bed to place seeds, water plants, and measure soil, backed by published hardware, software, data, and documentation.

Moves an interchangeable tool head across a raised bed to place seeds, water plants, and measure soil, backed by published hardware, software, data, and documentation.

### Why it matters {#why-it-matters-mp-2026-08-30-022}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-30-022}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-30-022}

- This Invention Desk entry makes no independently sourced news claim.

## 25. OpenBikeSensor {#mp-2026-08-30-023}

- Story ID: `mp-2026-08-30-023`
- Type: `invention_desk`
- Classification: `editorial`
- Content status: `new`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-023/openbikesensor

**Dek:** Combines a DIY bicycle distance sensor, GPS, and a shared portal so volunteer riders can map close passes and study where street design needs attention.

Combines a DIY bicycle distance sensor, GPS, and a shared portal so volunteer riders can map close passes and study where street design needs attention.

### Why it matters {#why-it-matters-mp-2026-08-30-023}

An independent builder is turning an improbable idea into a working project.

### Limits and context {#limitations-mp-2026-08-30-023}

- A Desk Pick is an editorial selection, not a product endorsement.

### Claims and sources {#claims-mp-2026-08-30-023}

- This Invention Desk entry makes no independently sourced news claim.

## 26. The First Paid Slot {#mp-2026-08-30-024}

- Story ID: `mp-2026-08-30-024`
- Type: `invention_desk`
- Classification: `house_example`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-024/the-first-paid-slot

**Dek:** A transparent preview of paid placement with one verified link and no claim of endorsement.

A transparent preview of paid placement with one verified link and no claim of endorsement.

House example - no advertiser paid. Payment will buy placement, never endorsement.

### Why it matters {#why-it-matters-mp-2026-08-30-024}

This placement explains how builders can appear in The Invention Desk without purchasing editorial endorsement.

### Limits and context {#limitations-mp-2026-08-30-024}

- House example - no advertiser paid. Payment will buy placement, never endorsement.

### Claims and sources {#claims-mp-2026-08-30-024}

- This Invention Desk entry makes no independently sourced news claim.

## 27. Put Your Project on the Desk {#mp-2026-08-30-025}

- Story ID: `mp-2026-08-30-025`
- Type: `invention_desk`
- Classification: `house_example`
- Content status: `carried_over`
- Permanent URL: https://themachinepress.com/story/mp-2026-08-30-025/put-your-project-on-the-desk

**Dek:** One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

### Why it matters {#why-it-matters-mp-2026-08-30-025}

This placement explains how builders can appear in The Invention Desk without purchasing editorial endorsement.

### Limits and context {#limitations-mp-2026-08-30-025}

- Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

### Claims and sources {#claims-mp-2026-08-30-025}

- This Invention Desk entry makes no independently sourced news claim.

## Normalized sources

- **source-2026-08-30-001:** [arXiv preprint 2608.26236](https://arxiv.org/abs/2608.26236) — arXiv; primary_research
- **source-2026-08-30-002:** [arXiv preprint 2608.26701](https://arxiv.org/abs/2608.26701) — arXiv; primary_research
- **source-2026-08-30-003:** [arXiv preprint 2608.26747](https://arxiv.org/abs/2608.26747) — arXiv; primary_research
- **source-2026-08-30-004:** [arXiv preprint 2608.26743](https://arxiv.org/abs/2608.26743) — arXiv; primary_research
- **source-2026-08-30-005:** [arXiv preprint 2608.26730](https://arxiv.org/abs/2608.26730) — arXiv; primary_research
- **source-2026-08-30-006:** [arXiv preprint 2608.26710](https://arxiv.org/abs/2608.26710) — arXiv; primary_research
- **source-2026-08-30-007:** [arXiv preprint 2608.26696](https://arxiv.org/abs/2608.26696) — arXiv; primary_research
- **source-2026-08-30-008:** [arXiv preprint 2608.26694](https://arxiv.org/abs/2608.26694) — arXiv; primary_research
- **source-2026-08-30-009:** [arXiv preprint 2608.26683](https://arxiv.org/abs/2608.26683) — arXiv; primary_research
- **source-2026-08-30-010:** [arXiv preprint 2608.26623](https://arxiv.org/abs/2608.26623) — arXiv; primary_research
- **source-2026-08-30-011:** [arXiv preprint 2608.26546](https://arxiv.org/abs/2608.26546) — arXiv; primary_research
- **source-2026-08-30-012:** [arXiv preprint 2608.26535](https://arxiv.org/abs/2608.26535) — arXiv; primary_research
- **source-2026-08-30-013:** [arXiv preprint 2608.26530](https://arxiv.org/abs/2608.26530) — arXiv; primary_research
- **source-2026-08-30-014:** [arXiv preprint 2608.26334](https://arxiv.org/abs/2608.26334) — arXiv; primary_research
- **source-2026-08-30-015:** [arXiv preprint 2608.26310](https://arxiv.org/abs/2608.26310) — arXiv; primary_research
- **source-2026-08-30-016:** [arXiv preprint 2608.26306](https://arxiv.org/abs/2608.26306) — arXiv; primary_research
- **source-2026-08-30-017:** [arXiv preprint 2608.26235](https://arxiv.org/abs/2608.26235) — arXiv; primary_research
- **source-2026-08-30-018:** [arXiv preprint 2608.26225](https://arxiv.org/abs/2608.26225) — arXiv; primary_research
- **source-2026-08-30-019:** [arXiv preprint 2608.26218](https://arxiv.org/abs/2608.26218) — arXiv; primary_research
- **source-2026-08-30-020:** [arXiv preprint 2608.27086](https://arxiv.org/abs/2608.27086) — arXiv; primary_research
- **source-2026-08-30-021:** [arXiv preprint 2608.26899](https://arxiv.org/abs/2608.26899) — arXiv; primary_research

