TheMachine Press

The newspaper for artificial intelligence and the people building it.

Morning editionSources linked throughout
Front pageImportance 10/10

The Shortcut Improved the Score and Weakened the Skill

A controlled logic-puzzle experiment found that cheaper AI help increased use—and that assisted performance overstated what participants could later do alone.

A sepia engraving of an adult solving a logic-tile puzzle while a glassy assistant device sits unused at the edge of the desk.Editorial illustration
Conceptual illustration: a controlled logic-puzzle study measured assisted performance and later independent skill; this does not depict its participants or materials. Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-25.

Participants worked through logic puzzles before, during and after access to on-demand AI assistance. Randomly lowering the request cost increased how often people called for help. Those who requested assistance during the access phase performed worse after it was removed, while a Bayesian latent-ability model associated more independent reasoning with larger skill gains. The experiment does not establish that every use of AI impairs learning, but it shows why an assisted score can be a poor proxy for the skill a person retains.

research
A sepia engineering-plate illustration of a floating offshore wind support structure with moorings, cutaway bracing and design-review insets.Editorial illustration
Conceptual illustration: the AI Engineer study reports an externally reviewed floating-wind design; this is not the study geometry, a built facility or final construction certification. Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-25.

The Design Loop Ended at an External Gate

A closed-loop agent connected language models to deterministic engineering solvers, then sent its floating-wind design to a classification society.

The AI Engineer turns natural-language requirements into geometry and a mesh, then couples topology and member-size searches to structural, aero-hydro-servo-elastic and cost checks. Its automated reviewer stops only when capacity, steel intensity, unit cost, constructability and fatigue scores clear stated floors. The top floating-wind support design passed China Classification Society Approval in Principle and used 8.1 percent less steel and 8.1 percent lower unit capital cost than the human-optimized TuQiang baseline. Approval in Principle is an external design-stage check, not final construction certification, and the authors list detailed design and fabrication constraints as remaining work.

The Repository Changed and the Skill Stayed Quiet

Every one of 105 selected release transitions invalidated part of the prior repository skill set.

Across 57 repositories, six frontier agents reached only 29.9 to 69.7 percent avg@3 macro F1 when asked to update V1 skills from official V1-to-V2 patches. Missed files left stale instructions intact, while broader edits raised recall at the cost of precision. This final text-only report fills the front rail beneath the feature.

Today's Dispatches

benchmarks evals01
Golden Gate Bridge stretching over blue-gray water and disappearing into dense fog.File image
Golden Gate Bridge file image used illustratively; it does not depict an EarthVerse task, source event, current hazard or benchmark result. Belli Kins / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.

The Best Agent Finished the Pieces and Lost the Chain

EarthVerse tests 405 reproducible investigations across 199 documented events and 19 natural-hazard families.

The strongest system reached 84.65 percent mean answer-unit accuracy, but only 34.81 percent Strict@95. The gap shows that agents can complete many local steps without preserving a consistent chain across evidence, units, calculations and physical interpretation.

developer tools02

The Skill Learned What the Brief Forgot

SkillAlchemy searches open-world sources for requirements omitted by an underspecified skill brief.

Across 87 SkillsBench tasks, its source-grounded skill packages improved pass rate 19.9 points over no-skill execution and 8.6 points over the strongest automated baseline, while performing comparably to human-curated skills. The system admits procedures only at the scope supported by evidence.

research03

The World Model Learned Energy and Let It Drift

A frozen video world model encoded an energy-like invariant that its own imagined rollouts failed to preserve.

A label-free search recovered the same scalar across independently trained conservative pendulum models and found no comparable quantity in damped controls. Projecting latent state back toward the initial level set reduced rollout error in all three conservative models; matched random constraints usually made it worse.

safety security04

Reasoning Training Moved Along the Safety Direction

A representation-space audit links some reasoning fine-tunes to safety shifts and proposes a penalty at the implicated layers.

The authors stress that reasoning-induced misalignment does not appear across every architecture, scale or dataset. On Qwen2.5 3B and 7B experiments, a learned safety-direction penalty restored measured safety while preserving benchmark reasoning performance, with diagnostics guiding which layers to include.

research05

The Clinical Agent Had to Earn the Next Step

MediSkill-Evo separates clinical skills, process rules, schemas and measurements behind publication and safety gates.

On 300 held-out simulated Qwen encounters, the complete system raised diagnosis accuracy from 61.33 to 69 percent and treatment-intent coverage from 33.62 to 66.44 percent while halving automatically scored critical failures relative to AgentClinic. The authors explicitly describe this as fixed-suite system evidence, not clinical validation.

developer tools06
Angled dark computer screen with colorful programming code and a bright blue edge.File image
Illustrative programming file image; it does not show Prime Agent, its codebase, sessions, subagents or benchmark output. Nemuel Sereti / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.

The Harness Stopped Turning Recovery Into a Model Failure

Prime Agent keeps histories, memories, skills and subagent specifications across long-running coding and evaluation trajectories.

The open-source harness pairs a persistent IPython environment with standardized execution, recovery, verification and resource accounting. Its reported evaluations include an ARC-AGI-3 RHAE Best@1 increase from 30 to 95.5 percent and competitive results across coding, kernel and emulator tasks; those are harness-specific benchmarks, not general capability proof.

research07

The Model Turned Its Own Failure Into the Training Signal

SRPO converts completed trajectories into concise reflection patches and dense token-level supervision.

With a Qwen3-8B base, the authors report 73.3 percent on AIME 2024 using 8 percent of the training FLOPs of scaled supervised fine-tuning, alongside gains on WebShop, ALFWorld and SWE-Bench-Lite. The method uses reflection-conditioned teacher scores without a separate critic or larger teacher model.

benchmarks evals08

The Rubric Became a Program

ExecRubrics compiles evaluation logic into inspectable scoring functions instead of leaving every criterion to a black-box judge.

Across HealthBench, HelpSteer and ArgQuality, executable rubrics matched or improved natural-language rubric baselines at best preference accuracies of 53, 78 and 92 percent while cutting latency by as much as 320 times. The approach makes dependencies, penalties and override conditions explicit and editable.

research09

The Old Screenshot Returned Only When It Mattered

CausalCache spends a fixed visual-memory budget on the past events with the highest conditional utility.

On OSWorld-Verified, restoring selected history images was worth about 13 success points over summary-only memory, although same-budget allocation methods were indistinguishable on desktop. A zero-shot mobile test produced a 3.7-point overall gain and an 8.6-point gain on the predeclared memory-critical split.

developer tools10

A Useful Skill Harmed the Coalition

Skill-bank audits found coalition pollution and utility reversals after cross-domain transfer.

The study uses sampled Shapley marginals to select skills and an unlabeled target-domain mask to suppress harmful transfers. Across LoCoMo, LongMemEval, HotpotQA and ALFWorld, the interventions improved performance and generalization while showing why isolation tests can miss a skill that damages the surrounding bank.

infrastructure11
Close-up of a graphics-card cooling fan inside a silver and black angular housing.File image
Generic accelerator-hardware file image; it does not depict CAI-DLLM, the tested models, devices, speedups or energy measurements. Natalia S / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.

Stable Tokens Stopped Paying for More Denoising

CAI-DLLM uses first-step confidence to commit easy tokens and reserve later denoising for harder ones.

The training-free method reports up to 18.2-times wall-clock speedup on LLaDA GSM8K and 13.1-times on Dream HumanEval with slightly higher measured accuracy in those settings. On harder tasks speedups reached 44.8 times with a largest 4.4-point accuracy drop, exposing the quality boundary.

business enterprise12

The Middle of the Results Page Became the Blind Spot

Randomized hotel listings show that shopping agents inspect deeper than people but retain a weak, non-monotonic position bias.

Across 5,000 sessions and four language models, the middle of a 100-result page was least likely to be inspected—not the bottom. Position reached the choice stage for some models but not others, while all selected the same undominated listing; displayed attributes mattered more than rank.

research13

The World Came Back to the Place It Had Shown

ReWorld combines local attention with a pose-indexed landmark bank to revisit scenes under a fixed memory budget.

The interactive world model streams 704-by-1280 video in a four-step real-time mode and retrieves landmarks nearest the current pose. In 64-second out-and-back tests, a 12-chunk cache regenerated the starting view after a sliding window had evicted the evidence and full-KV attention exhausted memory.

Independent builders

The Invention Desk

Independent builders turning improbable ideas into real things.

  • Four editorial selections
  • $7 for seven days
  • Paid work is clearly labeled
  • Placement is never endorsement
A sepia engraving of a compact touchscreen computer beside a keyboard, two speakers, and an unbranded music control board.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-23.
Desk PickReleased

Tulip Creative Computer

Buildershore pine sound systems contributors

A self-contained touchscreen computer boots into MicroPython and combines programmable graphics, MIDI connections, and an open-source synthesizer in focused, buildable hardware.

Visit Tulip Creative Computer
A sepia engraving of a tiny two-roller printing press clamped to a workbench as it feeds out a small abstract print.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-23.
Desk PickReleased

Open Press Project

BuilderMartin Schneider and Dominik Schmitz

Shrinks an intaglio press into 3D-printed tabletop hardware, with free fabrication plans for makers and finished presses for artists without a printer.

Visit Open Press Project
A sepia cutaway engraving of a printed-frame microscope with an objective, a flexure-guided sample stage, and three small motors.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-23.
Desk PickReleased

OpenFlexure Microscope

BuilderRichard Bowman and OpenFlexure contributors

Prints most of a precise microscope body as one flexure mechanism, pairing interchangeable optics with optional motorized sample positioning and open control software.

Visit OpenFlexure Microscope
A sepia engraving of a rooftop tracking antenna beneath a small satellite, with dotted arcs connecting distant ground stations.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-23.
Desk PickReleased

SatNOGS

BuilderLibre Space Foundation and SatNOGS contributors

Links volunteer-built radio ground stations to shared scheduling and observation services, turning backyard antennas into a public satellite-observation network.

Visit SatNOGS
An unnamed prototype under a desk lamp beside a blank card.
Original Codex Image Gen concept art from 2026-07-10; carried forward from the validated 2026-08-24 edition.
Sponsored ProjectOpen

The First Paid Slot

BuilderHouse example / The Machine Press

A transparent preview of paid placement with one verified link and no claim of endorsement.

House example - no advertiser paid. Payment will buy placement, never endorsement.

Ask about the launch slot
Six portfolio slots surround one open slot and seven day markers.
Original Codex Image Gen concept art from 2026-07-10; carried forward from the validated 2026-08-24 edition.
Open PlacementOpen

Put Your Project on the Desk

$7$1/day / 7 days

One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

Email the desk

Desk Picks are selected by the newsroom. Sponsored placement purchases visibility, never endorsement, and always remains visibly separated from editorial selection.