Across 200 FreeCAD tasks, the strongest tested computer-use agent passed 17.5%; expert reference work passed 87.0%.
Editorial illustration
Concept illustration: executable checks expose structural failure beneath a plausible CAD result; it is not a benchmark plot or documentary scene. Original editorial illustration generated with built-in Codex Image Gen for The Machine Press, 2026-09-16.
Computer-use benchmarks often stop when an interface appears to accept the right clicks. CADWorld checks the persistent engineering artifact instead. Its 200 FreeCAD tasks span 11 workflows, from sketches and assemblies to CAM, finite-element analysis, meshes and technical drawings. Executable checks inspect geometry, dimensions, parametric structure, constraints and downstream manufacturing state. Across seven current agents, the strongest passed 17.5% of the full suite versus 87.0% for an expert reference. Stronger agents increasingly reached plausible-looking outputs but failed structural and construction-process requirements, exposing a gap between GUI fluency and reliable engineering work.
Concept illustration: controlled dissipation reshapes a superconducting spin-chain simulation; it is not a device photograph or measured phase diagram. Original editorial illustration generated with built-in Codex Image Gen for The Machine Press, 2026-09-16.
A superconducting processor simulated 50 sites at depths up to 1,700 entangling gates and mapped 117 hardware points.
Open quantum systems can settle into steady states with no equilibrium counterpart, but their density matrices grow too quickly for controlled large-scale calculations. The team mapped a dissipative spin-1/2 Heisenberg chain onto 100 simultaneously active qubits on IBM's Kingston processor, using Stinespring dilation to represent 50 sites at entangling-gate depths up to 1,700. From 117 hardware measurements, they resolved ferromagnetic, antiferromagnetic, spin-density-wave and paramagnetic regions. Because the modeled dissipation continually erases some errors, hardware noise acts as a weaker competing bath. The result is a hardware study of one benchmark model, not a general claim of quantum advantage.
JWST separated reflected light from heat and measured transport becoming 7.1 ± 1.9 degrees more efficient per pressure decade.
A phase curve follows a planet through its orbit, mixing reflected starlight with thermal emission. JWST's NIRSpec PRISM separated those components for NGTS-10 Ab from 0.5 to 5.5 micrometers, producing joint temperature and reflectance maps. The two were anti-correlated: models best explained the pattern with micron-scale silicate clouds evaporating across the hottest eastern substellar region and weak atmospheric drag. The thermal offset changed by 7.1 ± 1.9 degrees per pressure decade, indicating stronger heat transport deeper in the atmosphere. The interpretation depends on retrieval and circulation models rather than direct images of clouds.
Generic processor-pin file image, used illustratively; it does not depict the tested networks, Fisher geometry or pruning results. Pixabay / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.
A Fisher-geodesic hierarchy beat magnitude and local-Fisher pruning across every tested architecture and dataset combination.
Setting a neural-network parameter to zero moves the model onto a constrained surface. This paper measures that displacement with the Fisher information geometry, then derives a hierarchy from simple magnitude pruning toward progressively more faithful geodesic approximations. Across fully connected networks and vision transformers on MNIST and CIFAR-10, over the complete zero-to-100% pruning range and five seeds, the proposed schemes outperformed magnitude and local-Fisher pruning on both accuracy and Matthews correlation. Intermediate approximations retained much of the gain with lower computational cost; the evidence remains limited to the reported models and datasets.
A 34-million-parameter module corrected 53.3% of tested errors while the matched evaluations showed no base-capability loss.
CRN v2 leaves a 4.65-billion-parameter language model frozen and learns a roughly 34-million-parameter logit correction layer on 83,400 pairs. On a 60-question domain exam it corrected 53.3% of base-model errors; a reworded version reached 43.3%. A matched LoRA baseline corrected more errors but lost 30% to 75% on the same small capability checks. Lowering the preservation penalty sharply reduced correction. The experiment supports frozen-base adjustment as a tradeoff, but its 60-question domain exam and small benchmark samples are not evidence of broad capability preservation.
A learned request router reached 0.864 mean goodput and matched round robin's seven-GPU result with six GPUs.
Disaggregated serving separates prompt prefill from token decoding, making request placement a latency and capacity problem. This study predicts completion time from prompt length, expected output, post-admission cache pressure and service class, then calibrates the scorer on eight NVIDIA A40 GPUs. Across three bursty traces, mean goodput reached 0.864 versus 0.835 to 0.847 for round robin, least-loaded and length heuristics. Calibration mattered: simulator-only constants erased 4.5 goodput points and much of the tail-latency gain. Under extreme scarcity, greedy routing concentrated work too aggressively, marking a clear operating boundary.
GPU, CPU and SSD tiering supported 73.02 times more sessions per GPU; predicted reuse and prefetching did not earn their cost.
Long-lived chats and agent loops fill expensive GPU memory with key-value cache blocks. A calibrated simulator compared recency, reuse frequency, predicted reuse and look-ahead prefetch across GPU HBM, CPU DRAM and SSD. The three-tier capacity model supported 73.02 times more concurrent sessions per GPU and cut cost per session 62.04 times, but those gains came from capacity rather than sophisticated placement. Recency minimized chat migration while reuse frequency led for agents and document work. Even an oracle prefetch policy failed to beat no prefetch on migration traffic, challenging recommendations that prediction alone improves this hierarchy.
A controller switched a frozen model among exploration, execution and reassessment using low-dimensional internal interventions.
Metacognitive Steering looks for process-level signals in scientist interaction traces rather than training only on finished scientific outputs. The authors identify a coordinated control surface across middle layers of a frozen mixture-of-experts model, then read its current regime and compose interventions for exploration, procedural convergence or critical reassessment. They report more sustained exploration, explicit pruning and evidence-responsive synthesis, and describe an autonomous system that reproduced eight BlueZ vulnerabilities and directed a rocket engineering project. Those demonstrations are author-reported case studies; independent replication is needed before treating the control method as reliable scientific judgment.
Generic code-screen file image, used illustratively; it does not depict BITCOS, ternary weights, kernels or benchmark results. Markus Spiske / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.
BITCOS reached 1.485 bits per weight on the sparsest tested model and improved decode throughput up to 1.27 times.
Ternary models are usually packed as if minus one, zero and plus one were equally likely. Measurements across 29 models found zeros reaching 51.5% of weights. BITCOS stores a presence bitmap plus compacted signs, costing two minus the zero density in bits per element. It beat five-trit packing in 26 of 29 tested models and reached 1.485 bits per weight on the sparsest. Optimized unpacking improved production kernel throughput up to 1.28 times and end-to-end decode up to 1.18 times on CPUs and 1.27 times on GPUs in the reported tests.
Executable ground truth improved expert-judge agreement 29% by MCC while reducing tokens 16% per test case.
Static answers go stale when a data-science agent is asked about live audience or business data. This framework encodes the expected result as a function that recomputes directly from the underlying database at evaluation time, then scores atomic facts regardless of whether the agent replies in prose, a table or HTML. In an author-reported human agreement study using a synthetic database shaped like a production system, ground-truth-as-code improved Matthews correlation with expert labels 29% over natural-language references and used 16% fewer tokens. A self-directed baseline without explicit ground truth was anti-correlated with human judgment.
Eight autonomous phantom resections produced negative margins across 77 cuts, with 1.61-millimeter mean absolute margin error.
Soft-tissue surgery changes the anatomy the robot is trying to track. This supervised research platform infers complete tumor, margin and kidney geometry from a single partial point cloud, updating its occupancy model as hydrogel tissue deforms and is cut. In eight consecutive open partial-nephrectomy phantom trials, dual robotic arms completed 77 electrosurgical cuts; every cut achieved a negative margin, with 1.61 ± 0.48 millimeters mean absolute margin error. These were patient-derived phantoms in a controlled open setting, not operations on people, and the results do not establish clinical autonomy.
Modifier-conditioned decoding improved force-direction following on a real whiteboard-wiping task while retaining speed control.
Contact-rich imitation learning usually reproduces an action without giving the operator a direct way to ask for slower, faster, gentler or firmer execution. Bi-MoDe injects a constrained modifier latent into every layer of a Transformer action decoder, allowing directives to alter action chunks. On a physical whiteboard-wiping task with combinations of temporal and force modifiers, it improved physical-directive following over the action-chunking baseline while maintaining comparable temporal control. The experiment demonstrates one task and robot setup, not a general natural-language safety interface.
Twenty operators reported less attention-shifting burden with registered mixed-reality overlays and similar torque regulation.
Underwater bilateral teleoperation can show reaction torque on a monitor, forcing the operator to shift attention between feedback and the workspace. MR-GLi spatially registers the torque indicator and wrist-camera view to the robot gripper. Twenty participants lifted and moved rigid and compliant objects in a counterbalanced comparison against the same feedback on a two-dimensional monitor. Torque-regulation performance remained similar, while subjective reports indicated less burden from shifting attention. The result supports information placement, not superior task performance or field readiness.
NASA OSAM-1 robotics file image, used illustratively; it does not depict AssemblyGrid, its simulated factory or reported results. Use does not imply NASA endorsement. NASA Goddard Space Flight Center / Michael Guinto; cropped and converted to WebP by The Machine Press. Use does not imply NASA endorsement.
AssemblyGrid separates task success from learning reward across flow, coalition and concurrency workloads.
Flexible production can require robots to route material, share resources, form temporary teams and work concurrently while each sees only local information. AssemblyGrid v1 combines those constraints with geometry-dependent feasibility in one repeatable benchmark. It defines three workload families at three scenario levels, with executable conformance checks and success metrics independent of any learning reward. Centralized references, structured decentralized controllers and IPPO, MAPPO and QMIX experiments all produced productive behavior. The benchmark is a controlled task-level environment, not evidence of autonomous deployment on a factory floor.
Without new demonstrations, geometric contracts reached 88.24% pick-and-place and 75.00% average functional-task success.
ManiSkillFormer replaces an end-to-end visuomotor policy with task-conditioned geometric contracts. Each skill declares the object keypoints, surface normals and other primitives perception must find; language-model agents generate contracts and motion templates inside human-defined skill structures. On a dual-arm robot, the authors report 88.24% zero-shot pick-and-place across eight object categories and 30 instances, 75.00% average success on unscrewing, pouring, pressing and folding, and 50% to 80% completion on three long-horizon tasks. The results are confined to the reported platform and skill library.
Stochastic consensus plus selective reconsideration cut mean displacement error 25.1% across 400 held-out aerial queries.
A vision-language waypoint planner normally emits one route without saying how reliable it is. UDAV samples multiple trajectories from aerial imagery, chooses the medoid and treats their spatial spread as uncertainty. When an interior waypoint crosses a threshold, it invokes a reconsideration stage. Across 400 held-out queries from two UAV flights, medoid selection cut mean average displacement error from 147.4 to 115.9 pixels; the complete planner reached 110.4 pixels, a 25.1% reduction, while returning valid trajectories for all queries. Pixel error on two flights is not evidence of safe autonomous navigation.
Builds a tactile digital camera around a Raspberry Pi Zero 2 W, an autofocus sensor, a small viewfinder, and mechanical-switch controls while publishing the software and printable hardware files.
Visit SATURNIXOriginal editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-09-13.
Combines a Raspberry Pi-class computer, an ultrasonic microphone, local storage, and a web interface into a modular recorder for active or unattended bat monitoring.
Visit WURB-2026Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-09-13.
BuilderOceanography for Everyone core team and contributors
Packages conductivity, temperature, and depth sensors with an Arduino-compatible controller, battery, and SD storage inside a user-built housing for nearshore research and education.
Visit OpenCTDOriginal editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-09-13.
Turns printable body parts, nineteen servomotors, a Raspberry Pi Zero 2 W, batteries, and shared control software into a compact humanoid platform for hands-on robotics work.
Visit MicrobanOriginal Codex Image Gen concept art from 2026-07-10; carried forward from the validated 2026-09-15 edition.
Desk Picks are selected by the newsroom. Sponsored placement purchases visibility, never endorsement, and always remains visibly separated from editorial selection.