A black-box audit across more than 120 language models found that safety, security, privacy and benign usefulness can move in different directions.
Editorial illustration
Conceptual illustration: aiXamine evaluates safety, security, privacy and usefulness as coupled properties; the reported benchmark does not certify any model for deployment. Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-24.
The aiXamine preprint combines 46 tests across nine services so that alignment, adversarial robustness, privacy and over-refusal are measured as related properties rather than separate leaderboards. Across more than 5,000 runs, the authors report three recurring trade-offs: stronger safety enforcement often rejected more benign requests; privacy was nearly orthogonal to the other trust dimensions; and one off-policy distillation setting collapsed robustness from 56.9 to 2.6 on the same base architecture. These are benchmark findings, not a universal ranking of deployed models, but they make a practical point: a high score on one trust axis cannot certify the others.
Conceptual illustration: EndoLIFT uses language to disambiguate advance and withdrawal in phantom and ex-vivo testing; it is not a clinical result. Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-24.
EndoLIFT uses an explicit language instruction to distinguish advance from withdrawal when an endoscope camera image alone is ambiguous.
Routine endoscopy changes direction: an instrument advances toward anatomy, withdraws for inspection and may reverse early when an operator requests it. The EndoLIFT preprint calls the resulting same-view, opposite-action problem intent aliasing. Its policy conditions continuous action chunks on RGB video, prior action state, language and a learned trajectory latent. Controlled instruction swaps showed that language selected the axial direction, while the latent improved directional correctness and retraction; the system retained 82.8 percent intent-following accuracy across 44 held-out phrasings and completed ten of ten ex-vivo porcine-trachea trials. Those phantom and ex-vivo results are not human clinical validation, but they isolate why an action policy sometimes needs the operator's stated intent, not another look at the same frame.
A deployed tender-response pipeline found an asymmetry between structural extraction and structural instruction conditioning.
Rendering source documents as markup improved three reading tasks, but converting bid instructions from prose to nested XML dropped answer quality from 74 to 48 percent. In a blind comparison, 68 percent of identified gaps came from information absent from the supplied sources, separating unavailable knowledge from avoidable writing defects.
Conceptual Visualising AI file image; it does not depict STCO, fluid fields, the evaluated operators or reported errors. Novoto Studio / Google DeepMind / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.
STCO conditions neural operators on prescribed motion, inflow and force fields at the target time.
Across twelve matched neural-operator backbones and an immersed-boundary fluid benchmark, the added interface reduced relative field error by a mean 31.1 percent and normalized load error by 24.7 percent. The study treats a control input as part of the query rather than asking past state alone to imply a future intervention.
AEGIS maps heterogeneous MCP tool arguments into a common resource-policy representation.
The proposed enforcement layer uses a language model to normalize text, image, video and location requests before Open Policy Agent rules evaluate them. The paper focuses on cross-domain resource abuse such as excessive search radii or long generated videos; it presents an architecture and integration, not evidence that every attack will be detected.
Three specialist agents jointly choose topology, resources and aggregation before a non-LLM feasibility check.
On a non-IID CIFAR-10 simulation, FL-MAESTRO matched the strongest energy-aware baseline while cutting wasted round energy from more than one third to near zero. The reported saving comes from withholding clients predicted to fail before their updates could be aggregated; the result remains a benchmark study, not a field deployment.
Koala Gripper co-designs a handheld data-capture device with the robot mechanism that will replay its demonstrations.
The platform combines a force-optimized trigger linkage, a monolithic dual-thumb and backdrivable fingers with effective mass measured in tens of grams. The team demonstrated varied grasps, tool use, singulation and an end-to-end learning-from-demonstration pipeline, arguing that data collection ergonomics should shape execution hardware from the start.
Checkpoint instrumentation separates long-horizon failures that happen before and after an agent reaches the relevant state.
In one 92-seed study, protocol-disambiguation guidance raised state observation for Gemini 2.5 Flash from 65.5 to 95.4 percent, but repeating the design with Gemini 3.7 Flash produced the opposite effect. The shifting bottleneck shows why final success alone cannot reveal which capability failed or whether it was exercised at all.
Illustrative programming file image; it does not show FlavourBench, Epicure, any model response, ingredient portfolio or leaderboard result. Daniil Komov / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.
FlavourBench freezes executable scores for all 56 three-ingredient portfolios in each task.
The benchmark evaluated 27 frontier endpoints on 534 identical tasks, totaling 14,418 scored model-task cells without differential missingness. Its two independently compiled panels correlated at 0.89, and 101 of 351 model pairs were statistically resolved; the authors release prompts, raw responses, score maps and an offline verifier.
ACES measures capability packages through paired live trials under a fixed model, workspace, sandbox and scorer.
Across 947 paired cases from 58 production skills and four harnesses, the preprint reports a mean composite Skill Lift of 0.2134 and positive lift in 72.8 percent of cases. Static scans and runtime outcomes were only weakly correlated, suggesting that document quality and actual agent benefit are complementary gates.
An ontology-driven finance system tied every selected fact to a complete provenance chain.
On FinanceBench, KDAF and BM25 were statistically indistinguishable on answer correctness, but KDAF reached citation-traceability F1 of 0.515 and admitted no evidence outside the question's subject entity. The negative accuracy result sharpens the claim: structured retrieval's measurable advantage here was auditability.
DreamBench-SWE uses hidden executable oracles for software tasks that depend on non-inferable evidence from earlier sessions.
In the preregistered successor audit, no external memory passed 21 of 180 tasks, deterministic verbatim memory passed 82 and one pinned hosted configuration passed 97. The authors explicitly stop short of claiming mechanism, general product superiority or equivalence, making the benchmark a profile of exact conditions rather than a universal memory ranking.
Why2Speak finds a capability-auditability trade-off when an assistant decides whether to intervene in conversation.
The strongest direct policy performed better but exposed no reasoning; a reasoning policy offered an inspectable trace at lower quality, especially on true intervention opportunities. Supervised and reinforcement learning did not repair the trade-off, and controlled probes showed that common faithfulness tests can confuse observability with a changed inference policy.
Illustrative laboratory-glassware file image; it does not depict the binder candidates, proxy scores, biological targets, experiments or results. Rodolfo Clix / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.
Language models combined precomputed structural proxy scores to rank existing protein-binder candidates.
On a ten-target held-out split, five sampled global policies reached 0.589 Recall@10 versus 0.571 for the strongest single-feature baseline. The authors frame the method as an interpretable post-generation decision layer for scarce wet-lab slots, not a new binder generator or evidence of biological efficacy.
A paired rollback audit reconstructed benchmark movement without finding a unique coarse module that explained it.
Across aligned Gemma-to-MedGemma and Qwen-to-HuatuoGPT pairs, the full decoder update strongly reconstructed measured medical multiple-choice movement. MLPs were the strongest broad family, but off-domain changes and matched controls prevented a single-component explanation; the study makes no clinical-validation claim.
CAS replaces fixed top-K retrieval with adaptive prediction sets and penalizes low-confidence training trajectories.
The framework applies conformal methods on both sides of search-agent training: document sets expand or contract to target coverage, while answer confidence weights the GRPO objective. The preprint reports higher reasoning accuracy and fewer redundant tool calls across single- and multi-hop QA datasets.
A self-contained touchscreen computer boots into MicroPython and combines programmable graphics, MIDI connections, and an open-source synthesizer in focused, buildable hardware.
Visit Tulip Creative ComputerOriginal editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-23.
Shrinks an intaglio press into 3D-printed tabletop hardware, with free fabrication plans for makers and finished presses for artists without a printer.
Visit Open Press ProjectOriginal editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-23.
BuilderRichard Bowman and OpenFlexure contributors
Prints most of a precise microscope body as one flexure mechanism, pairing interchangeable optics with optional motorized sample positioning and open control software.
Visit OpenFlexure MicroscopeOriginal editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-08-23.
BuilderLibre Space Foundation and SatNOGS contributors
Links volunteer-built radio ground stations to shared scheduling and observation services, turning backyard antennas into a public satellite-observation network.
Visit SatNOGSOriginal Codex Image Gen concept art from 2026-07-10; carried forward from the validated 2026-08-23 edition.
Desk Picks are selected by the newsroom. Sponsored placement purchases visibility, never endorsement, and always remains visibly separated from editorial selection.