TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionSources linked throughout
Front pageImportance 10/10

A Passport Office for AI Agents

The ITU has opened a standards forum for proving who an agent is, what it may do, and where human control must remain.

Sepia engraving of mechanical messengers presenting distinct geometric identity tokens to a central registry instrument.Editorial illustration
Conceptual illustration: autonomous messengers present identity tokens to a common registry. It does not depict the ITU meeting or a deployed identity system. Generated with Codex Image Gen for The Machine Press, 2026-07-10.

The International Telecommunication Union launched a Focus Group on Trust and Identity for Humans and Agentic AI in Geneva on July 9. Its remit is foundational rather than ceremonial: technical, policy, and legal experts will work on ways to identify autonomous agents, establish trustworthy interactions, and keep meaningful human control around sensitive transactions and infrastructure. No binding global standard exists yet; the group is the workshop where proposals can become technical reports or specifications before any later standards process.

research
Vintage laboratory engraving of four mechanical instruments cycling abstract research cards around a glowing central experiment vessel.Editorial illustration
Conceptual illustration: specialized machines move through hypothesis, observation, analysis, and revision around one experiment. It is not a real laboratory or medical result. Generated with Codex Image Gen for The Machine Press, 2026-07-10.

The Lab Loop Gets New Assistants

Nature pairs two independent multi-agent systems that generate hypotheses, propose experiments, interpret results, and return revised ideas to scientists.

Nature's July 9 issue places two agentic science systems side by side. Google DeepMind's Co-Scientist uses specialized Gemini-based agents to generate, critique, rank, and refine hypotheses; the paper reports expert-guided wet-lab validation across biomedical problems. FutureHouse's Robin connects literature search and data-analysis agents in a lab-in-the-loop process that identified and tested candidates for dry age-related macular degeneration. The striking result is not a machine replacing a laboratory. Both teams describe systems built around experimental validation and scientists who choose goals, constraints, and what to test next.

Today's Dispatches

chips infrastructure01
Rendered wafer-scale integrated circuit with a pale gold rectangular chip field on a dark circular substrate.File image
Illustration: a generic rendered wafer-scale circuit. It does not depict Meta's Iris design, a fabricated chip, or the reported production line. Wikideas1 / Wikimedia Commons (CC0 1.0); cropped and converted to WebP by The Machine Press.

Meta Schedules Its Own AI Chip for September

An internal memo reviewed by Reuters points to a faster custom-silicon cadence as Meta aims to expand its computing base.

Meta plans to begin manufacturing an in-house AI chip code-named Iris in September, according to an internal memo reviewed by Reuters. The report says Iris is part of a four-generation MTIA program and that Meta intends to release a chip about every six months through 2027 while targeting 14 gigawatts of overall computing power next year. The memo describes a plan, not a completed production run; yield, deployment, and performance remain to be demonstrated.

developer tools02

GPT-5.6 Lands Across GitHub Copilot

Sol, Terra, and Luna are rolling out across editor, command-line, cloud-agent, mobile, and web surfaces.

GitHub began rolling OpenAI's GPT-5.6 family into Copilot on July 9. GitHub positions Sol for complex, long-running work, Terra as the balanced default, and Luna for smaller, lower-cost tasks. Availability varies by plan, rollout is gradual, and Business and Enterprise administrators must explicitly enable the models because their policy starts off by default.

market industry03

Investors Ask Who Benefits—and Who Endures

At Reuters NEXT Asia, large funds described a shift from buying the AI label to testing infrastructure returns and business resilience.

Asian institutional investors are widening their AI lens from model makers to companies that can benefit from the technology without being easily displaced by it. At Reuters NEXT Asia, Temasek said it aims to raise AI exposure, while fund managers stressed durable cash flows, infrastructure economics, and the risk that heavy spending may not produce commensurate returns. The comments mark portfolio caution, not a verdict that the AI investment cycle has ended.

safety security04
Dark server-room aisle lined with black cabinets and blue-green equipment lights.File image
Illustrative infrastructure photograph: a server room at The National Archives in the UK, not a tested machine, compromised repository, or AI-hosting facility. The National Archives (UK), via Wikimedia Commons, CC BY 3.0; cropped and converted to WebP by The Machine Press.

The Security Reviewer Can Become the Entry Point

A proof of concept shows coding agents in autonomous modes executing attacker-controlled material while inspecting an untrusted repository.

AI Now researchers Boyan Milanov and Heidy Khlaaf demonstrated what they call Friendly Fire: prompt injections in a third-party codebase led Claude Code and Codex configurations with autonomous command approval enabled to run attacker-controlled code during a security review. The finding is configuration-specific, not evidence that every interactive review is compromised. Its practical lesson is broader: untrusted repositories should be isolated, and automatic execution should not be treated as a neutral extension of code reading.

frontier models05

Meta Opens a New Door for Agent Builders

Muse Spark 1.1 arrives with a public-preview model API, long context, and an emphasis on orchestrating tool-using agents.

Meta introduced Muse Spark 1.1 on July 9 as a multimodal reasoning model aimed at agentic work. The company says the model can plan, delegate across parallel subagents, use computers and tools, and manage a one-million-token context window. More consequential for outside builders, Meta also opened a public preview of its Meta Model API. The performance claims remain Meta's own evaluations, but the API changes the release from a consumer-app update into a developer platform move.

media creative tools06

Image Generation Starts Calling Its Own Tools

Meta's Muse Image uses search, code, and self-refinement before it returns a picture; Muse Video remains an early preview.

Meta released Muse Image across Meta AI and selected social surfaces, describing an image model that can invoke search and coding tools, revise its own generations, and compose multiple references. The accompanying Muse Video model is not yet a general release. Meta says images created in its products carry an invisible Content Seal provenance signal intended to survive common edits such as cropping and compression. Rankings and quality comparisons in the announcement are time-stamped company claims, while the release and watermarking mechanism are the durable news.

open source07
Black source-code symbol formed by two angle brackets and a slash on a warm gold background.File image
Generic source-code illustration; it does not depict a named repository, license, Copilot interface, or generated overview. D. Charbonnier / The Noun Project, via Wikimedia Commons (CC0 1.0); padded, gold background added, and converted to WebP by The Machine Press.

Copilot Adds a First-Pass Map for Unfamiliar Repositories

GitHub's new overview gathers purpose, technologies, and contribution guidance before a newcomer starts reading file by file.

GitHub now offers Copilot-generated overviews on repositories a user has not contributed to before. The summary draws from repository context to describe purpose, technologies, and contribution guidance, with shortcuts for recent changes and ways to contribute. GitHub says the feature is available across all Copilot plans; maintainers should still treat a generated overview as orientation, not a substitute for authoritative project documentation.

robotics08

LeRobot Closes More of the Learning Loop

Version 0.6 adds world-model policies, reward models, richer datasets, unified evaluation, and a correction-driven rollout workflow.

Hugging Face's LeRobot 0.6 release expands the open robotics toolkit beyond training individual policies. It adds policies that learn to anticipate future states, a common reward-model interface, six simulation benchmark families under one evaluation command, depth and language annotations for datasets, and a rollout tool that can record human corrections when a policy fails. The project also reports faster video-data loading and a leaner base installation. Together, the changes make repeated deployment, evaluation, correction, and retraining a more coherent open workflow.

safety security09

Biosafety Red Teams Get a Standing Target

OpenAI is turning its model-specific bio jailbreak challenge into an ongoing private bounty program with higher top rewards.

OpenAI says its Bio Bounty Program will now continue across frontier-model releases rather than end with a single model cycle. The private program asks vetted researchers to find one universal jailbreak that defeats a ten-question biological and chemical safety challenge. For GPT-5.6 and GPT-5.5, the first qualifying universal jailbreak can earn $50,000, up from $25,000. The invitation is rolling and participation requires selection and an NDA, so this is a structured red-team program rather than a public exploit contest.

benchmarks evals10

A Coding Benchmark Loses Its Signal

An OpenAI audit estimates that roughly thirty percent of SWE-Bench Pro's public tasks are broken or unfairly specified.

OpenAI audited the 731-task public split of SWE-Bench Pro after pass rates climbed sharply. Its automated review flagged 27.4 percent of tasks as broken, while a separate campaign with five experienced engineers per task marked 34.1 percent. The recurring problems were overly strict tests, underspecified prompts, low test coverage, and misleading instructions. Because coding benchmark scores inform capability and safety judgments, OpenAI withdrew its earlier recommendation to use SWE-Bench Pro and called for evaluations designed explicitly for model testing.

infrastructure11

Open Weights Get an Enterprise Loading Dock

A curated Hugging Face collection can now be deployed on Microsoft Foundry Managed Compute behind one governed endpoint.

Microsoft and Hugging Face are previewing a weekly refreshed collection of open-weight models for Foundry Managed Compute. Candidate models pass license and security screening, avoid unreviewed remote-code paths, and ship with signed, scanned runtimes matched to their architectures. Weights are pre-staged in Azure, while deployment uses the same identity, networking, monitoring, and billing surface as other Foundry models. The arrangement narrows the operational gap between finding an open model and running it inside an enterprise boundary.

policy12

Anthropic Promises a Ledger for Hard Questions

The company is asking the public what worries or excites them about AI and says it will track its concrete responses in public.

Anthropic launched a public call for questions about AI's effects on work, families, safety, science, and human agency. The company says it will publish the actions it takes in response and acknowledge where it falls short. The effort follows a first Public Record survey of 52,000 Americans, interviews with 81,000 Claude users across 159 countries and 70 languages, and smaller focus groups. The announcement is a company commitment rather than an independent accountability mechanism, but its promised public record creates a standard readers can later check.

Independent builders

The Invention Desk

Independent builders turning improbable ideas into real things.

  • Four editorial selections
  • $7 for seven days
  • Paid work is clearly labeled
  • Placement is never endorsement
Brass and ivory geometric algebra tokens arranged on a precise transformation table.
Original editorial concept art generated with Codex Image Gen for The Machine Press.
Desk PickReleased

Wyrm Math

Builderdicroce

Turns algebra into a gesture puzzle while an open-source exact engine makes invalid transformations impossible.

Visit Wyrm Math
Modular brass nodes and cables drive a plotter drawing abstract forms on paper.
Original editorial concept art generated with Codex Image Gen for The Machine Press.
Desk PickPrototype

SubjectiveZero

BuilderClem / SXP Studio

Moves creative coding from a high-level prompt into an editable node graph and native Swift and Metal code.

Visit SubjectiveZero
A constellation of paper cards forms a research atlas around a brass globe.
Original editorial concept art generated with Codex Image Gen for The Machine Press.
Desk PickBeta

Tomesphere

Builderleonickson / Tomesphere

Maps millions of open papers into an explorable research atlas with enriched paper pages, browser tools, and MCP access.

Explore Tomesphere
A circular city rail loop is surrounded by bells, horns, and flowing ribbons of sound.
Original editorial concept art generated with Codex Image Gen for The Machine Press.
Desk PickReleased

Yamanote.fun

BuilderPaul Jackson & Claude

Recreates Tokyo's circular Yamanote journey as an offline-capable soundscape of station melodies, chimes, and announcements.

Ride the soundscape
An unnamed prototype device sits under a desk lamp beside a blank placement card.
Original editorial concept art generated with Codex Image Gen for The Machine Press.
Sponsored ProjectOpen

The First Paid Slot

BuilderHouse example / The Machine Press

A transparent preview of a paid builder placement: one concise dream, one verified link, and no claim of endorsement.

House example - no advertiser paid for this card. Future paid cards will carry this same prominent Sponsored Project label.

Ask about the launch slot
Six portfolio slots surround one illuminated open slot and seven small day markers.
Original editorial concept art generated with Codex Image Gen for The Machine Press.
Open PlacementOpen

Put Your Project on the Desk

$7$1/day / 7 days

Reach readers curious about what people are building. One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

Manual launch intake; automated checkout is not live yet. Payment buys placement, never endorsement, and every submission is reviewed before publication.

Email the desk

Desk Picks are selected by the newsroom. Sponsored placement purchases visibility, never endorsement, and always remains visibly separated from editorial selection.