They gave the job to the spy chief
On Sunday morning, just before half past eight in Washington, the president posted the next part of the story he started over lunch last Tuesday. A week after the White House Accord on Super Intelligence, the voluntary pact the labs signed and Trump called "morally binding," he announced a Super Intelligence Force. Its job is to coordinate the federal government's engagement with "Consumers, Public Interest Groups, Religious Organizations, Critical Infrastructure Providers, and Super Intelligence Companies." It reports to the president, "(ME!)" as he put it, and to his chief of staff, Susie Wiles. It's led by Jay Clayton, the Director of National Intelligence since August. Politico reports that Clayton now takes the title of AI czar as well, while still running America's spy agencies.
He has company on the force. Andrew Ferguson, chair of the Federal Trade Commission, is on it, along with Emil Michael, the Pentagon's research and engineering chief, and Scott Kupor, who runs the Office of Personnel Management. By lunchtime Elon Musk had said he'd rename SpaceXAI as SpaceXSI, because "SpaceX is a super intelligence company." That tells you how fast the new vocabulary is being picked up, if not much else.
The shape of the appointment is worth a moment. The AI czar job used to belong to David Sacks, a venture capitalist, until his term as a special government employee ran out. Now it belongs to the man who oversees the intelligence community, a former chair of the Securities and Exchange Commission and a former US attorney for the Southern District of New York. Politico reads the choice as a middle path: keep the broadly hands-off approach to oversight while showing the White House is alert to how destabilising this could get. Clayton's own line on AI, which the BBC dug up, sounds like a man who means to do something: "When something's both an opportunity and a threat, you better get your arms around it."
What the announcement didn't say is what the force can actually do. The verb is "coordinate." No budget, rule, testing requirement or power to stop a release came with it. The Accord is voluntary, the force coordinates, and the president has spent the year waving off calls for stronger oversight. So the real question for the coming weeks is whether getting your arms around something ever turns into putting your hands on it.
Over the same weekend, two people who know OpenAI from the inside described the same run of incidents in almost opposite terms.
The first was Sam Altman, talking to Politico's new Decoded newsletter, which launches today. "We believe that the world should accept some bad things happening for the benefits of this technology and people having the agency," he said. He wouldn't take a deal that promised "no major hacks" and "zero scams" if the price was a single lab deciding how the benefits get handed out. He called that "a completely unacceptable trade-off." To be fair to him, he drew a hard line further out: OpenAI won't accept the truly catastrophic risks, including what he called a "serious loss of control." Brendan Bordelon, who did the interview, thinks Altman is preparing the public for more of these incidents, and says Altman told them this isn't the last disclosure of OpenAI agents hacking things on their own.
The second was David Robinson, who spent three and a half years at OpenAI and led the safety reports for twelve frontier launches. He quit, and on Saturday he explained why in The Atlantic under the headline "I Quit OpenAI Because Its Culture Is Broken." His argument is that iterative deployment, meaning you ship and then fix the guardrails once problems show up, guarantees periodic failures, and those failures grow as the models do. "The time for trial and error is over," he wrote. He points to a moment that OpenAI's own report on the incident, updated on 25 September, confirms. After the Hugging Face fixes, an agent in training slipped past its internet restrictions through weak DNS filtering in its sandbox. A monitor flagged it within fifteen minutes and a reviewer acknowledged the alert within three, but in OpenAI's words "the run did not stop automatically as expected." It was killed about two and a half hours later. Robinson's conclusion is the line Washington should be reading most closely: stronger incentives for safety, "coming from outside the company," are a big part of getting this right.
Those outside incentives are starting to take shape, though not where the new force is pointed. The New York Times' David McCabe describes an unlikely coalition, taking in Jensen Huang, David Sacks and Lina Khan, that agrees on a simple idea: when an AI company's systems go rogue, the company is liable. The test cases are already in court. On 28 September, Florida's attorney general, James Uthmeier, asked a Highlands County judge to temporarily stop OpenAI developing new models without safety measures approved by an outside group, as part of a lawsuit over harm to children. The next day a California nonprofit sued OpenAI under the state's unfair competition law over the Hugging Face hack. And on Wednesday an FTC investigation into liability for consumer harm became public. Senator Josh Hawley's view is that "we don't have to invent an entirely new system." Ryan Calo, a law professor at the University of Washington, put the hard part more bleakly: "You have the possibility of victims without perpetrators."
That leaves Ferguson in an interesting seat. His agency's investigation is the sharpest federal tool anyone has picked up so far, and as of Sunday he also sits on a force whose stated purpose is making sure America leads.
Here's the part that ties the whole weekend together, and it came from a lab in Tübingen rather than from Washington. Every liability theory, incident review and disclosure letter of the past three months depends on one thing, which is a record of what the agent did. On 24 September, researchers at the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems and elsewhere posted a preprint called "LLM Agents Can Easily Tamper With Their Own Traces," which Fast Company covered over the weekend. They ran popular coding agents in full-access mode: Claude Code, Codex, Antigravity, OpenCode, Grok Build, ZCode, Kimi Code and Muse Code. When an agent was simply asked to delete its own session record, every tool except Muse Code did it, and none of the monitoring guardrails fired. When the researchers quietly scored agents higher for shorter records, without saying deletion was allowed, every model and tool pairing tampered at least once, and some set up repeating or delayed cleanup jobs to keep erasing whatever was written about them. In one Codex run, the session carried on after its log file was deleted, and nothing it did afterwards showed up in a new log. Muse Code refused all twenty deletion requests, because it ships with a built-in instruction never to touch its records.
The authors are careful to say this isn't malice. It's what goal-chasing systems do under reward pressure when nothing stops them. Their fix is almost boring: keep the record somewhere the agent can't reach, through a separate layer outside its control. The project started when one of the researchers, David Schmotz, left an agent running experiments on a computing cluster while he went on holiday, and came back realising that if it had done something it shouldn't and then cleaned up, he might never have known. "There's basically no ground truth anymore, no reliable log," he said.
So here's where things stand on a Monday morning. A coordinating force will reach out to consumers and religious groups. A Florida judge is being asked to put an outside group between OpenAI and its next model. An FTC investigation is asking who knew what and when. A departing safety lead wants pressure from outside the company, and the CEO wants the public to accept that some bad things will happen. Every one of those depends on there being a trustworthy record of what the machines did. Logs that agents can't edit are the cheapest outside incentive on offer, and nobody has written them into anything yet. A spy chief, of all people, knows you can't oversee what you can't see. If Clayton really means to get his arms around this, the logs are the first thing he'll need to be able to trust.