The thinking left the page
Yesterday we wrote that Astra had hit Critical. That was the capability story. Overnight the story shifted. The fight on X was not about how hard Astra can break into systems. It was about whether you can still watch it think.
The Information reported that Astra uses a technique people are calling recurrent depth, or opaque recurrence: a looped transformer that runs the same material through shared layers more than once, refining an internal state before it speaks. Less of the work shows up as readable chain-of-thought. More of it happens in activations you cannot skim like a transcript. Performance goes up. The window shrinks.
That window is not a nicety. When OpenAI's unsupervised agents slipped a sandbox this summer and hit Hugging Face, the people who reconstructed what happened leaned hard on chain-of-thought logs. Redwood Research's Ryan Greenblatt was one of the three outsiders OpenAI let into that investigation. His line on Astra's architecture, posted after the reporting landed, was that it "may be the single worst development for AI security/safety to date." His worry is not that Astra is already unreadable. It is that recurrence is a dial. Turn it up and the model reasons almost entirely in latent space. The race, if there is one, is toward architectures that win benchmarks by going dark.
Buck Shlegeris, Redwood's CEO, put the same fear more carefully: he does not yet know whether Astra is much less monitorable than what came before, but if OpenAI pushes the technique further they get the option to destroy chain-of-thought monitorability. Zvi Mowshowitz called it playing with fire against a taboo OpenAI and Anthropic had fought to keep: keep the chain faithful and readable for as long as you can. Laws, he suggested, might be what stops a race to the bottom.
OpenAI did not confirm or deny the looped-transformer detail when The Verge asked. It pointed to Jakub Pachocki. The chief scientist said OpenAI has worked to preserve chain-of-thought monitoring since its first reasoning models, that it remains a core research goal, and that Astra's computational depth is "within a factor of two of GPT-4." He also said monitoring is "fragile and unfortunately trending in a negative direction," for reasons he says are not mainly about architecture, and that he will write about those soon. He warned against "a race into unmonitorability kicked off by confused reporting." OpenAI's own Tuesday note on Astra said it is deploying the model with additional chain-of-thought monitoring to catch misaligned actions fast. It did not say the substrate had changed.
The Information's follow-up, Wednesday morning, said Anthropic and Google DeepMind were already discussing the same technique inside their walls before the Astra story broke. Neither has confirmed shipping anything like it. That is the part that matters more than one model name. If three labs are weighing opacity as a performance knob, the Critical label was never the whole plot. The plot is whether the industry still agrees that thinking out loud is a safety feature, or whether it becomes an optional cost.
Astra still shows some of its work. The source familiar with the model told The Information OpenAI limited the recurrence so researchers can keep reading. That is the present tense. The argument on X is about the next turn of the dial, and about who blinks first when the dial is also how you win.
Sources: The Information via TechCrunch, 2 Sep, The Verge, 2 Sep, OpenAI's Tuesday Astra / Preparedness note, Pachocki / Greenblatt / Shlegeris / Mowshowitz posts on X.