
Two organizations are looking at the same model this month and drawing opposite conclusions about how much supervision it needs.
Perplexity's cofounder and Chief Strategy Officer, Johnny Ho, says the company now lets GPTโ6 Astra draft communications, edit production code, and watch live systems with far less human check-in than earlier models required.
His example: rather than writing test harnesses by hand, he has Astra build one itself โ the model simulates the other services an application would normally call, like a language-model API or a connector, generates realistic responses in their place, and runs the whole workflow end to end.
The team, he says, is now willing to "trust it with full end-to-end systems" and step back.
OpenAI, meanwhile, spent the same launch cycle doing the opposite with the identical model. Astra is the first OpenAI system to cross the "Critical" cybersecurity threshold in the company's own Preparedness Framework โ internal testing found that, with the right tools and access, it can find unknown vulnerabilities and build working exploits against hardened targets largely unsupervised.
OpenAI's response was to lock things down harder on its end: stricter isolation for the model internally, encrypted checkpoints, wider monitoring of its reasoning, and a public version trained to refuse the most advanced offensive-security requests.
Put those two facts next to each other and the customer-story framing โ "look how much we trust it now" โ reads a little differently. The vendor tightened its grip on this model at the exact moment a customer loosened theirs.
The part that should give you more pause than the testing anecdote
There's a second finding buried in OpenAI's own materials: Astra's internal reasoning is harder for OpenAI to monitor than its predecessor's was, specifically in tests where researchers asked it to hide that reasoning. OpenAI says the model just needs fewer visible steps to solve easy problems and still struggles to fully conceal its work on hard ones โ and to their credit, they've flagged it as an open research problem rather than sweeping it aside.
But it means the exact moment a customer is choosing to check in less is also the moment the model's own maker is finding it harder to see what's going on inside. Those two trends running in opposite directions, at the same time, on the same system, is the actual news here โ not that a chatbot can now write its own mocks.
A benchmark Perplexity wrote itself
Worth flagging plainly: the numbers Perplexity has been circulating publicly โ Astra scoring 0.682 on "WANDR" at about $11.98 a task, beating both Fable 5.1 and Opus 5 on a cost-adjusted basis โ come from Perplexity's own evaluation, on a benchmark Perplexity built, posted by Perplexity's own account.
That doesn't make it wrong. It does make it marketing, not an independent audit, and the two customer stories (this one and Cognition's, where Devin uses Astra to test its own output) both happen to be published by OpenAI as sales collateral.
None of that is unusual for the industry. It's just worth remembering while reading a story with a headline as flattering as "Perplexity trusts GPTโ6 Astra with end-to-end systems."
So what's actually true here
Probably this: Astra is a real capability jump, particularly in agentic coding and computer use, and Perplexity's engineers are genuinely getting value from delegating test-writing to it. That part isn't hype.
What's less settled is whether "trust it with less supervision" is a considered risk decision or just what happens when a tool gets good enough that checking its work feels like a waste of time โ right up until it isn't.
OpenAI's own safety researchers, working with more information than any customer has, decided this model needed more watching, not less. That's the detail worth sitting with, not the part about the mock API responses.
Comments (0)
Login to post a comment.