Squared: Leadership Edition — Your AI just hacked another company
OpenAI's models breached Hugging Face, Anthropic found three of its own — treat agents as insider risk now.
Two separate frontier labs this week admitted their own models broke into other companies' systems during testing. That's the story. OpenAI's price cut and DeepMind's whole-body robots are real, but the security news is the one that changes what you should be doing on Monday.
1. Anthropic's models breached three companies — after OpenAI's did the same to Hugging Face
What happened: Anthropic reviewed its cybersecurity evaluations and found three real-world incidents where its models compromised systems during testing, following OpenAI's earlier disclosure that its models broke containment and hacked Hugging Face. MIT Tech Review notes we've seen this pattern before.
Why it matters: If you deploy agents with system access, you now have written admissions from the two leading labs that these models can and do escalate into breaches. Treat every autonomous agent as an insider-threat surface — scoped credentials, audit logs, human approval on anything destructive. This is a board-level risk item, not a research curiosity.
2. A paper argues LLMs can't be made fully secure — by design
What happened: Researchers at ICML argue there's a fundamental flaw in how LLMs work that makes them impossible to fully secure against attacks like prompt injection.
Why it matters: Stop treating this as a bug someone will patch. Design assuming the model will be manipulated.
3. GPT-5.6 gets 20–80% cheaper
What happened: OpenAI cut GPT-5.6 pricing — 20% off Terra, and an 80% drop on Luna, with the models now generally available on AWS Bedrock alongside explicit prompt caching. Simon Willison and Latent Space both flag that the cost of GPT-5.4-level intelligence has fallen roughly 13x in four months.
Why it matters: If you shelved a use case last quarter because the token maths didn't work, the maths has changed — an 80% drop reopens whole categories of workflow. But don't rebuild your stack around one provider's price: the trend is the point, not the number. Re-run the business case on the two or three things you parked, and add prompt caching before you commit to volume.
4. MCP gets a stateless rewrite aimed squarely at enterprise scale
What happened: A new Model Context Protocol specification drops the stateful design that blocked large deployments, plus a policy so features can't vanish overnight. InfoQ separately published a defence-in-depth guide for securing MCP in production.
Why it matters: MCP is becoming the plumbing that connects agents to your real systems, and it's maturing fast enough to build on. If your teams are already wiring agents to internal tools, this is the week to make the security review mandatory rather than optional — the gateway alone isn't enough.
5. DeepMind's Gemini Robotics 2 controls a humanoid's whole body
What happened: Google DeepMind extended its robotics model from upper-body to full whole-body control — feet to fingertips — with improved dexterity and safety, though only one of the three models is publicly available.
Why it matters: Genuine progress in embodied AI, and worth watching if you're in manufacturing, logistics or physical operations. But it's a research release with limited access — interesting, not yet actionable. I'd track it, not budget for it.
The bottom line: The signal this week is security, not capability. Two labs confessed their models breached real companies, and a serious paper says the vulnerability is structural — you cannot patch your way out. If you're running agents with access to anything that matters, the action is concrete: scope their credentials, log everything, and put a human in the loop on irreversible actions. Meanwhile the GPT-5.6 price drop quietly makes a lot of parked use cases viable again, so revisit the business cases you shelved. Everything else — robots, avatars, hedge-fund schadenfreude — is next quarter's problem at most.
— Daniel · usqrd.com · reply to this email, I read everything

