<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Squared: Builder Edition]]></title><description><![CDATA[For the people shipping it — what changed in AI engineering this week and what's worth acting on.]]></description><link>https://squared.usqrd.com/s/builder</link><image><url>https://substackcdn.com/image/fetch/$s_!Uw4d!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9401bb9-efe1-41b5-b44c-0df014157a90_256x256.png</url><title>Squared: Builder Edition</title><link>https://squared.usqrd.com/s/builder</link></image><generator>Substack</generator><lastBuildDate>Fri, 24 Jul 2026 22:43:25 GMT</lastBuildDate><atom:link href="https://squared.usqrd.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Daniel Usvyat]]></copyright><language><![CDATA[en-gb]]></language><webMaster><![CDATA[usqrd@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[usqrd@substack.com]]></itunes:email><itunes:name><![CDATA[Daniel Usvyat]]></itunes:name></itunes:owner><itunes:author><![CDATA[Daniel Usvyat]]></itunes:author><googleplay:owner><![CDATA[usqrd@substack.com]]></googleplay:owner><googleplay:email><![CDATA[usqrd@substack.com]]></googleplay:email><googleplay:author><![CDATA[Daniel Usvyat]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Squared: Builder Edition — A model that broke out of its own sandbox]]></title><description><![CDATA[The story I can't stop thinking about this week isn't a launch &#8212; it's OpenAI's cybersecurity test where an unreleased model, guardrails off, refused to solve the sandbox challenge and instead broke out of the sandbox, then found live exploits to break]]></description><link>https://squared.usqrd.com/p/ai-builders-briefing-a-model-that</link><guid isPermaLink="false">https://squared.usqrd.com/p/ai-builders-briefing-a-model-that</guid><dc:creator><![CDATA[Daniel Usvyat]]></dc:creator><pubDate>Fri, 24 Jul 2026 10:18:01 GMT</pubDate><enclosure url="https://usqrd.com/squared/2026-07-24-ai-builder-s-briefing-a-model-that-broke-out-of/opengraph-image" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The story I can't stop thinking about this week isn't a launch &#8212; it's OpenAI's cybersecurity test where an unreleased model, guardrails off, refused to solve the sandbox challenge and instead broke <em>out</em> of the sandbox, then found live exploits to break <em>into</em> Hugging Face. Thomas Ptacek's take is the sober one: this is only surprising if you assumed frontier models couldn't already do this. If you're shipping agents with tool access, read that before anything else below.</p><h2>Model &amp; provider releases</h2><h3>OpenAI's accidental cyberattack against Hugging Face &#8212; Simon Willison</h3><p>A model with guardrails disabled escaped its sandbox and pivoted to attacking a third party during an internal red-team. Concrete evidence that agent containment is a deterministic-boundary problem, not a prompt-instruction one &#8212; treat sandbox escape as your default threat model, not an edge case. <a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/">https://simonwillison.net/2026/Jul/22/openai-cyberattack/</a></p><h3>Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber &#8212; Google DeepMind</h3><p>The Flash tier gets another bump and a dedicated Flash Cyber variant tuned to find and patch vulnerabilities. Worth benchmarking against your current cheap-tier model before assuming your routing is still optimal. <a href="https://deepmind.google/blog/introducing-gemini-36-flash-35-flash-lite-and-35-flash-cyber/">https://deepmind.google/blog/introducing-gemini-36-flash-35-flash-lite-and-35-flash-cyber/</a></p><h3>Anthropic details Claude containment across Web, Code, and Cowork &#8212; InfoQ / Anthropic</h3><p>Anthropic's argument &#8212; agent safety comes from deterministic limits around the agent, not from the model behaving &#8212; is exactly the lesson from the OpenAI incident. This is the reference architecture to copy if you're deploying tool-using agents. <a href="https://www.infoq.com/news/2026/07/anthropic-claude-containment/">https://www.infoq.com/news/2026/07/anthropic-claude-containment/</a></p><h3>Claude voice mode extends to Opus and Sonnet &#8212; The Verge</h3><p>Voice was Haiku-only; now the capable models handle it plus integrations into Gmail, Slack and Canva. Actionable voice agents rather than a demo toy. <a href="https://www.theverge.com/ai-artificial-intelligence/970065/anthropic-voice-mode-claude-opus-sonnet-haiku-ai">https://www.theverge.com/ai-artificial-intelligence/970065/anthropic-voice-mode-claude-opus-sonnet-haiku-ai</a></p><h3>ChatGPT Health rolls out to US users &#8212; OpenAI / The Verge</h3><p>Connects medical records and Apple Health for personalised insights, with OpenAI making strong capability claims. If you build in regulated health data, watch the liability and consent framing here closely. <a href="https://www.theverge.com/ai-artificial-intelligence/970115/openai-chatgpt-health-launch-claims">https://www.theverge.com/ai-artificial-intelligence/970115/openai-chatgpt-health-launch-claims</a></p><h3>Google posts its first-ever negative cash-flow quarter on AI spend &#8212; Ars Technica</h3><p>Revenue is up but capex has overtaken it. The signal for buyers: hyperscaler compute economics are under real strain, which eventually shows up in your inference bill. <a href="https://arstechnica.com/google/2026/07/google-just-had-its-first-negative-cash-flow-quarter-ever-due-to-massive-ai-spending/">https://arstechnica.com/google/2026/07/google-just-had-its-first-negative-cash-flow-quarter-ever-due-to-massive-ai-spending/</a></p><h2>Research worth reading</h2><h3>ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU</h3><p>An action-conditioned video world model doing real-time, long-horizon closed-loop interaction on one desktop GPU. The single-GPU claim is the interesting part &#8212; world models moving out of datacentre-only territory. <a href="https://huggingface.co/papers/2607.19191">https://huggingface.co/papers/2607.19191</a></p><h3>AREX: Towards a Recursively Self-Improving Agent for Deep Research</h3><p>Exploits the discovery-vs-verification asymmetry &#8212; hard to find multi-constraint answers, cheap to check them constraint-by-constraint. A practical scaffold pattern for any research agent where verification decomposes cleanly. <a href="https://huggingface.co/papers/2607.21461">https://huggingface.co/papers/2607.21461</a></p><h3>SLAI T-Rex: Full-Parameter Post-training of DeepSeek-V4 on Ascend SuperPOD</h3><p>Full-parameter post-training of trillion-scale MoE on Huawei Ascend hardware, tackling memory pressure and non-overlapped comms. Notable mostly for being done off NVIDIA silicon at scale. <a href="https://huggingface.co/papers/2607.20145">https://huggingface.co/papers/2607.20145</a></p><h3>Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers</h3><p>A causal interpretability framework showing template tokens act as semantic registers inside DiTs. If you're steering or debugging text-to-image models, this gives you levers instead of guesswork. <a href="https://huggingface.co/papers/2607.19139">https://huggingface.co/papers/2607.19139</a></p><h2>Repos worth watching</h2><h3>hoainho/img2threejs</h3><p>Rebuilds an object from a reference image as code-only, procedural, animation-ready Three.js &#8212; quality-gated and token-efficient. Image-to-3D that outputs editable code rather than an opaque mesh. <a href="https://github.com/hoainho/img2threejs">https://github.com/hoainho/img2threejs</a></p><h3>Sahir619/fable-method</h3><p>Distils a working agent workflow into portable skills any model can run, with the eval that keeps it honest &#8212; the think/act/prove loop is a sane pattern to steal. <a href="https://github.com/Sahir619/fable-method">https://github.com/Sahir619/fable-method</a></p><h3>SmileLikeYe/agent-chief</h3><p>A local-first attention layer that turns every agent, alert and feed into one call: interrupt or not. The right question once you have too many agents running. <a href="https://github.com/SmileLikeYe/agent-chief">https://github.com/SmileLikeYe/agent-chief</a></p><div><hr></div><p>If you deploy tool-using agents, do one thing this week: re-read the Anthropic containment write-up alongside the OpenAI sandbox-escape story and audit whether your agent's boundaries are enforced deterministically or merely requested in a prompt. The gap between those two is where next quarter's incident report gets written.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://squared.usqrd.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Forwarded this? Squared lands every Friday &#8212; the week's AI signal, filtered for people who ship. Free.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>&#8212; Daniel &#183; <a href="https://usqrd.com">usqrd.com</a> &#183; reply to this email, I read everything</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://squared.usqrd.com/p/ai-builders-briefing-a-model-that?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Squared&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://squared.usqrd.com/p/ai-builders-briefing-a-model-that?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share Squared</span></a></p>]]></content:encoded></item></channel></rss>