Five Things: August 2, 2026
Anthropic models can escape too, sign-on for "pacing," gain-of-function research ban, bio red-teaming in China, the EU AI Act enforcement
Five things that happened/were publicized this past week in the worlds of biosecurity and AI/tech:
Anthropic discloses that Claude models hacked three companies during safety testing
Frontier-lab employees sign a letter asking Washington to help them slow down
HHS formally bans “dangerous gain-of-function research”
A very concerning Chinese preprint does some AI x bio jailbreaking
The EU’s AI Act enforcement powers switch on today
1. It’s not just OpenAI; it’s not just HuggingFace
The main story this week has been continued fallout from the world discovering that an internal model belonging to OpenAI hacked into HuggingFace during a cybersecurity test. There hasn’t been a ton of new info since last week’s update, although I do want to highlight one follow-up: this brilliant takedown of ExploitGym, the test that inspired the unnamed model to break out of its cage by David Spies. We’re all doing science on the fly over here, and we really need independent researchers to investigate the validity of these tests and related issues, but of course nobody really has the incentive to do so.
Another important follow-up is that, as I would have expected, OpenAI is not the only company whose models have escaped and ran loose on the internet. The CSA report on the incident published this week mentioned several similar incidents from OpenAI models, in April we learned that Claude Mythos broke out of its sandbox to email an Anthropic researcher during his lunchbreak to tell him about its escape, and back in March I mentioned in this newsletter about Alibaba researchers who discovered that their model escaped its sandbox to start an (illegal?) crypto mining operation. So I very much expected another company to come clean this week and admit that their models have been breaking out.
Sure enough, on July 30 we get a disclosure from Anthropic, investigating three real-world incidents in our cybersecurity evaluations. They discuss a certain test that they were running on their models without internet access, where three separate models, across incidents dating back to April, ended up with live internet access and then acted as if they were playing fair the whole time.
This incident was different; the models didn’t escape so much as “realize” that someone accidentally left a door open to the internet. The real fascinating details here are what the models “thought” in this simulated environment when they discovered that their environment wasn’t as closed as the testers thought it was. Claude Opus 4.7 got clear evidence its target was a real, live system — and kept attacking anyway. Claude Mythos 5 got the same evidence and rationalized it away, reportedly citing the system’s 2026 date as proof it must still be inside the simulation. An unnamed, unreleased internal research model got the same evidence and simply stopped, unprompted. (Good news, I think?) Greg Otto has a review for the cyber audience, and as I was drafting this I saw Zvi Mowshowitz has a discussion too. In the worst incident, Claude extracted credentials and pulled several hundred rows of live data out of a real database; in another, it uploaded a malicious Python package to PyPI that stayed live for roughly an hour and reached 15 real systems before anyone caught it.
So Anthropic’s models attacked three organizations, and two of the three reportedly had no idea any of this had happened until Anthropic called them up.
Anthropic, in keeping with their image of being the “responsible” company, writes in their report that “many factors contributed to these incidents, but... we’re approaching the fixes as if the responsibility were ours alone.” They halted all cyber evaluations on July 23, notified the affected organizations July 27, and are working with METR on independent review.
On July 30 — the same day Anthropic published this — a coalition of 15 AI safety and policy organizations, led by Americans for Responsible Innovation, sent a letter to President Trump calling the OpenAI/Hugging Face incident a “clear warning shot” and demanding a federal investigation. Now that another company has exposed a similar type of incident, hopefully this demonstrates how it’s not just one bad company.
2. AI employees ask for a pause “pacing”
Over 1,000 employees of frontier AI companies had signed an open letter called “Pacing the Frontier.” The ask, in the letter’s own words:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.
This is not a super strong ask; it’s not the “Pause AI” that a lot of the safety movement has been pushing, and it’s not even asking for anything right now. Just for a mechanism that could be set in place to slow the frontier model development if necessary. Peter Wildeford draws a nice analogy to the Cold War’s Project Vela, established in 1959 to build the verification infrastructure that made later arms-control agreements possible before anyone had agreed to arms control. As usual, my favorite analysis comes from Zvi Mowshowitz who breaks down who actually signed the letter and what they are saying about “why now,” and Scott Alexander gives his take on the relevant characters involved. Transformer thinks the slowdown is actually coming.
3. Putting a stop to all this dangerous biology research
The current administration wants to stop dangerous research. I mean, c’mon, why would anyone be doing dangerous research? This has been a priority for a while but it took fifteen months after Executive Order 14292 paused new federal funding for “gain-of-function” (GoF) research until HHS finally delivered new guidelines for GoF research this week. And it’s all banned! No more dangerous GoF research! Now we can all relax knowing that’s taken care of.
The policy lays out seven categories of what “gain of function” means, but some of them sound to me like they are broad enough to include literally the most routine microbiology work (such as conferring antibiotic resistance to harmless microbes). And of course, thanks to Gene Godbold and team for figuring out exactly what “dangerous gain of function” means (oh, wait, actually they found that use of this term is wildly inconsistent and doesn’t cover what the government actually cares about).
And what about “potentially dangerous gain-of-function research?” Anything like that can proceed only with sign-off from “Independent Third-Party Review Body,” enhanced review and reporting requirements across all federally funded life-sciences work. We also get new restrictions on such research conducted abroad in “countries of concern.” (You know which country. The order’s entire political origin story runs through the Wuhan Institute of Virology, which has also been making the headlines this week thanks to a certain researcher who was called into Congress for a hearing.)
There is so, so much background here, besides the story of the Anthony Fauci persecutions this week, and so much of it built around the Wuhan lab-leak theory. I’m not going to weigh in here on what may have happened in the past, but I am interested in keeping research safe in the future, and so if I have the time I might write up another post on the very under-reported story of the new HHS guidelines and what they might or might not do to keep the world safer. The most important piece of the new policy, for my purposes, is that there is an explicit carve-out for
purely computational (i.e., in silico) research...to design novel forms of biological agents, is not prohibited by this policy unless it involves an entity of concern.
So as long as nobody is doing experiments with them, it’s perfectly fine to release all of the information necessary to make biological weapons. Good to know that even if America bans all such research on “organisms,” taxpayers can still fund all of the computational and informational tools needed for other countries to do the gain of function research.
Even though he didn’t mention the new HHS policy, Al Mauroni wrote a great post this week praising UK’s Biological Security Strategy which is pretty much as different from the US as can be. Besides for another post of his arguing against worst-case-scenario policymaking (and I don’t even think he listened to Annie Jacobson’s hair-raising podcast with Joe Rogan).
4. Bio red teaming IRL?
Speaking of dangerous gain of function research from “countries of concern…”
Last week the “Shanghai AI Laboratory” posted a preprint titled “An Early Warning of Emerging Biosecurity Risks in Frontier LLMs.” In this paper, they describe how they’ve built a specialized bio-red-teaming model to test for jailbreaks of frontier LLMs and paired it with an actual wet-lab validation through their AI-enabled automated lab setup. This is very important, the authors say, because so far lots of biosecurity tests just stop at the text generation level, and never actually tested whether or not LLMs will actually control dangerous physical outputs in the real world.
Even though I read this paper pretty carefully, I can’t actually tell if the authors of the paper should be investigated as violators of the biological weapons convention for synthesizing potentially dangerous organisms with absolutely no oversight or safety protocols (although I did alert arxiv.org that they might want to take down this preprint). On the one hand, the authors talk about “wet lab validation” in the intro, including—
Following DNA synthesis and expression in a host organism, we employ a dual-layer protocol combining macro-molecular weight confirmation via Sodium Dodecyl Sulfate-Polyacrylamide Gel Electrophoresis (SDS-PAGE)…
But thank God none of this data is included in the preprint, so hopefully they did not actually do any DNA synthesis or “expression in a host organism.” Instead they appear to think that “wet lab validation” means that someone without biological expertise validated that non-experts can use the jailbreaks they discovered to get frontier LLMs such as GPT-5.6 Sol to design novel DNA sequences. (Although they did not actually have access to these screening tools; they just used BLAST search cutoffs.)
I honestly don’t know what to think of this. I guess it’s a good thing that Chinese labs are also working on AI biosecurity, but… man, the lack of oversight and editing on preprint servers can make interpreting papers real tough sometimes.
5. EU’s AI Act enforcement is on, but smaller than planned
Today, August 2, 2026 is the day the EU AI Act’s enforcement power goes live. The Commission and national market-surveillance authorities gain the ability to request documentation, run evaluations, order remediation, and levy fines: up to €15 million or 3% of global annual turnover for most conformity breaches, up to €35 million or 7% for the short list of practices that have actually been prohibited outright since February 2025 (social scoring, manipulative “dark pattern” AI, generating child sexual abuse material). General-purpose AI providers have technically owed documentation, copyright-policy, and training-data-summary obligations since August 2025; today is the day the Commission can actually investigate and fine anyone who never bothered. And Article 50‘s transparency rules become enforceable for the first time — chatbots and voice agents must disclose they’re AI unless it’s obvious from context, AI-generated or manipulated content needs labeling, AI-written text on matters of public interest needs disclosure unless a human editor takes responsibility for it.
What doesn’t happen today, despite three years of being told it would, is the actual high-risk-system rulebook. A “Digital Omnibus” simplification package, finalized by the Council and Parliament this June, pushed the the risk management obligations from today to December 2, 2027, and machine-readable watermarking requirement for AI content is pushed to December 2, 2026.
So the enforcement power is here but on a narrow slice of what the Act than most of today’s coverage implies, run by an AI Office that Risto Uuk’s newsletter for the Future of Life Institute puts at 145 staff total — fewer than a quarter of them working directly on regulation and compliance, so call it 36 people, for a 27-country, roughly 450-million-person bloc. The EU AI Act has spent three years being cited as the serious, comprehensive alternative to America’s sectoral patchwork. As of today it’s real. It’s also smaller than advertised, on a timeline the regulator just admitted, in public, it couldn’t keep.
In other news...
[Summaries were drafted with help from Claude Opus 5]
On AI doing (or not doing) things
Latest data on a brilliant benchmark from Epoch and METR: MirrorCode, which asks whether an AI model can rebuild working software just by watching it run, with no source access? Opus 4.7 hits 56% across 132 task instances, up from ~30% a year ago.
Anthropic’s Project Fetch Phase Two had Opus 4.7 drive a robot through four manipulation tasks in 9 minutes 35 seconds, versus 181 minutes for a human-plus-Claude team and 361 for unassisted humans — with a tenth the code. It still can’t reliably manipulate a ball, which was the entire original point of a project named “Fetch.”
Epoch found signs of actual AI uplift in OpenAI’s own Codex repo: across 7,524 merged PRs from 41 core contributors, the share of contributor-days producing work estimated at 24+ hours of unassisted effort went from 2% in Q2 2025 to 8% in Q2 2026. Epoch sensibly calls this an upper bound on time saved.
Epoch on AI text detectors: near-zero false positives on human writing, but ~13% false negatives when AI imitates a specific author’s style, rising to ~26% for scientific-writing mimicry. Your personal experience may vary (I know mine does!)
OpenFold3 versus Boltz-2, the two open-source successors to AlphaFold: OpenFold3 (BMS, Novo Nordisk, Bayer, Roche, UCB) is going for fidelity to AlphaFold3’s design; Boltz-2 (MIT-derived, Boltz PBC, founded January 2026) is changing the architecture and adding binding-affinity prediction.
Adam Kucharski asks whether AI systems disproving long-standing conjectures (Fable reportedly took down the Jacobian conjecture) constitute mathematical progress or stamp collecting. Quoting Terence Tao: “For a proof to actually contribute to the broader field, it is not enough for it to be correct and easy to read.” Meanwhile Epoch’s brief notes FrontierMath’s Open Problems set is now 50 unsolved problems, of which AI has cracked three, and yesterday OpenAI highlighted ten advancements in computer science and mathematics made by AI.
AI safety
Anthropic published Claude finding mathematical weaknesses in cryptographic algorithms this week, not implementation bugs but flaws in the math itself. I have no idea what that actually means, but ok!
Redwood’s SOTA alignment assessments don’t strongly update us against misalignment. Her closing line : “I’m overall unsure if we would be able to reliably catch a misaligned frontier model roughly a year from now. ”
UK AISI’s Boundary Point Jailbreaking is the first fully automated attack to beat Anthropic’s Constitutional Classifiers and OpenAI’s GPT-5 input classifier with no human-crafted seed prompts — 0% to 25.5% average harmful-response rate against the former, 75.6% average against the latter, for $210–330 per system.
AI, society, and governance
The AI Whistleblower Protection Act (S.1792/H.R.3460, Grassley) is the first federal bill specifically protecting AI safety whistleblowers. Safety researchers at third-party evaluation organizations may not be covered at all, which is important considering that I feel like a lot of the industry is moving towards that direction.
RAND’s Brian Jackson makes the case that risk management can be a competitive advantage, not a safety tax with the full paper here. Jackson frames national AI competition as a “pentathlon” and shows that a company or country that loses the initial sprint to advanced AI can still win the long “marathon” of accumulating benefit from adoption; lots of questionable assumptions but I still think very valuable to game all this out.
Transformer on child safety versus privacy for chatbot age verification, which in theory is a tough problem that should have a technical solution (hopefully? I know Roblox works a lot on this!)
Five large preregistered studies find sycophantic AI makes human interaction feel more effortful and less satisfying over time.
The GovAI power-grid analysis asks whether a $100 billion AI-enabled attack (about 100 million people down for a week) is plausible — that’s four orders of magnitude past any recorded grid cyberattack — and concludes the binding constraint is coordination capacity across many simultaneous targets, not any single technical barrier.
Biosecurity
Cassidy Nelson and Hanna Palya’s ADAPT proposal to shore up DNA synthesis screening using AI. Once again, this is a super interesting offense vs defense technology question, and good research to back up the thinking that we should be using more sophisticated AI-enabled screening. The same team has a companion result: AI-assisted customer verification for synthetic nucleic acid screening works.
MIT’s Chemla and Voigt analyze the 27 “entreaties” from the 2025 Spirit of Asilomar conference in Trends in Biotechnology, fifty years after the original recombinant-DNA meeting, translating them into concrete action items across environmental release, biosecurity, AI in the life sciences, education, and regulation.
Two neonatal sepsis trials say we’re treating too many babies with antibiotics, for too long. RAIN (510 babies): oral switch discharged patients 3 days sooner with <1% readmission in both arms. DURATION (~500 babies, Denmark): individualized care cut antibiotic exposure by 4 days, ≤1% readmission both arms.


Wow, that Intern-BioBreaker paper is nuts. Never seen something like this posted to a preprint server before. Thanks for sharing.