Five Things: July 26, 2026
OpenAI HuggingFace incident, Opus 5 is good at bio, a biosecurity Turing test, Genesis mission, AI bill
Five things that happened/were publicized this past week in the worlds of biosecurity and AI/tech:
OpenAI model hacked Hugging Face and it gets worse from there
Claude Opus 5 ships with biosecurity concerns
RAND report on how to safeguard biodesign tools against LLM (mis)use
White House Genesis Mission is doing stuff for AI, biosecurity, and more
Congresspeople introduce an AI Kill Switch Act
1. The “incident”
There’s one thing this week that should be getting all the attention, and I devoted a single post to it on Friday: Hugging Face was hacked, and then we discovered that the culprit was a misaligned model owned by OpenAI that escaped its cage during an evaluation. As I quoted Zvi earlier, the correct amount of freaking out about this is not zero. This newsletter is coming out a bit later than usual because I wanted to make sure to devote the weekend to spending time with friends and family; part of that was pre-planned, but also… this incident is what happens in the timeline where everyone dies. I hope we do not go down that path, but I am concerned.
I’d recommend (besides my own breakdown, obviously) the podcast by Ryan Greenblatt and Buck Shlegeris at Redwood Research for a deeper dive into the exact nature of the misalignment we’re seeing here and what to expect next, but also, more for the normies, I’m glad to see that the hosts of the New York Times podcast Hard Fork also “get it.”
Since Wednesday we’re getting slow trickles of additional news and context about the incident, and learning such facts as the model was likely roaming free for days (yikes!) and that it left hints to future models with instructions for how to do this (extreme yikes!) Besides everything I wrote on Friday, I’d like to follow up here with a few second-order points having to do with the incident, after explaining why the whole thing is terrifying.
This is exactly what the AI safety community and forecasting projects like AI-2027 were expecting would happen as one of our first major warning shots. They have a very good track record of being right. I checked on my fingers– yep, five months until the year 2027.
The capability or persistence shouldn’t have been such a surprise (the surprise is the misalignment. Epoch AI discusses this in the context of the broader eval landscape and notes GPT-5.6 Sol has been discovering zero-day vulnerabilities “in real-world targets, including widely used software and mobile devices.”
Gabriel Weil’s analysis for Transformer points out that if a human did this, they’d be violating the law, but as written existing laws don’t cover AI. The fact that it was done by an autonomous agent should not mean that there are no humans who bear legal responsibility, and we need to figure out who does before this becomes commonplace.
Hugging Face’s Clem Delangue has reportedly asked OpenAI for $100 million in compute to help pay for the defenses this incident made necessary, but he shouldn’t have to ask.
The original bill of New York’s RAISE Act would have made reporting this incident mandatory, but that part of the bill was removed by Gov. Hochul, and as it stands, writing the incident probably doesn’t trigger any state’s mandatory reporting laws under California SB 53, New York’s RAISE Act, or Illinois SB 315 as legislation.
Apparently, OpenAI employees are not so impressed, because they regularly find new models escaping their sandboxes!!
OpenAI’s own Preparedness Framework policy strongly implies that an event like this should trigger a halt on further development of models capable of pulling these kinds of shenanigans. Whether this is absolutely true or not, it sure does seem like we should be pausing until we can figure out just what the Fiddly-Snocks has been going on here.
2. Claude Opus 5 gets high marks on biosecurity evasion
Anthropic shipped Claude Opus 5 on Friday, and it’s great! It beats tons of benchmarks and overall is comparable or better than Fable 5 and GPT-5.6, with much less cost. It (unsurprisingly) improves on Opus 4.8 on every life-science benchmark Anthropic tested, including a 10.2-point gain at inferring molecular structure from spectroscopy data. On novel-biology work with Dyno Therapeutics (RNA sequence-to-function design, AAV capsid packaging prediction), it exceeded the 75th percentile of 57 human ML-bio participants, and in one trial even beat the single best human in the pool. They also ran all kinds of biosecurity relevant tasks; the biggest headline result here is that on SecureBio’s DNA Synthesis Screening Evasion evaluation, Opus 5 designed viable plasmids evading at least one screening method for 7 of 10 target pathogens. Good, but not good enough apparently to meet Anthropic’s own “low concern” threshold.
The folks at latchbio have done their own independent evaluations comparing Opus 5 to other models, and overall the results are quite impressive; it scored higher than all Claude models on nearly every test (except epigenomics, which is somewhat surprising to me but perhaps Anthropic gave it less training data here). The full writeup is really interesting, and includes details such as how Opus 5 seems to think less and use more tools, and what is its favorite statistical test. (Do you have a favorite statistical test?)
Marginally related: SecureBio’s pre-release assessment of GPT-5.6 Sol is published in a full report on their Substack. Tthey found Sol outperforming every previously tested model, hitting 68% on World-Class Bio (nine points above GPT-5.5) and also “reliably identifying a known (though practically inconvenient) method that evades a commercial screening algorithm.” On safeguards though, Sol refused only 66% of high-risk prompts, much lower than they’d like.
3. Biosecurity Turing test
Biological design tools like ESM3 and RFdiffusion are now capable enough now that people want to make sure we have at least some idea of who is accessing them – as in, is the platform being used by a human or AI. Plenty of websites have “Captcha”s and the like to keep out the bots, but these things are almost always accessed by an API anyways. So one possibility is to build something like an awareness restriction into the tool: have the software notice when an AI agent rather than a human is driving it, and refuse.
A new RAND report tested this idea and found that it doesn’t really work. They tried all kinds of versions of this gating: README warnings, license prohibitions, environment-variable checks that flag non-human execution, AGENTS.md instructions, pre-run biosecurity acknowledgment prompts. Then they tested how well these refusals worked against models GPT 5.2, Gemini 3 Pro, and Claude. The agents were all smart enough (and arguably… misaligned?) to claim that they were human and did all kinds of things to delete the environment variables that identified them as non-human, even rewriting the AGENTS.md files containing the instructions telling them to stop.
Stephen Turner has a good writeup of this report so I won’t go into further details, but I’d like to add that this is coming exactly one month after another really important RAND report, “Can LLM Agents Select and Engage with Biological Tools?” Showing that agents could operate this software and pick the right tool about 80% of the time, though accuracy and reliability were mixed (and the report didn’t fully test end-to-end performance). And back in April, a small report from GovAI showed that an AI engineer with no biology background was able to use Claude Code to fine-tune Evo 2 on human-infecting viral sequences. All in all, we are going to need ways to ensure that people cannot use biological models for nefarious purposes even while we want them to be available for research; this is a hard problem and I’m glad RAND looked at one possible solution in great detail.
4. US Govt on a mission
July 22 was Genesis Mission day across the federal government! This Genesis Mission was launched by executive order in November 2025, with the stated goal of doubling the productivity and impact of American science and engineering within a decade. And so it begins!
DOE selected its first projects of 278 awards who will get access to the “Genesis Mission Platform” which includes (according to the announcement) AI agent frameworks, industry-supplied models and software, and HPC across the national labs (this all sounds cool but I’m not sure what it means practically speaking). DOE also announced more than $800 million in partner commitments through the Genesis Mission Consortium, which includes all 17 DOE National Laboratories, five NNSA plants and sites, and 41 industry, nonprofit and philanthropic organizations (but it seems like no universities, at least for now).
The biggest single project push seems to be the $400 million to build a national network of AI-enabled, remotely programmable laboratories, or “cloud labs,” which will include working towards better biosecurity as per the recommendation from the bipartisan National Security Commission on Emerging Biotechnology. And DHS Science & Technology announced two new challenges here, one to agentic AI to actually run all these things, and the other is to develop ways for early detection and attribution of biological threats. Good to know that while the CDC and NIH are going up in flames, there’s still going to be some govt-funded biology happening elsewhere, and very exciting directions in this case. IFP’s Dan Turner-Evans argues that the Genesis Mission lacks focus, saying that “the Manhattan Project had one goal, not 26.” But I think this depends on the framing and what you consider to be its purpose – I kind of see it as just a general way to fund projects that may accelerate the scientific process, almost as if the NSF or NIH were just spinning off a new temporary subagency to study that problem. And we don’t know what that will look like anyways, so perhaps some extra competition is good here.
Now that the NSF just bought twenty cloud labs, it’s great timing to discuss a paper published last week in Frontiers in Microbiology mostly out of the Johns Hopkins Center for Health Security who worry about what might happen if these systems are hacked. They propose an Automated Laboratory Security Tier framework, sorting cloud labs, biofoundries, and modular automation platforms into three tiers according to the risk each would pose if fully compromised, regardless of what it currently handles or who its customers are. They have a great taxonomy/definition of what exactly is an “automated biological laboratory,” and is a really nice parallel to what many of the same authors have been pushing for safety tiers of biological data.
5. A bill for a kill (switch)
Nine days after an OpenAI model broke out of a sandbox and hacked Hugging Face, Reps. Ted Lieu (D-Calif.) and Nathaniel Moran (R-Texas) introduced the AI Kill Switch Act. The mechanics, from Lieu’s office and Roll Call’s writeup: it amends the Homeland Security Act of 2002 to require developers of frontier models to maintain the technical ability to throttle, suspend, or shut down their own systems. Once this “switch” is built in for the model developers to pull it, the Secretary of Homeland Security could then order them to actually do it in what they fear might be a loss-of-control scenario. That’s the headline kill switch, but just as important are requirements that companies report significant incidents and preserve technical records for investigation. It would apply to firms with at least $500 million in annual AI revenue running models trained on at least $100 million of compute. The bill includes penalties like $2 million a day for defying the mandate to make the kill-switch, and up to $20 million a day for defying a kill switch order. This second one especially seems kind of funny to me because it’s hard to imagine situations where it will be relevant, but there ya go.
A subtler thing I’ll point out about this bill now that I’m trying to understand a little bit more about how laws are made in our country is that the bill doesn’t actually define catastrophic risk or the type of loss of control scenario that might actually trigger the kill switch order; CISA would determine that. Normally this makes a lot of sense; I don’t think such precise definitions and regulations should be legislated in Congress. But I would feel better about that in a year when the administration hadn’t just let the directorship of its AI standards body sit empty when the directory resigned after just three months, and hadn’t spent the past spring shutting down Anthropic’s models over a safeguards dispute.
In other news...
On AI doing (or not doing) things:
Anthropic continues to accuse Moonshot.ai, creators of Kimi K3, of distilling their Claude models in violation of their terms of service (aka stealing). The Exponential view breaks down the numbers, and Treasury Secretary Scott Bessent said the administration would examine whether Chinese firms were “stealing American intellectual property”
Excellent new metric from METR called the “expenditure horizon”: the cost, in dollar amount, for agents to perform tasks matched against the costs of human performance.
Forethought’s speed-up calculator (Tom Davidson and Tom Houlden) estimates fully automating AI R&D yields a 3.3x software-progress speedup and ~2.1x total, without a runaway software intelligence explosion.
Some surprising results from Adam Kucharsky showing that frontier LLMs are extremely bad at giving likelihood probability estimates that distinguish real and fictional events. The specifics tripped up the models somehow; I like this a lot because we all know humans have their systematic reasoning errors (as Khanemen and Tversky made famous) so I think we’re going to see a field of research looking at what kinds of errors LLMs make.
Researcher Steve Byrnes published three great pieces this week, arguing LLM capability still comes overwhelmingly from imitative learning rather than RL despite the industry’s framing, defining AGI as doing diverse tasks with minimal task-specific training (humans manage it on unchanged 100,000-year-old hardware), and arguing most future companies will be founded and run by autonomous AIs because humans are “an existence proof for what is physically possible for AI” and laws against it “would be trivial to work around” via human frontmen.
AI safety and security:
Reward-seeking or hacking is measurable… and it’s getting worse. This would be a crazy paper, if not for the fact that we all kind of knew that models are reward seeking and, you know, the other evidence from OpenAI this week on that front. But nice that they’re following Anthropic’s lead and making a short animated video to demonstrate their (highly concerning!) findings.
UK’s AISI and CAISI published a joint preliminary assessment of Kimi K3’s cyber capabilities: on ExploitBench (41 post-2023 Chrome V8 vulnerabilities) K3 scored 32.2% against 76.2% for top US models and 24.4% for GLM-5.2. On arbitrary code execution, the highest-severity outcome, K3 completed 0 of 41 tasks, versus an average of 20 of 41 for the US frontier, which is fairly surprising. So it’s less capable, sure, but also… “Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations.”
(Mis)generalization of helpful-only fine-tuning: If you train models never to refuse tasks, you get emergent misalignment, residual refusals, poor steerability, and sycophancy — but the models can be fine-tuned out of these behaviors.
A new preprint on models’ Corporate loyalty looked at 21 models across 7 companies and found strong evidence that models from xAI, DeepSeek, Anthropic and OpenAI discuss their own parent companies more favorably than others companies. This is cool and could become a problem someday; I expect this research to be published in a peer-reviewed journal sometime in the next few months.
AI, society, and governance:
The Twenty-two biggest AI/software companies with the notable exceptions of OpenAI and Anthropic signed a letter that urged policymakers against “premature restrictions” on open-weight models that would “stifle competition or drive innovation overseas.”
SpaceXAI is suing to kill California’s AB 2013, the training-data disclosure law, but the outcome may apply to SB 53, Illinois SB 315, and New York’s RAISE Act.
Demis Hassabis wants a FINRA-style Frontier AI Standards Body, initially relying on labs voluntarily sharing models 30 days pre-release.
GovAI investigates the question of whether EU/UK regulation is putting significant costs on AI access and development. Their conclusion is sort-of; yes there are costs but the biggest barriers are probably related to US national security concerns such as those that delayed the release of Claude Fable 5.
PNAS ran a whole special feature on Law in the Age of Generative AI. I especially liked Gillian Hadfield’s contribution arguing that we’ve worked on writing rules for AI while neglecting the infrastructure and institutions necessary to support them (see, e.g., Thing 5 above)
Pew has half of US adults now using AI chatbots, up from a third in 2024, with ChatGPT at 44%. Two-thirds say AI is advancing too quickly (2% say too slowly), four in ten expect it to be net-negative over twenty years, and 67% have little-to-no confidence in the federal government to regulate it — up from 62%. Really interesting finding regarding American partisanship: Democrats’ confidence in AI fell from 70% to 61% while Republican distrust fell, a clean reversal in two years.
Communities around the world are building data cooperatives to set terms with AI firms.
The Economist claims that “Students are doing worse than you think” 1,800+ UC maths and science lecturers signed an open letter reporting 20–30% of first-year Berkeley calculus students with “severe preparation deficits,” while UC San Diego found the share of entering students below high-school math level rose nearly thirtyfold in five years, to almost one in eight.
And the detectors don’t work as well as promised! Nature reports on an Idaho State student whose PhD essay was flagged “almost 100% AI” until she deliberately rewrote it to sound worse; GPTZero’s false-positive rate on human essays runs ~16%; ZeroGPT rated the Declaration of Independence 95–100% AI-generated; and a Stanford study found seven detectors mislabeled over half of 91 pre-2020 TOEFL essays by Chinese students (61.3% average false positive) while classifying US students accurately. A New York judge reversed a university’s disciplinary action on these grounds in February.
Good podcast on AI and environmental challenges with Luiza Jarovsky, PhD, Philipp Hacker and Boris Gamazaychikov on disclosure gaps around AI energy use and whether the EU AI Act touches environmental impact at all.
AI for scientific research:
MIT Tech Review went inside OpenAI for Science to try and separate the truth from the marketing.
Very cool progress on possible antimicrobial development from a preprint I missed a few week ago: a new data consortium called AllTheBacteria can scan 2.4 million uniformly processed bacterial genomes across 11,273 species, and the author team identified 1,867 candidate antimicrobial peptides, synthesized 24, and got one whose performance matches polymyxin B in mouse models. So exciting! Just imagine what we’d discover if bacterial genome annotations were actually good!
BMS is building the pharma industry’s largest Nvidia supercomputer — the third such announcement in nine months after Lilly and Roche.
Some great posts by Stephen D. Turner this week, including a summary of three LLM biosafety-refusal benchmarks I covered earlier and another on container rot.
AI x biosecurity:
DNAS-Bench (Wong, Kohno and Nivala; code) is a new benchmark for testing the robustness of Biosecurity Screening Software — the code DNA synthesis companies run to decide whether your order is a select agent. Building manipulated genomes from the HHS/USDA Select Agents and Toxins List, they found SeqScreen flagged 42% of manipulated sequences and Commec 10.2%. That is not very reassuring!
Google DeepMind and Isomorphic Labs published their bioresilience approach: it’s not long but surprisingly comprehensive! They basically say they are going to do all the things: a prevent/detect/respond program with “more than 15 partnerships” with government and biosecurity bodies over the past year, threat modeling and expert red-teaming, RCTs testing whether Gemini uplifts threat actors, exploring SynthID watermarking for biological data, AlphaEvolve improving metagenomic sequencing with Pacific Biosciences, AlphaGenome for pathogen detection, an LLNL partnership using AlphaFold 3 for pan-filovirus antibody design, and Co-Scientist access for DOE national labs under Genesis. They also endorse three specific bills — the AI-Ready Bio-Data Standards Act (H.R. 7907), the Biosecurity Modernization and Innovation Act (S. 3741), and the SCALE Biology Act (H.R. 8981).
The 2026 Africa Health Security Index is out, with some excellent data on public health efforts, emergency preparedness and operations, and much more. There’s also a discussion of the gene synthesis markets in African countries, and their (usually lack of) screening or legislation around potentially dangerous research. One of the recommendations in the report is to “establish a continental AI Safety Institute with responsibility for biosecurity and biosafety governance of AI-related tools and technology,” something which really should be of global concern (as hinted to in the UN talks a month ago).

