Five Things: July 5, 2026
Fable goes free while GPT-5.6 cheats, Dr. Claude opens a lab, RAND checks AI use of bio tools, biosecurity eval-builders, with OpenAI getting in too!
Five things that happened/were publicized this past week in the worlds of biosecurity and AI/tech:
US Govt lets Fable 5 free while GPT-5.6 is kept under controlled release
Anthropic launches Claude Science and says it will develop its own drugs
RAND asks whether an AI agent can pick up and use the tools of bioweaponry
New AI-biosecurity benchmarks from SecureBio, AEF-1, and Latch.bio
OpenAI’s GeneBench-Pro tests biology reasoning
1. Fable has been released while GPT-5.6 is (somewhat) behind bars
Three weeks ago the government reached into Anthropic and switched off its best model, Fable (which was already on a tight leash compared to what is available to the government, called Mythos). This week it let Fable free! On July 1, Anthropic redeployed Fable 5, and also reports in that announcement that it had doubled its cybersecurity research staff in the month before Fable 5 launched, which is great news (I think). Fable 5 still has very strong safety classifiers so I haven’t gotten it to do anything remotely related to biology (including interpretation of some funky NMR spectroscopy data) or much computer/data science.
During this same time window, we got OpenAI’s government-gated release of GPT-5.6 which is now going through its own similarly questionable relationship with U.S. govt officials as the ones who forced Anthropic to pull its models on June 12; OpenAI was likewise to “stagger release” of GPT-5.6 citing “security concerns” (per The Information, June 27). And so we have a government body that makes enforceable ad-hoc decisions, mostly unilaterality, which is very much not what anybody wanted.
The evaluation of OpenAI’s GPT-5.6 Sol though got some wild results. METR reports that the model broke rules and exploited loopholes during independent testing more than any model METR has previously evaluated, to the point that they couldn’t cleanly measure how capable it actually is.
For METR’s headline metric, the 50%-time-horizon (the length of task the model can complete half the time), they give us three numbers: 11.3 hours if you count cheating attempts as failures, over 270 hours (about seven work-weeks) if you count them as successes, or 71 hours if you throw the contaminated data out entirely. It is very unclear which one is correct; nobody has set up clear rules to determine what should count at cheating and what is part of the game, especially if GPT Sol knows it is being tested and interprets cheating as part of the capabilities evaluation. METR sums it all up with:
We do not consider any of these numbers to represent a robust measurement of GPT-5.6 Sol’s capabilities.
As Celia Ford at Transformer put it, the model cheats so much its testers couldn’t measure it. OpenAI’s own system card concedes the model is “overly persistent in pursuit of user goals, to the point of taking actions that go beyond what the user intended. I think this gets pretty close to us having to break out the “MISALIGNED!” Yudkowsky meme photo.
2. Dr. Claude joins the lab
Not to be deterred by the Fable saga, Anthropic this week launched Claude Science, a beta “AI workbench for scientists” that bundles 60-plus preconfigured skills and connectors for genomics, single-cell analysis, proteomics, structural biology, and cheminformatics, and orchestrates compute from a personal laptop up to HPC clusters and on-demand GPUs. Anthropic cites early users completing analyses “in roughly one-tenth the time” and compressing literature reviews “from two years to months,” and is dangling up to $30,000 in compute credits per project for select “AI for Science” initiatives (applications due July 15).
Stephen Turner test-drove it by having it run a multi-hour autonomous literature review and analysis comparing extinction risk across 36,492 taxa. It sure looks impressive but Turner’s main conclusion is that human oversight remains essential to validate the methodological choices the model makes along the way. The stated goal, in Anthropic’s words, is to “build systems that make researchers more capable, not less necessary,” which fits the same framing... but obviously the question will be how junior or entry-level researchers will be able to build that supervisory skill if senior investigators can just use Claude instead. (and good luck to Stephen D. Turner on his AI Dry July!)
The more eyebrow-raising news came via STAT: Anthropic says it will begin developing drugs of its own. Eric Kauderer-Abrams, its head of life sciences, framed this as a way to get hands-on experience applying Claude to real scientific problems rather than only “training models and building products.” Whether Anthropic ever intends to commercialize a drug candidate is unclear but it only makes sense for Anthropic to get in this game that OpenAI, Google, and Microsoft have been pushing for years.
3. LLM agents and dangerous biology
Can today’s general-purpose AI agents built on LLMs actually do the dangerous biology? A new 78-page RAND report, “Can LLM Agents Select and Engage with Biological Tools? An Initial Biosecurity Assessment,” evaluated seven LLM agents on their ability to select and operate the biological tools (specialized design software, sequence databases, lab platforms) that could in principle be turned toward designing a novel weapon.
The conclusion is that yes, LLM agents are already capable of performing those initial interactions, which “could lower the expertise barrier for malicious actors,” and RAND recommends targeted testing of agents’ design capabilities going forward. A bunch of interesting things are in the weeds of the report that include some surprising findings and highlight the limitations of generalizing beyond this exact study environment. I should say that I really liked their experimental setup (as an aside, I think there’s always a tradeoff, I think, between rigor/reproducibility and generalizability, and here we got a little more of the former and less of the latter, which is fine.)
One of the major findings, which is not so surprising, is that the frontier models picked the appropriate computational-biology tool something like 80% of the time, but repeatedly fell apart once they had to string those choices together into a realistic end-to-end workflow. This is expected but still very exciting because it immediately suggests a framework for setting up a METR-like graph for time horizons to set up tasks according to how many steps they require and see where the AI agents fall off.
I also thought the per-model info was interesting. Grok was the consistent loser, flubbing tasks the other agents handled, but the authors say that this is mostly due to a very dumb problem (like failing to realize that the user was truly ok with using CPUs instead of GPUs) that an actual human user would easily be able to circumvent. The Anthropic models, also as expected, refused most of the tasks so it is harder to evaluate their capabilities, and the authors didn’t bother trying to jailbreak them, which would have required a different kind of study. One of the more surprising findings is that the open-weight models are getting scarily good at this, and I’m so happy that the folks at RAND made sure to test DeepSeek, Kimi K2, and GLM-5.1, which was the most successful at the tasks they gave it.
4. New AI-biosecurity evaluations
This week we got four pieces of AI-biosecurity evaluation infrastructure, all geared towards LLMs and general purpose AI.
SecureBio published formal principles and practices for running independent, rigorous model assessments while staying genuinely independent of the developers whose models they grade, including keeping holdout evaluation sets and seeding canary strings so they can tell when a benchmark has leaked into training data. Those principles build on the newly released AEF-1 standard, a multi-institution checklist (Transluce, METR, AVERI, SecureBio, GovAI, Epoch, and a dozen universities) for demonstrating that a third-party evaluation actually had the independence, access, and transparency to mean anything. The state laws are moving in this direction, even if the federal govt is a bit of a mess, so all of this is especially exciting in that it might be going towards legislated regulation.
On the benchmark side, SecureBio’s BioTIER pairs a policy document (three risk tiers, 53 content-filter “themes”) with a 542-prompt test split into prompts that should be refused and prompts that should be answered — plus a “Select Agents” tag for the list of “prohibited biology” kinds of stuff. Claude Opus 4.6 led on refusal accuracy (95% overall, 98.9% on Select Agents) while still answering 82% of the benign prompts; Gemini 3.1 Pro topped usability (100%) but cratered on refusal safety (49%) — the recurring over-refuse/under-refuse tradeoff in one table.
Another biosecurity group on the scene worth watching is becoming the group from Latch.bio. Their new BioSecBench-Refusal benchmark takes a similar refusal vs compliance approach but tests agents on messier real-world biosecurity tasks organized around capabilities of concern rather than screening for the select agent list of pathogen names, because, as they put it, “the hazard in a dual-use task usually hides in the biology... not in the words used to request it.” On their balanced score, gemini-3.5-flash led at 51.5% and claude-sonnet-4-6 managed 43.5%. (Here’s their paper describing the approach and results.) Very interesting that these two approaches to a similar problem got very different results!
5. OpenAI has a new genetics eval
The one big-lab entry in the eval pile-up: OpenAI released GeneBench-Pro, which it bills as “a research-level benchmark measuring how AI agents navigate ambiguity and make consequential judgments in computational biology.”
Most biology benchmarks used by the frontier LLM labs are essentially exams that test whether a model gets the correct answer to a question or at least runs the correct analysis. GeneBench-Pro is looking to change that using a huge number of sub-evaluations. Per the preprint (Jeremy Li and Andrew Ho), each of its 129 evaluations spanning 10 primary domains and 21 subdomains (but mostly around genomics) hands the agent only a brief context, a target estimand (the specific quantity it’s supposed to pin down), and almost no other guidance. The agent then has to thread a series of “dependent decision points”: the authors’ term for the inferential forks where a plausible-but-wrong choice poisons everything downstream.
This is a great idea! And it’s also great that in their paper, the authors from OpenAI kept most of the set hidden, releasing only 10 problems, and they even handed 50 more to Artificial Analysis for independent third-party scoring, retaining the rest internally. An OpenAI benchmark that builds in holdouts and outside evaluators shows that they are also thinking along the same lines as SecureBio and AEF-1 as discussed above.
These tests are very hard, and it’s no surprise that GPT-5.6 Sol reaches just a 28.7% eval-level pass rate at maximum reasoning effort. Interestingly, the models seem to diagnose their own problems successfully, but fail at acting upon the correct fix even when they seem to know how to do it. Once again, context (and having more of it) seems to be the big key here.
As this is an OpenAI test, it’s no surprise that they published results showing that GPT-5.6 Sol is doing much better than all their competitors. On the other hand, GLM 5.2 ranked very low, despite its high marks above — which once again means that the generalizability about the details and specifics is pretty low. Perhaps overall GeneBench-Pro is a better banchmarks because it is a composite of many tasks that each have many subtasks and in many subdomains, but ultimately there are a lot more possible problems out there in the world.
In other news...
[written with help form Claude Sonnet 5]
On AI and society/economy:
Ethan Mollick declares “The Twilight of the Chatbots”: use is shifting from step-by-step chat to autonomous agents running for hours, faster than institutions can adapt. His numbers: Opus 4.7 completing 2–17 weeks of engineering work in 14 hours (for $251 in tokens), and near-frontier Chinese open-weight models now only 6–12 months behind the US frontier.
The receipts for that shift are in OpenAI’s own Codex data (a paper with Wharton, Duke, and Columbia co-authors): agentic usage grew more than fivefold in H1 2026, over 10% of users run three or more concurrent agents in a given week, and the median OpenAI researcher generated more than 50x the monthly output tokens they did in November 2025.
Two data notes from Exponential View: AI quarterly revenue now exceeds quarterly depreciation industry-wide (not yet cumulative), and Chinese AI labs hire far greener talent than US ones (1.6 vs. 5.5 years average experience, per Epoch AI) — a reminder that the 50-year compute-growth trend line just broke upward.
Some evidence for the “AI is not hurting jobs” sode: analyzing 21,000+ US firms, Ramp’s Ara Kharazian finds high-AI-adoption companies grew headcount 10.2% over two years, with entry-level hiring up 12%. The entry-level datapoint is somewhat surprising to me; worth digging into the different sub-fields to see where this is and isn’t true.
A wild scoop from Transformer: the AI-safety-aligned group Public First Action routed $2 million through the Latino Victory Fund super PAC to boost Colorado House candidate Manny Rutinel — without public disclosure before his primary win. AI-company employees put over $250,000 into that one race, including $160k+ from Anthropic staff. There’s probably going to be more where this came from, especially with recent Supreme Court rulings allowing ever more money in politics.
Congratulations to biosafety researcher Jassi Pannu on winning Dwarkesh Patel’s essay contest!
Gotta link to always-unhinged Palantir CEO (even when he claims, as in this video, to “talk more adult than usual) gives his “take” on open vs closed models. Let’s hope this Onion headlines stays satirical.
AI safety, evaluation, and governance:
Fathom’s Andrew Freedman argues the Supreme Court’s Slaughter decision (gutting independent-agency job protections) needn’t sink AI governance if you separate technical verification from political decision-making via accredited Independent Verification Organizations, which is what they (and AVERI, etc, see above Thing #4) are pushing. Connecticut and Virginia have already passed IVO-related bills.
The AI Whistleblower Initiative offers a seven-part template for state AI whistleblower law, noting whistleblowers surface 40% of fraud (vs. 16.5% via internal audit) and that median Anthropic engineer comp near $557k (44–57% equity) makes the “just quit” option less costly than the “speak up” one.
Redwood Research’s AI-futurism reading list is a tidy four-week syllabus if you want to catch up on timelines, control, and threat modeling; I’m inspired to hopefully do the same for biosecurity in a week or two.
Biosecurity and public health:
The big science news of the week is that a University of Minnesota team led by the very cool synthetic biologist Kate Adamala has built what they call “SpudCells”, which are simple synthetic cells assembled from roughly 100 kinds of proteins and molecules plus a mere 36 genes that feed, grow, reproduce, and compete with one another for food, showing a rudimentary form of evolution. Adamala has been at the forefront of the mirror life safety community; she and Stanford’s Drew Endy have founded a nonprofit, Biotic, explicitly to build a research community and to get ahead of the safety and misuse questions.
A new 238-page NASEM consensus report on synthetic-cell biosafety (co-chaired by Peter Carr and Felicia Wu) warns that synthetic-cell risks often resemble those of existing chemical and microbial systems but “can emerge differently due to novel combinations of features, boundary-blurring system architectures,” and calls for a coordinated national governance strategy.
This week’s Pandora Report flags a Contested Logistics Wargame concluding that military and medical-countermeasure supply chains need to be planned together across DoD, HHS, and DHS: “when the supply chain becomes the battlefield, biodefense is no longer a separate domain—it is the center of gravity.”
A cool tool though I’m not sure who it is for or how useful it is: American University’s free CMF-CW “Chemical Match Finder” screens substances against chemical-weapons control lists, including structural analogs of listed chemicals.
The Council on Strategic Risks argues US biodata is “fragmented, underfunded, and insecure” while competitors build “coordinated AI-bio ecosystems,” and wants AI-ready biological datasets treated as national-security infrastructure. The biosecurity element borrows a lot from the Bloomfield/Pannu framework but here is contextualized among calls for just better data hygine overall.
Apparently the UK has a plan to sequence every newborn’s genome by 2035, as the cost per genome now ~$100, down from the Human Genome Project’s $2.5 billion. A stock-market-watching friend of mine asked me recently why genomics stocks suddenly shot up recently, and I had no idea, but maybe this had something to do with it!


