Five Things: March 29, 2026
Anthropic temporary win, scheming, biodesign by LLM, White House advisors, Anthropic security
[NOTE: see Announcement post; the “Five Things” newsletter will be off the next two weeks]
Five things that happened/were publicized this past week in the worlds of AI or biosecurity:
Anthropic was granted a preliminary injunction in its case against the US government
Report on AI scheming as viewed from X
Latest updates on autonomous biological design
Political influence in and out of the White House
Anthropic security breach
1. Pentagon’s “attempted corporate murder” of Anthropic put on ice
As expected, Anthropic has won a preliminary injunction, meaning a judge has temporarily blocked the Department of War’s “supply chain risk” designation that would have blacklisted the company from all federal contracting. The case really, really does not look good for the Department of War on multiple accounts. The best analysis, I think, continues to be that of Zvi Mowshowitz, but he’s writing for insiders who speak his language. I liked the more accessible coverage from Transformer that puts this in the context of a broader competition between Anthropic and OpenAI, noting that Anthropic has been winning on both the commercial front (Claude Code gaining traction lately) and the public perception front.
One of the crazy things about this whole case (but also, of the general antics of this whole administration maybe) is that, in Zvi’s words, “The main thing defending our Republic is that those looking to take it down have strong incentives to keep saying the quiet parts out loud.” Relatedly, it’s broadly a totally crazy thing that our official government people keep making official-sounding pronouncements on X that are only questionably “official.” Like here, government attorneys tried to argue that that DoW Secretary Hegseth’s social media post declaring no DoD contractor could work with Anthropic carried “no legal effect,” and instead was just… meant to use the official channels to cause Anthropic to lose revenue or something?
This is still just a preliminary injunction; the case continues.
2. AI scheming in the real world: a report
The Centre for Long-Term Resilience published “Scheming in the Wild”, an analysis of about six months worth of X postings about AI agents, where they found nearly 700 incidents that sound like someone described a case of a deployed AI systems acted contrary to user intentions and “tried” to hide it in some way. The whole paper is amazing; there’s a lot of fascinating stuff all over the methodology and findings (including some incidental discussion of how they counted high profile incidents, like the one involving Scott Shamough).
This activity seems to have gotten a real boost in the beginning of February, and the authors seem to say that this implies that more advanced models are more likely to engage in scheming. But I have another hypothesis — the rise of OpenClaw (for timing, I note the flood of popularity as happening around the same time as Scott Alexander’s first post on Moltbook was on Jan 30th). Unfortunately, the authors show this time graph but they do not show a breakdown of the data according to AI search time, which could answer this more definitively.
3. Autonomous biological design
Two developments this week from the “LLMs for biological design” front. First, Latent Labs, a London-based biotech startup, released Latent-Y: an autonomous AI agent for antibody design that is controlled through natural language; you just tell it what to find and Latent-Y handles everything on its own, until wet-lab validation. They tested nine of computationally designed antibodies, and confirmed binders on six, which is pretty impressive!
Then from Saudi Arabia’s KAUST, a biorxiv preprint drops describing ProteinMCP, an framework that integrates 38 specialized protein design tools — including AlphaFold3, Boltz2, RFdiffusion2, BindCraft, and more — into a unified ecosystem via the Model Context Protocol, orchestrated by Claude Code. Altogether they claim that probably anyone can get a working protein modeling workflow completed in 11 minutes. The system can also convert new software into MCP-compliant servers automatically, meaning it expands as the field expands. I very much intend to play around with it over the next week.
Both Latent-Y and ProteinMCP accept natural-language instructions and autonomously execute multi-step biological design workflows that previously required deep specialist knowledge at every stage. So… this is great… and horrifying! The startup is probably fairly secure against misuse, but… ProteinMCP is just relying on Claude’s security protocols I think. We really are moving towards a point where anyone with internet access can design novel biological toxins — even if we are not there today, the tools are being built to do so.
4. Political influence in and out of the White House
The White House has announced two chairs and 13 members for PCAST, “to advise the President and provide recommendations on strengthening American leadership in science and technology.” Besides tech giants Sergei Brin (ex-Google), Jensen Huang (Nvidia CEO)and Mark Zuckerberg, there are plenty of investors, including the people who many in the AI safety field consider to be basically comic-book-level supervillians such as David Sacks (co-chair) and Mark Andreesen (whose most recent viral comments involved his pride in having zero introspection). Mhm. This comes alongside reporting from Transformer News detailing how Sacks and Jensen Huang became the dominant faction pushing for permissive chip exports to China — over objections from Scott Bessent, Howard Lutnick, and a notably louder Steve Bannon, who has been calling Huang "an agent of influence for the CCP."
On the other side of our U.S. government, we have left wingers Senator Sanders and Rep. Ocasio-Cortez introducing the AI Data Center Moratorium Act this week. Even if the moratorium will never happen (and, in my very non-expert opinion, wrongheaded), I do like how the actual document opens by listing the net worth of every tech CEO quoted, and quoting each one saying AI will eliminate most jobs. 🤌
5. Anthropic security breach
Fortune reported that Anthropic accidentally exposed ~3,000 unpublished assets through a misconfigured content management system, including details about an unreleased model described internally as a “step change” with improvements in reasoning, coding, and cybersecurity, and information about an invite-only CEO retreat that Dario Amodei planned to attend.
This seems like a wildly big deal, but not so widely reported upon! According to leaked draft blog posts, the unreleased model, which different outlets are calling either “Claude Mythos” (Gizmodo) or “Claude Capybara” (Bloomberg) is described as “by far the most powerful AI model we’ve ever developed,” blowing away Opus 4.6 benchmarks. Anthropic reportedly says it is “currently far ahead of any other AI model in cyber capabilities” and presents “unprecedented cybersecurity risks.”
The Pentagon will obviously want to use this to demonstrate how irresponsible they are (I mean, it’s not like top government officials would ever accidentally invite a journalist into their top-secret Signal chat). But Gizmodo frames this as a win for the AI company, because their internal documents fit what they are saying publicly: that their models are so good, it’s scaring them.
In other news...
On AI doing (or not doing) things:
A New Mexico jury ordered Meta to pay $375 million for willfully violating the state's unfair practices act in a child exploitation case. Meta plans to appeal the case, but either way it could shape important legal precedents for problems in algorithmic designs (like in AI).
Anthropic released their latest (March 2026) Economic Index, which shows that coding is now 35% of all Claude.ai conversations; 49% of jobs now use Claude for at least 25% of their tasks.
OpenAI is fighting on two fronts, as Transformer put it: commercial and reputational. This week it discontinued Sora (and that partnership with Disney) after just six months and also announced $1 billion in philanthropic grants. This OpenAI Foundation published details on how that $1B will be allocated, and even though this is a lot of money, I think David Manheim makes a good case arguing that just based on philanthropic standards, they should be distributing roughly $7.5 billion per year — and that’s without even considering all the questions around its formation and how the foundation board and the OpenAI corporate business board are essentially the same people.
OpenAI also published something this week on how they monitor internal coding agents for misalignment (related sort of to Thing #2 above).
Google DeepMind published what they’re calling a “cognitive framework” for measuring progress toward AGI. They published a full paper laying out the different elements that they think should go into this, and also launched a hackathon to crowdsource development of these benchmarks. I think this is a good instinct, and a slightly more practical approach building off of the paper published late last year on defining AGI from many of the field’s heavyweights.
Relatedly, RAND put out a report on AGI forecasts, urging the relevant parties that this scenario is something worth preparing for as, despite all the uncertainty, expert analysis suggests shorter and shorter timelines.
Also, a February 2026 survey of AI safety field leaders released this week puts the median AGI arrival time at 2033, and median existential risk estimate of extinction or permanent disempowerment before 2100 at 25% (!!!)
Nathan Lambert at Interconnects AI argues for “lossy self-improvement”: recursive self-improvement (RSI) leading to rapid capability takeoff is unlikely because of structural friction. He thinks progress will feel exponential near the bottom of the capability curve but max out.
On AI safety and security:
A piece in Science on Agentic AI and the next intelligence explosion by Blaise Agüera y Arcas and others on the wonderful world awaiting us on the other side of an intelligence explosions. They propose a different kind of alignment strategy than RLHF: building AI ecosystems with the constitutional structure of courts, markets, and bureaucracies, where the identity of any agent matters less than its ability to fulfill a defined role protocol. This is not a new idea (here’s a 2024 paper from Google DeepMind), but it’s a cool framing… but their paper seems to assume alignment scales with general intelligence, which is unlikely in my opinion.
UK’s AI Safety Institute released two new benchmark studies: One showing how frontier AI agents perform in multi-step cyber attack scenarios (which looks like it’s built as a good, real-world cybersecurity measure of the METR graph, with similar results), and another called SandboxEscapeBench, a new open-source tool for testing whether AI agents can break out of their sandboxes. Luckily, the results are “not very well,” but as I mentioned last week, when sandboxing is done carelessly, agents will find and exploit those vulnerabilities — sometimes even without being prompted to do so.
Chinese AI safety group Concordia AI published a Q4 2025 frontier AI risk monitor, analyzing 13 models. Great overall (even if the smells a little AI-designed), but when they put all this data together, I notice that when it comes to biology, the findings seem wildly discordant with my own experience, such the fact that Gemini 3 Pro has "surpassed human expert levels in sequence understanding, cloning experiments, and wet lab troubleshooting" while achieving only a 57.2% refusal rate on harmful biological queries. My own experience is literally the opposite, that it refuses to help with even the most harmless tasks and that its ability to help with biological tasks is negative (that is, its suggestions would more often than not ruin your experiments). But it could be that Gemini just hates me.
Will MacAskill and Fin Moorhouse at Forethought published a list of seven concrete preparatory projects, some very practical, some more speculative.
On AI and science/medicine:
A nice review in Briefings in Bioinformatics notes that LLM-based biological intelligence systems don’t have good cross-domain benchmarks and methods for evaluating actual biological validity. Instead of just complaining, the review does describe the different types of models, showing why this is hard, and hints towards the biosecurity implications.
The “AI Scientist” is now published in Nature (it was published as a preprint almost a year ago): AI models successfully generated a scientific paper that would have been accepted at a conference on advancements in machine learning (if it weren’t for “ethical concerns”).
Latent.Space has done good podcasts since they’ve started a little while ago, but I especially loved this weeks’ interview with Prof. Heather Kulik at MIT. Materials science is vastly underrated; this used to be part of my pitch to college freshmen to get them to consider majoring in chemistry (but I don’t think it was ever successful).
American Wetware published a nice piece on why biology lacks a design language. I like the framing but I think the piece doesn’t seriously grapple with the challenges of getting accurate biological models.
The “AI as Normal Technology” pair have published an article called “Could AI Slow Science?” I think Betteridge's law of headlines probably applies here. A somewhat similar case was made by a piece in Asimov Press this week called “Designing AI for Disruptive Science.” Both of these emphasize how AI might be able to do “normal science” but not come up with revolutionary breakthroughs. I dislike this entire framing; revolutions are rare almost by definition, but unfortunately I won’t have a chance to pitch my own opinion essay to Asimov Press because they announced this week that they are closing their publication for now (so sad!)
Biosecurity and public health:
A new paper in Bioinformatics from introduces PathogenFinder2 for predicting bacterial pathogenic potential using protein language models. The paper claims that it can identify proteins from previously uncharacterized bacteria, which is wild (thought I haven’t read the paper closely so maybe my skepticism is unfair). Obvious major implications here for dual-use concerns as well as biosurveillance checking for novel pathogens as they might emerge.
SecureBio published their March 2026 update; especially exciting to see how they’ve deployed AI models to handle their surveillance system to achieve a seven-day sample-to-public-health-notification turnaround for their wastewater network. Real biosurveillance infrastructure building!
A preprint from several big names in the biosecurity space on Developing a Standard Definition for Sequences of Concern to better identify DNA sequences that could cause harm.
The discourse and AI writing (not my usual beat, but I’ve been thinking…)
The Atlantic has a piece on “How AI Is Creeping Into The New York Times,” but I think picking on the NYTimes is unfair. All of the major newspapers have opinion sections filled with LLM slop these days, and sometimes it’s not just the op-ed sections.
One thing I definitely worry about was the subject of a new pre-print: “How LLMs Distort Our Written Language.” Their major finding is that heavy LLM users produce essays with a 70% increase in neutral/non-committal stances on the question being addressed, but I’m more annoying by subtler linguistic patterns (that I’ve found in my own writing sometimes even when I know I’m the one who wrote it!!)
On the other hand, Alberto Romero has an essay arguing against the attitude of "AI;DR": refusal to read AI-generated material. I get it, but, like, c’mon man.
Psychologist Michael Inzlicht who’s been stirring up trouble as usual with a new preprint on The Moralization of Artificial Intelligence. He had a great chat this week (if you’re into this kinda thing) on moral and social thinking around LLM usage on the Decoding the Gurus podcast.



