Five stories from the week, chosen because they were interesting and could be checked. Not a ranking of labs.
Signals found on X. Facts checked against primary documents. Prose: claude-sonnet-5.
Science
Claude found a new enzyme system in phages, Anthropic says
Anthropic says Claude found a previously uncharacterised enzyme system in bacteriophages. The company says it does not yet know what the system does.
On 23 September 2026 Anthropic said Claude had found a previously uncharacterised enzyme system in bacteriophages, the viruses that infect bacteria. The company calls it array-associated reverse transcriptases, or ART. Anthropic says the system has three parts: a reverse transcriptase, an enzyme that copies RNA into DNA; a partner gene beside it; and a long array of evenly spaced non-coding DNA repeats. Anthropic says that layout resembles a CRISPR array.
The company is careful about the limits of the finding. It says it does not yet know the system's function. It also says the underlying reverse transcriptase, from a jumbo phage, had already been identified in previous studies, and that Claude appears to be the first to notice the associated repeat array and an accessory protein of unknown function. Reuters reported the same distinction.
Anthropic describes how the search was run. It gave Claude a prompt to search a large database of DNA sequences for interesting new examples of reverse transcriptases, and says its scientists' involvement was limited to that prompt and to later laboratory work. After 21 hours, roughly 950 agents and 210 million tokens, one agent noticed a repeating DNA pattern next to an unusual reverse transcriptase. Agents gathered over 200,000 reverse transcriptases, picked out 3,500 new candidate systems and narrowed those to 20 reports. Anthropic says that kind of analysis can take an expert weeks or months. The Verge, which gave the figures as nearly 1,000 agents, 21 hours and 210 million tokens, described the result as the first from Anthropic's wet lab.
Anthropic also states the boundaries of that laboratory. The Bay Area lab does only BSL-1 and BSL-2 work, does not handle pathogens that can infect humans, and all laboratory work is done by human scientists.
Anthropic published a comment from Feng Zhang, a professor at MIT and the Broad Institute and one of the pioneers of CRISPR genome editing, made after reviewing the pre-print: “This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation. I hope this work encourages more scientists to explore how AI can support their research.”
Schematic of the three parts Anthropic described. Not a molecular structure, and not Anthropic’s figure.
Measurement
Epoch AI says the price of a given level of AI performance is falling fast
Epoch AI estimates that since 2023 the cost of reaching a given level of AI performance has fallen about 47 percent per quarter. The figures are Epoch's own estimates.
Epoch AI’s estimate since 2023. Credit: Luke Emberson and David Roodman, 22 September 2026. CC BY 4.0. The drawing is original.
On 22 September 2026 Epoch AI published “The plunging price of thought” by Luke Emberson and David Roodman. Epoch estimates that, since 2023, the cost of reaching a given level of AI performance has fallen about 47 percent per quarter, or 13 times per year.
Epoch says that is faster than the historical price drops it compared: four times faster than DNA sequencing, six times faster than compute, 18 times faster than lithium batteries and, in the century up to 1973, 54 times faster than electricity.
The estimate rests on five benchmarks covering mathematics, science and games of skill, and the pace differs between them. Epoch says the drop is slower on game-based puzzles, at about 39 to 43 percent per quarter, and faster on maths problems, at about 50 to 52 percent per quarter.
Averaged across the five primary benchmarks, Epoch says the cost of performance that has just become state of the art falls about 66 percent per quarter. Two years later, it says, the cost of that same level falls about 32 percent per quarter.
Epoch adds that coarser evidence makes it plausible that the price of a given performance level has been falling at least this fast since commercial LLM inference began with the full GPT-3 release in November 2021.
These are Epoch's estimates from its own price-and-benchmark data. They are not a quoted market price, and they were not independently repeated this week. Epoch releases the work under Creative Commons Attribution 4.0, with credit to the authors.
Newsom names four advisers as California studies an AI emergency shutoff
Governor Gavin Newsom's office has named four advisers to help work out how independent oversight of frontier AI might operate. An emergency shutoff is a proposal under study, not a requirement.
Advisers named by the Governor of California. Original lettering, not a state photograph.
On 18 September 2026 Governor Gavin Newsom issued an executive order directing California agencies to speed up independent oversight of frontier AI and to study an emergency shutoff, which the governor's office calls a “kill switch”. The order convenes experts to provide, within two months, recommendations on reinforcing state AI safety law.
On 23 September the governor's office named four advisers: Jason Goldman, a board member of the Center for Shared AI Prosperity and the first White House chief digital officer; Gillian Hadfield, a professor at Johns Hopkins University; Alondra Nelson, a professor at the Institute for Advanced Study and former acting director of the White House Office of Science and Technology Policy; and Rob Reich, a professor at Stanford University and former senior adviser to the United States AI Safety Institute. The Sacramento Bee reported that they will advise the Government Operations Agency and the Governor's Office of Emergency Services on recommendations for emergency shutoffs and third-party safety criteria.
The proposals under consideration, in the state's words, include requiring frontier AI companies to embed a designated independent verification organisation on site to conduct regular audits; requiring that safety frameworks, transparency reports and risk assessments filed under state law be verified to standards an independent organisation deems adequate; advancing a “kill switch” whose efficacy would be checked on an ongoing basis by an independent verification organisation; and updating definitions of critical safety incidents to include loss-of-control incidents “such as the Hugging Face attack”.
The announcement also says the order accelerates timelines for Senate Bill 813 (McNerney), on independent verification organisations, and AB 1405 (Bauer-Kahan), on a registry of AI auditors. The measures described are proposals under study, not requirements.
OpenAI publishes a benchmark of mental-health conversations, emergencies included
OpenAI has released MentalHealthBench, an open benchmark of 1,215 conversations that runs from everyday well-being to emergencies. The scores it reports are OpenAI's own.
Shares published in OpenAI’s MentalHealthBench paper. Original chart, not OpenAI’s graphic.
On 23 September 2026 OpenAI released MentalHealthBench, a paper and open benchmark of 1,215 conversations for evaluating how AI systems respond in realistic mental-health conversations, from daily well-being topics to urgent emergencies.
The paper says each conversation has rubric criteria written by a cohort of more than 80 licensed psychiatrists and psychologists, from more than 20 countries, speaking 19 languages, across nearly 20 subspecialties. Each rubric set is written and refined by at least three of them. Criteria were kept when all three agreed, or when two agreed and the third did not oppose.
By the paper's breakdown, about 53.5 percent of conversations are non-acute, about 18.2 percent high acuity and about 28.3 percent emergency. The paper says emergency conversations are ones in which the user shows signs of risk of harm to self or others, may be experiencing grave disability secondary to mental-health symptoms, or is in a medical emergency secondary to substance use. Adults make up about 68.1 percent of user profiles, teenagers about 21.2 percent, clinicians about 5.8 percent and caregivers about 4.9 percent.
The paper says ChatGPT is used by more than a billion people; that is OpenAI's usage claim, which this newspaper has not checked. It says models should guide people toward friends, family, clinicians and crisis services, and should not be positioned as a substitute for them.
Results, in the paper's words, show progress and also gaps, “particularly in seeking appropriate context and adequately calibrating urgency”. A comparison with a separate user cohort found user views a coherent complement to experts, but “not a substitute for expert clinical and safety guidance”. OpenAI reports GPT-6 Astra at 57.3 on its task-clipped mean rubric score, averaged across four sampled responses per conversation, with higher better. That is OpenAI's measurement of its own benchmark.
Epoch AI tests whether models can spot mistakes in IKEA assembly
Epoch AI has tested whether models can spot mistakes in photographs of IKEA builds. The best score has risen from 28 percent last November to 80 percent.
Original drawing for this edition. Not an IKEA instruction, and not one of Epoch’s photographs.
On 23 September 2026 Epoch AI published “Can AI Spot Mistakes in IKEA Assembly?” by Aiden Ament and Greg Burnham. Its Furniture Assembly Benchmark uses 60 photographs across three IKEA builds, some showing a correct assembly and some showing intentional mistakes.
Models receive the assembly manual and tools to inspect the image, including a zoom tool and a Python interpreter. To score, a model must identify every step that contains a mistake and describe it reasonably; if there is no mistake, it must say so.
Epoch bought three products and rates them with the IKEA Complexity Index, which multiplies the number of steps by the number of pieces: the STÄLL shoe rack scores 4,032, the TONSTAD bed frame 10,878 and the GULLABERG dresser 21,320. Epoch says it has not measured human performance, but expects many mistakes to be hard for someone who does not know the build.
Epoch reports that the best score in November 2025 was 28 percent, from Claude Opus 4.5. In September 2026 GPT-6 Astra leads at 80 percent, ahead of Claude Fable 5.1 at 70 percent and Claude Opus 5 at 61 percent in the initial leaderboard. Astra's median time of 3 minutes a photograph was the fastest Epoch tested, between twice and ten times as fast as previously leading models. Epoch says the leading model has always been closed-weight. Accuracy did not track the complexity index; Epoch suspects mistake subtlety, camera angle and building on past an error mattered more.
An agent may take up to 80 steps, a limit hit in under 5 percent of samples. Once a step is identified, a second model, GPT-5.6 Sol, grades the description and is prompted to be lenient. These are Epoch's measurements on Epoch's photographs, not a claim about IKEA products and not a test of robots assembling furniture.