When an AI system produces a scientific result, someone still has to decide whether it holds up, and three of this week's items answer that in different ways. In mathematics, nine researchers including Timothy Gowers and Edward Witten have formed an advisory group at the Institute for Advanced Study to help OpenAI decide how to release a batch of machine-produced results it has not yet published, which comes down to setting the terms on which those results enter the literature. ScientistTwo, a multi-agent system from Google, is judged the other way around: its papers are graded by AI reviewers and score higher than human papers accepted at ICLR, ICML, and NeurIPS, so the reviewer is itself a model. The twelve Millennium Problems for biology were chosen so that a human can check each answer in a standard wet lab within days to weeks, which keeps the test in a person's hands.
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
arXiv (Google), September 17 2026
A multi-agent framework that takes a research problem from a human expert and runs the full cycle autonomously, establishing baselines, forming hypotheses, running experiments, and drafting papers that score higher under automated AI review than human-authored papers accepted at ICLR, ICML, and NeurIPS, which means the primary evaluation is AI reviewing AI.
Claude discovers a novel enzyme system with CRISPR-like repeats
Anthropic, September 23 2026
Roughly 950 Claude agents searched a sequence database for 21 hours and identified a previously uncharacterized array-associated reverse transcriptase (ART) system in bacteriophages with CRISPR-like repeating arrays, with all biochemical characterization done by human scientists, who have not yet determined the system's function.
AgenticLS at NeurIPS 2026: Agentic AI for Biological Discovery
agenticls.github.io, September 22 2026
A non-archival workshop at NeurIPS 2026 in Sydney on December 11 covers building, benchmarking, and deploying agentic AI systems for life-science discovery, including how AI agents should interface with robotic labs and human scientists in closed-loop experimental pipelines, which establishes agentic biological discovery as a named subdiscipline at a major ML conference.
Millennium Problems for biology
Samuel G. Rodriques (Edison Scientific, FutureHouse), LinkedIn, September 19 2026
Twelve hard biology problems published at millenniumproblems.bio were selected on the explicit criterion that each must be checkable against an objective measurable endpoint in a standard wet lab within days to weeks, with anything lacking a clear validation endpoint excluded, and Rodriques frames the list as "the last reasonable eval for AI in biology," making the selection criteria themselves an argument about what makes a benchmark safe to construct.
A Breakdown of arXiv Submissions in Mathematics
Tammy Kolda and Daniel Appelo, Substack, September 21 2026
Mathematics arXiv submissions rose approximately 79 percent year over year from August 2025 to August 2026, the fastest growth of any field, concentrated in combinatorics and metric geometry, while the supply of expert referees able to evaluate the submissions has not changed.
Announcing the Advisory Group on Mathematics and Artificial Intelligence
Guest post on Terence Tao's blog, September 21 2026
Nine mathematicians, including Timothy Gowers and Edward Witten, have formed the Advisory Group on Mathematics and AI, hosted at the Institute for Advanced Study and operating without payment from AI companies, with the immediate task of advising OpenAI on how to coordinate the release of what the announcement describes as a large number of significant mathematical results produced by OpenAI's internal model that have not yet been publicly disclosed.
The AI discourse in academia is toxic
Ran Blekhman (University of Chicago), Substack, September 20 2026
After sharing a Nature paper on AI-enabled interactive scientific documents, Blekhman received hundreds of hostile replies including personal attacks and calls for him to lose funding, and he argues that the effect is selective: when a mildly AI-positive post draws a coordinated pile-on, researchers learn to stop posting, and the visible record of scientific opinion on AI becomes unrepresentative of what the field actually thinks.