The verification thread keeps tightening. Anthropic's Claude improved a Riemann zeta bound and formalized the proof in Lean. VibeMathed now catalogs 565 AI-solved math problems, 134 of them Lean-verified. And Litvak's planted-error experiment shows that AI peer reviewers are surprisingly uncorrelated. Pool five systems and you catch 93 of 100 errors, but every one of them misses omissions. The pattern across all three: individual AI outputs are unreliable, but ensembles with formal verification catch most of what slips through.
Peer review is overwhelmed. Can it survive in the AI era?
Ars Technica (Saima Sidik), August 10 2026
AI generates papers faster than humans can review them, piling onto a system already strained by reviewer shortages and rising workloads. AAAI 2026 drew roughly 30,000 submissions, and several conferences are piloting clearly-labeled AI reviews to supplement human reviewers.
How well does AI peer review work?
Paul Litvak, August 10 2026
An empirical test: 100 planted errors across 10 psychology papers, reviewed by Claude, GPT-5.5, Gemini, Reviewer3, and Refine.ink. The best single system caught 71, the worst caught 30, but pooling all of them caught 93, with the 7 undetected errors all being omissions, not insertions.
Company Offering '100% Human-Written, Never AI' Medical Research Is Entirely AI
404 Media (Emanuel Maiberg), August 11 2026
A company called Research Gold advertised "100% human-written, never AI" manuscript drafting and systematic reviews; its listed PhD methodologists were AI-generated, its communications were AI-generated, and when contacted its chatbot insisted it was a real person.
Learning more about Claude's mathematical capabilities
Anthropic, August 10 2026
An unreleased Claude improved the lower bound on Riemann zeta zeros on the critical line from 41.6% to 67.2% by connecting two existing papers across subfields, with the result peer reviewed by mathematicians (including Brian Conrey and Dan Goldston) and formalized in Lean. A first-party announcement, but the Lean formalization makes the claim independently checkable.
What sort of maths are LLMs good at?
Tim Gowers, August 12 2026
A Fields medalist draws the line: LLMs do well on problems solvable by broad search, brute-force exploration, and systematic example-checking, but humans keep the edge where a proof needs deep pruning, a surprising conceptual leap, or intuition about which directions are worth trying. A precise framework for the gap this newsletter has been mapping since edition 020.
VibeMathed: a community registry of AI-solved math problems
Rasmus Lindahl et al., ongoing
A community-curated database now tracking 565 problems, 416 resolved, 134 with Lean-verified proofs, tagged by field and significance. The zoom-out companion to the one-off results this newsletter tracks piece by piece, and the Lean-verified count is the number that matters most.
Latent Space (RJ Honicky, with Matthew McPartlon and Neil Patil of Chai Discovery), August 11 2026
Chai Discovery (roughly $4B, partnerships with Eli Lilly, Novartis, argenx) builds molecular-design tools that turn drug discovery from research into engineering, generating higher-quality candidates upfront, cutting lab iteration, and making previously impractical designs like bi-specific antibodies feasible.
Certainly! Generative AI and its Impact on Academic Writing (in Finance)
Thomas Walther and Marie Dutordoir, SSRN, June 2025
An analysis of over 41,000 finance articles finds that post-ChatGPT, publication quantity rose but readability declined, especially in lower-ranked journals and among non-native English speakers. A discipline-specific measurement of AI's effect on academic writing, sitting alongside the "Published Voice" thread from edition 020.