Back Original

The Hallway Track No. 021: AI peer review ensembles, Riemann zeta bound, VibeMathed registry

The verification thread keeps tightening. Anthropic's Claude improved a Riemann zeta bound and formalized the proof in Lean. VibeMathed now catalogs 565 AI-solved math problems, 134 of them Lean-verified. And Litvak's planted-error experiment shows that AI peer reviewers are surprisingly uncorrelated. Pool five systems and you catch 93 of 100 errors, but every one of them misses omissions. The pattern across all three: individual AI outputs are unreliable, but ensembles with formal verification catch most of what slips through.

Peer Review Under Pressure

AI in Mathematics

Research Practice