Back Original

The Hallway Track No. 023: Co-Scientist, Terminal-Bench Science, AI in peer review

Google claims its Co-Scientist ran full closed-loop autonomous research across three scientific domains. Terminal-Bench Science, released the same week, puts the best agent at a 30 percent pass rate on tasks drawn from real research workflows. Both can be true: narrow, well-defined problems are yielding to autonomous methods, while general scientific competence remains far off.

AI Agents as Researchers

Research Practice