TutorialsFrance
OpenMeasure AI agent resolution rates with Terminal-Bench-Science v0.1 across 70 expert scientific workflows
Practical guide to running Terminal-Bench-Science v0.1: reproduce the 70-task benchmark, compute resolution rates (top model ≈30%), and build a small audit-ready evaluation pipeline.