ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code
Published in Preprint (Zenodo); in preparation for NeurIPS Datasets and Benchmarks, 2026
Status: Preprint, in preparation for NeurIPS Datasets and Benchmarks.
ESM-bench tests whether AI agents understand the physics and code of Earth System Models, rather than producing plausible text and code without grounding.
What It Measures
- A 243-task benchmark drawn from Earth System Model physics and codebases
- A classification rubric scored with precision, recall, and F1
- Multi-model evaluation across several agents and models
- Leakage detection to guard against memorized answers
Why It Matters
AI agents can generate fluent code and explanations about Earth System Models without actually reasoning about the underlying physics or the structure of the codebase. ESM-bench provides a check on that gap.
Links
- Preprint (Zenodo): https://zenodo.org/records/19802836
- Blog post: ESM-bench: testing whether AI agents understand Earth System Model physics
Recommended citation: Wu, K., Cao, Y., & Mai, G. (2026). "ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code." Preprint, Zenodo. https://zenodo.org/records/19802836
Download Paper
