ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code

Published in Preprint (Zenodo); in preparation for NeurIPS Datasets and Benchmarks, 2026

Status: Preprint, in preparation for NeurIPS Datasets and Benchmarks.

ESM-bench tests whether AI agents understand the physics and code of Earth System Models, rather than producing plausible text and code without grounding.

What It Measures

  • A 243-task benchmark drawn from Earth System Model physics and codebases
  • A classification rubric scored with precision, recall, and F1
  • Multi-model evaluation across several agents and models
  • Leakage detection to guard against memorized answers

Why It Matters

AI agents can generate fluent code and explanations about Earth System Models without actually reasoning about the underlying physics or the structure of the codebase. ESM-bench provides a check on that gap.

Recommended citation: Wu, K., Cao, Y., & Mai, G. (2026). "ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code." Preprint, Zenodo. https://zenodo.org/records/19802836
Download Paper