Geepity Turned Benchmark Radar into an Interactive Reading
Published:
Geepity turned our paper Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation into a full interactive reading edition, walking through all nine chapters with live charts and a working BM25F demo.
Author: Koutian Wu; GitHub: ktwu01
The edition covers the finding problem, the two-stream architecture, the daily harvest, auditable BM25F ranking, the faceted taxonomy, the undated half of the catalog, the 82-of-790 comparability census, the Pareto frontier, and the prior-art search worked example. Each chapter pairs plain-English explanation with interactive widgets: drag the funnel gates, reweight the ranking formula, flip the counting mode, scrub a 577-model score history.
It closes by naming the paper, the v0.11.0 census, the dashboard at benchmark-radar.org, and the GitHub repository, framed as an independent interactive reading of arXiv:2609.11115.
An interactive edition reaches readers a PDF never meets: people who learn by dragging sliders, not by parsing equations. Nine chapters of it is the deepest third-party reading the paper has received.
See the Geepity interactive edition, or browse all Benchmark Radar coverage on my Media Coverage page.
