ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
Published in arXiv preprint; submitted to AAAI 2027, 2026
ResearchClawBench evaluates AI systems on end-to-end autonomous scientific research rather than isolated research subtasks. Its 40 tasks span 10 scientific domains, with each task grounded in a real published paper and evaluated using expert-curated multimodal rubrics.
Evaluation
- Seven autonomous-research agents are evaluated under a unified protocol.
- Seventeen native large language models are evaluated through ResearchHarness.
- Error analysis focuses on experimental-protocol mismatch, evidence mismatch, and missing scientific core.
Links
- Paper (arXiv): https://arxiv.org/abs/2606.07591
Recommended citation: Xu, W., et al. (including Wu, K.) (2026). "ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research." arXiv preprint. Submitted to AAAI 2027.
Download Paper
