AI Agent Knowledge Base Evaluation System
Multi-dimensional AI agent evaluation system with data anonymization and an LLM-as-a-judge architecture.
Multi-dimensional AI agent evaluation system with data anonymization and an LLM-as-a-judge architecture.
Full-stack LLM text processing system with polishing, AI detection, and plagiarism-reduction features.
A 243-task benchmark testing whether AI agents understand Earth System Model physics and code, with multi-model evaluation and leakage detection.
Advanced signal processing and data analysis of meteor-radar observations, with a novel double-Gaussian fitting algorithm.
A multi-expert AI agent framework for automated parameterization and validation of large-scale Fortran climate models.
NSF NCAR-funded project applying explainable AI to improve physics-based land-surface modeling of plant–rock–water interactions.
Unified resource platform for UT Austin campus services — 10,700+ visits, with cross-device optimization.
Published in Review of Geophysics and Planetary Physics, 2024
A comprehensive report on space physics practical education initiatives in 2022, documenting educational programs and outcomes in space science education.
Recommended citation: Wu, K.*, Xu, X., Jiang, J., & Shen, A. (2024). "A Summary Report on the Space Physics Practical Education in 2022." Review of Geophysics and Planetary Physics.
Download Paper
Published in JGR: Space Physics, 2024
This study investigates diurnal and seasonal variations of meteor speed and arrival angle using Mengcheng meteor radar observations, providing insights into meteoroid dynamics in the mesosphere and lower thermosphere.
Recommended citation: Wu, K., Yi, W.*, Xue, X.*, Reid, I., & Lu, M. (2024). "Diurnal and seasonal variations of meteor speed and arrival angle observed by Mengcheng meteor radar." JGR: Space Physics.
Download Paper
Published in Preprint (Zenodo); in preparation, 2025
Preprint. A multi-expert AI agent framework for automated parameterization and validation of large-scale Fortran climate models. Version 0.1, in preparation.
Recommended citation: Wu, K. (2025). "Noah-Agent: A Multi-Expert AI Agent Framework for Automated Parameterization and Validation of Large-Scale Fortran Climate Models (v0.1)." Preprint, Zenodo. https://zenodo.org/records/17862049
Download Paper
Published in Preprint (Zenodo); in preparation for NeurIPS Datasets and Benchmarks, 2026
Preprint. A 243-task benchmark testing whether AI agents understand Earth System Model physics and code, with multi-model evaluation, a classification rubric, precision/recall/F1 scoring, and leakage detection. In preparation for NeurIPS Datasets and Benchmarks.
Recommended citation: Wu, K., Cao, Y., & Mai, G. (2026). "ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code." Preprint, Zenodo. https://zenodo.org/records/19802836
Download Paper
Published in Geography According to Foundation Models, Vol. 422, IOS Press, 2026
Peer-reviewed book chapter reviewing key ethical issues in generative GeoAI. Wu authored Section 8, “Trust in AI and GeoAI Models,” covering geo-hallucination, uncertainty as an ethical requirement, and provenance-aware protocols.
Recommended citation: Mai, G., Lao, N., Zhang, J., Mao, L., Wang, Z., Wu, N., Janowicz, K., Wu, K., Rao, J., Gao, S., & Zhu, R. (2026). "On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust." In Geography According to Foundation Models, Vol. 422, pp. 215-232. IOS Press. DOI 10.3233/FAIA260483.
Download Paper
Published in arXiv preprint; submitted to AAAI 2027, 2026
A benchmark for end-to-end autonomous scientific research across 40 tasks from 10 scientific domains, with real-paper grounding and expert-curated multimodal rubrics.
Recommended citation: Xu, W., et al. (including Wu, K.) (2026). "ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research." arXiv preprint. Submitted to AAAI 2027.
Download Paper
Published in Submitted to AAAI 2027 Artificial Intelligence for Social Impact Track, 2027
A fitness-for-purpose survey of the progression from static personas to simulated users, submitted to the AAAI 2027 Artificial Intelligence for Social Impact Track.
Recommended citation: Liu, X., et al. (including Wu, K.) (2027). "From Personas to Simulated Users: A Fitness-for-Purpose Survey." Submitted to the AAAI 2027 Artificial Intelligence for Social Impact Track.
Published:
This oral presentation discussed the research findings on perturbations caused by the 2022 Hunga-Tonga volcano eruption in the mesosphere and lower thermosphere (MLT) region, combining WACCM-X simulation results with meteor radar observations.
Published:
Poster presentation at the 2023 AGU Fall Meeting investigating the atmospheric perturbations caused by the 2022 Hunga-Tonga volcano eruption in the mesosphere and lower thermosphere (MLT) region.
Published:
This poster presentation evaluated the Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS), presented at the Jackson School of Geosciences Research Symposium at UT Austin.
Published:
This poster presentation evaluated the Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS), presented at the Advancing Land Modeling Symposium.
Published:
This oral presentation showcased an open-source geospatial AI platform built on the AlphaEarth Foundation Model, winning 2nd place in the Geoscience Hackathon ‘25.
Published:
Poster presentation on evaluating the Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS), presented at the 106th American Meteorological Society Annual Meeting in Houston, Texas.
Published:
Oral presentation on ESM-bench at the 4th Annual Good Systems Smart Cities and AI Innovations Symposium at the University of Texas at Austin.
Graduate Teaching Assistant, University of Texas at Austin, Jackson School of Geosciences, 2024
Graduate Teaching Assistant for Earth in 2100 (GEO 303E), a course examining Earth system changes and climate projections for the year 2100.