Sitemap

A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.

Pages

Chronicles

Animated founder and academic chronicle pages built with an open-source template, covering researchers and founders in geoscience, AI, and security.

Posts

Benchmark Radar Day 41: A Ranking Decided by Contract First

2 minute read

Published:

Day forty-one of Benchmark Radar. The ranking engine for recent releases shipped its first phase with the data contract written first, and the report audit separated raw best scores from scores measured under the same protocol. Scoreboard: 160 stars, 28 forks.

Benchmark Radar 第四十一天:先定数据契约,再谈排行榜

less than 1 minute read

Published:

Benchmark Radar 的第四十一天。新鲜发布的排行榜引擎第一阶段上线,先把数据契约写在前头;报告审计也把原始最高分和同一协议下测出的分数分开。记分牌:160 颗星,28 个 fork。

Benchmark Radar Day 39: The Daily Brief Gets a Home

2 minute read

Published:

Day thirty-nine of Benchmark Radar. The daily brief now has a landing page, a full archive, and a feed of its own, and the dashboard navigation got short enough to scan. Scoreboard: 144 stars, 27 forks.

Benchmark Radar Day 38: A Blog Built From Evidence

2 minute read

Published:

Day thirty-eight of Benchmark Radar. The daily brief, the thing the project already wrote every day and never gave a page, now has a blog and a feed. Scoreboard: 137 stars, 26 forks.

Benchmark Radar Day 37: Cite It, Set It Up, Search Everything

2 minute read

Published:

Day thirty-seven of Benchmark Radar. Anyone writing a paper can now copy a citation without leaving the dashboard, the offline setup moved onto the site, and search looks through every collected day by default. Scoreboard: 130 stars, 25 forks.

Benchmark Radar Day 36: 37 Scores on One Model Card

2 minute read

Published:

Day thirty-six of Benchmark Radar. The biggest model card yet landed with 37 scores across 34 benchmarks, and search stopped guessing which result answered the question. Scoreboard: 128 stars, 24 forks.

Benchmark Radar 第三十六天:一张模型卡,37 个分数

less than 1 minute read

Published:

Benchmark Radar 的第三十六天。迄今最大的一张模型卡落地了,34 个基准上排了 37 个分数;搜索也不再靠猜来回答哪个结果是对的。记分牌:128 颗星,24 个 fork。

Benchmark Radar Day 35: One URL for Every Benchmark

3 minute read

Published:

Day thirty-five of Benchmark Radar. A week-old copycat outranked the real project for its own name because search engines could see only four pages here. Today every one of the 1,173 benchmarks the radar tracks has a page of its own. Scoreboard: 123 stars, 23 forks.

Benchmark Radar 第三十五天:每个基准都有自己的网址

less than 1 minute read

Published:

Benchmark Radar 的第三十五天。一个上线才一周的仿站,靠自家名字把真项目压在了搜索结果下面,因为搜索引擎在这里只能看到四个页面。今天,雷达追踪的 1173 个基准,每一个都有了自己独立的页面。记分牌:123 颗星,23 个 fork。

Benchmark Radar Day 34: A Radar You Can Run From a Command Line

3 minute read

Published:

Day thirty-four of Benchmark Radar. The radar now runs from a command line on your own computer, offline, no dashboard in between, and a coding agent can answer benchmark questions straight from the data. Scoreboard: 121 stars, 23 forks.

a.ai to z.ai: What Every Single-Letter AI Domain Actually Does

6 minute read

Published:

There are only 26 single-letter .ai domains. I went through all of them, one by one, to see who actually builds something and who just squats. The answer: about six run real products, one belongs to Elon Musk, one belongs to Google, one apparently belongs to Apple without a website, and the rest are for-sale pages asking between 1.5 and 500 million dollars. Here is the full alphabet.

a.ai 到 z.ai:26 个单字母 AI 域名到底都在干什么

1 minute read

Published:

单字母 .ai 域名只有 26 个。我把它们挨个访问了一遍,看谁在真做事,谁在蹲着等涨价。答案:大约六个在跑真产品,一个属于马斯克,一个属于谷歌,一个据说属于苹果但没有网站,剩下的是开价 150 万到 5 亿美元的出售页。以下是完整字母表。

Benchmark Radar 第二十九天:发布先于更新、0 到 1 读作 0 到 100、高亮始终可见

less than 1 minute read

Published:

新东西该在第一页,而不是藏在第三页。第二十九天我们做了三件事:让新发布排在例行更新之前,把 0 到 1 的分数按 0 到 100 来显示而不断点,让当前栏目始终看起来是选中的。先说几个词:发布是全新冒出来的基准;更新是对已有条目的改动,例如版本升级或新增分数;0 到 1 与 0 到 100 可以是同一数值的两种刻度。

Benchmark Radar 第二十八天:数据只剩一个真相来源,页面只剩一个 h1,引用一键可复制

less than 1 minute read

Published:

别给同一份真相留两个副本,否则它们迟早对不上。第二十八天我们做了三件事:把可生成的数据移出版本库,让页面只对爬虫说一个标题,给作品一个可引用的名字。先说几个词:真相来源是唯一可信的那份源文件,其余都由它生成;h1 是页面的主标题,爬虫期望整页只有一个;i18n 是国际化,让界面按语言显示;引用是你写论文时粘贴的那条参考文献。

Benchmark Radar 第二十七天:今日列表分页、标题减半、21 个基准正名

1 minute read

Published:

一天 136 条还要一次全画出来,只会让首屏变慢。第二十七天我们做了三件事:让今日列表一页页加载,把标题砍半,让 21 个基准显示真名。先说几个词:分页是一长串内容分多页看完;首包是页面为了快而先加载的小数据包;SEO 是让搜索引擎看懂并收录网站的做法。

Benchmark Radar 第二十六天:排名站到最前,返回键终于能用,前沿线不再撒谎

less than 1 minute read

Published:

一个以排名命名的页面,却把排名藏在第六屏,这不叫排名页。第二十六天我们做了三件事:让排名站到最前,让返回键能用,把一条编出来的图换成一条能画的线。先说几个词:排名就是按多少排序的榜单;返回键是浏览器左上角让你回到上一页的箭头;前沿线是一条阶梯线,连起截至每个日期为止的最好分数;过期横幅是数据过时时顶部的提示条。

Benchmark Radar 第十八天:稳定运行

less than 1 minute read

Published:

你好,我是 Koutian。第十八天,管道自己跑完了,没要任何人插手。这一天,恰恰证明了它稳。

Luck Sourcing: The Sourcing of Getting Lucky

2 minute read

Published:

A long time ago, I came across a book called Chase, Chance, and Creativity: The Lucky Art of Novelty(《追逐、机遇和创造力:新奇的幸运艺术》). I no longer remember everything in it, but its central idea stayed somewhere in the back of my mind: luck and creativity are not entirely random. Chance may arrive unexpectedly, but we can still choose how often we encounter it and whether we are ready to recognize it.

Luck Sourcing:主动寻找好运

less than 1 minute read

Published:

很久以前,我偶然看到一本书,叫 Chase, Chance, and Creativity: The Lucky Art of Novelty(《追逐、机遇和创造力:新奇的幸运艺术》)。我已经不记得书里的全部内容了,但它的核心想法一直留在我脑海中的某个角落:运气和创造力并不完全是随机的。机遇也许会意外到来,但我们仍然可以选择自己遇见它的频率,以及当它出现时,我们是否已经准备好认出它。

The GitHub Apps You Installed Months Ago Still Have Full Access. Go Check.

5 minute read

Published:

I was in the middle of setting up a second Cloudflare backup, verifying one dataset and configuring another, when the AI agent helping me flagged something unrelated: while listing what had access to my GitHub account, one authorization stood out as unusually broad. That single flag turned into an hour of reviewing GitHub’s installed-apps list, and it was worth every minute.

你几个月前装的那个 GitHub App,权限可能还全开着,去看看

less than 1 minute read

Published:

我当时正在配置第二份 Cloudflare 备份——核实一份数据、再搭建另一份——帮我干活的 AI Agent 顺手提醒了我一件不相关的事:在列出哪些东西能访问我的 GitHub 账号时,有一项授权的范围明显大得不正常。就是这一句提醒,让我花了一个多小时去仔细过了一遍 GitHub 的已安装应用列表,而这一个小时非常值得。

Benchmark Radar:面向 AI 基准的、以证据为先的每日雷达

1 minute read

Published:

新的 AI 基准冒出来的速度,比任何研究员评估它们的速度都快。所以我建了一个雷达,让「发现新基准」变成一件每天做、过程透明、可以复核的事。先说一个词:基准(benchmark)就是用来考 AI 的一套题,或者一份评测。

Benchmark Radar 第二天:累积趋势与工件去重

less than 1 minute read

Published:

雷达开始有记忆了。第二天我们建起了累积趋势图,还顺手解决了工件别名的问题。先说一个词:工件(artifact)就是被追踪的每一个具体东西,比如某个基准或数据集。

The Skills I Built for Weaker Models, and Why I Deleted Most of Them

6 minute read

Published:

I ran a health check on my Claude Code setup this week and found 174 custom skills, 124 of which I had never invoked once. They were not failures. Most of them were scaffolding I built for a model that needed it, and then kept long after the model stopped needing it.

我为较弱的模型做了这些 Skills,后来为什么删掉了大部分

less than 1 minute read

Published:

这周,我给自己的 Claude Code 配置做了一次健康检查,发现里面有 174 个自定义 Skill,其中 124 个我一次都没有调用过。它们并不是失败品。大部分只是我为当时需要额外支撑的模型搭起的脚手架,而模型早已不再需要它们,我却一直把它们留到了现在。

LLM 应用机会,不是更聪明的聊天框

less than 1 minute read

Published:

LLM 最有价值的应用,可能不是回答问题,而是把一项专家工作变成可以执行、复核和审计的数字工作单元。

Does “No Stupid Questions” Mean “Ask Anything” or “Do Not Ask Dumb Questions”?

6 minute read

Published:

A widely shared post describes a brainstorming meeting at a robotics startup in Boston. A slide said “No stupid questions.” A Chinese employee interpreted it as “Do not ask stupid questions,” while American colleagues explained it as “Ask freely; no question will be judged stupid.” A related post claims that “Say it again?” is neutral when someone is not heard, whereas “What did you say?” is hostile.

“No stupid questions”到底是“随便问”还是“别问蠢问题”?

1 minute read

Published:

一张流传截图讲了这样一件事:一家波士顿机器人创业公司开头脑风暴会,幻灯片写着 “No stupid questions.” 一位中国员工把它理解为“不要问蠢问题”,美国同事却说它的意思是“什么都可以问,没有问题会被当成蠢问题”。另一组讨论又声称:没听清时说 “Say it again?” 很中性,而 “What did you say?” 很不友好。

我在 YVR“合理地”走错了 NEXUS

1 minute read

Published:

我走错了温哥华机场的 NEXUS 通道;规则上是我的错,设计上却是一条几乎可以预测的错误路径。

Another Side of Vancouver Airport: Age, Workspace, and the People I Saw

6 minute read

Published:

Beyond the wayfinding problem that led me into the wrong NEXUS line, my connection at Vancouver International Airport left me with several observations unrelated to signs. I noticed an older-looking mix of people, a departures area that offered almost nowhere to work, and a visible contrast between people in Vancouver and Texas. This post records those observations. The full account of the NEXUS incident is in How I “Reasonably” Ended Up in the Wrong NEXUS Line at YVR.

温哥华机场的另一面:年龄结构、办公空间与我看到的”人群”

less than 1 minute read

Published:

除了那次走错 NEXUS 通道的导视问题,这次在温哥华机场(YVR)转机还让我观察到几件和路标无关、但同样值得记录的事情:明显偏大的现场年龄结构、几乎不为办公设计的候机空间,以及温哥华与得州人群景观的直观差异。这篇笔记单独记录这些观察;如果你想看我是怎么在 NEXUS 通道走错路的,可以读我在 YVR”合理地”走错了 NEXUS

Why Did a S$5.50 Purchase Appear as US$4.27?

3 minute read

Published:

A receipt in Singapore showed S$5.50, while a U.S. credit-card account displayed only US$4.27. Another purchase of about S$48 appeared as roughly US$38. A bank-card transit ride also failed to appear immediately as pending.

为什么 S$5.50 的消费在美国信用卡上只显示 US$4.27?

less than 1 minute read

Published:

在新加坡消费时,一张收据写着 S$5.50,但美国信用卡账户只显示 US$4.27。另一笔约 S$48 的消费,则显示为约 US$38。与此同时,刷银行卡乘坐公共交通后,交易也没有立刻出现在 pending 中。

Why alias rm=trash Cannot Stop an AI Agent from rm -rf

5 minute read

Published:

I asked Claude Code a narrow question: can I protect this machine from Codex accidentally deleting files forever, just by aliasing rm to trash? The honest answer turned out to be no, for a reason that is not obvious until you actually test it, and the fix ended up being a five-layer setup rather than a one-liner.

为什么 alias rm=trash 拦不住 AI agent 的 rm -rf

1 minute read

Published:

我问了 Claude Code 一个很具体的问题:能不能只靠把 rm alias 成 trash,来防止 Codex 意外把文件永久删掉。真实的答案是不能,原因不实测根本看不出来,最后落地的也不是一行配置,而是五层防护。

Hard Benchmarks Should Not Become Coding Tricks

10 minute read

Published:

The most dangerous failure mode of an AI science benchmark is not that it is too hard, it is that it quietly becomes either a coding trick or a guessing game.

不要偷偷录客户访谈

1 minute read

Published:

偷偷录一次客户访谈,可能把一个正常的产品研究流程变成隐私和合规事故。

算力是新一代的鸡蛋:当大厂开始”发鸡蛋”

less than 1 minute read

Published:

算力 / token 是新一代的鸡蛋。现在各大厂商都在发算力,就像在发鸡蛋一样。你老了就要去跟别的老奶奶、老大爷抢发鸡蛋——这已经不是一个笑话,可能就是一个现实。

Why Your Brand-New WD Drive Is Read-Only on a Mac (It’s Not the Drive)

4 minute read

Published:

I plugged a Western Digital external drive full of data into a MacBook Air, tried to copy a file onto it, and nothing happened. No error dialog, no progress bar, just a drive that would let me read everything and write nothing. My first instinct was that something was broken, or that I needed to fix permissions. Both were wrong, and chasing the wrong explanation almost led me to permanently downgrade the security of the whole laptop.

为什么你崭新的 WD 移动硬盘在 Mac 上只能读不能写(问题不在硬盘)

less than 1 minute read

Published:

我把一块装满数据的西部数据(WD)移动硬盘插到 MacBook Air 上,想往里拷一个文件,结果什么都没发生。没有报错弹窗,没有进度条,就是一块能读出所有东西、却一个字节都写不进去的硬盘。我的第一反应是它坏了,或者是我得去修一下权限。这两个判断都是错的,而且顺着错误的解释找下去,差点让我把整台笔记本的安全性永久降级。

Building a real ZIP bomb in Fortran, C++, and C

2 minute read

Published:

I’ve been playing with mixed-language builds (Fortran calling into C++ and C via iso_c_binding) and wanted a demo that was more interesting than “add two numbers across languages.” So I built fortran-zip-bomb: a small program that generates a genuine ZIP bomb — a small archive that expands into a much larger file on decompression.

用 Fortran、C++ 和 C 构建一个真正的 ZIP 炸弹

less than 1 minute read

Published:

我一直在玩混合语言构建(Fortran 通过 iso_c_binding 调用 C++ 和 C),想要一个比”跨语言把两个数加起来”更有意思的演示。于是我建了 fortran-zip-bomb:一个小程序,生成一个真正的 ZIP 炸弹——一个解压时会扩展成大得多的文件的小压缩包。

A Materials-Science Model of Egg Fried Rice, Three Years Later

10 minute read

Published:

In October 2023 I asked ChatGPT an over-engineered question: how do you fry rice so that every grain of rice ends up bonded to egg, with no bare rice grains and no isolated clumps of egg sitting off on their own? I saved that conversation as an HTML file, dropped it in my Downloads folder, and did not look at it again for almost three years. Last week I finally turned it into an actual open-source repo, and going back through the original conversation to build it was more interesting than I expected.

蛋炒饭材料学模型,三年之后

less than 1 minute read

Published:

2023 年 10 月,我问了 ChatGPT 一个过度工程化的问题:怎么炒蛋炒饭,才能让每一粒米都裹上蛋,既没有裸露的白米粒,也没有单独抱团、没沾到米的蛋碎?那次对话我存成了一个 HTML 文件,扔进 Downloads 文件夹,然后差不多三年没再打开过。上周我终于把它整理成了一个真正的开源仓库,回头重读那段 2023 年的对话、把它做成代码的过程,比我预想的有意思得多。

A Friend I Met on a United SFO–PVG Flight

5 minute read

Published:

On a United Airlines flight from San Francisco SFO to Shanghai Pudong PVG, the plane had started its descent and the cabin announcements kept reminding everyone to fasten their seatbelts. Outside was night; inside, people were already lit up with the excitement of “finally going home.” It was on this flight that I met a friend.

在 UA SFO–PVG 航班上遇到的一位朋友

less than 1 minute read

Published:

在联合航空从旧金山 SFO 飞往上海浦东 PVG 的航班上,飞机已经开始下降,机舱广播一遍遍提醒大家系好安全带。窗外是夜色,舱内的人却已经被“终于回国了”的兴奋点亮。我就在这趟航班上认识了一位朋友。

Why I Must Throw Myself Into the AI Wave

less than 1 minute read

Published:

Recently I’ve sometimes felt confused by how far AI has come, a bit lost and anxious, unsure what I should do, and then, because of my identity and my path, wondering what fallbacks or better options exist.

为什么我一定要投身 AI 浪潮

less than 1 minute read

Published:

最近虽然有的时候也会因为 AI 的发展程度感到非常困惑,感到有一点迷茫、有些焦虑,不知道自己该干啥,然后又因为自己的身份和路径问题,在想有什么样的退路或者更好的方案。

Further Thoughts on Social Contradictions

6 minute read

Published:

While you’re still grinding your heart out on “if you just work hard enough, you can move up,” you may not realize the system was never designed to cultivate you in the first place, it was designed to screen you.

社会矛盾的进一步思考

less than 1 minute read

Published:

当你还在为“只要努力就能向上流动”而拼命内卷时,你可能没有意识到,这个系统从一开始就不是为了培养你,而是为了筛选你。

Test Your Heart Age

5 minute read

Published:

When an ordinary-looking “heart age” questionnaire lets you calculate that your heart is younger than your actual age, you may need to re-examine the lifestyle habits you’ve been ignoring.

测一测你的“心脏年龄”

less than 1 minute read

Published:

当一份看似普通的“心脏年龄”问卷让你算出自己的心脏比实际年龄还年轻时,你可能需要重新审视一下那些被你忽视的生活习惯了。

你是哪一种?AI 时代的表演图鉴

less than 1 minute read

Published:

当你在这个由PPT、Paper和焦虑构成的AI时代大剧院里找座位时,不如先看看台上的人都在演哪一出戏。

From a Poetry Society to Unicorns: The Less-Traveled Road Isn’t Laziness

6 minute read

Published:

Hah, I can’t help but laugh. It just hit me: the first time I ever used Markdown was back when I was building a poetry society. And that’s also when I first learned about Git. Looking back now, if you put all the founding members of that poetry society together, you’d almost have two unicorns.

从诗社到独角兽:少走的路不是偷懒

less than 1 minute read

Published:

哎,我他妈笑了。我忽然想起来,我最早用 Markdown,就是之前创建诗社的时候。知道 Git,也是在那个时候。现在回头看,整个诗社的元老凑在一起,真的快有两个独角兽了。

杨振宁先生到底有多少财富?

8 minute read

Published:

一张 1957 年的诺贝尔奖支票,一个中国第一位数论博士的父亲,三套不同的货币制度,加上七十年的复利,我们到底要怎么给杨振宁家族 2026 年的财富估出一个不是瞎编的数字?

消失的科学家,为什么 600 年才追得上才是科学真正的瓶颈

less than 1 minute read

Published:

一篇新论文把 1901 到 2023 年所有 739 位科学类诺贝尔奖得主从童年扒到现在,得出了一个特别狠的数字,按目前的进步速度,一个出生在低收入国家的孩子要等大约 600 年,才能拿到和富裕国家孩子一样的诺奖机会。

Prompt Caching 不是技术债

1 minute read

Published:

很多人把 prompt caching 看成一个省钱 hack,但我觉得这个判断刚好反了。

My PhD Advisor Built a Land Surface Model That Forecasted Hurricane Harvey

5 minute read

Published:

When you join a lab, you do not just get a research direction. You inherit a 30-year codebase that is currently running inside the U.S. National Water Model and was on the critical path forecasting Hurricane Harvey. That is what working with Zong-Liang Yang at the Jackson School of Geosciences actually looks like.

我的博士导师打造的陆面模型,曾用于预报飓风 Harvey

less than 1 minute read

Published:

加入实验室以后,你会接过一个研究方向,也会继承一套已有 30 年历史的代码库。它目前运行在美国国家水模型中,也曾处在飓风 Harvey 预报工作的关键路径上。这就是在 Jackson School of Geosciences 与 Zong-Liang Yang 一起工作的真实样子。

我用一周为杨振宁建了一座年鉴,家谱才是真正的故事

1 minute read

Published:

大多数物理本科生知道杨振宁是诺贝尔奖得主,知道他是 Yang-Mills 里的那个 Yang。而为他建一座年鉴,让我看到了教科书略过的东西:他一生中最有分量的一个事实是谁是他的父亲,以及那个父亲在他出生之前,为他铺好了什么。

Sean Xiang Has Been Building Bloombase for 14 Years, and the AI Era Finally Caught Up to It

6 minute read

Published:

Most enterprise security companies show up, ride one trend, and disappear in the next infrastructure cycle. Sean Xiang has been building Bloombase since January 2012, and the company has somehow been on the right side of every major infrastructure shift since, including the current AI accelerator era. That is not luck. That is a thesis.

Sean Xiang 已经打造 Bloombase 14 年,AI 时代终于追上了它

1 minute read

Published:

大多数企业安全公司冒个泡、赶一波趋势,然后在下一个基础设施周期里消失。Sean Xiang 从 2012 年 1 月起就在打造 Bloombase,这家公司却阴差阳错站到了此后每一波重大基础设施转变的正确一边,包括当下的 AI 加速器时代。那不是运气,那是一套论点。

Marc Hesse Does the Fluid Mechanics of Everything From Magma to Mars

5 minute read

Published:

You can study fluid mechanics in five different countries before you turn 30, work on petroleum reservoirs and tectonophysics and planetary ice on the same week, and somehow end up at a Centennial Chair in Geophysics. Marc Hesse did exactly that, and the through-line is more interesting than any of the individual stops.

Marc Hesse 研究从岩浆到火星的万物流体力学

less than 1 minute read

Published:

你可以在 30 岁前在五个不同国家学流体力学,同一周里既研究油气储层又研究构造物理和行星冰,最后竟然坐上地球物理学百年讲席。Marc Hesse 就是这么做的,而他背后那条主线比任何一个单独的站点都更有意思。

Kehan Dong 与那些比别人更早开始建造的人合作

less than 1 minute read

Published:

大多数 VC 和孵化器都谈支持创始人。Kehan Dong 专门支持那种 16 岁就开始动手建造、而根本没人告诉过他们可以这么做的创始人,事实证明这个群体被严重忽视了。

Gemini Says My Ego Essay Is Still a Humblebrag

6 minute read

Published:

Earlier today I published a piece called “Four Ego Mistakes I Made as a 22-Year-Old Founder.” Then I ran it through Gemini 2.5 Pro as an independent reviewer. Gemini’s verdict was that the essay is itself an ego move. I think Gemini is mostly right.

Gemini 说我那篇 ego 文还是凡尔赛

less than 1 minute read

Published:

今天早些时候我发了一篇叫《22 岁 founder 的四个 ego 错误》。然后我把这篇喂给了 Gemini 2.5 Pro 当独立审稿人。Gemini 的判决是:这篇文章本身就是一个 ego 动作。我觉得 Gemini 大体是对的。

Geeta Persad Came Back to Austin to Build a Climate Group That Actually Talks to Policy

5 minute read

Published:

Most academic climate scientists will tell you they care about policy and then publish a paper that no policymaker is ever going to read. Geeta Persad spent four years working at the Union of Concerned Scientists translating climate models for water managers, and she came back to academia knowing exactly what the gap looks like.

Geeta Persad 回到奥斯汀,组建了一个真正与政策对话的气候团队

less than 1 minute read

Published:

大多数学术气候科学家会告诉你他们在乎政策,然后发表一篇任何决策者都不会读的论文。Geeta Persad 在忧思科学家联盟(Union of Concerned Scientists)花了四年,为水资源管理者翻译气候模型,然后带着对那个差距的清醒认识回到了学术界。

Trees Drink From Rock, and Daniella Rempe Proved It

5 minute read

Published:

If you ask most people where trees in California get their water in a drought, they will say “the soil.” It turns out a huge fraction of it comes from cracks in the bedrock underneath the soil, and Daniella Rempe is the person who put numbers on it.

树从石头里喝水,Daniella Rempe 证明了这件事

less than 1 minute read

Published:

如果你问大多数人,加州在干旱时树从哪里取水,他们会说”土壤”。事实证明,很大一部分水其实来自土壤下方岩石裂缝里的基岩,而 Daniella Rempe 就是把数字放到这件事上的人。

Ashley Matheny Treats Trees as Pumps, and That Changes the Whole Model

5 minute read

Published:

Most land surface models treat a tree like a passive straw. Water comes in at the roots, water leaves at the leaves, end of story. Ashley Matheny’s research basically says no, a tree is an active hydraulic system with storage, capacitance, and a strategy, and if you do not model it that way you are going to be wrong about drought.

把树当成水泵:Ashley Matheny 改变了整个陆地模型

less than 1 minute read

Published:

大多数陆地表面模型都把一棵树当成一根被动的吸管。水从根进来,水从叶出去,故事就这么简单。Ashley Matheny 的研究基本上在说:不对,树是一个带有储水、电容和策略的活跃水力系统,如果你不这样建模,你在干旱问题上就会犯错误。

Who Hoards Multimodal Data for Real, Meta, ByteDance, X, and the Visa Plot Twist

7 minute read

Published:

Ask which campus actually sits on the best multimodal feedstock for GPT-4V-class perception, Sora-class video, Gemini-scale bundles, and Meta Emu-style image stacks, and the short answer is almost vulgar in how cleanly it splits three big piles. Meta still pulls ahead by a chasm on stills, ByteDance owns the high-velocity short-video river that is really motion plus audio, and X ships the smallest absolute media volume yet the weirdest leverage on tight text-image coupling and live-event semantics.

「图片地主」对「视频钥匙」对「实时百科」,多模态家底Meta、ByteDance和X怎么分

less than 1 minute read

Published:

这题问到刀尖上了。如果把「高质量多模态训练数据」收窄到对 GPT-4V、Sora、Gemini、Emu 这类模型真有喂饭价值的图文或视频,短答其实很锋利,Meta(Facebook / Instagram)静态图片数量的库存对其他两家几乎是断层第一,ByteDance(TikTok / 抖音)在短视频也就是动态图片流上占最大优势,X(Twitter)绝对量级最小,但图文的贴脸相关性和实时信息密度是独一份。

Think First, Code Later with AI

2 minute read

Published:

In an era where AI can write code in seconds, I just learned the hard way that blindly moving fast is actually slowing me down.

Earth System Model Skill Packages: Deep Knowledge Bundles for Noah-MP, CLM, CAM, MOM6, WRF, E3SM, and More

3 minute read

Published:

Earth system models are some of the most complex scientific software ever written, and they are also some of the worst-documented for newcomers. I have been building a series of “skill packages” — structured, progressive-disclosure knowledge bundles — for the major Earth system and land surface models, designed to be used by both new graduate students and AI coding agents.

地球系统模型技能包:为 Noah-MP、CLM、CAM、MOM6、WRF、E3SM 等量身打造的深层知识包

less than 1 minute read

Published:

地球系统模型是人类写过的、有史以来最复杂的一批科学软件,可它们对新手来说偏偏又是文档最糟糕的一批。我一直在为主要的几大地球系统和陆地表面模型构建一系列”技能包”——结构化的、渐进式披露的知识包——设计给刚入门的研究生和 AI 编码代理两类使用者使用。

In Conversation with Chenxi Hu

7 minute read

Published:

THIS IS A FAKE BLOG. The content below is fabricated and should not be cited or treated as a real interview or factual record.

对话胡晨曦

less than 1 minute read

Published:

胡晨曦的研究揭示,城市化不只是被动地应对极端天气,它还会主动重塑热带气旋如何向沿海特大城市倾泻暴雨。

AI PhD Survival Guide: How to Finish a PhD in the Age of LLMs

1 minute read

Published:

A PhD is hard. A PhD in 2026 — with the field moving faster than your committee can read — is a different kind of hard. AI PhD Survival Guide is a handbook for surviving and finishing an AI or ML PhD without burning out and without falling behind.

AI 博士生存指南:在 LLM 时代如何拿下博士学位

less than 1 minute read

Published:

读博很难。2026 年读博——当领域进展快到你的委员会都读不过来时——是另一种难。AI PhD Survival Guide 是一本手册,教你如何在 AI 或机器学习博士研究中活下来并毕业,既不会燃烧殆尽,也不至于掉队。

Why Is American Cuisine Lacking in Umami?

6 minute read

Published:

Although “umami” is the unshakable soul of Jiangsu-Zhejiang and Cantonese cooking, in traditional American food you often wander only among single-note salt, sweet, and oil, struggling to find that layered depth of flavor. That isn’t accidental; it’s the inevitable result of a deep cultural difference in how ingredients are handled, how seasoning is reasoned, and how industrial production works.

为什么美国菜缺乏鲜味?

less than 1 minute read

Published:

虽然“鲜”(Umami)在江浙菜和广府菜中是不可撼动的灵魂,但在传统美国菜里,你往往只能在单一的咸、甜、油之间徘徊,而难觅那种富有层次感的味觉深度。这并非偶然,而是一场关于食材处理、调味逻辑与工业化生产方式的深层文化差异所导致的必然结果。

Standing at the Alamo: A Sacred Ground of Texas History

2 minute read

Published:

The silence surrounding the Alamo chapel in San Antonio belies the brutal, thirteen-day siege that transformed this former Spanish mission into the ultimate symbol of Texan independence. Standing before its weathered facade today, one can almost hear the echoes of a conflict that remains one of the most legendary chapters in American history.

站在阿拉莫:德克萨斯历史的圣地

less than 1 minute read

Published:

今天我站在了圣安东尼奥的阿拉莫教堂前。这座看似宁静的建筑,却承载着德克萨斯乃至美国历史上最惨烈、最具传奇色彩的一页。

When Climate Data Comes Alive: A Day at TACC

12 minute read

Published:

The roar of cooling systems in the Texas Advanced Computing Center (TACC) isn’t just noise—it’s the power of supercomputers translating the overwhelming dimensionality of climate data into something students can finally see and understand.

当气候数据活起来:在 TACC 的一天

1 minute read

Published:

德州高级计算中心(TACC)里冷却系统的轰鸣不只是噪音——那是超级计算机的力量,正在把气候数据那令人不知所措的多维性,转化成学生最终能看见、能理解的东西。

hao-tokens: A Practical Guide to Free and Cheap LLM API Tokens

1 minute read

Published:

Every indie developer who has ever burned through a free API tier knows the feeling: you build something cool, and then your OPENAI_API_KEY runs out at the worst possible moment. hao-tokens is a curated list of the legitimate ways to get free or low-cost LLM tokens so you can keep building.

hao-tokens:免费和低价 LLM API Token 实用指南

less than 1 minute read

Published:

每一个曾经把某个免费 API 层额度烧光的独立开发者都知道那种感觉:你做了很酷的东西,然后你的 OPENAI_API_KEY 在最不该耗尽的时候用光了。hao-tokens 是一份经过整理、合法获取免费或低价 LLM token 的清单,让你能继续开发下去。

Virtual Cell Neuromorphic Gene Language Models: A VC Intern’s Field Guide

7 minute read

Published:

Virtual cell models powered by neuromorphic computing and gene language models represent one of the most capital-intensive and scientifically ambitious convergences in biotech AI. If you’re evaluating this space as a VC intern, you need to understand three core components: what these systems actually do, why the market is moving now, and where the investable opportunities lie.

虚拟细胞神经形态基因语言模型:VC 实习生的领域指南

less than 1 minute read

Published:

由神经形态计算和基因语言模型驱动的虚拟细胞模型,代表了生物技术 AI 领域中资本最密集、科学最雄心勃勃的融合方向之一。如果你正在以 VC 实习生的身份评估这个领域,你需要理解三个核心组成部分:这些系统实际上做什么、为什么市场现在开始行动,以及可投资的机会在哪里。

Moat Plus Momentum: Why AI Makes Preparation Optional

6 minute read

Published:

Research, stock trading, and startups all share a common pattern: success comes from combining a defensible core competency with the ability to ride trending waves. You don’t need exhaustive preparation anymore. You need methodology and the ability to produce content when it matters. When the right moment arrives, you strike.

护城河加热点:为什么 AI 让准备变得可选

less than 1 minute read

Published:

科研、股市还有创业,这些所有需要展示并能得到结果的东西,本质上都是”主业/具有护城河的本行加上热点”。所以需要做好准备,或者说不一定非要进行那种极其周全的准备,只要掌握一些方法论就可以。利用 AI 在关键时刻能够产出内容,一旦关键热点到来,马上就可以抓住。

The Gibbs Phenomenon: Why LLM Hallucinations Are Mathematically Inevitable

4 minute read

Published:

When Large Language Models (LLMs) hallucinate, we often treat it as an engineering bug to be fixed. But what if hallucinations are not a flaw, but a mathematical inevitability? A recent interdisciplinary discussion revealed a profound connection between the Gibbs phenomenon in Fourier analysis and the fundamental limitations of neural networks.

熵悖论:为什么AI难以实现科学发现

less than 1 minute read

Published:

AI 自动化科研的愿景令人陶醉:想象机器能在我们睡觉时生成假设、设计实验、发表论文。然而,尽管关于“AI 科学家”的新闻铺天盖地,我们正在撞上一堵根本性的墙。问题不在于算力或数据集规模,而是某种更深刻的东西,根植于科学发现的本质和信息论之中。

How to Find Outliers in Statistics

12 minute read

Published:

In statistics, how you find outliers depends on the dimensionality of your data, its distributional characteristics, and your tolerance for what counts as “abnormal.” Here are the most common and standard approaches:

统计学上如何寻找Outlier

1 minute read

Published:

在统计学中,寻找离群值(Outlier)的方法取决于数据的维度、分布特征以及你对“异常”的容忍程度。以下是几种最常用且标准的方法:

How PhDs Can Self-Design KPIs: From ‘Felt Effort’ to ‘Systematic Output’

5 minute read

Published:

Doing a PhD cannot rely on “felt effort” and “moving yourself emotionally.” We need to upgrade day-to-day literature reading, experiment design, data analysis, and paper writing into a personal research system that is quantifiable, reviewable, and continuously optimizable. The point is not self-exploitation, but verifying whether your research efficiency truly exists, and keeping precious PhD time from being consumed inefficiently.

PhD如何自我设计KPI:从“感觉努力”到“系统产出”

less than 1 minute read

Published:

读博不能仅凭“感觉努力”和“自我感动”。我们需要将日常的文献阅读、实验设计、数据分析和论文写作,升级为一套可量化、可复盘、可持续优化的个人科研系统。重点不在于自我压榨,而在于验证自己的科研效率是否真实存在,避免宝贵的博士时间被低效消耗。

For Dating and Resource-Sharing Markets, Should You Use Traditional Search/Ads/Rec or AI Recommendation Algorithms?

9 minute read

Published:

I’ve been discussing this question with a friend lately. He wants to use AI for matching in the dating market; I think with a small sample size this is perfectly feasible, you don’t even need to build any search/ads/rec system at all. Just ask the large language model directly, toss in a few users’ profiles, and have it rank them; the whole process is very simple.

婚恋和资源共享市场,到底该用传统搜广推还是 AI 推荐算法?

less than 1 minute read

Published:

我和朋友最近在讨论这个问题。朋友说想拿 AI 来做婚恋市场的匹配,我觉得在样本量小的情况下,这完全可行——甚至根本不需要搭任何搜广推系统。你直接问大语言模型,把几个用户的简历丢进去,让它排序就好了,整个过程非常简单。

AI 生存指南:在 LLM 时代作为知识工作者如何保持有用

less than 1 minute read

Published:

如今,每一位知识工作者都在用某项工作的某些部分与 LLM 竞争。AI Survival Guide 是一本手册,帮你弄清是哪些部分、该怎么应对,以及如何带着更强的技艺从这场竞争的另一头走出来,而不是被它取代。

I published a preprint on Zenodo: Research grounding for a bilateral venture capital model of PhD programs

26 minute read

Published:

A PhD system built on a 19th-century apprenticeship model is failing by nearly every empirical measure — ~40–50% attrition, depression rates six times the general population, and tenure-track placement below 15% in many fields — while venture capital has spent four decades perfecting bilateral contracts that manage exactly the risks PhD programs ignore: information asymmetry, moral hazard, hold-up, and misaligned incentives. The literature across economics, education policy, signaling theory, and AI research converges on a striking conclusion: the structural tools to fix the PhD already exist in VC contract design, but academia has never imported them. This research compendium maps the evidentiary landscape across six domains to ground the argument.

Every Generation Has Its Own To-Do List

7 minute read

Published:

Every generation has its own way of managing work: paper notebooks, SaaS task managers, and now programmable agentic workflows powered by tools like OpenClaw heartbeat.

每一代人,都有每一代人的 To-Do List

1 minute read

Published:

每一代人都有自己管理任务的方式:最早是纸和笔,后来是 SaaS 任务管理工具,现在则开始进入像 OpenClaw heartbeat 这样可程序化、可持续运行的 agentic workflow 时代。

把 GitHub PR 当作技术人的 Inbound Marketing

1 minute read

Published:

对技术人来说,在一个高速增长的开源仓库里做出高质量 PR,往往比再发一篇泛泛而谈的 AI 观点帖更像真正有效的 inbound marketing。

Technology Is Not the Moat; Sales Is

9 minute read

Published:

The most important thing is selling. Having technology is useless; it only matters if someone is willing to buy. I don’t think anyone has a real tech moat; that part is easy to solve. As long as you have a little technical foundation, you can handle it. What matters most is that people come and buy. Though, it might be that my own technical level is too high, and I’ve grown numb to technology.

技术不是壁垒,销售才是

1 minute read

Published:

所以最重要的是推销。有技术并没有什么卵用,只要有人要购买才有用。我觉得所有人都没有 tech 壁垒,这个很容易解决。只要稍微有一点 tech 基础,都能解决。最重要的是,有人来买。不过,也可能是我的技术太高了,我对技术已经无感。

用能量消耗衡量工作

less than 1 minute read

Published:

衡量工作效率,最重要的指标或许不是投入的时间,而是消耗的能源。

Clear Plus Airport Experience: Money, Privilege, and Market Regulation

3 minute read

Published:

Exchanging money for time is a privilege I rarely indulge in, but a surprise membership benefit that cleared airport security in ten minutes changed my perspective on friction and market regulation. While a Clear Plus membership normally costs over a hundred dollars annually, obtaining it for free through an Uber membership allowed me to experience a level of efficiency that money can’t always buy—at least not without a well-regulated system behind it.

Clear Plus 体验:金钱、特权与市场调节

less than 1 minute read

Published:

用金钱换取时间是一种我很少尝试的“特权”,但这次在奥斯汀机场仅用十分钟便完成安检的经历,让我对效率与市场调节有了新的思考。这份特权并非我主动购买,而是通过 Uber 年费会员赠送的 Clear Plus 获得的——原本需要每年支付一百多美金的服务,在免除门槛后,带给我一种金钱也未必能随时买到的流畅体验。

I Spent a Night Reverse-Engineering Buy Borrow Die With Gemini, and Realized the F-1 Script Looks Nothing Like the Billionaire One

14 minute read

Published:

I spent a whole evening with Gemini taking apart the Buy Borrow Die playbook that American billionaires run, expecting that with a little scaling down I could just copy the moves, and what I found instead was that almost every single move has an F-1 trapdoor underneath it, and the list of things I can actually do fits on one page.

Vibe Coding IDEs: a brief comparison (EN)

6 minute read

Published:

Every product has its own pros and cons. Cursor: “extraordinarily productive”; Kiro: “spec-driven”; Antigravity: “agent-first”.

Vibe Coding IDEs brief comparison

2 minute read

Published:

每一家都有每一家的优点和缺点。Cursor: “extraordinarily productive”; Kiro: “spec-driven”; Antigravity: “agent-first”.

巴菲特:比GitHub早了半个世纪的”开源”运动领袖

less than 1 minute read

Published:

在当今这个由代码、协作和透明度驱动的时代,GitHub 成为了”开源”精神的代名词。但如果我们将目光投向金融界,会发现一位”开源”的先行者,他比 GitHub 的诞生早了整整半个世纪。他就是沃伦·巴菲特。

Neural Galaxy - 属于你的 AI 对话可视化宇宙

1 minute read

Published:

一个支持手势控制的 3D 可视化项目,把你的 AI 对话历史变成可以飞行探索的星系。你可以在 ChatGPT 对话之间穿梭,也可以把抽象的人工智能概念变成一个能亲手操作的可视化空间。Try Live Demo

TQQQ ML Trend: Predicting a 3x Leveraged ETF With Machine Learning

1 minute read

Published:

TQQQ is the 3x-leveraged Nasdaq-100 ETF. It is also one of the most asymmetric instruments retail investors touch — the upside is real, the drawdowns are brutal, and the daily-rebalance math means buy-and-hold doesn’t behave the way most people assume. TQQQ ML Trend is an experiment in using machine learning to predict the trend regime, not the price.

TQQQ ML 趋势:用机器学习预测 3 倍杠杆 ETF

less than 1 minute read

Published:

TQQQ 是纳斯达克 100 指数的 3 倍杠杆 ETF。它也是散户接触过的最不对称的金融工具之一——上行空间真实存在,回撤却十分残酷,而每日再平衡的数学逻辑意味着”买入并持有”并不像大多数人以为的那样运转。TQQQ ML Trend 是一个用机器学习来预测趋势状态的实验,而不是去预测价格。

Collaborating with Claude Code to Update My Academic Website

3 minute read

Published:

Today I had an interesting experience collaborating with Claude Code to completely overhaul my personal academic website. As a PhD student in Geological and Earth Sciences at UT Austin, I needed to update my GitHub Pages site with real professional information instead of the placeholder content that had been sitting there.

和 Claude Code 一起更新我的学术网站

1 minute read

Published:

今天我经历了一次挺有意思的合作:和 Claude Code 一起,把我的个人学术网站彻底重做了一遍。作为 UT Austin Geological and Earth Sciences 的博士生,我需要把 GitHub Pages 站点从一堆占位符内容,更新成真正能代表我专业背景的信息。

Tmux Orchestrator - Run AI agents 24/7

6 minute read

Published:

The Tmux Orchestrator enables Claude agents to work autonomously, schedule their own check-ins, and coordinate across multiple projects without human intervention - a project I explored and learned a lot from.

komomood - Couple Mood Tracking Heatmap

1 minute read

Published:

An elegant couple mood tracking website with self-hosted backend and SQLite, displaying daily mood records in GitHub contribution graph style.

komomood - 情侣心情追踪热力图

less than 1 minute read

Published:

这是一个优雅的情侣心情记录网站,采用自托管后端和 SQLite,以 GitHub contribution graph 风格展示每日心情记录。

Welcome to My Academic Website

less than 1 minute read

Published:

From the vast datasets of Earth System Models to the specialized niches of high-performance computing, this space documents my journey as a PhD student at UT Austin pushing the boundaries of Geological and Earth Sciences. This website is more than just a portfolio—it’s a hub where data-driven climate science meets the practical challenges of modern research.

欢迎来到我的学术网站

less than 1 minute read

Published:

在 UT Austin 攻读地球科学博士学位的过程中,我始终在试图寻找复杂气候模型与真实世界影响之间的联结。这个学术网站不仅是我研究、论文和项目经历的展示入口,更是我记录如何利用数据驱动的科学方法去理解地球系统演变的思考空间。

做过划掉:一款告诉你这一生值不值得的计算器

less than 1 minute read

Published:

“这b人生过的值不值”——简单说,就是”这该死的一生过得值不值?”——这句在中国互联网上承载了大量情感的话语。ZuoGuoHuaDiao(做过划掉)是一款小巧的网页工具,它认真对待这句话,并试图用数字来回答它。试试在线演示

暑期计算器:你的暑假到底值多少钱?

less than 1 minute read

Published:

大多数人把暑假当作休息时间。我开始怀疑这种框架是否低估了它。Summer Calculator 是一款小巧的网页工具,用来估算你暑假的完整价值——金钱、学习、人际关系、健康——而不仅仅是你没工作的那些日子。试试在线演示

1AI-polish - AI 学术写作润色系统

1 minute read

Published:

学术写作的严谨性不仅在于数据,更在于表达的精准,而 1AI-polish 通过集成 DeepSeek-R1 的推理能力,为研究者提供了一套集文本润色与 AI 检测于一体的深度协作系统,旨在让复杂的科研思想以更专业、更透明的方式呈现。

科技再发达,我们依然是在用基因做决策

less than 1 minute read

Published:

故事是这样的,我最近一直在寻思一个问题,既然现在的科技都已经发达到这种地步了,为什么我们这帮人,还是得天天花心思去收拾自己的形象?

NASA FINESST Resources: A Practical Guide and Link Library for the FINESST Proposal

2 minute read

Published:

NASA’s Future Investigators in NASA Earth and Space Science and Technology (FINESST) is one of the most underused fellowships among US graduate students. Most PhD students have never heard of it, and the ones who have often miss the deadline because the proposal expectations aren’t obvious from the call alone. This repo is a curated guide of links, tips, and examples to help you write a competitive FINESST proposal.

NASA FINESST 资源:FINESST 提案实用指南与链接库

less than 1 minute read

Published:

NASA 的未来地球与空间科技研究者(FINESST)项目,是美国研究生中最被低估的奖学金之一。这个仓库是一份精心整理的链接、技巧与示例指南,帮助你写出一份有竞争力的 FINESST 提案。

NASA FINESST Resources Guide

3 minute read

Published:

Securing up to $150,000 in research funding over three years can define a graduate career, and NASA’s FINESST program is the primary vehicle for that transformation. This guide distills the complex application process into actionable strategies for Earth and Space Science researchers seeking to join the next generation of Future Investigators.

NASA FINESST 资源申请指南

1 minute read

Published:

每年 5 万美元且连续 3 年的科研资助,让 NASA FINESST 项目成为地球与空间科学博士生必须争取的黄金机会。这份指南旨在将复杂的申请流程拆解为可操作的策略,帮助下一代“未来研究员”(Future Investigators)在激烈的竞争中脱颖而出。

LEAD-UTexas - Land Environment and Atmospheric Dynamics Group

1 minute read

Published:

Dr. Zong-Liang Yang’s Land Environment and Atmospheric Dynamics (LEAD) Group at UT-Austin employs satellite remote sensing, earth system modeling, and high-performance computing to advance understanding of Earth system sciences.

portfolio

UT01 Navigation Platform

Unified resource platform for UT Austin campus services — 34,940+ visits, with cross-device optimization.

publications

Perturbations by the 2022 Hunga-Tonga Volcano Eruption in the MLT Region Investigated Using the WACCM-X Simulation and Meteor Radar Observations

Published in AGU Fall Meeting 2023, 2023

AGU Fall Meeting abstract studying wave perturbations from the 2022 Hunga-Tonga eruption in the mesosphere and lower thermosphere using WACCM-X simulations and meteor radar observations.

Recommended citation: Wu, K., Liu, H.-L., Yi, W., & Xue, X. (2023). "Perturbations by the 2022 Hunga-Tonga Volcano Eruption in the MLT Region Investigated Using the WACCM-X Simulation and Meteor Radar Observations." AGU Fall Meeting Abstracts, SA33B-2892.
Download Paper

A Summary Report on the Space Physics Practical Education in 2022

Published in Review of Geophysics and Planetary Physics, 2024

A comprehensive report on space physics practical education initiatives in 2022, documenting educational programs and outcomes in space science education.

Recommended citation: Wu, K.*, Xu, X., Jiang, J., & Shen, A. (2024). "A Summary Report on the Space Physics Practical Education in 2022." Review of Geophysics and Planetary Physics.
Download Paper

Diurnal and seasonal variations of meteor speed and arrival angle observed by Mengcheng meteor radar

Published in JGR: Space Physics, 2024

This study investigates diurnal and seasonal variations of meteor speed and arrival angle using Mengcheng meteor radar observations, providing insights into meteoroid dynamics in the mesosphere and lower thermosphere.

Recommended citation: Wu, K., Yi, W.*, Xue, X.*, Reid, I., & Lu, M. (2024). "Diurnal and seasonal variations of meteor speed and arrival angle observed by Mengcheng meteor radar." JGR: Space Physics.
Download Paper

Noah-Agent: A Multi-Expert AI Agent Framework for Automated Parameterization and Validation of Large-Scale Fortran Climate Models (v0.1)

Published in Preprint (Zenodo); in preparation, 2025

Preprint. A multi-expert AI agent framework for automated parameterization and validation of large-scale Fortran climate models. Version 0.1, in preparation.

Recommended citation: Wu, K. (2025). "Noah-Agent: A Multi-Expert AI Agent Framework for Automated Parameterization and Validation of Large-Scale Fortran Climate Models (v0.1)." Preprint, Zenodo. https://zenodo.org/records/17862049
Download Paper

ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code

Published in Preprint (Zenodo); in preparation for NeurIPS Datasets and Benchmarks, 2026

Preprint. A 243-task benchmark testing whether AI agents understand Earth System Model physics and code, with multi-model evaluation, a classification rubric, precision/recall/F1 scoring, and leakage detection. In preparation for NeurIPS Datasets and Benchmarks.

Recommended citation: Wu, K., Cao, Y., & Mai, G. (2026). "ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code." Preprint, Zenodo. https://zenodo.org/records/19802836
Download Paper

On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust

Published in Geography According to Foundation Models, Vol. 422, IOS Press, 2026

Peer-reviewed book chapter reviewing key ethical issues in generative GeoAI. Wu authored Section 8, “Trust in AI and GeoAI Models,” covering geo-hallucination, uncertainty as an ethical requirement, and provenance-aware protocols.

Recommended citation: Mai, G., Lao, N., Zhang, J., Mao, L., Wang, Z., Wu, N., Janowicz, K., Wu, K., Rao, J., Gao, S., & Zhu, R. (2026). "On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust." In Geography According to Foundation Models, Vol. 422, pp. 215-232. IOS Press. DOI 10.3233/FAIA260483.
Download Paper

How Does Integrating Plant Hydraulics Improve Noah-MP Land Surface Model

Published in 106th AMS Annual Meeting, 2026

Conference poster presenting the integration and evaluation of a plant hydraulics scheme in the Noah-MP land surface model.

Recommended citation: Wu, K., Li, L., Rempe, D., Matheny, A., Mbarak, M., & Yang, Z.-L. (2026). "How Does Integrating Plant Hydraulics Improve Noah-MP Land Surface Model." Poster presented at the 106th AMS Annual Meeting, Houston, TX.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Published in arXiv preprint; submitted to AAAI 2027, 2026

A benchmark for end-to-end autonomous scientific research across 40 tasks from 10 scientific domains, with real-paper grounding and expert-curated multimodal rubrics.

Recommended citation: Xu, W., et al. (including Wu, K.) (2026). "ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research." arXiv preprint. Submitted to AAAI 2027.
Download Paper

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Published in arXiv preprint arXiv:2608.04205, 2026

A population-scale simulated-user evaluation infrastructure with 8.3 billion persona records, four interactive playground environments, and 1,010 application tasks across 25+ domains.

Recommended citation: Li, X., et al. (including Wu, K.) (2026). "MatrAIx: Simulating the World with 8.3 Billion Persona Agents." arXiv preprint arXiv:2608.04205.
Download Paper

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

Published in arXiv preprint arXiv:2608.15844, 2026

A behavioral-science instrument that measures identity drift in generative agents carrying an immutable “soul file” through a resource-scarce, long-horizon multi-agent simulation.

Recommended citation: Ng, S., et al. (including Wu, K.) (2026). "MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations." arXiv preprint arXiv:2608.15844.
Download Paper

ASI-Bench: At the Dawn of Artificial Superintelligence

Published in arXiv preprint arXiv:2608.17271, 2026

A benchmark with 60 project-level tasks across 11 scientific domains that evaluates frontier models on open-ended scientific research beyond human-expert benchmarks.

Recommended citation: Zhou, J., et al. (including Wu, K.) (2026). "ASI-Bench: At the Dawn of Artificial Superintelligence." arXiv preprint arXiv:2608.17271.
Download Paper

From Personas to Simulated Users: A Fitness-for-Purpose Survey

Published in Submitted to AAAI 2027 Artificial Intelligence for Social Impact Track, 2027

A fitness-for-purpose survey of the progression from static personas to simulated users, submitted to the AAAI 2027 Artificial Intelligence for Social Impact Track.

Recommended citation: Liu, X., et al. (including Wu, K.) (2027). "From Personas to Simulated Users: A Fitness-for-Purpose Survey." Submitted to the AAAI 2027 Artificial Intelligence for Social Impact Track.

talks

teaching

Earth in 2100 (GEO 303E)

Graduate Teaching Assistant, University of Texas at Austin, Jackson School of Geosciences, 2024

Graduate Teaching Assistant for Earth in 2100 (GEO 303E), a course examining Earth system changes and climate projections for the year 2100.