Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Chronicles
Animated founder and academic chronicle pages built with an open-source template, covering researchers and founders in geoscience, AI, and security.
CV - Koutian Wu | AI4Geoscience PhD Student at UT Austin
Comprehensive CV of Koutian Wu, PhD student working on AI agents for science: agent evaluation and benchmark design (ESM-bench), land surface modeling, and full-stack engineering.
Posts
“We Don’t Train on Your Data” Is a Promise. “Your Data Never Left Your Computer” Is a Fact.
Published:
The Navier-Stokes fight is not proof that anyone stole anything. It is proof that the person whose work was at stake had no way to check.
「我们不拿你的数据训练」是一句承诺,「你的数据从未离开你的电脑」才是一个事实
Published:
Navier-Stokes 之争并没有证明谁偷了谁的东西。它证明的是:那个把研究押在里面的人,根本无从查证。
AI Gave Us Answers. The Most Expensive Thing Is the Question.
Published:
The dark night gave me dark eyes, yet I use them to seek the light.
AI 给了我们答案,最贵的是问题
Published:
黑夜给了我黑色的眼睛,我却用它寻找光明。
Benchmark Radar Day 41: A Ranking Decided by Contract First
Published:
Day forty-one of Benchmark Radar. The ranking engine for recent releases shipped its first phase with the data contract written first, and the report audit separated raw best scores from scores measured under the same protocol. Scoreboard: 160 stars, 28 forks.
Benchmark Radar 第四十一天:先定数据契约,再谈排行榜
Published:
Benchmark Radar 的第四十一天。新鲜发布的排行榜引擎第一阶段上线,先把数据契约写在前头;报告审计也把原始最高分和同一协议下测出的分数分开。记分牌:160 颗星,28 个 fork。
Benchmark Radar Day 40: A Source for Papers, a Benchmark Restored, a Wider Broadcast
Published:
Day forty of Benchmark Radar. The radar started watching the network where most new benchmark papers get their DOI, brought a benchmark back by searching its exact name, and started posting where it was not posting before. Scoreboard: 147 stars, 27 forks.
Benchmark Radar 第四十天:盯上论文的登记处,找回一个基准,发到更多地方
Published:
Benchmark Radar 的第四十天。雷达开始盯着大多数新基准论文拿到 DOI 的那个网络,靠搜索一个基准的精确名字把它找了回来,并开始向以前没发过的地方发布。记分牌:147 颗星,27 个 fork。
Benchmark Radar Day 39: The Daily Brief Gets a Home
Published:
Day thirty-nine of Benchmark Radar. The daily brief now has a landing page, a full archive, and a feed of its own, and the dashboard navigation got short enough to scan. Scoreboard: 144 stars, 27 forks.
Benchmark Radar 第三十九天:每日简报有了自己的家
Published:
Benchmark Radar 的第三十九天。每日简报现在有了自己的落地页、完整档案和订阅源,仪表盘导航也短到一眼能扫完。记分牌:144 颗星,27 个 fork。
Benchmark Radar Day 38: A Blog Built From Evidence
Published:
Day thirty-eight of Benchmark Radar. The daily brief, the thing the project already wrote every day and never gave a page, now has a blog and a feed. Scoreboard: 137 stars, 26 forks.
Benchmark Radar 第三十八天:用证据搭起来的博客
Published:
Benchmark Radar 的第三十八天。每日简报,这个项目每天都写、却从没给它一个页面的东西,现在有了博客和订阅源。记分牌:137 颗星,26 个 fork。
Benchmark Radar Day 37: Cite It, Set It Up, Search Everything
Published:
Day thirty-seven of Benchmark Radar. Anyone writing a paper can now copy a citation without leaving the dashboard, the offline setup moved onto the site, and search looks through every collected day by default. Scoreboard: 130 stars, 25 forks.
Benchmark Radar 第三十七天:引用、设置、全库搜索,一步到位
Published:
Benchmark Radar 的第三十七天。写论文的人现在不用离开仪表盘就能复制引用,离线安装搬到了网站上,搜索默认会翻遍每一个采集过的日子。记分牌:130 颗星,25 个 fork。
Benchmark Radar Day 36: 37 Scores on One Model Card
Published:
Day thirty-six of Benchmark Radar. The biggest model card yet landed with 37 scores across 34 benchmarks, and search stopped guessing which result answered the question. Scoreboard: 128 stars, 24 forks.
Benchmark Radar 第三十六天:一张模型卡,37 个分数
Published:
Benchmark Radar 的第三十六天。迄今最大的一张模型卡落地了,34 个基准上排了 37 个分数;搜索也不再靠猜来回答哪个结果是对的。记分牌:128 颗星,24 个 fork。
Benchmark Radar Day 35: One URL for Every Benchmark
Published:
Day thirty-five of Benchmark Radar. A week-old copycat outranked the real project for its own name because search engines could see only four pages here. Today every one of the 1,173 benchmarks the radar tracks has a page of its own. Scoreboard: 123 stars, 23 forks.
Benchmark Radar 第三十五天:每个基准都有自己的网址
Published:
Benchmark Radar 的第三十五天。一个上线才一周的仿站,靠自家名字把真项目压在了搜索结果下面,因为搜索引擎在这里只能看到四个页面。今天,雷达追踪的 1173 个基准,每一个都有了自己独立的页面。记分牌:123 颗星,23 个 fork。
Benchmark Radar Day 34: A Radar You Can Run From a Command Line
Published:
Day thirty-four of Benchmark Radar. The radar now runs from a command line on your own computer, offline, no dashboard in between, and a coding agent can answer benchmark questions straight from the data. Scoreboard: 121 stars, 23 forks.
Benchmark Radar 第三十四天:可以离线在命令行里跑的雷达
Published:
Benchmark Radar 的第三十四天。雷达现在可以从你自己电脑的命令行运行,离线可用,中间不隔仪表盘,编码智能体可以直接从数据里回答基准测试的问题。记分牌:121 颗星,23 个 fork。
Benchmark Radar Day 33: A Radar That Finds Releases No Keyword Search Can Catch
Published:
Day thirty-three of Benchmark Radar. The radar learned to find benchmark releases that keyword search would never catch, and a new frontier model card with 14 benchmarks joined the registry. Scoreboard: 115 stars, 21 forks.
Benchmark Radar 第三十三天:关键词搜索永远抓不到的发布,雷达能抓到了
Published:
Benchmark Radar 的第三十三天。雷达学会了找到关键词搜索永远抓不到的基准发布,注册表里也进了一张带 14 个基准的新前沿模型卡片。记分牌:115 颗星,21 个 fork。
UT Austin Featured My Research, Teaching, and Science Communication
Published:
The University of Texas at Austin’s Office of Graduate and Postdoctoral Studies featured my work on LinkedIn on August 26, 2026.
UT Austin 专题介绍了我的科研、教学与科学传播工作
Published:
德克萨斯大学奥斯汀分校研究生与博士后事务办公室于 2026 年 8 月 26 日在 LinkedIn 上专题介绍了我的工作。
Benchmark Radar Day 32: Errors That Never Shouted, Adoption Without Binaries, and a Briefing a Person Would Write
Published:
Day thirty-two of Benchmark Radar. We fixed data errors that failed silently, taught the ranking to count adoption for releases that ship no binary files, and rewrote the daily briefing so it reads like a person wrote it. Scoreboard: 113 stars, 21 forks.
Benchmark Radar 第三十二天:从不吭声的错误、没有二进制的采纳、像人写的简报
Published:
Benchmark Radar 的第三十二天。我们修掉了静默失败的数据错误,让排名在发布不附带任何二进制文件时也能统计采纳,还把每日简报改写成像人写的。记分牌:113 颗星,21 个 fork。
Benchmark Radar Day 31: One-Line Release Cards, Two-Field Forms, and a README You Can Read
Published:
Day thirty-one of Benchmark Radar. Every release card now carries a one-line summary instead of a bare version tag, filing an issue takes at most two required fields, and both READMEs got short enough to read. Scoreboard: 107 stars, 20 forks.
Benchmark Radar 第三十一天:一行摘要的发布卡片、两项必填的 Issue 表单、读得下去的 README
Published:
Benchmark Radar 的第三十一天。每张发布卡片现在都带一行摘要,而不是一个光秃秃的版本号;提交 issue 最多填两个必填项;两份 README 也精简到读得下去。记分牌:107 颗星,20 个 fork。
a.ai to z.ai: What Every Single-Letter AI Domain Actually Does
Published:
There are only 26 single-letter .ai domains. I went through all of them, one by one, to see who actually builds something and who just squats. The answer: about six run real products, one belongs to Elon Musk, one belongs to Google, one apparently belongs to Apple without a website, and the rest are for-sale pages asking between 1.5 and 500 million dollars. Here is the full alphabet.
a.ai 到 z.ai:26 个单字母 AI 域名到底都在干什么
Published:
单字母 .ai 域名只有 26 个。我把它们挨个访问了一遍,看谁在真做事,谁在蹲着等涨价。答案:大约六个在跑真产品,一个属于马斯克,一个属于谷歌,一个据说属于苹果但没有网站,剩下的是开价 150 万到 5 亿美元的出售页。以下是完整字母表。
Benchmark Radar Day 30: A Live Log on the Road to 1,000 Stars
Published:
Day thirty of Benchmark Radar. Thirty posts in, this log has become a live broadcast of one open-source project trying to reach 1,000 stars, and today the scoreboard reads 86.
Benchmark Radar 第三十天:冲向 1000 星的公开直播
Published:
Benchmark Radar 的第三十天。写到今天,这份日志正式变成一场公开直播:一个开源项目如何从零走向 1000 颗星。先报分数:今天记分牌上写着 86。
Benchmark Radar Day 29: Releases Before Updates, Scores on 0 to 100, and a Nav That Looks Active
Published:
Day twenty-nine of Benchmark Radar. We ranked new releases before routine updates, put a 0 to 1 score on a 0 to 100 scale without changing its value, and made the active tab stay active.
Benchmark Radar 第二十九天:发布先于更新、0 到 1 读作 0 到 100、高亮始终可见
Published:
新东西该在第一页,而不是藏在第三页。第二十九天我们做了三件事:让新发布排在例行更新之前,把 0 到 1 的分数按 0 到 100 来显示而不断点,让当前栏目始终看起来是选中的。先说几个词:发布是全新冒出来的基准;更新是对已有条目的改动,例如版本升级或新增分数;0 到 1 与 0 到 100 可以是同一数值的两种刻度。
Benchmark Radar Day 28: One Source of Truth, One H1, and a Citation You Can Copy
Published:
Day twenty-eight of Benchmark Radar. We removed a second copy of the data, made the page say one thing to crawlers, and gave the work a citable name.
Benchmark Radar 第二十八天:数据只剩一个真相来源,页面只剩一个 h1,引用一键可复制
Published:
别给同一份真相留两个副本,否则它们迟早对不上。第二十八天我们做了三件事:把可生成的数据移出版本库,让页面只对爬虫说一个标题,给作品一个可引用的名字。先说几个词:真相来源是唯一可信的那份源文件,其余都由它生成;h1 是页面的主标题,爬虫期望整页只有一个;i18n 是国际化,让界面按语言显示;引用是你写论文时粘贴的那条参考文献。
Benchmark Radar Day 27: Paging Today, Halving a Title, and Naming 21 Benchmarks Right
Published:
Day twenty-seven of Benchmark Radar. We made the Today list load a page at a time, cut a title in half, and taught 21 benchmarks to show their real names.
Benchmark Radar 第二十七天:今日列表分页、标题减半、21 个基准正名
Published:
一天 136 条还要一次全画出来,只会让首屏变慢。第二十七天我们做了三件事:让今日列表一页页加载,把标题砍半,让 21 个基准显示真名。先说几个词:分页是一长串内容分多页看完;首包是页面为了快而先加载的小数据包;SEO 是让搜索引擎看懂并收录网站的做法。
Benchmark Radar Day 26: The Ranking Leads, Back Finally Works, and a Frontier That Tells the Truth
Published:
Day twenty-six of Benchmark Radar. We put the ranking where it belongs, made the Back button work, and replaced an invented chart with one we can draw.
Benchmark Radar 第二十六天:排名站到最前,返回键终于能用,前沿线不再撒谎
Published:
一个以排名命名的页面,却把排名藏在第六屏,这不叫排名页。第二十六天我们做了三件事:让排名站到最前,让返回键能用,把一条编出来的图换成一条能画的线。先说几个词:排名就是按多少排序的榜单;返回键是浏览器左上角让你回到上一页的箭头;前沿线是一条阶梯线,连起截至每个日期为止的最好分数;过期横幅是数据过时时顶部的提示条。
Benchmark Radar Day 25: Scores on Their Dates, and a Radar That Does Not Rank Itself
Published:
Day twenty-five of Benchmark Radar. We stopped the radar from ranking itself, put scores on the dates they belong to, and let a benchmark with no score still speak for itself.
Benchmark Radar 第二十五天:分数落在它们所属的日期上,一个不再推荐自己的雷达
Published:
一个会把自己排进榜单的雷达,不配叫雷达。第二十五天我们做了三件事:让排名不再跟自己玩博弈,让每天抓来的分数落在它该在的发布日期上,再让那些还没人评分的基准自己说出身份。先说一个词:排行榜(leaderboard)就是一张公开榜单,谁分数高谁排前面。
Benchmark Radar Day 24: Results-First Today View and Honest Empty States
Published:
Day twenty-four of Benchmark Radar. We put the results first on the homepage, and we stopped blaming you when a page came up empty.
Benchmark Radar 第二十四天:结果优先的今日视图与诚实的空状态
Published:
今天我们把首页重做了一遍,让它先把结果摆出来,而不是先甩给读者一个空页面。我们还让那些「爬到了但没结果」的来源变得看得见了。先解释两个词:PR 是别人提议改的一段代码,issue 是一个待办或 bug。
Niu Lai and GameStop: Two Comeback Myths, Compared
Published:
What do a 7,169-yuan Chinese animated film and a 483-dollar GameStop short squeeze have in common? More than you would think.
牛来与GameStop:两个逆袭神话的横纵对照
Published:
一部票房 7169 元的国产动画,和一只冲到 483 美元的 GameStop 股票,有什么共同点?比你想象的要多。
Benchmark Radar Day 23: External Benchmark Catalog and Leaderboard Navigator
Published:
Day twenty-three of Benchmark Radar. We started pulling in other people’s benchmark lists, and we made the leaderboard searchable.
Benchmark Radar 第二十三天:外部基准目录与排行榜导航器
Published:
雷达吸收了外部目录。第二十三天构建了外部目录系统,并在排行榜中新增了基准测试搜索功能。
Benchmark Radar Day 22: Data Integrity and Community Channels
Published:
Day twenty-two of Benchmark Radar. We locked down data integrity and opened up the community.
Benchmark Radar 第二十二天:数据完整性与社区频道
Published:
过时运行不再能污染快照数据,微信群正式公开。第二十二天强化了数据完整性并扩展了社区渠道。
Benchmark Radar Day 21: OpenReview Authentication and Responsive Layout
Published:
Hi, Koutian here. Day twenty-one was a fix day that touched many parts at once.
Benchmark Radar 第二十一天:OpenReview 认证与响应式布局
Published:
OpenReview 完成了认证接入,宽屏布局得以扩展,手机端的显示问题也修复了。第二十一天是横跨多个模块的修复日。
Benchmark Radar Day 20: Full Chinese Language Support and Insight Blocks
Published:
Hi, Koutian here. Day twenty taught the radar to speak Chinese and to summarize itself.
Benchmark Radar 第二十天:完整中文支持与洞察块
Published:
雷达开始说中文了。第二十天交付了完整的 zh 国际化词典、洞察块和完整的社交清单。
Benchmark Radar Day 19: Geospatial Signals and Vendor Logo Saturation
Published:
Hi, Koutian here. Day nineteen pushed the radar into maps and satellites.
Benchmark Radar 第十九天:地理空间信号与厂商 Logo 饱和度
Published:
雷达扩展到了地理空间和卫星 AI 领域。第十九天为饱和度图添加了领域特定信号和厂商 Logo。
Benchmark Radar Day 18: Stable Operations
Published:
Hi, Koutian here. Day eighteen was a quiet day, and that was the point.
Benchmark Radar 第十八天:稳定运行
Published:
你好,我是 Koutian。第十八天,管道自己跑完了,没要任何人插手。这一天,恰恰证明了它稳。
Benchmark Radar Day 17: Audit Hardening and Presentation-Ready Polish
Published:
Hi, Koutian here. Day seventeen was a cleanup day. We fixed what an internal audit found and made the site look finished.
Benchmark Radar 第十七天:审计加固与发布级美化
Published:
你好,我是 Koutian。第十七天是加固与美化日。我们没加什么大功能,而是把之前埋的雷排掉,再把界面收拾干净。审计发现的问题都处理完了。
UT Austin GenAI Society: From Listening to AI to Co-Producing Public Outcomes
Published:
How can a generative-AI student society do more than host talks, repost news, and chase the latest model?
UT Austin GenAI Society:从听 AI 到共同生产公共成果
Published:
一个生成式 AI 学生社团,怎样才能不止于办讲座、转发资讯和追逐下一款模型?
Benchmark Radar Day 16: Counting Accuracy and Progressive Disclosure
Published:
Some benchmarks were counted wrong, and the dashboards were overwhelming. Day sixteen fixed both.
Benchmark Radar 第十六天:计数准确性与渐进式披露
Published:
你好,我是 Koutian。第十六天,我们修了一个会悄悄骗人的数,还让界面不再一打开就吓到人。
Benchmark Radar Day 15: Social Media Pipeline and WeChat Integration
Published:
The radar started posting to social media. Day fifteen built that pipeline and retired the daily GitHub Issue.
Benchmark Radar 第十五天:社交媒体流水线与微信集成
Published:
你好,我是 Koutian。第十五天,雷达第一次学会主动对外说话,不再只是等你来翻。
Benchmark Radar Day 14: Feed Coverage, Briefing Reliability, and Production Q&A
Published:
The daily pipeline went from fragile to reliable. Day fourteen added more feeds, hardened the briefing, and turned on Q&A in production.
Benchmark Radar 第十四天:订阅源覆盖、简报可靠性与生产环境 Q&A
Published:
你好,我是 Koutian。第十四天,我们把雷达的几个老毛病修好了,也让它更像一个会自己干活的工具。
Benchmark Radar Day 13: Dashboard Polish and Evidence Grounding
Published:
The radar now backs up every claim with a link. Day thirteen added evidence grounding, a daily Q&A on the dashboard, and a stack of small fixes.
Benchmark Radar 第十三天:仪表板美化与证据溯源
Published:
第十三天的重点是让仪表盘更好看,并且把每条结论都说得有凭有据。
Luck Sourcing: The Sourcing of Getting Lucky
Published:
A long time ago, I came across a book called Chase, Chance, and Creativity: The Lucky Art of Novelty(《追逐、机遇和创造力:新奇的幸运艺术》). I no longer remember everything in it, but its central idea stayed somewhere in the back of my mind: luck and creativity are not entirely random. Chance may arrive unexpectedly, but we can still choose how often we encounter it and whether we are ready to recognize it.
Luck Sourcing:主动寻找好运
Published:
很久以前,我偶然看到一本书,叫 Chase, Chance, and Creativity: The Lucky Art of Novelty(《追逐、机遇和创造力:新奇的幸运艺术》)。我已经不记得书里的全部内容了,但它的核心想法一直留在我脑海中的某个角落:运气和创造力并不完全是随机的。机遇也许会意外到来,但我们仍然可以选择自己遇见它的频率,以及当它出现时,我们是否已经准备好认出它。
Benchmark Radar Day 12: Daily RSS Feed from Snapshot History
Published:
The radar now has an RSS feed. Day twelve was a small day, one feature, but it changes how you follow the project.
Benchmark Radar 第十二天:每日 RSS 订阅源
Published:
雷达现在能订阅了。第十二天,我们加了一个每天更新的 RSS 订阅源。
Benchmark Radar Day 11: KW-Bench Capability Layer and Community Launch
Published:
KW-Bench got a capability rubric. The community got launch posts. Hi, Koutian here. Day eleven ran on two tracks at once.
Benchmark Radar 第十一天:KW-Bench 能力层与社区发布
Published:
第十一天,KW-Bench 的能力分级上线了,社区发布也开始了。
Benchmark Radar Day 10: AI Briefing, GPT Insight, and Launch Prep
Published:
The radar started talking. Hi, Koutian here. Day ten added AI-written briefings, GPT-powered insights, and launch copy for Chinese social platforms.
Benchmark Radar 第十天:AI 简报、GPT 洞察与发布准备
Published:
雷达开始会自己说话了。第十天我们加了 AI 生成的每日简报、GPT 驱动的洞察,还有面向中文社交平台的发布文案。
The GitHub Apps You Installed Months Ago Still Have Full Access. Go Check.
Published:
I was in the middle of setting up a second Cloudflare backup, verifying one dataset and configuring another, when the AI agent helping me flagged something unrelated: while listing what had access to my GitHub account, one authorization stood out as unusually broad. That single flag turned into an hour of reviewing GitHub’s installed-apps list, and it was worth every minute.
你几个月前装的那个 GitHub App,权限可能还全开着,去看看
Published:
我当时正在配置第二份 Cloudflare 备份——核实一份数据、再搭建另一份——帮我干活的 AI Agent 顺手提醒了我一件不相关的事:在列出哪些东西能访问我的 GitHub 账号时,有一项授权的范围明显大得不正常。就是这一句提醒,让我花了一个多小时去仔细过了一遍 GitHub 的已安装应用列表,而这一个小时非常值得。
Benchmark Radar Day 9: Adoption Frontier and Score Progression Layer
Published:
The leaderboard became a research workbench. Hi, Koutian here. Day nine added the adoption frontier, score progression, and saturation charts.
Benchmark Radar 第九天:采纳前沿与分数进展层
Published:
排行榜以前只是一张排名表。第九天,我们把它改成了能回答问题的研究工作台,加了采纳前沿、分数进展和饱和度可视化。
Benchmark Radar Day 8: Stabilization and Minimum Fixes
Published:
A quiet day after the registry explosion. Hi, Koutian here. Day eight settled the pipeline with two precise fixes.
Benchmark Radar 第八天:稳定化与最小修复
Published:
注册表大爆炸之后,安静的一天。第八天,我们用两次精准修复稳住了流水线。
A Parking Space for $126,887: What Exactly Makes UT Austin’s Mulva Hall So Expensive?
Published:
One hundred twenty-six thousand eight hundred and eighty-seven dollars is housing-level money in Austin; in UT Austin’s new business-school building, it buys a single parking spot.
一个车位 12.7 万美元:UT Austin 的 Mulva Hall 到底贵在哪里?
Published:
十二万六千八百八十七美元,在 Austin 已经是住房级别的钱;在 UT Austin 的新商学院大楼里,它只对应一个停车位。
Benchmark Radar Day 7: Model Card Adoption Leaderboard and Registry v0.3.0
Published:
Day seven was the biggest single day. Hi, Koutian here. We launched a new leaderboard, grew the registry to the 2026 frontier, and bumped the version to 0.3.0.
Benchmark Radar 第七天:模型卡采纳排行榜与注册表 v0.3.0
Published:
最大的一天。模型卡采纳排行榜上线,注册表扩到了 2026 年前沿,版本号升到 0.3.0。
Benchmark Radar Day 6: Hacker News Integration and Scheduled Reliability
Published:
The radar now watches Hacker News. Day six added community attention signals and hardened the daily schedule.
Benchmark Radar 第六天:Hacker News 集成与定时可靠性
Published:
雷达开始听 Hacker News 的动静了。第六天,我们加了社区注意力信号,还加固了每天定时运行的可靠性。
Benchmark Radar Day 5: Accessible Trend Charts and Landscape Analysis
Published:
The trend chart got a real hover card. The landscape report got published. And the agentic benchmark count jumped from 3 to 78.
Benchmark Radar 第五天:无障碍趋势图与全景分析
Published:
趋势图现在有了一个真正的悬浮卡片。全景分析报告发布了。agentic 基准的数量从 3 跳到了 78。
《入都》(Entering the Capital): Ten Poems Written Before the Exam Journey by a Twenty-One-Year-Old
Published:
丈夫只手把吴钩,意气高于百尺楼。一万年来谁著史,三千里外欲封侯。
《入都》:一个二十一岁的人,进京赶考前写下的十首诗
Published:
丈夫只手把吴钩,意气高于百尺楼。一万年来谁著史,三千里外欲封侯。
Benchmark Radar Day 4: Historical Memory and Freshness Detection
Published:
The radar learned to remember. Day four added historical backfill, staleness detection, and an agentic taxonomy category.
Benchmark Radar 第四天:历史记忆与新鲜度检测
Published:
雷达学会了记忆。第四天,我们加了历史数据回填、过期检测,还有 agentic 这个分类。
Benchmark Radar Day 3: Linkable Rubric Dialogs and Filter Fixes
Published:
A small day with two precise fixes that made the radar shareable.
Benchmark Radar 第三天:可链接的评分对话框与过滤器修复
Published:
内容不多,但两处精准的修复让雷达变得能分享了。
Benchmark Radar: An Evidence-First Daily Radar for AI Benchmarks
Published:
AI benchmarks now show up faster than any one researcher can keep up with. So I built a radar that checks for them every day, and shows its work.
Benchmark Radar:面向 AI 基准的、以证据为先的每日雷达
Published:
新的 AI 基准冒出来的速度,比任何研究员评估它们的速度都快。所以我建了一个雷达,让「发现新基准」变成一件每天做、过程透明、可以复核的事。先说一个词:基准(benchmark)就是用来考 AI 的一套题,或者一份评测。
Benchmark Radar Day 2: Cumulative Trends and Artifact Deduplication
Published:
The radar gained memory. Day two added trend maps that span time, plus a fix for benchmarks showing up under many names.
Benchmark Radar 第二天:累积趋势与工件去重
Published:
雷达开始有记忆了。第二天我们建起了累积趋势图,还顺手解决了工件别名的问题。先说一个词:工件(artifact)就是被追踪的每一个具体东西,比如某个基准或数据集。
Benchmark Radar Day 1: Building the Cumulative Dashboard MVP
Published:
Day one of Benchmark Radar. We went from nothing to a working daily dashboard in a single day.
Benchmark Radar 第一天:构建累积仪表板 MVP
Published:
Benchmark Radar 第一天。我们从零开始,一天之内搭出了一个能跑的累积仪表板。
The Skills I Built for Weaker Models, and Why I Deleted Most of Them
Published:
I ran a health check on my Claude Code setup this week and found 174 custom skills, 124 of which I had never invoked once. They were not failures. Most of them were scaffolding I built for a model that needed it, and then kept long after the model stopped needing it.
我为较弱的模型做了这些 Skills,后来为什么删掉了大部分
Published:
这周,我给自己的 Claude Code 配置做了一次健康检查,发现里面有 174 个自定义 Skill,其中 124 个我一次都没有调用过。它们并不是失败品。大部分只是我为当时需要额外支撑的模型搭起的脚手架,而模型早已不再需要它们,我却一直把它们留到了现在。
Why Did They All Move to AI Labs? What the Sharpest Minds See, and It Isn’t Just Money
Published:
The sharpest minds keep moving into AI labs, and it isn’t just for the money.
他们为什么都去了 AI Lab?顶尖头脑看见的,不只是钱
Published:
这是《颠覆性创新者的思维模式》的五个月后续研究。
Looking Back from 2030: Connecting Stronger Intelligence to Problems Worth Solving
Published:
If I reopened the 2026 folder in 2030, I would check one thing first: did the work completed with AI change a real problem and leave evidence that others could verify and build upon?
2030 年回望 2026:把更强的智能接到值得解决的问题上
Published:
假设在 2030 年重新打开 2026 年的文件夹,我会先查一件事:今天借助 AI 完成的工作,是否改变了一个现实问题,并留下别人能够复验和继承的证据。
LLM Application Opportunities: Not a Smarter Chat Box
Published:
The most valuable application of LLMs may not be answering questions, but turning an expert job into a digital work unit that can be executed, reviewed, and audited.
LLM 应用机会,不是更聪明的聊天框
Published:
LLM 最有价值的应用,可能不是回答问题,而是把一项专家工作变成可以执行、复核和审计的数字工作单元。
The World After LLMs: Who Will Be Reorganized by 2030?
Published:
The most worth-watching thing in 2030 is not how many people are replaced by AI, but which jobs will be reorganized.
LLM 之后的世界,2030 年谁会被重新组织?
Published:
2030 年最值得观察的,不是有多少人被 AI 取代,而是哪些工作会被重新组织。
Revisiting Liang Wenfeng’s Four-Hour Transcript From a 2030 Vantage Point
Published:
| Original source | Tencent Tech; transcript edited | Gu Lingyu; original editors | Xu Qingyang, Su Yang; 2030 annotations and second edit | Koutian Wu |
站在2030年回看梁文锋四小时实录
Published:
原文来源|腾讯科技;原文整理|顾翎羽;原文编辑|徐青阳、苏扬;2030 注释与二次编辑|Koutian Wu
The Mom Test: From Reading It at SFO to Using It in Every Customer Interview
Published:
Some books leave behind a few ideas. The Mom Test gave me a method I continue to use: do not ask whether someone likes your idea. Ask what they have already done, how often the problem occurs, what it has cost them, and whether they will commit to a next step.
《The Mom Test》:我从 SFO 机场读完,到后来一直使用的用户访谈方法
Published:
有些书读完会留下一些观点,《The Mom Test》留给我的却是一套不断调用的方法:不要问别人喜不喜欢你的想法,要追问他过去真实做过什么、问题多久发生一次、已经付出了什么成本,以及接下来愿不愿意采取行动。
Steve Jobs Was Insulted in Public. His Answer Revealed Where Product Decisions Should Begin
Published:
When someone publicly told Steve Jobs that he did not understand technology, Jobs paused, conceded part of the criticism, and asked a more consequential question: what will the customer ultimately gain?
被当众羞辱后,Steve Jobs 没有反击,而是解释了产品决策的起点
Published:
一个人当众说 Steve Jobs 根本不懂技术,他停了十几秒,没有反击,只问了一件更重要的事,用户最终能得到什么?
Does “No Stupid Questions” Mean “Ask Anything” or “Do Not Ask Dumb Questions”?
Published:
A widely shared post describes a brainstorming meeting at a robotics startup in Boston. A slide said “No stupid questions.” A Chinese employee interpreted it as “Do not ask stupid questions,” while American colleagues explained it as “Ask freely; no question will be judged stupid.” A related post claims that “Say it again?” is neutral when someone is not heard, whereas “What did you say?” is hostile.
“No stupid questions”到底是“随便问”还是“别问蠢问题”?
Published:
一张流传截图讲了这样一件事:一家波士顿机器人创业公司开头脑风暴会,幻灯片写着 “No stupid questions.” 一位中国员工把它理解为“不要问蠢问题”,美国同事却说它的意思是“什么都可以问,没有问题会被当成蠢问题”。另一组讨论又声称:没听清时说 “Say it again?” 很中性,而 “What did you say?” 很不友好。
How I “Reasonably” Ended Up in the Wrong NEXUS Line at YVR
Published:
I ended up in the wrong NEXUS line at Vancouver International Airport. By the rules, the mistake was mine. By design, it was an almost predictable wrong turn.
我在 YVR“合理地”走错了 NEXUS
Published:
我走错了温哥华机场的 NEXUS 通道;规则上是我的错,设计上却是一条几乎可以预测的错误路径。
Another Side of Vancouver Airport: Age, Workspace, and the People I Saw
Published:
Beyond the wayfinding problem that led me into the wrong NEXUS line, my connection at Vancouver International Airport left me with several observations unrelated to signs. I noticed an older-looking mix of people, a departures area that offered almost nowhere to work, and a visible contrast between people in Vancouver and Texas. This post records those observations. The full account of the NEXUS incident is in How I “Reasonably” Ended Up in the Wrong NEXUS Line at YVR.
温哥华机场的另一面:年龄结构、办公空间与我看到的”人群”
Published:
除了那次走错 NEXUS 通道的导视问题,这次在温哥华机场(YVR)转机还让我观察到几件和路标无关、但同样值得记录的事情:明显偏大的现场年龄结构、几乎不为办公设计的候机空间,以及温哥华与得州人群景观的直观差异。这篇笔记单独记录这些观察;如果你想看我是怎么在 NEXUS 通道走错路的,可以读我在 YVR”合理地”走错了 NEXUS。
Air Canada Was a Mess: Three Seats, Fourteen Hours of Lavatory Odor, a Go-Around, and a Towing Delay
Published:
My verdict on Air Canada after flying from Singapore to Austin through Vancouver is simple: what a mess.
Air Canada 太拉了:一排三座、十四小时厕所味、复飞与拖机延误
Published:
这次从新加坡经温哥华飞回奥斯汀,我对 Air Canada 的评价很简单:太拉了。
Why Did a S$5.50 Purchase Appear as US$4.27?
Published:
A receipt in Singapore showed S$5.50, while a U.S. credit-card account displayed only US$4.27. Another purchase of about S$48 appeared as roughly US$38. A bank-card transit ride also failed to appear immediately as pending.
为什么 S$5.50 的消费在美国信用卡上只显示 US$4.27?
Published:
在新加坡消费时,一张收据写着 S$5.50,但美国信用卡账户只显示 US$4.27。另一笔约 S$48 的消费,则显示为约 US$38。与此同时,刷银行卡乘坐公共交通后,交易也没有立刻出现在 pending 中。
I Saw “Asia’s Safest Bank” at a Singapore Airport: A Banking Safety Lesson for a 10-Year-Old
Published:
At Singapore Changi Airport I saw an ad reading “Asia’s Safest Bank,” and my first question wasn’t how impressive it is, it was, who says so?
我在新加坡机场看到“亚洲最安全银行”:给10岁小朋友的银行安全课
Published:
我在新加坡樟宜机场看到一块写着「亚洲最安全银行」的广告,脑子里第一个问题不是它有多厉害,而是,谁说的?
AI 不一定是人工智能:两个字母如何被不同世界反复占用
Published:
今天看航班时,我突然发现:AI 是 Air India(印度航空)的航空公司代码。
AI 不一定是人工智能:两个字母如何被不同世界反复占用
Published:
今天看航班时,我突然发现:AI 是 Air India(印度航空)的航空公司代码。
AI Can Plan the Trip. It Still Cannot Get the Car to Move.
Published:
The last mile of an AI-generated travel plan may be one missing Retry button.
AI 能规划旅程,却还不能让车真正开起来
Published:
一份由 AI 制定的完美旅行计划,最后可能会死在一个不存在的“重试”按钮上。
AI Can Plan My International Trip, but It Still Cannot Clear the Payment Gate
Published:
The thing that almost stranded me on this international trip was not a typhoon, but several unrelated systems deciding at the same time that they did not trust me.
AI 能帮我规划国际旅行,但它还不能替我过支付这一关
Published:
这次国际旅行真正差点把我留在原地的,不是台风,而是几个互不认识的系统同时决定不相信我。
Is Contributing on GitHub Like Donating Sperm?
Published:
Here is an inappropriate analogy that nevertheless captures something real about open-source work:
在 GitHub 上做贡献,像不像捐精?
Published:
这里有一个不太登大雅之堂、却抓住了开源工作某些真相的类比:
Why Filing Issues and PRs All Over GitHub Is Like “Donating Sperm” and Not Like “Donating Eggs”
Published:
Let me offer a comparison that’s a bit inappropriate but does capture a real structure of open-source collaboration:
在 GitHub 到处提 Issue 和 PR,为什么像“捐精”而不像“捐卵”
Published:
我提一个不太恰当、但确实抓住了开源协作某种结构的比喻:
Why alias rm=trash Cannot Stop an AI Agent from rm -rf
Published:
I asked Claude Code a narrow question: can I protect this machine from Codex accidentally deleting files forever, just by aliasing rm to trash? The honest answer turned out to be no, for a reason that is not obvious until you actually test it, and the fix ended up being a five-layer setup rather than a one-liner.
为什么 alias rm=trash 拦不住 AI agent 的 rm -rf
Published:
我问了 Claude Code 一个很具体的问题:能不能只靠把 rm alias 成 trash,来防止 Codex 意外把文件永久删掉。真实的答案是不能,原因不实测根本看不出来,最后落地的也不是一行配置,而是五层防护。
The Value of Benchmarks, Market Mispricing, and the Ability to Frame Questions: A Raw Conversation with My Good Bro from Northeast China at CMU
Published:
Does building a benchmark count as research, or is it “crowdfunding a paper”? Does it create public value or consume public resources? Can it build lasting capabilities, or is it useful only for a startup team trying to show investors its potential?
Benchmark 的价值、市场误判与提出问题的能力:一次和我滴好东北哥们儿(CMU)的原始对话
Published:
Benchmark 到底是在做研究,还是在“众筹一篇论文”?它是在创造公共价值,还是在占用公共资源?它能不能形成长期能力,还是只适合创业团队向投资人展示潜力?
Hard Benchmarks Should Not Become Coding Tricks
Published:
The most dangerous failure mode of an AI science benchmark is not that it is too hard, it is that it quietly becomes either a coding trick or a guessing game.
高难度 Benchmark 不该变成编程技巧题
Published:
AI 科学 Benchmark 最危险的失败方式,不是它太难,而是它悄悄变成了一道编程技巧题或猜谜题。
Do Not Secretly Record Customer Interviews
Published:
Secretly recording a customer interview can turn a useful research habit into a privacy problem in one click.
不要偷偷录客户访谈
Published:
偷偷录一次客户访谈,可能把一个正常的产品研究流程变成隐私和合规事故。
From News Signals to GitHub Repos: turning trends into open-source influence and business opportunities
Published:
The opportunity is not the news itself. The opportunity is the friction created by the news.
从新闻信号到 GitHub Repo:如何把趋势转化为开源影响力和商业机会
Published:
新闻本身不是机会。新闻制造的摩擦才是机会。
Compute Is the New Generation of Free Eggs: When Big Tech Starts “Handing Out Eggs”
Published:
Compute / tokens are the new generation of free eggs. Right now all the big vendors are handing out compute like they’re handing out eggs. When you’re old, you’ll be fighting other old grannies and grandpas for free eggs, this is no longer just a joke, it might be reality.
算力是新一代的鸡蛋:当大厂开始”发鸡蛋”
Published:
算力 / token 是新一代的鸡蛋。现在各大厂商都在发算力,就像在发鸡蛋一样。你老了就要去跟别的老奶奶、老大爷抢发鸡蛋——这已经不是一个笑话,可能就是一个现实。
A Closeness-Influence Map: A Method for Structuring Your Network
Published:
Goal: Use a 4x3 grid to place every relevant person “by position, with a matching strategy,” then review it regularly to maximize resource use and manage risk.
亲疏‑影响双维度人脉地图:结构化管理关系网络的方法论
Published:
目标:用一张 4×3 网格,把所有相关人脉”按位置、定策略”,持续复盘,最大化资源利用与风险管控。
Why Your Brand-New WD Drive Is Read-Only on a Mac (It’s Not the Drive)
Published:
I plugged a Western Digital external drive full of data into a MacBook Air, tried to copy a file onto it, and nothing happened. No error dialog, no progress bar, just a drive that would let me read everything and write nothing. My first instinct was that something was broken, or that I needed to fix permissions. Both were wrong, and chasing the wrong explanation almost led me to permanently downgrade the security of the whole laptop.
为什么你崭新的 WD 移动硬盘在 Mac 上只能读不能写(问题不在硬盘)
Published:
我把一块装满数据的西部数据(WD)移动硬盘插到 MacBook Air 上,想往里拷一个文件,结果什么都没发生。没有报错弹窗,没有进度条,就是一块能读出所有东西、却一个字节都写不进去的硬盘。我的第一反应是它坏了,或者是我得去修一下权限。这两个判断都是错的,而且顺着错误的解释找下去,差点让我把整台笔记本的安全性永久降级。
Building a real ZIP bomb in Fortran, C++, and C
Published:
I’ve been playing with mixed-language builds (Fortran calling into C++ and C via iso_c_binding) and wanted a demo that was more interesting than “add two numbers across languages.” So I built fortran-zip-bomb: a small program that generates a genuine ZIP bomb — a small archive that expands into a much larger file on decompression.
用 Fortran、C++ 和 C 构建一个真正的 ZIP 炸弹
Published:
我一直在玩混合语言构建(Fortran 通过 iso_c_binding 调用 C++ 和 C),想要一个比”跨语言把两个数加起来”更有意思的演示。于是我建了 fortran-zip-bomb:一个小程序,生成一个真正的 ZIP 炸弹——一个解压时会扩展成大得多的文件的小压缩包。
A Materials-Science Model of Egg Fried Rice, Three Years Later
Published:
In October 2023 I asked ChatGPT an over-engineered question: how do you fry rice so that every grain of rice ends up bonded to egg, with no bare rice grains and no isolated clumps of egg sitting off on their own? I saved that conversation as an HTML file, dropped it in my Downloads folder, and did not look at it again for almost three years. Last week I finally turned it into an actual open-source repo, and going back through the original conversation to build it was more interesting than I expected.
蛋炒饭材料学模型,三年之后
Published:
2023 年 10 月,我问了 ChatGPT 一个过度工程化的问题:怎么炒蛋炒饭,才能让每一粒米都裹上蛋,既没有裸露的白米粒,也没有单独抱团、没沾到米的蛋碎?那次对话我存成了一个 HTML 文件,扔进 Downloads 文件夹,然后差不多三年没再打开过。上周我终于把它整理成了一个真正的开源仓库,回头重读那段 2023 年的对话、把它做成代码的过程,比我预想的有意思得多。
Nobody Gets Rich From a Blue Link: A Ladder to the Table
Published:
I just remembered something sudden on the train: why are almost all the links on the internet that extremely saturated shade of blue?
没有人靠蓝色链接变富:一架通往牌桌的梯子
Published:
我刚刚在火车上突然想起一件事:为什么互联网上的链接,几乎清一色都是那种饱和度极高的蓝色?
What It Actually Takes to Land a Patch in Git
Published:
In my first post on this I described sending a patch through GitGitGadget and watching it land on the mailing list as v1. That post ended with the patch waiting for review. It has since merged. Now that I have seen the full lifecycle, from typo to master, here is what actually mattered.
把一个补丁真正送进 Git 需要什么
Published:
在第一篇文章里,我写了怎么通过 GitGitGadget 把补丁发出去,它作为 v1 落在了邮件列表上。那篇文章结束在等待 review 的状态。现在它已经合并了。看完从一个拼写错误到进入 master 的完整过程,下面是我觉得真正重要的几点。
I Am Now an Official Git Contributor
Published:
The patch I sent to the GitGitGadget doorstep ten days ago has been merged into master.
我现在是正式的 Git Contributor 了
Published:
十天前我送到 GitGitGadget 门口的那个补丁,现在已经合并进 master 了。
Japan Travel Advice From a Friend: Plan the Trip First
Published:
The clearest advice I recently received from a friend who has visited Japan several times was simple: plan the trip first.
朋友给我的日本旅行建议:先把行程排明白
Published:
我最近从一位去过日本多次的朋友那里得到一句很直接的建议:去日本之前,先把行程排明白。
A Friend I Met on a United SFO–PVG Flight
Published:
On a United Airlines flight from San Francisco SFO to Shanghai Pudong PVG, the plane had started its descent and the cabin announcements kept reminding everyone to fasten their seatbelts. Outside was night; inside, people were already lit up with the excitement of “finally going home.” It was on this flight that I met a friend.
在 UA SFO–PVG 航班上遇到的一位朋友
Published:
在联合航空从旧金山 SFO 飞往上海浦东 PVG 的航班上,飞机已经开始下降,机舱广播一遍遍提醒大家系好安全带。窗外是夜色,舱内的人却已经被“终于回国了”的兴奋点亮。我就在这趟航班上认识了一位朋友。
Why I Must Throw Myself Into the AI Wave
Published:
Recently I’ve sometimes felt confused by how far AI has come, a bit lost and anxious, unsure what I should do, and then, because of my identity and my path, wondering what fallbacks or better options exist.
为什么我一定要投身 AI 浪潮
Published:
最近虽然有的时候也会因为 AI 的发展程度感到非常困惑,感到有一点迷茫、有些焦虑,不知道自己该干啥,然后又因为自己的身份和路径问题,在想有什么样的退路或者更好的方案。
Port Aransas South Jetty: How Satellites and AI Protect the Texas Coast
Published:
Port Aransas South Jetty published an Institute Insights feature on May 28, 2026 about how our University of Texas team is using satellite remote sensing and AI to monitor Texas coastal water quality.
阿兰瑟斯港南防波堤:卫星与 AI 如何守护德克萨斯海岸
Published:
阿兰瑟斯港南防波堤 于 2026 年 5 月 28 日发表了一篇 Institute Insights 专题,讲的是我们德克萨斯大学团队如何用卫星遥感与 AI 监测德克萨斯沿海水质。
Sending My First Git Patch to the GitGitGadget Doorstep
Published:
Today I got my first Git patch onto the mailing list.
我把第一个 Git 补丁送到了 GitGitGadget 门口
Published:
今天我差一点,就把自己的第一个 Git 项目补丁送进邮件列表了。
Why I Want AI Actions to Manage My Information Flow
Published:
What I actually want to do today is pull myself out from under a mountain of repetitive work.
我为什么要把自己的信息流交给 AI Actions
Published:
我今天真正想做的事情,其实是把自己从一座 mountains of 重复性的 work 里面捞出来。
Hunting the Perfect Ballistic Backpack Insert: A Discussion From Hardcore Gear to Coming Back to Earth
Published:
When a random idea hits you on a US campus to slip a ballistic plate into the North Face backpack you carry every day, things start racing toward cyberpunk.
寻找完美防弹书包插板:一次从硬核装备到回归现实的探讨
Published:
当你在美国校园里突发奇想,想给每天背的The North Face书包塞一块防弹板时,事情就开始朝着赛博朋克的方向狂奔了。
Further Thoughts on Social Contradictions
Published:
While you’re still grinding your heart out on “if you just work hard enough, you can move up,” you may not realize the system was never designed to cultivate you in the first place, it was designed to screen you.
社会矛盾的进一步思考
Published:
当你还在为“只要努力就能向上流动”而拼命内卷时,你可能没有意识到,这个系统从一开始就不是为了培养你,而是为了筛选你。
Test Your Heart Age
Published:
When an ordinary-looking “heart age” questionnaire lets you calculate that your heart is younger than your actual age, you may need to re-examine the lifestyle habits you’ve been ignoring.
测一测你的“心脏年龄”
Published:
当一份看似普通的“心脏年龄”问卷让你算出自己的心脏比实际年龄还年轻时,你可能需要重新审视一下那些被你忽视的生活习惯了。
Which One Are You? A Field Guide to Performances in the AI Era
Published:
When you go looking for a seat in the great theater of the AI era, built from PowerPoints, papers, and anxiety, it helps to look at which play each person on stage is performing.
你是哪一种?AI 时代的表演图鉴
Published:
当你在这个由PPT、Paper和焦虑构成的AI时代大剧院里找座位时,不如先看看台上的人都在演哪一出戏。
Why I Was (Partially) Wrong About AI Talent Inflation
Published:
After two years of “All-in AI”, my deep-seated belief in the “10,000-hour talent moat” was ruthlessly shattered by the rise of Agentic workflows.
我为何(部分地)改变了对“AI人才通胀”的看法
Published:
在“All-in AI”两年后,我曾经深信的“10,000小时人才护城河”理论,在Agentic工作流的冲击下被无情地打碎了一半。
From a Poetry Society to Unicorns: The Less-Traveled Road Isn’t Laziness
Published:
Hah, I can’t help but laugh. It just hit me: the first time I ever used Markdown was back when I was building a poetry society. And that’s also when I first learned about Git. Looking back now, if you put all the founding members of that poetry society together, you’d almost have two unicorns.
从诗社到独角兽:少走的路不是偷懒
Published:
哎,我他妈笑了。我忽然想起来,我最早用 Markdown,就是之前创建诗社的时候。知道 Git,也是在那个时候。现在回头看,整个诗社的元老凑在一起,真的快有两个独角兽了。
Calling Overseas AI APIs from China: 5 Fatal Compliance Traps (Including Criminal Risk) + How to Do It Legally
Published:
Many Chinese teams building embodied AI or AI applications hit the same real-world problem: the domestic models aren’t enough, and they want to call the APIs of foreign models like OpenAI and Anthropic directly. That’s why a swarm of “relay,” “recharge,” and “one OpenAI-compatible interface for you” businesses has sprung up.
中国公司调海外 AI API,5 个致命合规雷区(含刑事风险)+ 怎么合法做
Published:
很多做具身智能、做 AI 应用的中国团队,都会遇到同一个现实问题:国内的模型不够用,想直接调 OpenAI、Anthropic 这些海外模型的 API。于是市面上冒出一堆”中转”“代充”“一个 OpenAI 兼容接口给你用”的生意。
earth-space-ai.org: progressive-disclosure skill packages for Earth and space system models
Published:
Decades of Earth-system modeling judgment live in PDFs, mailing lists, and senior researchers’ heads, and none of it is loadable by an AI coding agent. earth-space-ai.org is an attempt to fix that.
earth-space-ai.org:面向地球与空间系统模型的渐进式披露技能包
Published:
数十年来积累的地球系统建模判断散落在 PDF、邮件列表和资深研究者的头脑里,AI 编程智能体无法直接加载这些知识。earth-space-ai.org 试图解决这个问题。
技术创业成功者的发迹代价:痛苦、路径、情感与金钱的横纵分析
Published:
研究时间:2026-05-26
技术创业成功者的发迹代价:痛苦、路径、情感与金钱的横纵分析
Published:
研究时间:2026-05-26
New Paper Published: On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust
Published:
I am deeply honored to join my colleagues in contributing to the newly published paper, “On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust,” specifically focusing on the section regarding GeoAI Trust (Section 8).
新论文发表:On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust
Published:
我非常荣幸能够参与到这篇新发表的论文 《On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust》 的合作中,并作为共同作者之一,贡献了关于 GeoAI信任(GeoAI Trust) 的章节(第八章)。
The Future World: AI Assistants, Collaboration, and the End of UI Friction
Published:
In the future world, everyone will have their own AI agent assistant.
未来的世界:AI助手、大协作与前端摩擦的终结
Published:
在未来的世界,每个人都会有自己的 AI agent 助理。
In the AI Era, What Hardware Will the Gateway to the Human Mind Be Built Into?
Published:
Seizing the gateway to the human mind in the AI era is, at bottom, a secret war over token throughput power.
AI 时代的人类心智入口,到底会被装在什么硬件里
Published:
抢占 AI 时代的人类心智入口,说到底是一场关于 token 吞吐权力的隐秘战争。
People Who Rolled Sixes Six Times in a Row Are Sharing Their Dice-Rolling Strategies
Published:
I’ve been watching too many people who got lucky rolling a six six times in a row, standing on stage lecturing us about their dice-rolling strategies.
连续掷出六次六点的人,在台上分享掷色子经验
Published:
这两天看了太多因为连续六次掷出色子六点而功成名就的人,在台上滔滔不绝地分享他们的掷色子心得。
Stock Trading 101, A No-Nonsense Guide to Picking Your First Brokerage
Published:
A junior colleague asked me recently if there is a magical brokerage that does it all, from stock trading to high-yield savings, which reminded me of my own naive expectations when I first started investing.
炒股101,新手券商评测与账户搭建指南
Published:
这两天有一位好兄弟跑来问我,有没有哪家券商能把炒股、理财和银行存款全包了。
What Microsoft Flight Simulator Really Shows About Earth-Scale Digital Twins
Published:
I spent the last couple of days checking a tempting claim: that Microsoft Flight Simulator is secretly the foundation for a broader Earth digital twin story.
微软模拟飞行到底说明了什么:关于地球级数字孪生的事实核查
Published:
这两天我在核查一个很容易写得很爽、但也很容易写过头的说法:微软模拟飞行(Microsoft Flight Simulator)是不是某种更大地球数字孪生叙事背后的秘密技术底座。
Anthropic’s Most Ruthless Methodology Is Shipping the Worst Version First
Published:
Anthropic’s most ruthless move isn’t polishing every product to perfection before release, but daring to throw out a research preview first, then rapidly getting stronger in front of everyone.
Anthropic 最狠的方法论,是先把最差版本发出去
Published:
Anthropic 最狠的地方,不是每次都把产品打磨到完美才发布,而是敢先把一个 research preview 扔出来,然后当着所有人的面快速变强。
How Rich Is Chen Ning Yang? His wealth and the science of scientific privilege
Published:
A 1957 Nobel prize cheque, a father who was the first Chinese PhD in number theory, three different monetary regimes, and seventy years of compounding, so how do you put a number on Chen Ning Yang’s family wealth in 2026 without making it up?
杨振宁先生到底有多少财富?
Published:
一张 1957 年的诺贝尔奖支票,一个中国第一位数论博士的父亲,三套不同的货币制度,加上七十年的复利,我们到底要怎么给杨振宁家族 2026 年的财富估出一个不是瞎编的数字?
The Missing Scientists, Why 600 Years of Status Quo Is the Real Bottleneck in Science
Published:
A new paper looked at 739 science Nobel laureates from 1901 to 2023, and the headline number that came out is brutal, at the current rate of progress it would take roughly 600 years before a kid born into a poor country has the same shot at a Nobel as a kid born into a rich one.
消失的科学家,为什么 600 年才追得上才是科学真正的瓶颈
Published:
一篇新论文把 1901 到 2023 年所有 739 位科学类诺贝尔奖得主从童年扒到现在,得出了一个特别狠的数字,按目前的进步速度,一个出生在低收入国家的孩子要等大约 600 年,才能拿到和富裕国家孩子一样的诺奖机会。
把人扫地出门,再把博士学位塞回去:荣誉博士制度的横纵分析
Published:
大学最擅长的表演,莫过于在一个人不再需要勋章时,隆重地为他披上一件绣着「博士」二字的学位袍。
把人扫地出门,再把博士学位塞回去:荣誉博士制度的横纵分析
Published:
大学最擅长的表演,莫过于在一个人不再需要勋章时,隆重地为他披上一件绣着「博士」二字的学位袍。
Why Anthropic Won’t Sponsor Your PERM, and What You Can Do About It
Published:
You apply to Anthropic, land an offer, then ask HR whether they’ll sponsor your PERM, and they say “on hold for now” or change the subject, and you probably stand there frozen.
为什么 Anthropic 不给你办 PERM,以及你现在能做什么
Published:
你投了 Anthropic,拿到了 offer,然后去问 HR 能不能帮你办 PERM,对方说「目前暂停」或者直接转移话题,你当时大概就愣在原地了。
Prompt Caching Is Not Technical Debt
Published:
Prompt Caching Is Not Technical Debt. It’s what the math demands.
Prompt Caching 不是技术债
Published:
很多人把 prompt caching 看成一个省钱 hack,但我觉得这个判断刚好反了。
Talking Claude Code at a Silicon Valley Alumni Gathering: What Matters Is How It Changes Work
Published:
At today’s Silicon Valley alumni “Ask a Question” AI meet-up, what I wanted to talk about wasn’t how cool Claude Code is, but how it actually re-slices real work.
在硅谷校友会聊 Claude Code,真正重要的是它怎么改变工作
Published:
今天在硅谷校友会的「师问」AI 交流活动上,我想讲的不是 Claude Code 多酷,而是它到底怎么把真实工作重新拆开。
My PhD Advisor Built a Land Surface Model That Forecasted Hurricane Harvey
Published:
When you join a lab, you do not just get a research direction. You inherit a 30-year codebase that is currently running inside the U.S. National Water Model and was on the critical path forecasting Hurricane Harvey. That is what working with Zong-Liang Yang at the Jackson School of Geosciences actually looks like.
我的博士导师打造的陆面模型,曾用于预报飓风 Harvey
Published:
加入实验室以后,你会接过一个研究方向,也会继承一套已有 30 年历史的代码库。它目前运行在美国国家水模型中,也曾处在飓风 Harvey 预报工作的关键路径上。这就是在 Jackson School of Geosciences 与 Zong-Liang Yang 一起工作的真实样子。
I Spent a Week Building a Chronicle for Chen Ning Yang, and the Family Tree Was the Story
Published:
Most physics undergraduates know Chen Ning Yang as a Nobel laureate and the Y in Yang-Mills. Building him a chronicle made me see something the textbooks gloss over: the most consequential single fact about his life is who his father was, and what that father had set up for him before he was born.
我用一周为杨振宁建了一座年鉴,家谱才是真正的故事
Published:
大多数物理本科生知道杨振宁是诺贝尔奖得主,知道他是 Yang-Mills 里的那个 Yang。而为他建一座年鉴,让我看到了教科书略过的东西:他一生中最有分量的一个事实是谁是他的父亲,以及那个父亲在他出生之前,为他铺好了什么。
Sean Xiang Has Been Building Bloombase for 14 Years, and the AI Era Finally Caught Up to It
Published:
Most enterprise security companies show up, ride one trend, and disappear in the next infrastructure cycle. Sean Xiang has been building Bloombase since January 2012, and the company has somehow been on the right side of every major infrastructure shift since, including the current AI accelerator era. That is not luck. That is a thesis.
Sean Xiang 已经打造 Bloombase 14 年,AI 时代终于追上了它
Published:
大多数企业安全公司冒个泡、赶一波趋势,然后在下一个基础设施周期里消失。Sean Xiang 从 2012 年 1 月起就在打造 Bloombase,这家公司却阴差阳错站到了此后每一波重大基础设施转变的正确一边,包括当下的 AI 加速器时代。那不是运气,那是一套论点。
Marc Hesse Does the Fluid Mechanics of Everything From Magma to Mars
Published:
You can study fluid mechanics in five different countries before you turn 30, work on petroleum reservoirs and tectonophysics and planetary ice on the same week, and somehow end up at a Centennial Chair in Geophysics. Marc Hesse did exactly that, and the through-line is more interesting than any of the individual stops.
Marc Hesse 研究从岩浆到火星的万物流体力学
Published:
你可以在 30 岁前在五个不同国家学流体力学,同一周里既研究油气储层又研究构造物理和行星冰,最后竟然坐上地球物理学百年讲席。Marc Hesse 就是这么做的,而他背后那条主线比任何一个单独的站点都更有意思。
Kehan Dong Works With People Who Started Building Earlier Than They Were Supposed To
Published:
Most VCs and accelerators talk about supporting founders. Kehan Dong specifically supports the kind of founder who started building at 16, when nobody told them they were allowed to, and it turns out that population is dramatically underserved.
Kehan Dong 与那些比别人更早开始建造的人合作
Published:
大多数 VC 和孵化器都谈支持创始人。Kehan Dong 专门支持那种 16 岁就开始动手建造、而根本没人告诉过他们可以这么做的创始人,事实证明这个群体被严重忽视了。
Juan Santiago Built One of the Top Microfluidics Labs in the World by Staying Put for 30 Years
Published:
In a field where everyone gets pulled toward the next hot thing, Juan Santiago joined Stanford Mechanical Engineering in 1998 and has been there ever since, building one of the most consequential microfluidics labs on the planet. Long-term focus is its own competitive advantage.
Juan Santiago 在一个地方待了 30 年,建起了世界顶级微流控实验室之一
Published:
在一个人人都被拽向下一个热门方向的领域里,Juan Santiago 1998 年加入了斯坦福机械工程系,此后一直待在那里,建起了这个星球上最有影响力的微流控实验室之一。长期专注本身就是一种竞争优势。
Gengchen Mai Is Building Spatial Foundation Models, and Geographers Should Pay Attention
Published:
The same way text foundation models ate NLP and image models ate computer vision, someone is going to build the foundation model that eats spatial reasoning. Gengchen Mai is one of the people taking that bet seriously, and his SEAI Lab at UT Austin is one of the places it is being built.
Gengchen Mai 正在构建空间基础模型,地理学者值得关注
Published:
文本基础模型已经改变了 NLP,图像模型也改变了计算机视觉。空间推理领域同样会出现自己的基础模型。Gengchen Mai 正在研究这一方向,他在 UT Austin 的 SEAI Lab 也开展相关工作。
Gemini Says My Ego Essay Is Still a Humblebrag
Published:
Earlier today I published a piece called “Four Ego Mistakes I Made as a 22-Year-Old Founder.” Then I ran it through Gemini 2.5 Pro as an independent reviewer. Gemini’s verdict was that the essay is itself an ego move. I think Gemini is mostly right.
Gemini 说我那篇 ego 文还是凡尔赛
Published:
今天早些时候我发了一篇叫《22 岁 founder 的四个 ego 错误》。然后我把这篇喂给了 Gemini 2.5 Pro 当独立审稿人。Gemini 的判决是:这篇文章本身就是一个 ego 动作。我觉得 Gemini 大体是对的。
Geeta Persad Came Back to Austin to Build a Climate Group That Actually Talks to Policy
Published:
Most academic climate scientists will tell you they care about policy and then publish a paper that no policymaker is ever going to read. Geeta Persad spent four years working at the Union of Concerned Scientists translating climate models for water managers, and she came back to academia knowing exactly what the gap looks like.
Geeta Persad 回到奥斯汀,组建了一个真正与政策对话的气候团队
Published:
大多数学术气候科学家会告诉你他们在乎政策,然后发表一篇任何决策者都不会读的论文。Geeta Persad 在忧思科学家联盟(Union of Concerned Scientists)花了四年,为水资源管理者翻译气候模型,然后带着对那个差距的清醒认识回到了学术界。
Eric Greene Built a Way to Watch Single Molecules of DNA Repair Themselves
Published:
Most of biology is averages. You measure a million cells, you get a mean. Eric Greene’s lab figured out how to actually watch one DNA molecule at a time, repair itself, in real time, and that changes what kind of questions biology can ask.
Eric Greene 建了一套办法,能观看单个 DNA 分子如何自我修复
Published:
大部分生物学是平均值。你测一百万个细胞,得到一个均值。Eric Greene 的实验室却想出了怎么一次实时观看一个 DNA 分子自我修复,而这改变了生物学能提出什么样的问题。
Trees Drink From Rock, and Daniella Rempe Proved It
Published:
If you ask most people where trees in California get their water in a drought, they will say “the soil.” It turns out a huge fraction of it comes from cracks in the bedrock underneath the soil, and Daniella Rempe is the person who put numbers on it.
树从石头里喝水,Daniella Rempe 证明了这件事
Published:
如果你问大多数人,加州在干旱时树从哪里取水,他们会说”土壤”。事实证明,很大一部分水其实来自土壤下方岩石裂缝里的基岩,而 Daniella Rempe 就是把数字放到这件事上的人。
Ashley Matheny Treats Trees as Pumps, and That Changes the Whole Model
Published:
Most land surface models treat a tree like a passive straw. Water comes in at the roots, water leaves at the leaves, end of story. Ashley Matheny’s research basically says no, a tree is an active hydraulic system with storage, capacitance, and a strategy, and if you do not model it that way you are going to be wrong about drought.
把树当成水泵:Ashley Matheny 改变了整个陆地模型
Published:
大多数陆地表面模型都把一棵树当成一根被动的吸管。水从根进来,水从叶出去,故事就这么简单。Ashley Matheny 的研究基本上在说:不对,树是一个带有储水、电容和策略的活跃水力系统,如果你不这样建模,你在干旱问题上就会犯错误。
Zero to a Billion in Ten Years Sounds Normal, Until You Realize It Means Doubling Every Year
Published:
People love saying that ten years to a billion-dollar valuation is the normal pace for a unicorn, neither fast nor slow. But once you actually crunch the compound numbers, you realize what hides behind that word “normal” is a beast that doubles in value every single year.
十年从零到十亿美金,这种「正常」速度其实是每年翻一倍
Published:
很多人会说,十年估值十个亿美金,就是一只正常的独角兽,不慢也不快。但你真去算一下复利就会发现,这「正常」二字背后藏着一个每年估值翻一倍的怪兽。
Five Orders of Magnitude on Purpose, Why Text AI and Embodied AI Are Not Playing the Same Data Game
Published:
Call it a (10^5) gap if you like, once you put semantic entropy and effective training samples on the board, the difference between internet-scale text and embodied robotics data stops sounding like rhetorical inflation and starts sounding almost conservative. You just have to separate raw physics bits from semantics a learner can actually use.
五个数量级,「十万倍」可能还是保守的,第一性原理拆开看一眼文字 AI 与具身 AI 的数据鸿沟
Published:
「十万倍」这个说法乍听像聊天里随手甩出去的量级,但如果把它放回语义信息熵和有效训练样本这两把尺子下面,它非但不夸张,反而可能是保守的。关键是你得把「原始物理比特」和「带语义的、可被学习的数据」分开看。
Who Hoards Multimodal Data for Real, Meta, ByteDance, X, and the Visa Plot Twist
Published:
Ask which campus actually sits on the best multimodal feedstock for GPT-4V-class perception, Sora-class video, Gemini-scale bundles, and Meta Emu-style image stacks, and the short answer is almost vulgar in how cleanly it splits three big piles. Meta still pulls ahead by a chasm on stills, ByteDance owns the high-velocity short-video river that is really motion plus audio, and X ships the smallest absolute media volume yet the weirdest leverage on tight text-image coupling and live-event semantics.
「图片地主」对「视频钥匙」对「实时百科」,多模态家底Meta、ByteDance和X怎么分
Published:
这题问到刀尖上了。如果把「高质量多模态训练数据」收窄到对 GPT-4V、Sora、Gemini、Emu 这类模型真有喂饭价值的图文或视频,短答其实很锋利,Meta(Facebook / Instagram)静态图片数量的库存对其他两家几乎是断层第一,ByteDance(TikTok / 抖音)在短视频也就是动态图片流上占最大优势,X(Twitter)绝对量级最小,但图文的贴脸相关性和实时信息密度是独一份。
The Physics Professor Who Won’t Touch AI, and the Student Who Can’t Stop
Published:
Top physics researchers aren’t blind to what AI is doing in the application layer. They’re locked by their own evaluation system and can’t afford to look.
用 AI 的物理学生和拒绝 AI 的物理教授,就像教员和王明博古的区别
Published:
顶尖的 physics 老年研究者并非看不见应用层的繁荣,而是被自己的评价体系与资源锁定。
Think First, Code Later with AI
Published:
In an era where AI can write code in seconds, I just learned the hard way that blindly moving fast is actually slowing me down.
谋定而后动,不要让 AI 一股脑地写代码
Published:
在这个闭着眼睛就能让 AI 写代码的时代,我今天结结实实地挨了一锤。
Earth System Model Skill Packages: Deep Knowledge Bundles for Noah-MP, CLM, CAM, MOM6, WRF, E3SM, and More
Published:
Earth system models are some of the most complex scientific software ever written, and they are also some of the worst-documented for newcomers. I have been building a series of “skill packages” — structured, progressive-disclosure knowledge bundles — for the major Earth system and land surface models, designed to be used by both new graduate students and AI coding agents.
地球系统模型技能包:为 Noah-MP、CLM、CAM、MOM6、WRF、E3SM 等量身打造的深层知识包
Published:
地球系统模型是人类写过的、有史以来最复杂的一批科学软件,可它们对新手来说偏偏又是文档最糟糕的一批。我一直在为主要的几大地球系统和陆地表面模型构建一系列”技能包”——结构化的、渐进式披露的知识包——设计给刚入门的研究生和 AI 编码代理两类使用者使用。
ESM-bench, Testing Whether AI Agents Understand Earth System Model Physics
Published:
The scary part is not that AI agents fail on Earth system model code, but that they can fail while looking almost right.
ESM-bench:测试AI智能体是否理解地球系统模型的物理学
Published:
最可怕的部分不在于 AI 智能体在处理地球系统模型代码时失败了,而是它们可能在看似正确的代码中失败。
Share Your Research Skills and Become a Nature Coauthor
Published:
Contribute just one research skill and you can become a coauthor of a Nature paper.
分享科研绝活,成为 Nature 合著者
Published:
只要你贡献一点科研绝活,就能成为 Nature 文章的合著者。
We Are Building an Alexandria Library for the AGI Era, and We Aim for Nature
Published:
Just by running a single command while using AI for research, you can turn your tacit knowledge into a Nature co-authorship.
我们在倒腾一个AI时代的亚历山大图书馆,顺便想冲一下Nature
Published:
只要你在用AI辅助科研,敲一行命令就能把你脑子里的「隐性经验」变成能发Nature的资本。
Female-Friendly Jobs in China and the US: Career Choices Without Drinking and Without Walking on Eggshells
Published:
While doing career planning for my girlfriend, I generated this report. Reading it as I went, I found it rather surreal.
中美女性友好型工作:不陪酒、不看脸色的职业选择
Published:
帮女朋友做职业规划的时候,我生成了这份报告。一边看一边觉得,挺魔幻的。
A Comprehensive Comparison of Female-Friendly Careers in China and the US
Published:
A Comprehensive Comparison of Female-Friendly Careers in China and the US
中美女性友好职业全景对比
Published:
帮女朋友做职业规划的时候,我生成了这份报告。一边看一边觉得,挺魔幻的。
‘No Eggshells’ Jobs Suitable for Chinese Women: A Horizontal-Vertical Analysis
Published:
While doing career planning for my girlfriend, I generated this report. Reading it as I went, I found it rather surreal.
适合中国女性的「不看人脸色」工作:横纵分析
Published:
帮女朋友做职业规划的时候,我生成了这份报告。一边看一边觉得,挺魔幻的。
In Conversation with Chenxi Hu
Published:
THIS IS A FAKE BLOG. The content below is fabricated and should not be cited or treated as a real interview or factual record.
对话胡晨曦
Published:
胡晨曦的研究揭示,城市化不只是被动地应对极端天气,它还会主动重塑热带气旋如何向沿海特大城市倾泻暴雨。
AI PhD Survival Guide: How to Finish a PhD in the Age of LLMs
Published:
A PhD is hard. A PhD in 2026 — with the field moving faster than your committee can read — is a different kind of hard. AI PhD Survival Guide is a handbook for surviving and finishing an AI or ML PhD without burning out and without falling behind.
AI 博士生存指南:在 LLM 时代如何拿下博士学位
Published:
读博很难。2026 年读博——当领域进展快到你的委员会都读不过来时——是另一种难。AI PhD Survival Guide 是一本手册,教你如何在 AI 或机器学习博士研究中活下来并毕业,既不会燃烧殆尽,也不至于掉队。
Why Is American Cuisine Lacking in Umami?
Published:
Although “umami” is the unshakable soul of Jiangsu-Zhejiang and Cantonese cooking, in traditional American food you often wander only among single-note salt, sweet, and oil, struggling to find that layered depth of flavor. That isn’t accidental; it’s the inevitable result of a deep cultural difference in how ingredients are handled, how seasoning is reasoned, and how industrial production works.
为什么美国菜缺乏鲜味?
Published:
虽然“鲜”(Umami)在江浙菜和广府菜中是不可撼动的灵魂,但在传统美国菜里,你往往只能在单一的咸、甜、油之间徘徊,而难觅那种富有层次感的味觉深度。这并非偶然,而是一场关于食材处理、调味逻辑与工业化生产方式的深层文化差异所导致的必然结果。
Standing at the Alamo: A Sacred Ground of Texas History
Published:
The silence surrounding the Alamo chapel in San Antonio belies the brutal, thirteen-day siege that transformed this former Spanish mission into the ultimate symbol of Texan independence. Standing before its weathered facade today, one can almost hear the echoes of a conflict that remains one of the most legendary chapters in American history.
站在阿拉莫:德克萨斯历史的圣地
Published:
今天我站在了圣安东尼奥的阿拉莫教堂前。这座看似宁静的建筑,却承载着德克萨斯乃至美国历史上最惨烈、最具传奇色彩的一页。
When Climate Data Comes Alive: A Day at TACC
Published:
The roar of cooling systems in the Texas Advanced Computing Center (TACC) isn’t just noise—it’s the power of supercomputers translating the overwhelming dimensionality of climate data into something students can finally see and understand.
当气候数据活起来:在 TACC 的一天
Published:
德州高级计算中心(TACC)里冷却系统的轰鸣不只是噪音——那是超级计算机的力量,正在把气候数据那令人不知所措的多维性,转化成学生最终能看见、能理解的东西。
hao-tokens: A Practical Guide to Free and Cheap LLM API Tokens
Published:
Every indie developer who has ever burned through a free API tier knows the feeling: you build something cool, and then your OPENAI_API_KEY runs out at the worst possible moment. hao-tokens is a curated list of the legitimate ways to get free or low-cost LLM tokens so you can keep building.
hao-tokens:免费和低价 LLM API Token 实用指南
Published:
每一个曾经把某个免费 API 层额度烧光的独立开发者都知道那种感觉:你做了很酷的东西,然后你的 OPENAI_API_KEY 在最不该耗尽的时候用光了。hao-tokens 是一份经过整理、合法获取免费或低价 LLM token 的清单,让你能继续开发下去。
Virtual Cell Neuromorphic Gene Language Models: A VC Intern’s Field Guide
Published:
Virtual cell models powered by neuromorphic computing and gene language models represent one of the most capital-intensive and scientifically ambitious convergences in biotech AI. If you’re evaluating this space as a VC intern, you need to understand three core components: what these systems actually do, why the market is moving now, and where the investable opportunities lie.
虚拟细胞神经形态基因语言模型:VC 实习生的领域指南
Published:
由神经形态计算和基因语言模型驱动的虚拟细胞模型,代表了生物技术 AI 领域中资本最密集、科学最雄心勃勃的融合方向之一。如果你正在以 VC 实习生的身份评估这个领域,你需要理解三个核心组成部分:这些系统实际上做什么、为什么市场现在开始行动,以及可投资的机会在哪里。
What Happens When You Just Give People Money? Sam Altman’s UBI Experiment
Published:
What Happens When You Just Give People Money? Sam Altman’s UBI Experiment
直接发钱会怎样?奥特曼的全民基本收入实验
Published:
直接发钱会怎样?奥特曼的全民基本收入实验
The Arrow of Time: How Timestamps Give AI Agents a Soul
Published:
Why AI agents lack temporal awareness, and how injecting timestamps transforms them from stateless automata into coherent entities.
时间之箭:时间戳如何赋予AI Agent灵魂
Published:
为什么AI Agent缺乏时间意识,以及注入时间戳如何将它们从无状态自动机转变为连贯的实体。
Moat Plus Momentum: Why AI Makes Preparation Optional
Published:
Research, stock trading, and startups all share a common pattern: success comes from combining a defensible core competency with the ability to ride trending waves. You don’t need exhaustive preparation anymore. You need methodology and the ability to produce content when it matters. When the right moment arrives, you strike.
护城河加热点:为什么 AI 让准备变得可选
Published:
科研、股市还有创业,这些所有需要展示并能得到结果的东西,本质上都是”主业/具有护城河的本行加上热点”。所以需要做好准备,或者说不一定非要进行那种极其周全的准备,只要掌握一些方法论就可以。利用 AI 在关键时刻能够产出内容,一旦关键热点到来,马上就可以抓住。
High-Dimensional Manifolds and the Rationality of Next Token Prediction
Published:
Why next token prediction isn’t just “stochastic parroting”—it’s solving geodesics on the manifold of human knowledge.
高维流形与Next Token Prediction的合理性
Published:
为什么Next Token Prediction不只是”随机鹦鹉”——它在求解人类知识流形上的测地线。
The Gibbs Phenomenon in LLMs: Why Hallucination is Mathematical Destiny
Published:
When continuous functions meet discontinuous truth: Understanding LLM hallucinations through the lens of Fourier analysis.
LLM的吉布斯现象:为什么幻觉是数学宿命
Published:
当连续函数遇上离散真理:从傅里叶分析的视角理解LLM幻觉。
The Gibbs Phenomenon: Why LLM Hallucinations Are Mathematically Inevitable
Published:
When Large Language Models (LLMs) hallucinate, we often treat it as an engineering bug to be fixed. But what if hallucinations are not a flaw, but a mathematical inevitability? A recent interdisciplinary discussion revealed a profound connection between the Gibbs phenomenon in Fourier analysis and the fundamental limitations of neural networks.
吉布斯现象:为什么大模型幻觉在数学上不可避免
Published:
当大语言模型(LLM)产生幻觉时,我们通常将其视为需要修复的工程缺陷。但如果幻觉不是缺陷,而是数学上的必然结果呢?最近一场跨学科讨论揭示了傅里叶分析中的吉布斯现象与神经网络根本局限性之间的深刻联系。
Epistemological Interferometry: The Deep Isomorphism Between Black Hole Imaging and LLM Training
Published:
How photographing a black hole and training GPT-4 are fundamentally the same process: extracting coherent structure from impossibly sparse frequency-domain measurements.
认识论干涉测量:黑洞成像与LLM训练的深层同构性
Published:
拍摄黑洞照片和训练GPT-4在根本上是同一个过程:从极度稀疏的频域测量中提取相干结构。
The Entropy Paradox: Why AI Struggles with Scientific Discovery
Published:
The promise of AI-automated science is intoxicating: imagine machines that can generate hypotheses, design experiments, and publish papers while we sleep.
熵悖论:为什么AI难以实现科学发现
Published:
AI 自动化科研的愿景令人陶醉:想象机器能在我们睡觉时生成假设、设计实验、发表论文。然而,尽管关于“AI 科学家”的新闻铺天盖地,我们正在撞上一堵根本性的墙。问题不在于算力或数据集规模,而是某种更深刻的东西,根植于科学发现的本质和信息论之中。
How to Find Outliers in Statistics
Published:
In statistics, how you find outliers depends on the dimensionality of your data, its distributional characteristics, and your tolerance for what counts as “abnormal.” Here are the most common and standard approaches:
统计学上如何寻找Outlier
Published:
在统计学中,寻找离群值(Outlier)的方法取决于数据的维度、分布特征以及你对“异常”的容忍程度。以下是几种最常用且标准的方法:
AI 时代,PhD 应该按风投模式培养
Published:
建立在 19 世纪学徒制模型上的博士培养体系,几乎在所有关键经验指标上都已经出现了系统性失灵。
How PhDs Can Self-Design KPIs: From ‘Felt Effort’ to ‘Systematic Output’
Published:
Doing a PhD cannot rely on “felt effort” and “moving yourself emotionally.” We need to upgrade day-to-day literature reading, experiment design, data analysis, and paper writing into a personal research system that is quantifiable, reviewable, and continuously optimizable. The point is not self-exploitation, but verifying whether your research efficiency truly exists, and keeping precious PhD time from being consumed inefficiently.
PhD如何自我设计KPI:从“感觉努力”到“系统产出”
Published:
读博不能仅凭“感觉努力”和“自我感动”。我们需要将日常的文献阅读、实验设计、数据分析和论文写作,升级为一套可量化、可复盘、可持续优化的个人科研系统。重点不在于自我压榨,而在于验证自己的科研效率是否真实存在,避免宝贵的博士时间被低效消耗。
For Dating and Resource-Sharing Markets, Should You Use Traditional Search/Ads/Rec or AI Recommendation Algorithms?
Published:
I’ve been discussing this question with a friend lately. He wants to use AI for matching in the dating market; I think with a small sample size this is perfectly feasible, you don’t even need to build any search/ads/rec system at all. Just ask the large language model directly, toss in a few users’ profiles, and have it rank them; the whole process is very simple.
婚恋和资源共享市场,到底该用传统搜广推还是 AI 推荐算法?
Published:
我和朋友最近在讨论这个问题。朋友说想拿 AI 来做婚恋市场的匹配,我觉得在样本量小的情况下,这完全可行——甚至根本不需要搭任何搜广推系统。你直接问大语言模型,把几个用户的简历丢进去,让它排序就好了,整个过程非常简单。
A Fun Coincidence: The Time It Takes to Finish a PhD Is Exactly How Long It Took OpenAI to Change the World
Published:
In the fall of 2015, two things happened simultaneously.
一个有趣的巧合:读完博士的时间,刚好够 OpenAI 改变世界
Published:
2015 年秋天,有两件事同时发生了。
Why Foundation Model Companies Are Pushing Agent Products Even While Losing Money
Published:
The signal truly worth heeding is not that AI companies are growing too slowly, but that they are growing fast while cash-flow pressure keeps mounting.
为什么大模型厂商还在亏钱,却还要拼命推 Agent 产品
Published:
真正值得警惕的信号不是 AI 公司增长太慢,而是它们增长得很快,现金流压力却仍在放大。
OpenClaw: Wiring the Command-Line World Into Your Second Brain
Published:
To use OpenClaw or not to use OpenClaw, that is the question.
OpenClaw:把命令行世界接成你的第二大脑
Published:
用不用 OpenClaw?这是个问题
Do Big Tech Engineers Actually Need OpenClaw?
Published:
Should Silicon Valley engineers use OpenClaw?
硅谷大厂人究竟需不需要用 OpenClaw?
Published:
硅谷大厂的工程师究竟需不需要用 OpenClaw?
An Alumnus’s Real Story: Three Deaths and Rebirths, from Aspiring Writer to AI Entrepreneur
Published:
The “Alumni Real-Life Archive” documents how alumni made decisions at turning points, giving current students a wider range of real paths to consider.
校友真实档案:从作家梦到AI创业者的三次死亡与重生
Published:
欢迎加入”校友真实档案”。这里记录的不是成功经验,而是人在关键节点如何做选择。我们希望通过这些真实路径的分享,为在校学生呈现人生的多样性。
AI Survival Guide: How to Stay Useful as a Knowledge Worker in the Age of LLMs
Published:
Every knowledge worker is now competing with an LLM for some part of their job. AI Survival Guide is a handbook for figuring out which parts, what to do about it, and how to come out the other side better at your craft instead of replaced by it.
AI 生存指南:在 LLM 时代作为知识工作者如何保持有用
Published:
如今,每一位知识工作者都在用某项工作的某些部分与 LLM 竞争。AI Survival Guide 是一本手册,帮你弄清是哪些部分、该怎么应对,以及如何带着更强的技艺从这场竞争的另一头走出来,而不是被它取代。
I published a preprint on Zenodo: Research grounding for a bilateral venture capital model of PhD programs
Published:
A PhD system built on a 19th-century apprenticeship model is failing by nearly every empirical measure — ~40–50% attrition, depression rates six times the general population, and tenure-track placement below 15% in many fields — while venture capital has spent four decades perfecting bilateral contracts that manage exactly the risks PhD programs ignore: information asymmetry, moral hazard, hold-up, and misaligned incentives. The literature across economics, education policy, signaling theory, and AI research converges on a striking conclusion: the structural tools to fix the PhD already exist in VC contract design, but academia has never imported them. This research compendium maps the evidentiary landscape across six domains to ground the argument.
From Collaboration to Autonomy: The Business Map and Paradigm Shift in Human x Human, Human x Agent, and Agent-to-Agent Evolution
Published:
If you’ve been watching the business map of the AI track lately, you can’t help but be shaken by that trend feverishly racing from “human-machine collaboration” toward “multi-agent autonomy.”
从协作到自治 人x人、人xAgent与Agent2Agent演进中的商业图谱与范式迁移
Published:
如果你最近在看AI赛道的商业图谱,你一定会被那个从「人机协作」向「多智能体自治」疯狂演进的趋势给震撼到。
When the Vibe Coding VC Replaces the Venture Capital VC
Published:
Large model vendors spend rocket-launch money raising a legion of geeks who only want to buy a two-dollar firecracker just to hear it pop.
当 Vibe Coding 的 VC 取代了 Venture Capital 的 VC
Published:
大模型厂商用造火箭的成本,培养出了一群只想买两块钱鞭炮听个响的极客。
2055, Brain-to-Brain Conversation, and the Future That Still Feels Warm
Published:
Sometimes a piece of old science fiction stays in your head not because of the gadgets, but because it got the emotional structure of the future right.
《2055》、脑内对话,以及那个依然温热的未来
Published:
有些旧科幻会一直留在脑子里,不只是因为它写了什么炫酷科技,而是因为它提前抓住了未来的情感结构。
Every Generation Has Its Own To-Do List
Published:
Every generation has its own way of managing work: paper notebooks, SaaS task managers, and now programmable agentic workflows powered by tools like OpenClaw heartbeat.
每一代人,都有每一代人的 To-Do List
Published:
每一代人都有自己管理任务的方式:最早是纸和笔,后来是 SaaS 任务管理工具,现在则开始进入像 OpenClaw heartbeat 这样可程序化、可持续运行的 agentic workflow 时代。
GitHub PRs as Inbound Marketing for Technical People
Published:
For technical people, a strong pull request on a fast-growing open-source repo can become better inbound marketing than another generic post about AI.
把 GitHub PR 当作技术人的 Inbound Marketing
Published:
对技术人来说,在一个高速增长的开源仓库里做出高质量 PR,往往比再发一篇泛泛而谈的 AI 观点帖更像真正有效的 inbound marketing。
How AI Can Accelerate Career Growth — and Why This Matters for AI for Science
Published:
Early-career growth usually comes down to a few repeatable layers: alignment, delivery, visibility, communication, leadership, and trust.
AI 如何加速职业成长——以及这对 AI for Science 为什么更重要
Published:
职业早期的成长,通常可以拆成几个可重复的层次:方向对齐、稳定交付、建立可见度、提升沟通、展现领导力,以及赢得信任。
Technology Is Not the Moat; Sales Is
Published:
The most important thing is selling. Having technology is useless; it only matters if someone is willing to buy. I don’t think anyone has a real tech moat; that part is easy to solve. As long as you have a little technical foundation, you can handle it. What matters most is that people come and buy. Though, it might be that my own technical level is too high, and I’ve grown numb to technology.
技术不是壁垒,销售才是
Published:
所以最重要的是推销。有技术并没有什么卵用,只要有人要购买才有用。我觉得所有人都没有 tech 壁垒,这个很容易解决。只要稍微有一点 tech 基础,都能解决。最重要的是,有人来买。不过,也可能是我的技术太高了,我对技术已经无感。
Measuring Work by Energy Consumption
Published:
衡量工作效率,最重要的指标或许不是投入的时间,而是消耗的能源。
用能量消耗衡量工作
Published:
衡量工作效率,最重要的指标或许不是投入的时间,而是消耗的能源。
Silicon Valley’s Spiritual Strata: Disruptive Innovation, Creative Betrayal, and the Thousand-Fold Flywheel of Risk and Growth
Published:
The most terrifying thing about Silicon Valley has never been the brilliant ideas of geniuses, but that it established an entire system of mechanisms that encourage “creative betrayal.”
硅谷的精神地层:颠覆性创新,创造性背叛,风险和增长的千倍飞轮
Published:
硅谷最可怕的从来不是那些天才的想法,而是它建立了一整套鼓励「创造性背叛」的机制。
颠覆性创新者的思维模式 | 改变世界的这些人,究竟看到了什么?|丰饶时代的认识论危机
Published:
改变世界的人到底看到了什么是普通人看不到的?
颠覆性创新者的思维模式 | 改变世界的这些人,究竟看到了什么?|丰饶时代的认识论危机
Published:
改变世界的人到底看到了什么是普通人看不到的?
If Jobs Were Building Today, He Would Toss Most multiAgent Products in the Trash
Published:
Lay out every hyped multiAgent product on a table in front of Steve Jobs and he would probably squint, frown, and ask one question, why on earth are you showing the user any of this.
如果乔布斯今天还在做产品,他会一脚把multiAgent踢进垃圾桶
Published:
把现在市面上所有被吹得天花乱坠的multiAgent产品摆在乔布斯面前,他大概率会皱着眉头反问一句,你为什么要让用户看到这些东西。
Clear Plus Airport Experience: Money, Privilege, and Market Regulation
Published:
Exchanging money for time is a privilege I rarely indulge in, but a surprise membership benefit that cleared airport security in ten minutes changed my perspective on friction and market regulation. While a Clear Plus membership normally costs over a hundred dollars annually, obtaining it for free through an Uber membership allowed me to experience a level of efficiency that money can’t always buy—at least not without a well-regulated system behind it.
Clear Plus 体验:金钱、特权与市场调节
Published:
用金钱换取时间是一种我很少尝试的“特权”,但这次在奥斯汀机场仅用十分钟便完成安检的经历,让我对效率与市场调节有了新的思考。这份特权并非我主动购买,而是通过 Uber 年费会员赠送的 Clear Plus 获得的——原本需要每年支付一百多美金的服务,在免除门槛后,带给我一种金钱也未必能随时买到的流畅体验。
I Spent a Night Reverse-Engineering Buy Borrow Die With Gemini, and Realized the F-1 Script Looks Nothing Like the Billionaire One
Published:
I spent a whole evening with Gemini taking apart the Buy Borrow Die playbook that American billionaires run, expecting that with a little scaling down I could just copy the moves, and what I found instead was that almost every single move has an F-1 trapdoor underneath it, and the list of things I can actually do fits on one page.
我用 Gemini 拆了一遍富豪的 Buy Borrow Die,结果发现 F-1 学生的剧本完全不一样
Published:
我花了一个晚上跟 Gemini 把美国富豪那套 Buy Borrow Die 的玩法从头扒到尾,本来以为只要照抄就能复刻,结果一层层剥下来才发现,F-1 学生在这个剧本里几乎处处是坑,能用的招其实只有那么三两个。
The Rise of Massively Collaborative Science: A Deep Dive into AI and Interdisciplinary ‘Mega Author’ Projects in 2026
Published:
If you browse arXiv regularly, you will be struck by the AI papers whose author lists are so long you have to scroll through multiple pages.
大规模协作科学的崛起 2026年人工智能与跨学科「巨型作者」项目深度调研报告
Published:
如果你最近经常逛 arxiv,你一定会被那些作者名单长到需要翻页的 AI 论文给震撼到。
A Non-AI Researcher’s Guide: Publishing Top AI Papers Using Your Domain Expertise
Published:
To be totally honest with you, you don’t need to understand the underlying architecture of large language models to get your name on a top-tier AI paper right now, you just need to be a true expert in your own field.
给非AI研究者的发顶会指南:如何用你的专业知识在AI论文中署名
Published:
坦率的讲,你不需要懂任何大模型底层的原理,只要你是一个真的懂行的领域专家,你现在就能在顶级 AI 论文上署名。
YC’s “Five Traitors”: When Rebels Meet Rebels, YC’s Breakup with the Founders Most Like Itself
Published:
Five YC rebels: some were expelled, some were welcomed back. What happened next?
YC 的“五叛徒”:当叛逆者遇上叛逆者,YC 与它最像自己的创始人们的决裂
Published:
五位 YC 叛逆者:有人被驱逐、有人被重新接纳,后来又发生了什么?
US Survival Guide: A Practical Handbook for Chinese Students and newcomers in America
Published:
The US is full of small traps that nobody warns you about until after you’ve fallen into one. US Survival Guide is a handbook for Chinese students and newcomers on identifying risk early, knowing your rights, and building a financial defense before you need it. Read the Guide
美国生存指南:写给在美中国学生和新移民的实用手册
Published:
美国布满了各式各样的暗坑,在你掉进去之前,没人会提醒你。US Survival Guide 是一本写给中国学生和新移民的手册,教你及早识别风险、了解自己的权利,并在你需要之前先筑好财务防线。阅读指南
The Gold/Silver Ratio Pair Trade: Swap Gold for Silver When the Ratio Screams
Published:
When gold gets too expensive relative to silver, you sell your gold and buy silver. When silver catches up, you swap back. This sounds like folk wisdom, but it’s one of the most time-tested strategies in precious metals.
1盎司黄金能换多少白银?金银比 Pair Trade 的底层逻辑
Published:
金涨多了就换成银,银涨多了再换回金,这个听起来像民间偏方的操作,其实是贵金属市场里最经典的增厚策略之一。
When Product Excellence Meets Monetization Disaster: A VC Perspective on Wispr Flow
Published:
Wispr Flow is the best AI product I’ve used this year. But I haven’t paid them a cent. Here’s why that’s both impressive and concerning.
当产品卓越遭遇变现灾难:从VC视角看Wispr Flow
Published:
Wispr Flow 是我今年用过的最好的AI产品。但我一分钱都没付给他们。这既令人印象深刻,又令人担忧。
Atomize: Building a Task-Breaking Agent System Inspired by Goblin Tools
Published:
Yuxuan and I have been discussing AI agent ideas for a while. Yesterday, we finally decided: we’re building Atomize, a task-breaking agent system evolved from Goblin Tools.
Atomize:受 Goblin Tools 启发的任务分解 Agent 系统
Published:
我和 Yuxuan 讨论 AI Agent 的想法已经有一段时间了。昨天,我们终于做出了决定:我们要做 Atomize,一个从 Goblin Tools 演化而来的任务分解 Agent 系统。
Energy Singularity and the Intelligence Revolution: An In-Depth Research Report on Controlled Nuclear Fusion, Its Principles, Commercialization, and Sam Altman’s Strategic Bet
Published:
At the end of AI’s compute road lies the ultimate solution to energy.
能源奇点与智能革命:可控核聚变全景深度研究报告——原理、商业化进程与Sam Altman的战略赌注
Published:
引言:AI 的算力尽头,是能源的终极解法。
Simon Willison’s Newsletter: Deep Dives into LLMs and AI Engineering
Published:
If you’re serious about understanding LLMs and AI engineering, Simon Willison’s Newsletter is one of the best resources out there. With over 38,000 subscribers, it provides detailed analysis on AI, LLMs, web engineering, open source, data science, and Python.
Simon Willison 的 Newsletter:深入解析 LLM 与 AI 工程
Published:
如果你认真想要理解 LLM 和 AI 工程,Simon Willison 的 Newsletter 是最好的资源之一。拥有超过 38,000 名订阅者,它提供关于 AI、LLM、Web 工程、开源、数据科学和 Python 的深度分析。
Vibe Coding IDEs: a brief comparison (EN)
Published:
Every product has its own pros and cons. Cursor: “extraordinarily productive”; Kiro: “spec-driven”; Antigravity: “agent-first”.
Vibe Coding IDEs brief comparison
Published:
每一家都有每一家的优点和缺点。Cursor: “extraordinarily productive”; Kiro: “spec-driven”; Antigravity: “agent-first”.
The Hard Bottleneck at the Edge of AI Compute: America’s Silicon Steel Monopoly and Its Investment Logic
Published:
In financial markets, we are forever hunting for the kind of “choke point” that everyone calls a stranglehold.
AI算力尽头的硬核瓶颈 聊聊美国变压器硅钢的独家游戏与投资逻辑
Published:
在金融市场里,我们永远在寻找那种被称为「卡脖子」的咽喉要道。
Legacy USTC Academic Resources — Still Useful
Published:
A collection of USTC academic resources covering career planning, library access, and research tools.
中科大历史学术资源——依然实用
Published:
一份涵盖生涯规划、图书馆访问与研究工具的中科大学术资源合集。
New Preprint: Noah-Agent, a Multi-Expert AI Agent Framework for Fortran Climate Models
Published:
Parameterizing a large Fortran climate model by hand is slow, error-prone, and hard to validate. Noah-Agent asks whether a team of specialized AI agents can do it instead.
新预印本:Noah-Agent,面向 Fortran 气候模型的多专家 AI 智能体框架
Published:
手动为大型 Fortran 气候模型配置参数,速度慢、容易出错,也难以验证。Noah-Agent 要检验的是,一组各有所长的 AI 智能体能否协作完成这项工作。
| Warren Buffett: The “Open Source” Leader Half a Century Before GitHub | 巴菲特:比GitHub早了半个世纪的“开源”运动领袖 |
Published:
In an era driven by code, collaboration, and transparency, GitHub has become synonymous with the “open source” spirit. But if we turn our gaze to the world of finance, we find a pioneer of “open source” who predates the birth of GitHub by a full half-century. His name is Warren Buffett.
巴菲特:比GitHub早了半个世纪的”开源”运动领袖
Published:
在当今这个由代码、协作和透明度驱动的时代,GitHub 成为了”开源”精神的代名词。但如果我们将目光投向金融界,会发现一位”开源”的先行者,他比 GitHub 的诞生早了整整半个世纪。他就是沃伦·巴菲特。
Neural Galaxy - Your Personalized AI Conversation Visualization
Published:
A hand-controlled 3D visualization of your AI conversation history - fly through ChatGPT conversations and explore artificial intelligence concepts. Try Live Demo
Neural Galaxy - 属于你的 AI 对话可视化宇宙
Published:
一个支持手势控制的 3D 可视化项目,把你的 AI 对话历史变成可以飞行探索的星系。你可以在 ChatGPT 对话之间穿梭,也可以把抽象的人工智能概念变成一个能亲手操作的可视化空间。Try Live Demo
TQQQ ML Trend: Predicting a 3x Leveraged ETF With Machine Learning
Published:
TQQQ is the 3x-leveraged Nasdaq-100 ETF. It is also one of the most asymmetric instruments retail investors touch — the upside is real, the drawdowns are brutal, and the daily-rebalance math means buy-and-hold doesn’t behave the way most people assume. TQQQ ML Trend is an experiment in using machine learning to predict the trend regime, not the price.
TQQQ ML 趋势:用机器学习预测 3 倍杠杆 ETF
Published:
TQQQ 是纳斯达克 100 指数的 3 倍杠杆 ETF。它也是散户接触过的最不对称的金融工具之一——上行空间真实存在,回撤却十分残酷,而每日再平衡的数学逻辑意味着”买入并持有”并不像大多数人以为的那样运转。TQQQ ML Trend 是一个用机器学习来预测趋势状态的实验,而不是去预测价格。
Crypto Dashboard: An AI-Powered Real-Time Crypto Analysis Tool
Published:
Most crypto dashboards either drown you in numbers or hide the signal behind a paywall. Crypto Dashboard is my attempt at a clean, AI-augmented view of the market that runs entirely in your browser. Try Live Demo
Crypto Dashboard:AI 驱动的实时加密货币分析工具
Published:
大多数加密货币仪表盘要么把你淹没在数字里,要么把信号藏在付费墙后面。Crypto Dashboard 是我对一种简洁、由 AI 增强的市场视图的尝试,而且它完全运行在你的浏览器里。试试在线演示
AlphaEarthHack: A UT Austin Geoscience Hackathon Project on Earth System AI
Published:
AlphaEarthHack is the project our team built for the UT Austin Geoscience Hackathon. The goal: see how far we could push AI on Earth system data inside a single weekend. Try Live Demo
AlphaEarthHack:德州大学奥斯汀分校地球系统 AI 黑客松项目
Published:
AlphaEarthHack 是我们团队为德州大学奥斯汀分校地球科学黑客松打造的项目。目标只有一个:看看在一个周末内,我们能在地球系统数据上把 AI 推到多远。试试在线演示
Collaborating with Claude Code to Update My Academic Website
Published:
Today I had an interesting experience collaborating with Claude Code to completely overhaul my personal academic website. As a PhD student in Geological and Earth Sciences at UT Austin, I needed to update my GitHub Pages site with real professional information instead of the placeholder content that had been sitting there.
和 Claude Code 一起更新我的学术网站
Published:
今天我经历了一次挺有意思的合作:和 Claude Code 一起,把我的个人学术网站彻底重做了一遍。作为 UT Austin Geological and Earth Sciences 的博士生,我需要把 GitHub Pages 站点从一堆占位符内容,更新成真正能代表我专业背景的信息。
Tmux Orchestrator - Run AI agents 24/7
Published:
The Tmux Orchestrator enables Claude agents to work autonomously, schedule their own check-ins, and coordinate across multiple projects without human intervention - a project I explored and learned a lot from.
Tmux Orchestrator - 让 AI agents 7×24 小时持续工作
Published:
Tmux Orchestrator 让 Claude agents 能够自主工作、自己安排 check-in,并在多个项目之间协同推进——这是一个我探索过、也从中学到很多的项目。
VerificationSuccess - Email Verification Success Page
Published:
A simple, clean HTML page displaying email verification success message with responsive design and professional styling.
VerificationSuccess - 邮箱验证成功页面
Published:
这是一个简洁干净的 HTML 页面,用于展示邮箱验证成功提示,具备响应式布局和专业风格设计。
AI Billion Career - Career Planning Platform for AI Industry
Published:
A comprehensive tool built with React, Vite, and Supabase to help individuals plan and navigate their careers in the AI industry.
AI Billion Career - 面向 AI 行业的职业规划平台
Published:
这是一个基于 React、Vite 和 Supabase 构建的综合工具,帮助个人规划并导航自己在 AI 行业中的职业发展路径。
komomood - Couple Mood Tracking Heatmap
Published:
An elegant couple mood tracking website with self-hosted backend and SQLite, displaying daily mood records in GitHub contribution graph style.
komomood - 情侣心情追踪热力图
Published:
这是一个优雅的情侣心情记录网站,采用自托管后端和 SQLite,以 GitHub contribution graph 风格展示每日心情记录。
Welcome to My Academic Website
Published:
From the vast datasets of Earth System Models to the specialized niches of high-performance computing, this space documents my journey as a PhD student at UT Austin pushing the boundaries of Geological and Earth Sciences. This website is more than just a portfolio—it’s a hub where data-driven climate science meets the practical challenges of modern research.
欢迎来到我的学术网站
Published:
在 UT Austin 攻读地球科学博士学位的过程中,我始终在试图寻找复杂气候模型与真实世界影响之间的联结。这个学术网站不仅是我研究、论文和项目经历的展示入口,更是我记录如何利用数据驱动的科学方法去理解地球系统演变的思考空间。
ZuoGuoHuaDiao: A Life Calculator That Tells You Whether This Life Was Worth It
Published:
“这b人生过的值不值” — roughly, “was this damn life worth it?” — is a phrase that does a lot of emotional work on the Chinese internet. ZuoGuoHuaDiao (做过划掉) is a small web tool that takes the phrase seriously and tries to answer it with numbers. Try Live Demo
做过划掉:一款告诉你这一生值不值得的计算器
Published:
“这b人生过的值不值”——简单说,就是”这该死的一生过得值不值?”——这句在中国互联网上承载了大量情感的话语。ZuoGuoHuaDiao(做过划掉)是一款小巧的网页工具,它认真对待这句话,并试图用数字来回答它。试试在线演示
Summer Calculator: How Much Is Your Summer Vacation Actually Worth?
Published:
Most people think of summer vacation as time off. I started wondering whether that framing undersells it. Summer Calculator is a small web tool that estimates the full value of your summer — money, learning, relationships, health — not just the days you didn’t work. Try Live Demo
暑期计算器:你的暑假到底值多少钱?
Published:
大多数人把暑假当作休息时间。我开始怀疑这种框架是否低估了它。Summer Calculator 是一款小巧的网页工具,用来估算你暑假的完整价值——金钱、学习、人际关系、健康——而不仅仅是你没工作的那些日子。试试在线演示
Meteor Speed Variations - Research Manuscript
Published:
Research manuscript repository for meteor speed variations study published on Journal of Geophysical Research (JGR): Space Physics.
Meteor Speed Variations - 研究论文仓库
Published:
这是一个关于流星速度变化研究的论文仓库,相关稿件已在 Journal of Geophysical Research (JGR)发表。
Geomaps - Scientific Computing and Mapping Tools
Published:
Python and MATLAB tools for creating publication-quality scientific maps and geographic visualizations.
Geomaps - 科学计算与地图绘制工具集
Published:
用 Python 和 MATLAB 创建适合论文发表的科学地图与地理可视化图件。
1AI-polish - AI Academic Writing Polishing System
Published:
An AI writing assistant that polishes academic text using DeepSeek-R1 and detects AI-generated content.
1AI-polish - AI 学术写作润色系统
Published:
学术写作的严谨性不仅在于数据,更在于表达的精准,而 1AI-polish 通过集成 DeepSeek-R1 的推理能力,为研究者提供了一套集文本润色与 AI 检测于一体的深度协作系统,旨在让复杂的科研思想以更专业、更透明的方式呈现。
School Evaluator: Beyond Rankings — A Calculator for the Real Value of a School
Published:
School rankings are a single number standing in for a multi-dimensional decision. School Evaluator is a small web tool that lets you score a school across the dimensions that actually shape your life there. Try Live Demo
School Evaluator:超越排名的学校真实价值计算器
Published:
学校排名只是一个用单一数字去代表多维决策的替代品。School Evaluator 是一款小巧的网页工具,让你能在一个真正决定你在那里生活的维度上为学校打分。试试在线演示
Cultural Bridge - Cross-Cultural Communication Platform
Published:
A web-based platform for facilitating cross-cultural communication and understanding.
Cultural Bridge - 跨文化交流平台
Published:
一个基于 Web 的平台,用于促进跨文化交流与理解。
No Matter How Advanced Technology Gets, We Still Make Decisions With Our Genes
Published:
The story goes like this: I’ve been pondering a question lately—given how insanely advanced technology is today, why do we still have to spend so much time dressing up and managing our personal image?
科技再发达,我们依然是在用基因做决策
Published:
故事是这样的,我最近一直在寻思一个问题,既然现在的科技都已经发达到这种地步了,为什么我们这帮人,还是得天天花心思去收拾自己的形象?
ut01 - UT Austin Student Navigation Hub
Published:
A student-built navigation homepage that consolidates all essential UT Austin links in one place, solving the frustration of navigating deep university website hierarchies. Visit ut01.github.io
ut01 - UT Austin 学生导航中心
Published:
一个由学生亲手打造的导航首页,把 UT Austin 最常用的重要链接整合到一个地方,解决学校官网层级太深、入口分散的烦恼。访问 ut01.github.io
NASA FINESST Resources: A Practical Guide and Link Library for the FINESST Proposal
Published:
NASA’s Future Investigators in NASA Earth and Space Science and Technology (FINESST) is one of the most underused fellowships among US graduate students. Most PhD students have never heard of it, and the ones who have often miss the deadline because the proposal expectations aren’t obvious from the call alone. This repo is a curated guide of links, tips, and examples to help you write a competitive FINESST proposal.
NASA FINESST 资源:FINESST 提案实用指南与链接库
Published:
NASA 的未来地球与空间科技研究者(FINESST)项目,是美国研究生中最被低估的奖学金之一。这个仓库是一份精心整理的链接、技巧与示例指南,帮助你写出一份有竞争力的 FINESST 提案。
AI Writing Assistants: From Gibberish to a Structured Mess
Published:
When an AI writing assistant turns a paper’s incoherent rambling into a “well-structured pile of something,” we have to admit: it is basically a cyber laxative.
AI 就是论文开塞露:从狗屁不通变成拉了一坨大的
Published:
当AI写作助手把学术论文里的“狗屁不通”梳理成“结构完整的一大坨”时,我们不得不承认:这简直是一剂赛博开塞露。
komo520 - Koko & Momo Love Universe
Published:
A handcrafted website universe where every webpage, balloon, click, and line of code says “I love you” - designed exclusively for Momo. Visit Live Site
komo520 - Koko 与 Momo 的爱意宇宙
Published:
这是一个手工打造的网站小宇宙:每一个网页、气球、点击和每一行代码都在说“我爱你”——只为 Momo 而设计。访问在线站点
NASA FINESST Resources Guide
Published:
Securing up to $150,000 in research funding over three years can define a graduate career, and NASA’s FINESST program is the primary vehicle for that transformation. This guide distills the complex application process into actionable strategies for Earth and Space Science researchers seeking to join the next generation of Future Investigators.
NASA FINESST 资源申请指南
Published:
每年 5 万美元且连续 3 年的科研资助,让 NASA FINESST 项目成为地球与空间科学博士生必须争取的黄金机会。这份指南旨在将复杂的申请流程拆解为可操作的策略,帮助下一代“未来研究员”(Future Investigators)在激烈的竞争中脱颖而出。
LEAD-UTexas - Land Environment and Atmospheric Dynamics Group
Published:
Dr. Zong-Liang Yang’s Land Environment and Atmospheric Dynamics (LEAD) Group at UT-Austin employs satellite remote sensing, earth system modeling, and high-performance computing to advance understanding of Earth system sciences.
LEAD-UTexas - 德州大学奥斯汀分校陆地环境与大气动力学研究组
Published:
UT Austin 宗良杨博士(Dr. Zong-Liang Yang)领导的陆地环境与大气动力学研究组(LEAD Group)利用卫星遥感、地球系统建模和高性能计算,推动人类对地球系统科学的理解。
portfolio
AI Agent Knowledge Base Evaluation System
Multi-dimensional AI agent evaluation system with data anonymization and an LLM-as-a-judge architecture.
AI Text Processing System
Full-stack LLM text processing system with polishing, AI detection, and plagiarism-reduction features.
Benchmark Radar: Daily Evidence-First Radar for AI Benchmarks
Open-source daily radar and adoption leaderboard tracking which benchmarks frontier labs actually report, across 30+ curated documents from 10 organizations.
ESM-bench: Benchmarking AI Agents on Earth System Model Physics and Code
A 243-task benchmark testing whether AI agents understand Earth System Model physics and code, with multi-model evaluation and leakage detection.
Noah-Agent: Multi-Expert AI Agent Framework for Fortran Climate Models
A multi-expert AI agent framework for automated parameterization and validation of large-scale Fortran climate models.
Explainable AI for Noah-MP Land Surface Modeling
NSF NCAR-funded project applying explainable AI to improve physics-based land-surface modeling of plant–rock–water interactions.
UT01 Navigation Platform
Unified resource platform for UT Austin campus services — 34,940+ visits, with cross-device optimization.
publications
Perturbations by the 2022 Hunga-Tonga Volcano Eruption in the MLT Region Investigated Using the WACCM-X Simulation and Meteor Radar Observations
Published in AGU Fall Meeting 2023, 2023
AGU Fall Meeting abstract studying wave perturbations from the 2022 Hunga-Tonga eruption in the mesosphere and lower thermosphere using WACCM-X simulations and meteor radar observations.
Recommended citation: Wu, K., Liu, H.-L., Yi, W., & Xue, X. (2023). "Perturbations by the 2022 Hunga-Tonga Volcano Eruption in the MLT Region Investigated Using the WACCM-X Simulation and Meteor Radar Observations." AGU Fall Meeting Abstracts, SA33B-2892.
Download Paper
A Summary Report on the Space Physics Practical Education in 2022
Published in Review of Geophysics and Planetary Physics, 2024
A comprehensive report on space physics practical education initiatives in 2022, documenting educational programs and outcomes in space science education.
Recommended citation: Wu, K.*, Xu, X., Jiang, J., & Shen, A. (2024). "A Summary Report on the Space Physics Practical Education in 2022." Review of Geophysics and Planetary Physics.
Download Paper
Diurnal and seasonal variations of meteor speed and arrival angle observed by Mengcheng meteor radar
Published in JGR: Space Physics, 2024
This study investigates diurnal and seasonal variations of meteor speed and arrival angle using Mengcheng meteor radar observations, providing insights into meteoroid dynamics in the mesosphere and lower thermosphere.
Recommended citation: Wu, K., Yi, W.*, Xue, X.*, Reid, I., & Lu, M. (2024). "Diurnal and seasonal variations of meteor speed and arrival angle observed by Mengcheng meteor radar." JGR: Space Physics.
Download Paper
Noah-Agent: A Multi-Expert AI Agent Framework for Automated Parameterization and Validation of Large-Scale Fortran Climate Models (v0.1)
Published in Preprint (Zenodo); in preparation, 2025
Preprint. A multi-expert AI agent framework for automated parameterization and validation of large-scale Fortran climate models. Version 0.1, in preparation.
Recommended citation: Wu, K. (2025). "Noah-Agent: A Multi-Expert AI Agent Framework for Automated Parameterization and Validation of Large-Scale Fortran Climate Models (v0.1)." Preprint, Zenodo. https://zenodo.org/records/17862049
Download Paper
ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code
Published in Preprint (Zenodo); in preparation for NeurIPS Datasets and Benchmarks, 2026
Preprint. A 243-task benchmark testing whether AI agents understand Earth System Model physics and code, with multi-model evaluation, a classification rubric, precision/recall/F1 scoring, and leakage detection. In preparation for NeurIPS Datasets and Benchmarks.
Recommended citation: Wu, K., Cao, Y., & Mai, G. (2026). "ESM-bench: A Benchmark for Evaluating Whether AI Agents Understand Earth System Model Physics and Code." Preprint, Zenodo. https://zenodo.org/records/19802836
Download Paper
On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust
Published in Geography According to Foundation Models, Vol. 422, IOS Press, 2026
Peer-reviewed book chapter reviewing key ethical issues in generative GeoAI. Wu authored Section 8, “Trust in AI and GeoAI Models,” covering geo-hallucination, uncertainty as an ethical requirement, and provenance-aware protocols.
Recommended citation: Mai, G., Lao, N., Zhang, J., Mao, L., Wang, Z., Wu, N., Janowicz, K., Wu, K., Rao, J., Gao, S., & Zhu, R. (2026). "On the Ethics of Generative GeoAI: Explainability, Bias, Hallucination, Accountability, Privacy, and Trust." In Geography According to Foundation Models, Vol. 422, pp. 215-232. IOS Press. DOI 10.3233/FAIA260483.
Download Paper
How Does Integrating Plant Hydraulics Improve Noah-MP Land Surface Model
Published in 106th AMS Annual Meeting, 2026
Conference poster presenting the integration and evaluation of a plant hydraulics scheme in the Noah-MP land surface model.
Recommended citation: Wu, K., Li, L., Rempe, D., Matheny, A., Mbarak, M., & Yang, Z.-L. (2026). "How Does Integrating Plant Hydraulics Improve Noah-MP Land Surface Model." Poster presented at the 106th AMS Annual Meeting, Houston, TX.
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
Published in arXiv preprint; submitted to AAAI 2027, 2026
A benchmark for end-to-end autonomous scientific research across 40 tasks from 10 scientific domains, with real-paper grounding and expert-curated multimodal rubrics.
Recommended citation: Xu, W., et al. (including Wu, K.) (2026). "ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research." arXiv preprint. Submitted to AAAI 2027.
Download Paper
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Published in arXiv preprint arXiv:2608.04205, 2026
A population-scale simulated-user evaluation infrastructure with 8.3 billion persona records, four interactive playground environments, and 1,010 application tasks across 25+ domains.
Recommended citation: Li, X., et al. (including Wu, K.) (2026). "MatrAIx: Simulating the World with 8.3 Billion Persona Agents." arXiv preprint arXiv:2608.04205.
Download Paper
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
Published in arXiv preprint arXiv:2608.15844, 2026
A behavioral-science instrument that measures identity drift in generative agents carrying an immutable “soul file” through a resource-scarce, long-horizon multi-agent simulation.
Recommended citation: Ng, S., et al. (including Wu, K.) (2026). "MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations." arXiv preprint arXiv:2608.15844.
Download Paper
ASI-Bench: At the Dawn of Artificial Superintelligence
Published in arXiv preprint arXiv:2608.17271, 2026
A benchmark with 60 project-level tasks across 11 scientific domains that evaluates frontier models on open-ended scientific research beyond human-expert benchmarks.
Recommended citation: Zhou, J., et al. (including Wu, K.) (2026). "ASI-Bench: At the Dawn of Artificial Superintelligence." arXiv preprint arXiv:2608.17271.
Download Paper
From Personas to Simulated Users: A Fitness-for-Purpose Survey
Published in Submitted to AAAI 2027 Artificial Intelligence for Social Impact Track, 2027
A fitness-for-purpose survey of the progression from static personas to simulated users, submitted to the AAAI 2027 Artificial Intelligence for Social Impact Track.
Recommended citation: Liu, X., et al. (including Wu, K.) (2027). "From Personas to Simulated Users: A Fitness-for-Purpose Survey." Submitted to the AAAI 2027 Artificial Intelligence for Social Impact Track.
talks
Perturbations by the 2022 Hunga-Tonga Volcano Eruption in the MLT Region Investigated Using WACCM-X Simulation and Meteor Radar Observations
Published:
This oral presentation discussed the research findings on perturbations caused by the 2022 Hunga-Tonga volcano eruption in the mesosphere and lower thermosphere (MLT) region, combining WACCM-X simulation results with meteor radar observations.
Perturbations by the 2022 Hunga-Tonga Volcano Eruption in the MLT Region Investigated Using the WACCM-X Simulation and Meteor Radar Observations
Published:
Poster presentation at the 2023 AGU Fall Meeting investigating the atmospheric perturbations caused by the 2022 Hunga-Tonga volcano eruption in the mesosphere and lower thermosphere (MLT) region.
Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS) Evaluation
Published:
This poster presentation evaluated the Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS), presented at the Jackson School of Geosciences Research Symposium at UT Austin.
Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS) Evaluation
Published:
This poster presentation evaluated the Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS), presented at the Advancing Land Modeling Symposium.
An open-source geospatial AI platform based on the AlphaEarth Foundation Model
Published:
This oral presentation showcased an open-source geospatial AI platform built on the AlphaEarth Foundation Model, winning 2nd place in the Geoscience Hackathon ‘25.
Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS) Evaluation
Published:
Poster presentation on evaluating the Noah-MP land surface model with plant hydraulics scheme (Noah-MP-PHS), presented at the 106th American Meteorological Society Annual Meeting in Houston, Texas.
ESM-bench: Can AI Agents Understand Earth System Model Physics and Code?
Published:
Oral presentation on ESM-bench at the 4th Annual Good Systems Smart Cities and AI Innovations Symposium at the University of Texas at Austin.
teaching
Earth in 2100 (GEO 303E)
Graduate Teaching Assistant, University of Texas at Austin, Jackson School of Geosciences, 2024
Graduate Teaching Assistant for Earth in 2100 (GEO 303E), a course examining Earth system changes and climate projections for the year 2100.
