<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://koutian.is-a.dev/feed.xml" rel="self" type="application/atom+xml" /><link href="https://koutian.is-a.dev/" rel="alternate" type="text/html" /><updated>2026-08-14T04:43:41+00:00</updated><id>https://koutian.is-a.dev/feed.xml</id><title type="html">koutian.is-a.dev</title><subtitle>AI4Geoscience PhD student at UT Austin specializing in explainable AI for physics-based land surface modeling, Noah-MP research, and Earth system science. Expert in Python, MATLAB, HPC, and machine learning for atmospheric and geoscience applications.</subtitle><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><entry><title type="html">UT Austin GenAI Society：从听 AI 到共同生产公共成果</title><link href="https://koutian.is-a.dev/zh/posts/2026/08/ut-austin-genai-society/" rel="alternate" type="text/html" title="UT Austin GenAI Society：从听 AI 到共同生产公共成果" /><published>2026-08-11T00:00:00+00:00</published><updated>2026-08-11T00:00:00+00:00</updated><id>https://koutian.is-a.dev/zh/posts/2026/08/ut-austin-genai-society-zh</id><content type="html" xml:base="https://koutian.is-a.dev/zh/posts/2026/08/ut-austin-genai-society/"><![CDATA[<p>一个生成式 AI 学生社团，怎样才能不止于办讲座、转发资讯和追逐下一款模型？</p>

<blockquote>
  <p>作者：<a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>；<a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>GenAI Society at UT Austin 成立于 2024 年，公开定位横跨研究探索、产业创业与社会文化讨论。这份横纵分析把它放回 UT Austin 的 AI 发展脉络，也将它与校内外的 AI 学生组织并置，追问一个更具体的问题：它如何把短期注意力转化为成员持续参与、校园真实使用和可公开复用的成果？</p>

<p>报告提出一个可执行的方向：以 4 至 6 周为一期，组织跨专业学生围绕真实、低风险的校园问题完成原型、风险卡和复盘。评价标准不是报名人数，而是成果是否有人使用、能否复现，以及成员是否愿意继续参与。</p>

<p>完整 HTML 研究报告：<a href="/research/hv-analysis/UT_Austin_GenAI_Society_横纵分析报告.html">UT Austin GenAI Society：从「听 AI」走向「共同生产公共成果」</a></p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="artificial-intelligence" /><category term="student-organization" /><category term="ut-austin" /><category term="community" /><category term="hv-analysis" /><summary type="html"><![CDATA[一个生成式 AI 学生社团，怎样才能不止于办讲座、转发资讯和追逐下一款模型？]]></summary></entry><entry><title type="html">Luck Sourcing：主动寻找好运</title><link href="https://koutian.is-a.dev/zh/posts/2026/08/luck-sourcing/" rel="alternate" type="text/html" title="Luck Sourcing：主动寻找好运" /><published>2026-08-07T00:00:00+00:00</published><updated>2026-08-07T00:00:00+00:00</updated><id>https://koutian.is-a.dev/zh/posts/2026/08/luck-sourcing-zh</id><content type="html" xml:base="https://koutian.is-a.dev/zh/posts/2026/08/luck-sourcing/"><![CDATA[<p>很久以前，我偶然看到一本书，叫 <em>Chase, Chance, and Creativity: The Lucky Art of Novelty</em>（《追逐、机遇和创造力：新奇的幸运艺术》）。我已经不记得书里的全部内容了，但它的核心想法一直留在我脑海中的某个角落：运气和创造力并不完全是随机的。机遇也许会意外到来，但我们仍然可以选择自己遇见它的频率，以及当它出现时，我们是否已经准备好认出它。</p>

<blockquote>
  <p>作者：<a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>；<a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>昨天和 Anne 聊天时，我为这个想法找到了一个说法：<strong>luck sourcing</strong>——主动寻找好运的来源。</p>

<p>我们通常谈起运气时，会把它说得好像只是碰巧发生在我们身上。但也许运气是可以被主动寻找的：通过我们认识的人、参与的对话、待过的地方、不断提出的问题，以及那些我们没有太快否定的奇怪想法。我们无法命令一个幸运的结果发生，但我们也许可以创造更多让运气出现的地方。</p>

<p>这就是我想在这里探索的想法。</p>

<h1 id="luck-sourcing">Luck Sourcing</h1>

<p>我大概六年前开始“扩列”。</p>

<p>当时没有什么明确目标，只是觉得：认识更多人，总会发生一些有意思的事情。于是我会利用各种碎片时间加人、聊天、建立联系。</p>

<p>有一次，我在火车上连续加了几十个人。聊得太投入，结果不知不觉坐过站了。</p>

<p>后来我做 VC intern，负责 deal sourcing。理论上，我应该不断寻找项目、研究创业者、建立投资线索。</p>

<p>但实际工作中，我经常做到一半，就开始和已经认识的朋友聊天。表面上看，这像是在摸鱼；但很多新的项目、合作和信息，恰恰来自这些没有明确目的的交流。</p>

<p>我后来意识到，这种行为可以像 deal sourcing 一样取一个名字。</p>

<p>就应该叫 <strong>luck sourcing</strong>。</p>

<p>Deal sourcing，是主动寻找 deal。
Talent sourcing，是主动寻找 talent。
Luck sourcing，则是主动寻找 luck。</p>

<p>很多人把运气理解为随机事件：等一个机会出现，等一个贵人联系自己，等某件好事刚好发生。</p>

<p>但很多所谓的运气，其实可以被增加。</p>

<p>你认识的人越多，交流越频繁，参与的社区越多，暴露在新信息中的次数越多，意外机会出现的概率也越高。</p>

<p>Luck sourcing 不是保证成功。</p>

<p>它只是让自己更频繁地出现在“好运可能发生”的地方。</p>

<p>You are not waiting for luck.</p>

<p>You are sourcing it.</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="luck" /><category term="creativity" /><category term="ideas" /><summary type="html"><![CDATA[很久以前，我偶然看到一本书，叫 Chase, Chance, and Creativity: The Lucky Art of Novelty（《追逐、机遇和创造力：新奇的幸运艺术》）。我已经不记得书里的全部内容了，但它的核心想法一直留在我脑海中的某个角落：运气和创造力并不完全是随机的。机遇也许会意外到来，但我们仍然可以选择自己遇见它的频率，以及当它出现时，我们是否已经准备好认出它。]]></summary></entry><entry><title type="html">Luck Sourcing: The Sourcing of Getting Lucky</title><link href="https://koutian.is-a.dev/posts/2026/08/luck-sourcing/" rel="alternate" type="text/html" title="Luck Sourcing: The Sourcing of Getting Lucky" /><published>2026-08-07T00:00:00+00:00</published><updated>2026-08-07T00:00:00+00:00</updated><id>https://koutian.is-a.dev/posts/2026/08/luck-sourcing</id><content type="html" xml:base="https://koutian.is-a.dev/posts/2026/08/luck-sourcing/"><![CDATA[<p>A long time ago, I came across a book called <em>Chase, Chance, and Creativity: The Lucky Art of Novelty</em>（《追逐、机遇和创造力：新奇的幸运艺术》）. I no longer remember everything in it, but its central idea stayed somewhere in the back of my mind: luck and creativity are not entirely random. Chance may arrive unexpectedly, but we can still choose how often we encounter it and whether we are ready to recognize it.</p>

<blockquote>
  <p>Author: <a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>; <a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>Yesterday, while talking with Anne, I found a phrase for this idea: <strong>luck sourcing</strong>—the sourcing of getting lucky.</p>

<p>We usually speak about luck as if it simply happens to us. But perhaps luck can be sourced: through the people we meet, the conversations we enter, the places we spend time, the questions we keep asking, and the strange ideas we decide not to dismiss too quickly. We cannot command a lucky outcome, but we may be able to create more places for luck to come from.</p>

<p>That is the idea I want to explore here.</p>

<h1 id="luck-sourcing">Luck Sourcing</h1>

<p>About six years ago, I started “expanding my contact list.”</p>

<p>At the time, I did not have any clear goal. I just felt that if I knew more people, interesting things would always happen. So I would use all kinds of spare moments to add people, chat, and build connections.</p>

<p>Once, I added dozens of people in a row on a train. I got so absorbed in chatting that I passed my stop without realizing it.</p>

<p>Later, I worked as a VC intern and was responsible for deal sourcing. In theory, I should have been constantly looking for projects, researching founders, and building investment leads.</p>

<p>But in practice, I would often get halfway through my work and start chatting with friends I already knew. On the surface, this looked like slacking off; but many new projects, collaborations, and pieces of information came precisely from these conversations with no clear purpose.</p>

<p>I later realized that this behavior could be given a name, just like deal sourcing.</p>

<p>It should be called <strong>luck sourcing</strong>.</p>

<p>Deal sourcing is actively looking for deals.
Talent sourcing is actively looking for talent.
Luck sourcing is actively looking for luck.</p>

<p>Many people understand luck as a random event: waiting for an opportunity to appear, waiting for a benefactor to contact them, waiting for something good to happen by chance.</p>

<p>But much of what we call luck can actually be increased.</p>

<p>The more people you know, the more frequently you communicate, the more communities you participate in, and the more often you are exposed to new information, the higher the probability that unexpected opportunities will appear.</p>

<p>Luck sourcing does not guarantee success.</p>

<p>It simply lets you appear more frequently in places where “good luck might happen.”</p>

<p>You are not waiting for luck.</p>

<p>You are sourcing it.</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="luck" /><category term="creativity" /><category term="ideas" /><summary type="html"><![CDATA[A long time ago, I came across a book called Chase, Chance, and Creativity: The Lucky Art of Novelty（《追逐、机遇和创造力：新奇的幸运艺术》）. I no longer remember everything in it, but its central idea stayed somewhere in the back of my mind: luck and creativity are not entirely random. Chance may arrive unexpectedly, but we can still choose how often we encounter it and whether we are ready to recognize it.]]></summary></entry><entry><title type="html">你几个月前装的那个 GitHub App，权限可能还全开着，去看看</title><link href="https://koutian.is-a.dev/zh/posts/2026/08/github-app-permissions-quietly-pile-up/" rel="alternate" type="text/html" title="你几个月前装的那个 GitHub App，权限可能还全开着，去看看" /><published>2026-08-04T00:00:00+00:00</published><updated>2026-08-04T00:00:00+00:00</updated><id>https://koutian.is-a.dev/zh/posts/2026/08/github-app-permissions-quietly-pile-up-zh</id><content type="html" xml:base="https://koutian.is-a.dev/zh/posts/2026/08/github-app-permissions-quietly-pile-up/"><![CDATA[<p>我当时正在配置第二份 Cloudflare 备份——核实一份数据、再搭建另一份——帮我干活的 AI Agent 顺手提醒了我一件不相关的事：在列出哪些东西能访问我的 GitHub 账号时，有一项授权的范围明显大得不正常。就是这一句提醒，让我花了一个多小时去仔细过了一遍 GitHub 的已安装应用列表，而这一个小时非常值得。</p>

<blockquote>
  <p>作者：<a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>；<a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>好在没有任何东西被滥用，没有数据泄露，也没有东西被删除。但我发现那些安安静静躺在那里、早已获得授权的东西，是一个足够普遍的问题，值得写清楚：GitHub 上有两类不同的第三方访问权限，每一类当初都是因为一个具体理由被授予一次，但每一类如今持有的权限，都远远超出了当初那个理由本身。</p>

<h2 id="两种不同的授权方式区别为什么重要">两种不同的授权方式，区别为什么重要</h2>

<p>GitHub 有两种相关但本质不同的机制，用来让外部工具接触你的账号，而这个区别，直接决定了每一种的危险程度。</p>

<p><strong>GitHub App 安装（Installed App）</strong> 的范围限定在你选择的具体仓库上，权限也是具体的，比如”写入仓库内容”或”读取 Pull Request”。你可以在 <code class="language-plaintext highlighter-rouge">github.com/settings/installations</code> 查看并收紧这份清单。</p>

<p><strong>OAuth App 授权</strong> 则不同，范围也更大。它不是被限定在一个或几个仓库上，而是可以让某个服务”像你本人一样”行动，使用你账号本身的权限，触达你账号能到达的一切地方，包括你所在的组织。这类授权要单独在 <code class="language-plaintext highlighter-rouge">github.com/settings/applications</code> 查看。</p>

<p>大多数人，包括我自己，偶尔会瞥一眼第一份清单，但几乎从不会去想第二份，因为它在体验上更像是一次性的”登录”步骤，而不是一份持续生效的授权。</p>

<h2 id="我发现了什么">我发现了什么</h2>

<p>我平时用的一个 AI 编程助手，持有一个 OAuth 授权，允许它以我的名义在四个不同的 GitHub 账号上行动，而不只是一个。其中两个是我并不完全掌控的组织。这个授权是我几个月前，为了一个很具体、很窄的任务而给出的一次性授权。而它就这样一直保留着完整的访问范围，安安静静地，因为日常使用这个工具的过程中，从来没有任何东西会主动把”它现在到底还能做什么”摆到我面前。</p>

<p>另外，一个大约六个月前为了某次一次性部署配置而安装的集成应用，至今仍然对它能看到的每一个仓库，持有仓库管理设置、Webhook 和 Pull Request 的完整读写权限。我从来没有回头检查过，这个权限范围是否还符合我现在实际需要它做的事情，因为没有任何东西提醒我这么做。</p>

<p>这两件事都不是因为什么恶意行为造成的。它们只是安静地待在那里，什么坏事都没做——就像抽屉里一把备用钥匙，早已过了当初配它的理由，却还一直躺在那儿。真正的风险，不在于这两个应用主动做了什么出格的事，而在于：在我真正去看之前，我完全不知道，究竟有什么东西能够操作我的账号，以及这个操作范围到底有多大。</p>

<h2 id="为什么这在-ai-agent-时代更值得重视">为什么这在 AI Agent 时代更值得重视</h2>

<p>这种模式其实在 AI Agent 出现之前就一直存在：过度宽泛的 OAuth 授权范围、长期未被重新审视的应用安装，从来都是真实存在的风险。AI Agent 真正改变的是你积累这些授权的速度，以及你有多容易停止注意到它们。</p>

<p>当你借助 Agent 快速推进工作、接入各种集成、连接各种服务、搭建自动化部署时，”授予权限”往往只是完成另一件事过程中一次摩擦很小的点击。你为了让任务继续推进而授权了某个应用，任务完成了，而这份授权本身，却比它存在的理由活得更久。把这个过程乘以你在一两年时间里连接过的每一个集成应用，”已安装应用”和”已授权应用”这两个页面，就变成了一片安静却在不断扩大的暴露面，没有人在主动盯着它。</p>

<p>一个帮你快速搭建东西的 Agent，同样也会在你让它这么做的时候，毫不犹豫地替你点掉一个 OAuth 授权确认页面，因为从它的角度看，这只是为了让下一步能继续推进。它没有任何独立的理由去提醒你：”顺便说一句，这个授权六个月后可能还会原封不动地留在这里，没人重新审视过。”这恰恰是人该做的事，也恰恰是最容易被跳过的一步。</p>

<h2 id="到底该检查什么">到底该检查什么</h2>

<p>这不需要多深的安全专业知识，只需要两个页面，外加几分钟时间。</p>

<ol>
  <li>打开 <code class="language-plaintext highlighter-rouge">github.com/settings/installations</code>。逐一检查每个已安装的应用能访问哪些仓库、持有哪些权限。凡是超出这个应用实际用途所需的范围，都应该收紧。</li>
  <li>打开 <code class="language-plaintext highlighter-rouge">github.com/settings/applications</code>。这是大多数人会跳过的一页。检查每一个 OAuth 授权到底覆盖了多少个账号和组织，而不只是仓库。凡是你不清楚记得最近授予过的、以及授权范围明显超出当初那个具体任务的，都应该撤销。</li>
  <li>把”这个我几个月前就配好了，一直用得好好的”当作一个”该去检查一下”的信号，而不是”可以放着不管”的理由。”能正常工作”和”权限范围仍然合适”是两件不同的事，而其中只有一件，在你不去看的情况下是看不出来的。</li>
</ol>

<p>在一个我自认为已经相当留意的账号上，我在不到一个小时内，就在这两类授权里都找到了真实、可以修正的过度授权问题。工具本身不是问题所在，一个从未被重新审视过的默认设置才是。</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="AI Agents" /><category term="GitHub" /><category term="Security" /><category term="Cloud Backup" /><summary type="html"><![CDATA[我当时正在配置第二份 Cloudflare 备份——核实一份数据、再搭建另一份——帮我干活的 AI Agent 顺手提醒了我一件不相关的事：在列出哪些东西能访问我的 GitHub 账号时，有一项授权的范围明显大得不正常。就是这一句提醒，让我花了一个多小时去仔细过了一遍 GitHub 的已安装应用列表，而这一个小时非常值得。]]></summary></entry><entry><title type="html">The GitHub Apps You Installed Months Ago Still Have Full Access. Go Check.</title><link href="https://koutian.is-a.dev/posts/2026/08/github-app-permissions-quietly-pile-up/" rel="alternate" type="text/html" title="The GitHub Apps You Installed Months Ago Still Have Full Access. Go Check." /><published>2026-08-04T00:00:00+00:00</published><updated>2026-08-04T00:00:00+00:00</updated><id>https://koutian.is-a.dev/posts/2026/08/github-app-permissions-quietly-pile-up</id><content type="html" xml:base="https://koutian.is-a.dev/posts/2026/08/github-app-permissions-quietly-pile-up/"><![CDATA[<p>I was in the middle of setting up a second Cloudflare backup, verifying one dataset and configuring another, when the AI agent helping me flagged something unrelated: while listing what had access to my GitHub account, one authorization stood out as unusually broad. That single flag turned into an hour of reviewing GitHub’s installed-apps list, and it was worth every minute.</p>

<blockquote>
  <p>Author: <a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>; <a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>Nothing had been misused. No data was leaked, nothing was deleted. But what I found sitting there, quietly authorized, is common enough that it is worth writing down plainly: two separate categories of GitHub third-party access, each granted once for a specific reason, each still holding far more reach than that reason justified.</p>

<h2 id="two-different-kinds-of-access-and-why-the-difference-matters">Two different kinds of access, and why the difference matters</h2>

<p>GitHub has two related but distinct mechanisms for letting outside tools touch your account, and the distinction matters a lot for how dangerous each one is.</p>

<p>A <strong>GitHub App installation</strong> is scoped to specific repositories you choose, with specific permissions like “write to repository contents” or “read pull requests.” You can review and narrow this list at <code class="language-plaintext highlighter-rouge">github.com/settings/installations</code>.</p>

<p>An <strong>OAuth App authorization</strong> is different and broader. Rather than being scoped to one or more repositories, it can grant a service the ability to act as you, using your own account’s permissions, across everywhere your account reaches, including organizations. You review this separately at <code class="language-plaintext highlighter-rouge">github.com/settings/applications</code>.</p>

<p>Most people, myself included, glance at the first list occasionally and rarely think about the second at all, because it is presented as a one-time login step rather than an ongoing grant of access.</p>

<h2 id="what-i-found">What I found</h2>

<p>An AI coding assistant I use had an OAuth authorization that let it act on my behalf across four separate GitHub accounts, not one. Two were organizations I do not solely control. I had granted this once, for a narrow, specific task, months earlier. It had kept that full reach the entire time since, silently, because nothing about using the tool day to day ever surfaced the scope of what it could still do.</p>

<p>Separately, a deployment integration I had installed roughly six months prior, for a specific one-time setup, still held full read-and-write access to repository administration settings, webhooks, and pull requests, on every repository it could see. I had never gone back to check whether that scope still matched what I actually needed from it, because nothing prompted me to.</p>

<p>Neither of these was caused by anything malicious. They were both just sitting there, doing nothing wrong, the way a spare key sits in a drawer long after the reason you cut it has passed. The risk was not that either app was actively misbehaving. The risk was that I genuinely did not know, until I looked, what could act on my account and how far that reach extended.</p>

<h2 id="why-this-matters-more-with-ai-agents-in-the-loop">Why this matters more with AI agents in the loop</h2>

<p>This same pattern predates AI agents entirely; over-broad OAuth scopes and stale app installations have always been a real risk. What changes with AI agents is the pace at which you accumulate these grants, and how easy it becomes to stop noticing.</p>

<p>When you are moving fast with an agent helping you wire up integrations, connect services, and automate deployments, granting access becomes a small, low-friction click in the middle of getting something else done. You authorize an app to keep the task moving, the task finishes, and the authorization simply outlives its reason for existing. Multiply that by every integration you have ever connected over a year or two of building things, and the two installed-apps pages become a quiet, growing surface that nobody is actively watching.</p>

<p>An agent helping you build things fast is also an agent that will happily click through an OAuth consent screen if you tell it to, because from its point of view that is just unblocking the next step. It has no independent reason to flag “and by the way, this grant will still be here in six months, un-reviewed.” That is squarely a human’s job, and it is an easy one to skip.</p>

<h2 id="what-to-actually-check">What to actually check</h2>

<p>This does not require deep security expertise, just two pages and a few minutes.</p>

<ol>
  <li>Go to <code class="language-plaintext highlighter-rouge">github.com/settings/installations</code>. For each installed app, check what repositories it can reach and what permissions it has. Narrow anything broader than the app actually needs for what you use it for.</li>
  <li>Go to <code class="language-plaintext highlighter-rouge">github.com/settings/applications</code>. This is the one people skip. Check every OAuth authorization for how many accounts and organizations it reaches, not just repositories. Revoke anything you do not clearly remember granting recently, and anything whose reach is broader than the task you originally granted it for.</li>
  <li>Treat “I set this up months ago and it still works” as a reason to check it, not a reason to leave it alone. Working fine and being appropriately scoped are two different things, and only one of them is visible without looking.</li>
</ol>

<p>I found real, correctable over-reach in both categories in under an hour, on an account I thought I paid reasonable attention to. The tools themselves were not the problem. The unreviewed default was.</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="AI Agents" /><category term="GitHub" /><category term="Security" /><category term="Cloud Backup" /><summary type="html"><![CDATA[I was in the middle of setting up a second Cloudflare backup, verifying one dataset and configuring another, when the AI agent helping me flagged something unrelated: while listing what had access to my GitHub account, one authorization stood out as unusually broad. That single flag turned into an hour of reviewing GitHub’s installed-apps list, and it was worth every minute.]]></summary></entry><entry xml:lang="zh"><title type="html">一个车位 12.7 万美元：UT Austin 的 Mulva Hall 到底贵在哪里？</title><link href="https://koutian.is-a.dev/zh/posts/2026/08/mulva-hall-126887-dollar-parking-space/" rel="alternate" type="text/html" title="一个车位 12.7 万美元：UT Austin 的 Mulva Hall 到底贵在哪里？" /><published>2026-08-02T00:00:00+00:00</published><updated>2026-08-02T00:00:00+00:00</updated><id>https://koutian.is-a.dev/zh/posts/2026/08/mulva-hall-126887-dollar-parking-space-zh</id><content type="html" xml:base="https://koutian.is-a.dev/zh/posts/2026/08/mulva-hall-126887-dollar-parking-space/"><![CDATA[<p>十二万六千八百八十七美元，在 Austin 已经是住房级别的钱；在 UT Austin 的新商学院大楼里，它只对应一个停车位。</p>

<blockquote>
  <p>作者：<a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>；<a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>这不是我把混凝土、白线和一块车牌随手估成了天价。数字来自 <a href="https://www.utsystem.edu/sites/default/files/offices/board-of-regents/board-meetings/board-minutes/11-2024Meeting1248.pdf">UT System Board of Regents 2024 年 11 月会议纪要</a>：Mulva Hall 项目有 164 个停车位，停车库的 building cost 是 20,809,541 美元，官方给出的 <strong>building cost per car 是 126,887 美元</strong>。</p>

<p>计算也完全对得上：</p>

\[\frac{20{,}809{,}541}{164}=126{,}887.45\]

<p>问题不是除法，而是：为什么一所公立大学会为 164 个车位配置超过 2,000 万美元的建设成本？</p>

<h2 id="这不是普通的地下车库比较贵">这不是普通的“地下车库比较贵”</h2>

<p>地下停车库当然不能和郊区地面停车场直接比较。它需要基坑开挖、支护、防水、排水、通风、消防、照明、坡道、电梯和楼梯。Mulva Hall 的停车库还是两层地下结构；<a href="https://construction.utexas.edu/construction-advisory/USB-Underground-Infrastructure">UT Austin 的施工说明</a>也显示，项目周围需要接入新的给排水、电力、通信和热力管线。</p>

<p>但 UT 自己的基准表已经把“地下车库很贵”这个解释推到了极限。会议纪要给出的比较是：</p>

<table>
  <thead>
    <tr>
      <th>项目或基准</th>
      <th style="text-align: right">每车位建设成本</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Mulva Hall 停车库</td>
      <td style="text-align: right">$126,887</td>
    </tr>
    <tr>
      <td>Dallas 地区中位数</td>
      <td style="text-align: right">$29,042</td>
    </tr>
    <tr>
      <td>Houston 地区中位数</td>
      <td style="text-align: right">$29,518</td>
    </tr>
    <tr>
      <td>UT System 其他项目高四分位</td>
      <td style="text-align: right">$34,845</td>
    </tr>
    <tr>
      <td>全美项目高四分位</td>
      <td style="text-align: right">$61,365</td>
    </tr>
  </tbody>
</table>

<p>也就是说，Mulva Hall 的数字大约是 Dallas 中位数的 <strong>4.37 倍</strong>、Houston 中位数的 <strong>4.30 倍</strong>、UT System 其他项目高四分位的 <strong>3.64 倍</strong>，甚至是全美项目高四分位的 <strong>2.07 倍</strong>。</p>

<p>这不是拿地面停车场去碰瓷地下车库。最关键的异常，正是由 UT System 自己选择的比较组和官方表格呈现出来的。</p>

<h2 id="2081-million-美元里到底装了什么">20.81 million 美元里到底装了什么？</h2>

<p>一种可能的解释是，这个“停车库成本”不只服务于停车。</p>

<p>地下结构可能承担整栋高层建筑的荷载；基础、柱网、转换结构、施工支护或机电空间，也可能在会计上分配到停车库包件。如果 20.81 million 美元包含大量塔楼基础和非停车空间，那么“每个车位 126,887 美元”会夸大真正为了停车功能付出的边际成本。</p>

<p>但公开的项目成本表目前无法证明这一点。相反，它还把许多相关费用另列为独立科目：</p>

<ul>
  <li>Site Development：$11.62 million；</li>
  <li>Institutionally Managed Work：$20.22 million；</li>
  <li>Architectural/Design Services：$17.81 million；</li>
  <li>Other Professional Fees：$10.03 million；</li>
  <li>Project Contingency：$12.75 million。</li>
</ul>

<p>因此，不能笼统地说管线迁移、场地施工、设计费和所有杂费都塞进了停车库的 20.81 million 美元。至少在董事会批准的预算表面上，它们有自己的预算科目。</p>

<p>这也意味着，对这个数字最诚实的表述不是“每画一个停车格花了十二万美元”，而是：</p>

<blockquote>
  <p>UT 把 20.81 million 美元的停车库 building cost 分摊到 164 个车位后，每个车位承担 126,887 美元；公开资料尚不足以判断其中有多少实际属于塔楼基础或其他非停车功能。</p>
</blockquote>

<h2 id="这笔钱不是靠停车费轻松赚回来的">这笔钱不是靠停车费轻松赚回来的</h2>

<p>如果把每个车位当成一项独立投资，并且荒谬地忽略利息、保安、电力、通风、维修、保险和空置率，那么静态回收期是：</p>

<table>
  <thead>
    <tr>
      <th>每月停车收入</th>
      <th style="text-align: right">收回 $126,887 所需时间</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>$200</td>
      <td style="text-align: right">约 52.9 年</td>
    </tr>
    <tr>
      <td>$300</td>
      <td style="text-align: right">约 35.2 年</td>
    </tr>
    <tr>
      <td>$500</td>
      <td style="text-align: right">约 21.1 年</td>
    </tr>
  </tbody>
</table>

<p>现实中的回收期只会更长。停车库显然不是一个靠停车费独立盈利的普通项目，而是整栋旗舰教学楼的配套设施。</p>

<p>资金结构也说明了这一点。Mulva Hall 总项目成本是 <strong>425 million 美元</strong>：150 million 美元来自捐赠，50 million 美元来自 Available University Fund，另外 225 million 美元来自 Revenue Financing System 债券。董事会文件明确写着，这 225 million 美元债务将由 <strong>designated tuition 和 parking revenues</strong> 偿还。McCombs 自己的<a href="https://news.mccombs.utexas.edu/news/mulva-hall-project-underway/">项目介绍</a>也确认了 425 million 美元总价和 225 million 美元融资规模。</p>

<p>这不代表每一美元学费都会直接流入车库，也不代表停车收入只用于偿还车库。但它说明，车库不是与学校其余财务隔离的私人投资。成本最终进入了大学整体的融资和资源配置。</p>

<h2 id="最值得追问的不是有没有人偷钱">最值得追问的不是“有没有人偷钱”</h2>

<p>一个离群值不是腐败证据。它可能来自合理但尚未公开的工程与会计分配，也可能来自一个极其昂贵、但被决策者认为值得的设计选择。</p>

<p>真正需要公开的是：</p>

<ol>
  <li>两层地下空间的总面积，以及其中真正用于停车的面积；</li>
  <li>车库每平方英尺造价，而不只是每个车位造价；</li>
  <li>基坑、基础、主体结构、防水和机电各自的金额；</li>
  <li>哪些成本只服务停车，哪些同时服务上方建筑；</li>
  <li>Construction Manager-at-Risk 合同下的 GMP 和 schedule of values；</li>
  <li>学校为何选择建设 164 个如此昂贵的地下车位，以及评估过哪些替代方案。</li>
</ol>

<p>替代方案并不神秘：少建车位、使用现有车库、在校园外围提供低成本停车、加强班车与公共交通补贴，或者把地下空间留给更难在别处替代的教学和设备用途。每一种方案都有代价，但至少应该和 20.81 million 美元放在同一张表上比较。</p>

<h2 id="结论">结论</h2>

<p>“一个车位 12.7 万美元”听起来像一句笑话，但它确实是 UT System 自己公布的预算指标。</p>

<p>它不自动证明有人偷钱，也不能被简化为两条白线值十二万美元；但它已经足以构成一个明确的成本离群值。尤其当官方数字比全美高成本项目的高四分位还高一倍时，“地下停车比较贵”不再是充分解释。</p>

<p>UT 应该公布成本分配，让公众看清楚：这 20.81 million 美元究竟在购买 164 个停车位，还是在替整栋 Mulva Hall 承担一部分没有被正确命名的地下工程。</p>

<p>在答案出现以前，最准确的问题不是“为什么一条白线这么贵”，而是：<strong>为什么大学把足以进入 Austin 低价住房价格区间的资本，配置给了每一个地下车位？</strong></p>

<h2 id="主要资料">主要资料</h2>

<ul>
  <li><a href="https://www.utsystem.edu/sites/default/files/offices/board-of-regents/board-meetings/board-minutes/11-2024Meeting1248.pdf">UT System Board of Regents，2024 年 11 月会议纪要：项目资金、成本明细与每车位基准</a></li>
  <li><a href="https://construction.utexas.edu/construction-advisory/USB-Underground-Infrastructure">UT Austin Planning, Design and Construction：Mulva Hall 施工与地下基础设施说明</a></li>
  <li><a href="https://news.mccombs.utexas.edu/news/mulva-hall-project-underway/">McCombs School of Business：Mulva Hall Project Underway</a></li>
</ul>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="UT Austin" /><category term="Mulva Hall" /><category term="parking" /><category term="public finance" /><category term="cost analysis" /><summary type="html"><![CDATA[十二万六千八百八十七美元，在 Austin 已经是住房级别的钱；在 UT Austin 的新商学院大楼里，它只对应一个停车位。]]></summary></entry><entry xml:lang="zh"><title type="html">《入都》：一个二十一岁的人，进京赶考前写下的十首诗</title><link href="https://koutian.is-a.dev/zh/posts/2026/07/li-hongzhang-rudu-poem-startup-parallel/" rel="alternate" type="text/html" title="《入都》：一个二十一岁的人，进京赶考前写下的十首诗" /><published>2026-07-30T00:00:00+00:00</published><updated>2026-07-30T00:00:00+00:00</updated><id>https://koutian.is-a.dev/zh/posts/2026/07/li-hongzhang-rudu-poem-startup-parallel-zh</id><content type="html" xml:base="https://koutian.is-a.dev/zh/posts/2026/07/li-hongzhang-rudu-poem-startup-parallel/"><![CDATA[<p>丈夫只手把吴钩，意气高于百尺楼。一万年来谁著史，三千里外欲封侯。</p>

<blockquote>
  <p>作者：<a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>；<a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>道光二十三年（1843 年），李鸿章 21 岁，那年入选优贡，奉父亲李文安之命，从安徽老家入京，准备第二年的顺天乡试。入京路上，他写下《入都》十首。我最近反复想起其中的第一首和第二首，觉得跟我现在的状况很像。</p>

<h2 id="第一首意气">第一首：意气</h2>

<blockquote>
  <p>丈夫只手把吴钩，意气高于百尺楼。
一万年来谁著史，三千里外欲封侯。
定将捷足随途骥，那有闲情逐水鸥。
笑指泸沟桥畔月，几人从此到瀛洲？</p>
</blockquote>

<p>21 岁的人，从安徽走三千里去北京，脑子里想的是”一万年来谁著史”。这句话的野心大到有点可笑，可正是这种可笑，才配得上一个刚要出发的年轻人。他还没有任何战功，没有任何官职，甚至连乡试都还没考。他有的只是”意气”，以及一句”定将捷足随途骥，那有闲情逐水鸥”：我要跟着最快的马跑，没空理会水面上悠闲的鸥鸟。</p>

<p>这句话我读了很多遍。跟着快马跑的人，是没有资格去羡慕水鸥的悠闲的。这两种活法不能既要又要。</p>

<h2 id="第二首疲惫和自我怀疑">第二首：疲惫和自我怀疑</h2>

<blockquote>
  <p>频年伏枥困红尘，悔煞驹光二十春。
马足出群休恋栈，燕辞故垒更图新。
遍交海内知名士，去访京师有道人。
即此可求文字益，胡为抑郁老吾身！</p>
</blockquote>

<p>同一个人，同一次进京，第二首的调子却完全变了。”频年伏枥困红尘，悔煞驹光二十春”：这些年困在尘世里、像伏在槛里的马一样，二十年的光景，让他后悔。</p>

<p>这才是真实的年轻人。不是只有第一首的意气风发，还有第二首的自我怀疑：我是不是浪费了前二十年？我现在才出发，是不是已经晚了？</p>

<p>但紧接着，他给自己一个答案：”马足出群休恋栈，燕辞故垒更图新。”马要跑出群，就不能恋着旧的马厩；燕子要换新窝，就要先辞别旧的巢。然后是”遍交海内知名士，去访京师有道人”：去结交天下有名的人，去拜访京城真正有道行的人。最后一句是自己给自己的当头棒喝：”即此可求文字益，胡为抑郁老吾身！”这条路本身就能带来长进，为什么还要抑郁到老？</p>

<h2 id="为什么这跟我现在的状况很像">为什么这跟我现在的状况很像</h2>

<p>我现在同时在读 PhD、做创业、转专业。这几件事放在一起，很像李鸿章 21 岁那年：手里没有任何已经兑现的战功，只有”意气”和一个”一万年来谁著史”式的野心。第一首里那种”定将捷足随途骥，那有闲情逐水鸥”，我太熟悉了。跟着快马跑的时候，确实没有力气去羡慕别人悠闲的生活。</p>

<p>但第二首才是我更需要记住的部分。”频年伏枥困红尘，悔煞驹光二十春”，这种情绪我也有。会怀疑自己是不是走得太慢，是不是该更早开始，是不是已经把一些年份浪费在了不该浪费的地方。</p>

<p>李鸿章给自己的解法，不是回头去弥补过去，而是”马足出群休恋栈，燕辞故垒更图新”：既然要跑出群，就不要恋着旧马厩；既然要换巢，就干脆去建新的。然后去交真正的人，去访真正有道行的人，把这条新路本身当作长进的来源。</p>

<p>这十首诗写在他 21 岁，还什么都没有做成的时候。后来的事，历史书里都有，不需要我在这里重复。我更在意的是这十首诗写下的那个时间点：一个人还没有任何证明，却已经决定”意气高于百尺楼”，同时也已经在自我怀疑”悔煞驹光二十春”。这两种情绪同时存在，并不矛盾。</p>

<p>意气和怀疑，本来就应该同时出现在一个正在出发的人身上。区别只在于，出发之后，是回头去恋旧马厩，还是”休恋栈”，继续往前跑。</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="reflection" /><category term="startup" /><category term="创业" /><category term="history" /><category term="poetry" /><summary type="html"><![CDATA[丈夫只手把吴钩，意气高于百尺楼。一万年来谁著史，三千里外欲封侯。]]></summary></entry><entry><title type="html">Benchmark Radar: An Evidence-First Daily Radar for AI Benchmarks</title><link href="https://koutian.is-a.dev/posts/2026/07/benchmark-radar/" rel="alternate" type="text/html" title="Benchmark Radar: An Evidence-First Daily Radar for AI Benchmarks" /><published>2026-07-28T00:00:00+00:00</published><updated>2026-07-28T00:00:00+00:00</updated><id>https://koutian.is-a.dev/posts/2026/07/benchmark-radar</id><content type="html" xml:base="https://koutian.is-a.dev/posts/2026/07/benchmark-radar/"><![CDATA[<p>AI benchmarks are now appearing faster than any researcher can evaluate them, so I built a radar that makes discovery daily, transparent, and auditable.</p>

<blockquote>
  <p>Author: <a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>; <a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p><a href="https://github.com/ktwu01/benchmark-radar">Benchmark Radar</a> is an open-source system that looks for newly released AI benchmarks, evaluation methods, datasets, leaderboards, and data-quality research every day. It gathers records from primary and structured sources, removes duplicates, classifies them with a visible taxonomy, ranks them with explainable signals, and publishes both a daily GitHub Issue and a <a href="https://ktwu01.github.io/benchmark-radar/">cumulative dashboard</a>.</p>

<p>I am about to begin a new adventure, and an important part of my job will be searching for exactly this kind of information. I need to know which benchmarks, evaluation methods, datasets, and data-quality ideas are emerging—and I need a reliable way to keep up with them. That is why I needed Benchmark Radar, and why I built it.</p>

<p>The project addresses a simple problem: following AI evaluation work has become a research task of its own.</p>

<p>A new benchmark may first appear as an arXiv paper, an OpenReview submission, a GitHub repository, or a Hugging Face dataset. Its leaderboard may arrive later. The same artifact may then be discussed by several secondary sources, each with a slightly different title and description. A normal feed gives all of these records equal visual weight and leaves the reader to reconstruct what is new, what is duplicated, and what has real evidence behind it.</p>

<p>Benchmark Radar is my attempt to make that process more systematic.</p>

<h2 id="what-the-radar-watches">What the radar watches</h2>

<p>The default taxonomy covers four connected areas:</p>

<ul>
  <li>new AI and LLM benchmarks, challenge sets, and evaluation suites;</li>
  <li>evaluation frameworks, judge models, safety and capability evaluations, and leaderboards;</li>
  <li>public datasets, preference data, synthetic data, and other data releases;</li>
  <li>work on contamination, leakage, provenance, deduplication, annotation quality, and related data-quality problems.</li>
</ul>

<p>The collector queries arXiv, OpenReview, Hugging Face, GitHub, GitHub Releases, Semantic Scholar, OpenAlex, and Brave Search. Optional sources can fail or be unavailable without stopping the daily report. When that happens, the source-health table shows the missing coverage instead of quietly pretending that the run was complete.</p>

<p>That detail matters. A trend line built from changing source coverage can look like momentum even when it is only a collection artifact.</p>

<h2 id="evidence-before-novelty">Evidence before novelty</h2>

<p>The central design choice is that every ranked item carries four visible component scores:</p>

<ul>
  <li><strong>Relevance</strong> measures how closely the record matches the benchmark, evaluation, dataset, and data-quality taxonomy.</li>
  <li><strong>Evidence</strong> rewards primary or structured sources, authorship signals, and corroborating artifacts.</li>
  <li><strong>Recency</strong> captures how recently the work was published or materially updated.</li>
  <li><strong>Adoption</strong> uses signals such as stars, downloads, likes, or citations on a logarithmic scale.</li>
</ul>

<p>The default priority score is:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>0.35 relevance + 0.20 evidence + 0.20 recency + 0.25 adoption
</code></pre></div></div>

<p>The result is reported on a 0–100 scale. The weights and component bands live in the same code that scores the records, and that rubric is also exported to the dashboard. Clicking a priority score shows the record’s component values and the weighted calculation behind its rank.</p>

<p>This does not make the ranking objectively correct. It makes it inspectable.</p>

<p>Benchmark Radar is a triage system, not a scientific quality judge. It cannot tell whether a benchmark measures the capability its authors claim, whether its test set is contaminated, or whether a leaderboard result will reproduce. What it can do is show why an artifact reached the daily list and give the reader a cleaner evidence trail for deciding what deserves a closer look.</p>

<h2 id="counting-artifacts-instead-of-mentions">Counting artifacts instead of mentions</h2>

<p>Deduplication is one of the less visible but more important parts of the system.</p>

<p>The cumulative corpus resolves entities from exact identifiers such as DOI, arXiv, OpenReview, GitHub, and Hugging Face IDs. It does not silently merge two similarly titled projects with fuzzy matching. Observations remain connected to their discovered artifacts, and the dashboard can expand multiple records without turning every mention into a new benchmark.</p>

<p>The same caution applies to historical comparisons. Trend calculations only compare snapshots collected with the same report limit and the same connector-coverage signature. Incomplete days remain visible and are labeled rather than being smoothed away.</p>

<p>That sounds conservative because it is. A radar is useful only if an increase in the chart means more than “the crawler behaved differently today.”</p>

<h2 id="a-daily-research-artifact">A daily research artifact</h2>

<p>The entire workflow runs in GitHub Actions at 12:15 UTC and can also be triggered manually. A run collects and scores records, validates a versioned daily snapshot, updates the date-filtered GitHub Issue, and rebuilds the cumulative dashboard.</p>

<p>Each snapshot records the selection funnel from fetched records to deduplicated, qualified, and published items. It also preserves retrieval time, parser version, and a fingerprint of the upstream payload without publishing credentials or raw API responses.</p>

<p>The repository therefore contains more than a generated webpage. It contains the code, configuration, schemas, and versioned observations needed to inspect how the feed was produced.</p>

<p>You can run it locally with Python 3.11 or later:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>git clone https://github.com/ktwu01/benchmark-radar.git
<span class="nb">cd </span>benchmark-radar
python3 <span class="nt">-m</span> venv .venv
<span class="nb">source</span> .venv/bin/activate
python <span class="nt">-m</span> pip <span class="nb">install</span> <span class="nt">-e</span> <span class="s1">'.[dev]'</span>
benchmark-radar
</code></pre></div></div>

<p>The main outputs are a Markdown report, a machine-readable evidence snapshot, a dated corpus snapshot, and the browser-ready data used by the dashboard.</p>

<h2 id="why-i-built-it-in-public">Why I built it in public</h2>

<p>AI evaluation needs better discovery infrastructure, but it also needs skepticism toward the infrastructure doing the discovery.</p>

<p>Ranking systems easily hide editorial decisions inside a score. Data pipelines easily hide missing sources behind a clean interface. Cumulative dashboards easily count attention as evidence and repeated mentions as new activity. Benchmark Radar does not eliminate those risks, but it tries to expose them in the product itself: visible components, source-health warnings, versioned schemas, explicit limits, deterministic rebuilds, and links back to the discovered records.</p>

<p>The project is inspired by <a href="https://github.com/duanyytop/agents-radar">agents-radar</a>, with its sources and scoring redesigned for benchmark and AI-data research. It is released under the MIT License, and both the <a href="https://github.com/ktwu01/benchmark-radar">source code</a> and <a href="https://ktwu01.github.io/benchmark-radar/">live dashboard</a> are public.</p>

<p>The goal is not to produce one more leaderboard. It is to make the fast-moving landscape around benchmarks a little easier to inspect—and a little harder to misread.</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="AI" /><category term="Benchmarks" /><category term="Evaluation" /><category term="Datasets" /><category term="Open Source" /><summary type="html"><![CDATA[AI benchmarks are now appearing faster than any researcher can evaluate them, so I built a radar that makes discovery daily, transparent, and auditable.]]></summary></entry><entry><title type="html">我为较弱的模型做了这些 Skills，后来为什么删掉了大部分</title><link href="https://koutian.is-a.dev/zh/posts/2026/07/skills-i-built-for-weaker-models/" rel="alternate" type="text/html" title="我为较弱的模型做了这些 Skills，后来为什么删掉了大部分" /><published>2026-07-26T00:00:00+00:00</published><updated>2026-07-26T00:00:00+00:00</updated><id>https://koutian.is-a.dev/zh/posts/2026/07/skills-i-built-for-weaker-models-zh</id><content type="html" xml:base="https://koutian.is-a.dev/zh/posts/2026/07/skills-i-built-for-weaker-models/"><![CDATA[<p>这周，我给自己的 Claude Code 配置做了一次健康检查，发现里面有 174 个自定义 Skill，其中 124 个我一次都没有调用过。它们并不是失败品。大部分只是我为当时需要额外支撑的模型搭起的脚手架，而模型早已不再需要它们，我却一直把它们留到了现在。</p>

<blockquote>
  <p>作者：<a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>；<a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>在 Claude Code 里，一个 Skill 就是一个文件夹，里面有一份 Markdown 文件，告诉 Agent 如何完成某项具体任务。每次会话开始时，所有已安装 Skill 的名称和描述都会进入模型的 Context，让 Agent 知道自己可以调用什么。只有在 Skill 被调用时，它的正文才会载入。这种设计很好，因为一百个 Skill 占用的是一百条描述，而不是一百套完整流程。</p>

<p>问题在于，描述也不是免费的。在我动手清理前，这份 Skill 列表已经增长到大约 13,800 个 Token。Claude Code 为 Skill 列表预留的空间约占 Context Window 的百分之一，超过以后就会截断内容，路由效果也会下降。我的列表已经达到预算的约七倍。那些我真正使用的 Skill，反而因为大量不用的 Skill 而变得更难被 Agent 找到。</p>

<h2 id="这些-skills-当初在补偿什么">这些 Skills 当初在补偿什么</h2>

<p>回头阅读那些从未调用过的 Skill，我看到了一个规律。它们可以分成几组，而且几乎每一组都在绕过某种具体的能力短板。</p>

<p>最大的一组是流程拆解。<code class="language-plaintext highlighter-rouge">paper-plan</code>、<code class="language-plaintext highlighter-rouge">research-pipeline</code>、<code class="language-plaintext highlighter-rouge">experiment-plan</code> 和 <code class="language-plaintext highlighter-rouge">dse-loop</code> 之类的 Skill，负责把一个多步骤流程固定下来。先写大纲，再写方法，然后写结果，最后对照论点检查图表。我编写这些 Skill，是因为早期模型常常开局很好，到了第四步就丢了主线。Skill 相当于一份外部记忆，替模型保存它自身无法持续掌握的流程。</p>

<p>第二组用来约束输出形态，包括 <code class="language-plaintext highlighter-rouge">formal-docs-affirmative-prose</code>、<code class="language-plaintext highlighter-rouge">sci-writing-prose-rules</code> 和 <code class="language-plaintext highlighter-rouge">figure-never-annotation-caption</code>。它们编码了那些我已经厌倦反复强调的规则：不要含糊其词，不要把 Caption 放进图里，不要在方法部分使用营销语言。每一个 Skill 的存在，都是因为如果没有东西把模型约束住，它就会漂回一种通用的表达方式。</p>

<p>第三组是验证脚手架，包括 <code class="language-plaintext highlighter-rouge">result-to-claim</code>、<code class="language-plaintext highlighter-rouge">sci-validation-checklist</code> 和 <code class="language-plaintext highlighter-rouge">sci-geoscience-metric-claims</code>。它们迫使 Agent 把论文里的一个数字一路追溯到生成该数字的代码。我曾被那些看似可信、其中数字却经不起核查的图表坑过，于是构建了这些 Skill。</p>

<p>第四组与能力完全无关。<code class="language-plaintext highlighter-rouge">dbs-*</code>、<code class="language-plaintext highlighter-rouge">nature-*</code> 和 <code class="language-plaintext highlighter-rouge">ios-*</code> 来自外部 Skill Pack。我安装了它们，试过其中一两个，之后就把其余部分留在那里。它们从来没有补偿过任何能力，只是杂物，因此也最容易被删除。</p>

<h2 id="那些不再值得占据位置的-skills">那些不再值得占据位置的 Skills</h2>

<p>第一组和第二组体现了最有意思的变化。现在的模型无需一份逐步清单，也能规划多步骤工作；只要告诉它采用什么表达风格，它也能保持这种风格。当我今天调用 <code class="language-plaintext highlighter-rouge">paper-plan</code> 时，大多数情况下只是让模型按照我十八个月前选定的顺序，去做它原本就会做的事。</p>

<p>这正是人们容易忽略的成本。一个过时的 Skill 并非中性。它是一条固定指令，会和模型自身的判断竞争，而且它会胜出，因为我把它写成了指令。如果 Skill 里编码的流程还不如模型在没有指令时会采用的方案，那么它一边显得更加严谨，一边主动让输出变差。</p>

<p><code class="language-plaintext highlighter-rouge">result-to-claim</code> 不是在补偿推理能力不足，而是在应对一个事实：Agent 无法独立核实某个数字。它必须查看代码。这是结构性缺口，不是能力缺口。模型能力的提升无法弥合这个缺口，因为再深入的推理也无法验证涉及我数据的结论。因此，这些 Skill 会留下。</p>

<p>基于同样的理由，我还保留了另外几个 Skill。<code class="language-plaintext highlighter-rouge">gemini-review</code> 会把一份产物交给另一个模型评审，它的价值恰恰来自评审者不是产生产物的同一个模型。<code class="language-plaintext highlighter-rouge">wkt</code> 会在实施工作开始前创建 Git Worktree，这是我对仓库的策略选择，不是推理辅助。<code class="language-plaintext highlighter-rouge">atomic-commit</code> 编码了我希望 Git 历史呈现的样子。这些都不是脚手架，而是偏好和结构性事实。模型变得更强，并不会让它们过时。</p>

<h2 id="我最后采用的判断方法">我最后采用的判断方法</h2>

<p>我最终用来区分它们的问题是：一个 Skill 补偿的，是模型无法做到的事情，还是它过去暂时做不到的事情？</p>

<p>编码了外部访问能力的 Skill 应该保留，例如读取文件、调用另一个模型、对照来源核查数字，以及遵循我的仓库特有的约定。模型变得更聪明，也无法自行推导出这些东西。</p>

<p>如果一个 Skill 编码的是模型现在可以自行推断的流程，它就应该被移除。检验这一点的诚实办法，是删掉这个 Skill，再在没有它的情况下完成一次任务。如果输出相同，这个 Skill 只是仪式。如果输出变差，说明它确实承担了工作，应该放回来。</p>

<p>有两件事让这个过程比听上去容易。首先，禁用是可逆的，所以判断错一次只需要改回一项设置。其次，使用次数让第一轮筛选变得机械。一千多次会话以后，调用次数仍然为零的 Skill，不需要再做主观判断。</p>

<h2 id="实际发生了什么变化">实际发生了什么变化</h2>

<p>我禁用了 124 个从未使用的 Skill。列表从大约 13,800 个 Token 降到约 5,700 个。剩下的大部分是内置 Skill，以及一个我仍然启用、但并不适用于当前仓库的 Plugin。</p>

<p>我还把一大段项目特有的评审标准，从一个始终会加载的指令文件移进了只在调用时才加载的 Skill。这是同一条原则的另一个方向：这部分内容值得保留，但没有必要常驻在每一次会话中。</p>

<p>真正值得推广的不是 Token 数量，而是我已经积累了一整年的变通方案，却从未重新检查当初需要绕过的问题是否仍然存在。围绕模型局限构建的工具，应该在局限发生变化时重新审视。记录实际调用情况的计数器一直都在配置文件里。</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="Claude Code" /><category term="AI Agents" /><category term="Tooling" /><category term="Context Engineering" /><category term="Developer Experience" /><summary type="html"><![CDATA[这周，我给自己的 Claude Code 配置做了一次健康检查，发现里面有 174 个自定义 Skill，其中 124 个我一次都没有调用过。它们并不是失败品。大部分只是我为当时需要额外支撑的模型搭起的脚手架，而模型早已不再需要它们，我却一直把它们留到了现在。]]></summary></entry><entry><title type="html">The Skills I Built for Weaker Models, and Why I Deleted Most of Them</title><link href="https://koutian.is-a.dev/posts/2026/07/skills-i-built-for-weaker-models/" rel="alternate" type="text/html" title="The Skills I Built for Weaker Models, and Why I Deleted Most of Them" /><published>2026-07-26T00:00:00+00:00</published><updated>2026-07-26T00:00:00+00:00</updated><id>https://koutian.is-a.dev/posts/2026/07/skills-i-built-for-weaker-models</id><content type="html" xml:base="https://koutian.is-a.dev/posts/2026/07/skills-i-built-for-weaker-models/"><![CDATA[<p>I ran a health check on my Claude Code setup this week and found 174 custom skills, 124 of which I had never invoked once. They were not failures. Most of them were scaffolding I built for a model that needed it, and then kept long after the model stopped needing it.</p>

<blockquote>
  <p>Author: <a href="https://www.linkedin.com/in/ktwu01/">Koutian Wu</a>; <a href="https://github.com/ktwu01/">GitHub: ktwu01</a></p>
</blockquote>

<p>A skill, in Claude Code, is a folder with a markdown file that tells the agent how to do a specific task. The name and description of every installed skill sit in the model’s context at the start of every session, so the agent knows what it can reach for. The body loads only when the skill is invoked. That design is good: it means a hundred skills cost you a hundred descriptions, not a hundred procedures.</p>

<p>The catch is that descriptions are not free. Mine had grown to roughly 13,800 tokens of listing before I touched anything. Claude Code budgets the skill listing at about one percent of the context window, and past that it truncates and routing degrades. I had gone seven times over. The skills I actually used were getting harder for the agent to find because of the ones I did not.</p>

<h2 id="what-the-skills-were-compensating-for">What the skills were compensating for</h2>

<p>Reading back through the ones I never called, a pattern shows up. They fall into a few groups, and almost every group is a workaround for a specific weakness.</p>

<p>The largest group is procedural decomposition. Skills like <code class="language-plaintext highlighter-rouge">paper-plan</code>, <code class="language-plaintext highlighter-rouge">research-pipeline</code>, <code class="language-plaintext highlighter-rouge">experiment-plan</code>, and <code class="language-plaintext highlighter-rouge">dse-loop</code> exist to hold a multi-step process in place. Write the outline, then the methods, then the results, then check the figures against the claims. I wrote them because earlier models would start strong and lose the thread by step four. The skill was an external memory for a process the model could not hold on its own.</p>

<p>The second group is output-shape enforcement. <code class="language-plaintext highlighter-rouge">formal-docs-affirmative-prose</code>, <code class="language-plaintext highlighter-rouge">sci-writing-prose-rules</code>, <code class="language-plaintext highlighter-rouge">figure-never-annotation-caption</code>. These encode rules I was tired of restating. Do not hedge. Do not put the caption inside the figure. Do not use marketing language in a methods section. Each one existed because the model would drift back to a generic register unless something held it.</p>

<p>The third group is verification scaffolding: <code class="language-plaintext highlighter-rouge">result-to-claim</code>, <code class="language-plaintext highlighter-rouge">sci-validation-checklist</code>, <code class="language-plaintext highlighter-rouge">sci-geoscience-metric-claims</code>. These force the agent to trace a number in a paper back to the code that computed it. I built them after being burned by plausible-sounding figures whose numbers did not survive checking.</p>

<p>The fourth group is not about capability at all. <code class="language-plaintext highlighter-rouge">dbs-*</code>, <code class="language-plaintext highlighter-rouge">nature-*</code>, <code class="language-plaintext highlighter-rouge">ios-*</code> came in as external skill packs. I installed them, tried one or two, and left the rest. They were never compensating for anything. They were just clutter, and they are the easiest to justify removing.</p>

<h2 id="the-ones-that-stopped-earning-their-place">The ones that stopped earning their place</h2>

<p>Groups one and two are where the interesting change is. Current models plan multi-step work without a checklist telling them the steps, and they hold a register once you tell them the register. When I invoke <code class="language-plaintext highlighter-rouge">paper-plan</code> today, I am mostly telling the model to do what it was already going to do, in an order I picked eighteen months ago.</p>

<p>That is the cost people miss. A stale skill is not neutral. It is a fixed instruction competing with the model’s own judgment, and it wins, because I wrote it as an instruction. If the procedure encoded in the skill is worse than what the model would do unprompted, the skill actively makes the output worse while appearing to add rigor.</p>

<p>Group three did not stop earning its place, and I want to be precise about why. <code class="language-plaintext highlighter-rouge">result-to-claim</code> is not compensating for weak reasoning. It is compensating for the fact that an agent has no independent access to whether a number is true. It has to go read the code. That is a structural gap, not a capability gap, and no amount of model improvement closes it, because the model cannot verify a claim about my data by thinking harder about it. Those skills stay.</p>

<p>The same logic keeps a few others. <code class="language-plaintext highlighter-rouge">gemini-review</code> routes an artifact to a different model for critique, which is valuable precisely because it is not the same model grading itself. <code class="language-plaintext highlighter-rouge">wkt</code> creates a git worktree before implementation work, which is a policy choice about my repo, not a reasoning aid. <code class="language-plaintext highlighter-rouge">atomic-commit</code> encodes how I want history to look. None of these are scaffolding. They are preferences and structural facts, and a better model does not make them obsolete.</p>

<h2 id="the-test-i-ended-up-with">The test I ended up with</h2>

<p>The distinction I landed on is whether a skill compensates for something the model cannot do, or for something it merely could not do yet.</p>

<p>Skills that encode access to the world stay. Reading a file, calling another model, checking a number against its source, following a convention specific to my repository. The model cannot derive any of these by being smarter.</p>

<p>Skills that encode a procedure the model could now infer should go, and the honest way to test that is to delete the skill and do the task without it. If the output is the same, the skill was ceremony. If it degrades, the skill was carrying something real and belongs back.</p>

<p>Two things make this easier than it sounds. Disabling is reversible, so a wrong call costs one settings edit. And usage counters make the first pass mechanical: a skill with zero invocations across more than a thousand sessions is not a judgment call.</p>

<h2 id="what-actually-changed">What actually changed</h2>

<p>I disabled the 124 unused skills. The listing dropped from roughly 13,800 tokens to about 5,700, and most of what remains is built-in skills and one plugin I had left enabled in a repository where it does not apply.</p>

<p>I also moved a long block of project-specific review criteria out of an always-loaded instructions file and into a skill that loads only when invoked. That is the same principle from the other direction: the content was worth keeping and did not need to be resident in every session.</p>

<p>The part worth generalizing is not the token count. It is that I had accumulated a year of workarounds and never revisited whether the thing being worked around was still there. Tooling built against a model’s limitations should be re-examined when the limitations move, and the counters that tell you which tools you actually reach for are sitting in a config file the whole time.</p>]]></content><author><name>Koutian Wu</name><email>ktwugoat@gmail.com</email><uri>https://ktwu01.github.io</uri></author><category term="Claude Code" /><category term="AI Agents" /><category term="Tooling" /><category term="Context Engineering" /><category term="Developer Experience" /><summary type="html"><![CDATA[I ran a health check on my Claude Code setup this week and found 174 custom skills, 124 of which I had never invoked once. They were not failures. Most of them were scaffolding I built for a model that needed it, and then kept long after the model stopped needing it.]]></summary></entry></feed>