Looking Back from 2030: Connecting Stronger Intelligence to Problems Worth Solving

13 minute read

Published:

If I reopened the 2026 folder in 2030, I would check one thing first: did the work completed with AI change a real problem and leave evidence that others could verify and build upon?

Author: Koutian Wu; GitHub: ktwu01

After writing “Why Did They All Go to AI Labs?”, the camera turns from them back to us.

The previous essay explained why top talent is concentrating in frontier labs. Those labs have expensive compute, internal models, a high concentration of talent, and feedback from deployment. Joining such an organization lets someone run experiments that were previously impossible. It also places them closer to the source of a new generation of cognitive infrastructure.

This essay asks how to put that capability to work in settings where reality can prove its results wrong. Proximity to the source of capability only gives someone more room to act. Deciding how to use that room still requires a different kind of judgment.

I move the calendar forward to 2030 for one reason: to add an acceptance criterion to today’s choices. Four years later, what did they leave behind?

By 2030, the model leaderboards will be forgotten first

If we reopen the 2026 folder in 2030, much of what consumed our attention at the time will be hard to distinguish. One model led for two weeks. One prompting technique gained a few percentage points. One company announced a new capability at a launch event. That information was useful at the time, but it had a short life.

Capabilities spread. Interfaces change. A feature that requires repeated debugging today may become a default setting tomorrow. If someone possesses only a few months of informational advantage, the next model update may erase it.

What remains is more modest: a dataset others still use, a reproducible experiment, a tool that keeps running, or a record that clearly explains a failure. These can outlast the distinction of being first to learn a new term.

Company names and job titles affect access to resources. They show which room someone entered at the time, but they do not show what that person changed inside it. By 2030, we can inspect whether reality changed as a result and whether those who came later could continue the work.

Two axes for judging whether work has entered reality

We can begin judging an AI project along two axes.

The vertical axis measures the strength of the intelligence it can access. Moving upward means getting closer to internal checkpoints, large-scale training, scarce compute, and top research teams. Talent moves rapidly up this axis by joining a frontier lab.

The horizontal axis measures proximity to real-world feedback. Can an experiment disprove the conclusion? Do users actually change their behavior? Does the system make fewer errors after running for several months? Can the cause be traced when something goes wrong? Feedback like this moves the work to the right.

 Weak real-world feedbackStrong real-world feedback
Access to stronger intelligenceImpressive demos, internal metrics, self-evaluationReproducible experiments, continuously operating systems, traceable external outcomes
Access to weaker intelligenceContent repackaging, short-lived tool wrappersHuman baselines, long-term observation, reliable processes

The upper-left corner creates the easiest sense of awe. The model is powerful and the demonstration looks polished, but the outside world has not yet had a chance to prove it wrong. The lower-right corner often lacks prestige. Someone may use an ordinary model to improve a review process and record false decisions for six months. That result may last longer than a popular demo.

The upper-right corner is where this migration of talent holds the most promise. Powerful models enter scientific experiments or production workflows, reality exposes their errors, and the organization accepts responsibility for the results. Joining an AI lab only moves someone upward. The meaning of the work depends on whether it can also move to the right.

As capability gets cheaper, choosing the problem gets more expensive

In 2026, the cost of generating text and code is still falling. One person can test more ideas in a day, and a small team can complete work that once required a large organization. That speed is tempting, but it also magnifies the cost of choosing the wrong problem.

AI can quickly produce one hundred answers. It cannot decide on its own which question deserves years of a person’s life. Nor can it conjure experimental results, user rejection, or the constraints of the physical world. Models can help people reason through possibilities. Reality is responsible for saying “no.”

I therefore care more about problems that machines cannot score by themselves. They may begin with an anomalous measurement, a research workflow that no one has maintained for years, or a user who repeatedly works around a product. A person has to go to the scene, admit that the original assumption failed, and return new evidence to the system.

The selection criteria can be stricter. First find the people or systems already bearing a visible cost. Then confirm that you can reach evidence capable of overturning the proposed solution. Finally, ask whether the problem will still exist after the model changes. If these conditions remain unclear, treat the work as an exploration before claiming to solve the problem.

The benefit of speed should shorten this loop. If it only multiplies output by ten, we get more content that looks complete. If we use the saved time to verify results and revise assumptions, we may arrive at more reliable knowledge.

Bring AI back to the problems you have encountered over time

In June 2026, I wrote in “Why I Have to Commit to the AI Wave” that I could not see a future in geoscience, so I had to commit to the AI wave.

Looking at it now, that sentence was only half right.

I should still understand AI and enter the new ways of working that it is creating. But if committing to AI means abandoning the domain problems I have encountered over time and joining everyone else in chasing the same model metrics, then I also abandon the place where I am closest to real-world feedback.

Frontier labs know how to make models stronger. Earth system science still needs people who understand observations, physical models, and production code to define its problems. If domain researchers all leave those problems to compete for positions at the center of model development, the center becomes more crowded while reality loses the people who can identify the models’ mistakes.

For me, ESM-bench is one concrete starting point. It tests whether AI agents can modify production Earth system model code such as Noah-MP while preserving the physics encoded in that code. Existing code and expert-defined tasks provide the reference. If a patch edits the wrong subroutine, reverses the sign of a flux, or breaks conservation, it fails even if it compiles.

I want it to leave reproducible task data, evaluation programs, and failed patches in 2026. Model versions can change, while these records can still test whether the next generation of agents repeats the same errors less often.

This also explains why a domain researcher without an AI title may leave more durable work by 2030. That researcher did not train the largest model, but made a new capability pass a test imposed by reality.

Power appears after the model enters a workflow

The previous essay discussed the relationship between AI and power. Looking back from 2030, that judgment needs greater precision.

One model output can already influence a decision. When model outputs repeatedly enter hiring, research funding, diagnosis, or production scheduling, that influence hardens into the power to allocate opportunities and resources over time. A model can make a recommendation, while an organization decides whether to accept it. The organization also decides what data the model can read and who can stop it after an error.

Building a stronger model is only one part of the chain of power. The person who defines the objective holds some power. The person who approves deployment holds some as well. If the people affected have no route to appeal, power concentrates inside invisible interfaces.

If an AI system is still worthy of trust in 2030, it should have retained failure records in 2026, specified when to return a case to a human process, and allowed the person giving final approval to explain the decision. The stronger the capability, the earlier these arrangements need to be built.

Individuals face the same problem. We can delegate research, first drafts, and routine experiments to agents, but we still need to keep the standards of judgment in our own hands. If someone cannot explain why they trust a result, stronger intelligence only lets errors spread faster.

Assets that the next generation of models will not erase

Specific prompts depreciate. Simple wrappers get absorbed by platforms. Model brands may change. Problem definitions, sourced data, and evaluation methods tied to real consequences usually last longer.

Decision records can survive model updates. By 2030, a model may be able to rewrite a piece of code from 2026. It cannot automatically reconstruct why the team abandoned another approach, which batch of sensors had drifted, or which attractive metric later misled an experiment. Recording that context lets the next generation of tools continue from a real starting point.

Public writing can also become such an asset. An essay submits a judgment for others to inspect, records the evidence available to the author at the time, and lets later readers see where an error began. This kind of writing can outlast the traffic it received on the day it was published.

Meaningful work does not have to be large. A reusable script, a dataset with documented failure conditions, or a test suite that reduces a team’s false judgments may lower the starting cost for those who follow. If someone can still begin from that point four years later, time has not erased the work.

Trust is a slow-moving variable

The faster technology changes, the harder it becomes to replace trust between people.

AI can generate introductory emails, maintain contact lists, and remind us to follow up on time. It cannot manufacture a history of sharing risk or fulfill a commitment on someone’s behalf. By 2030, the weight of a relationship can be measured by one concrete result: how many people are willing to work with you a second time?

This gives renewed weight to actions that look slow. Replying thoughtfully to a young researcher, publicly crediting a collaborator, and correcting a mistaken judgment quickly will not appear on a model leaderboard. They will determine who is willing to bring you a problem that has not yet been made public several years later.

You can gain a high concentration of talent by joining an organization. A high concentration of trust can only accumulate through repeated collaboration. The former gives someone faster access to capability. The latter determines whether that person can carry the capability beyond a single position.

Four receipts to carry into 2030

I want every project from 2026 to leave four receipts.

The first comes from reality. The project has a baseline before it begins and an external outcome after it ends. The experiment may fail, and users may decline to use the result, but the world must have an opportunity to say “no.”

The second comes from inheritance. If I leave tomorrow, can someone else understand the data, reproduce the experiment, and continue maintaining it? If the model provider changes, does the validation method still hold?

The third comes from judgment. I need to record why I chose this route, what evidence would overturn it, and what changed after new evidence arrived. A correct result is worth preserving. The ability to locate an error is worth preserving too.

The fourth comes from collaboration. Are the people who shared responsibility for the result willing to work with me again? If the answer is no, even high short-term efficiency may come at the expense of future collaboration.

Connect limited time to a verifiable future

From “The Mindset of Disruptive Innovators”, to the question of why top talent enters AI labs, to this essay, I have kept returning to the same question: how can a person see farther and bring what they see into reality?

Foresight cannot be proven by guessing the name of a model. It eventually becomes a set of actions that can begin today: choose a problem worth amplifying, let reality continue correcting the answer, and leave the process behind for others to inspect.

By 2030, I hope the record lets me answer four concrete questions. What changed in reality? Where is the evidence? Can those who came later continue the work? Are the people who shared responsibility willing to collaborate again?

Those answers will tell my future self whether the intelligence available in 2026 entered reality or merely increased the speed of the moment.