Contents

中文

Model vendors love to show off long-horizon task capability with claims like “over 20 hours without human intervention,” implying that long-horizon tasks are fully automatic. But in reality, how many of our tasks can simply be left alone — running for a long time with no intake of external information at all? I saw Dr. Kong mention in a WeChat post that he ran a small survey and found that real-world tasks lasting over an hour don’t seem to be that common.

In real work, we often find partway through that we need to supplement information — we have to talk to other people to get it, digest the new information, and then process it further.

To run unattended — to have an agent execute a very long task — means the operator has to provide complete information upfront. That is extremely hard.

Moreover, after a few hours, if the person is responsible and thinks diligently, they may well have new ideas about the problem themselves, and they need to feed those new ideas back into the agent’s process.

Some tasks do fit the bill — proving a mathematical conjecture, for example, or migrating a programming project from one language to another. These problems are hard, take a long time to execute, and the information can be fully provided at the start.

To break through on long-horizon tasks, model vendors have improved models’ ability to proactively verify closed loops and to handle long contexts. With these abilities, an agent can work on a task for dozens of hours. But given the way humans deliver information described above, I think this has become impractical. Besides, when a task is too long, even if the agent delivers, the workload is so large that it’s hard for a human to review the results well — it would be better to split the work into smaller issues and solve them separately. And that brings in another question: an issue-based agent management mechanism.

After these breakthroughs, what models lack when executing tasks is no longer capability, but information. This is exactly where the Matt skills bundle is useful: it solves the information problem. Information always has to be provided, because the model can’t possibly know my requirements. The model is no mind reader — it doesn’t know what my words mean. Different people saying the same word can mean different things. The model cannot know what the person giving it requirements actually means; it has to ask me from multiple angles, and only then can we align and minimize misunderstanding.

This is the idea behind Matt skills.

There is a fantasy that a company can simply accumulate its knowledge base and then do Deep Research.

I don’t know how many people have actually tried Deep Research. By my standards, the accuracy of Deep Research output is still not high enough.

Very often the problem is “search quotient”: the model doesn’t know what it doesn’t know. It searches and constructs according to its own ideas. Frequently the information it holds is already outdated, but it doesn’t know that — it keeps searching based on the old assumptions it has internalized, finds old information, and everything comes out correct and self-consistent. In the end it produces a “carving the boat to seek the sword” report — one built on stale premises. In my scenarios, 50% accuracy would already be decent.

The information-coverage problem of “not knowing what you don’t know” is very hard to solve, and it is a difficulty every agentic long-horizon task must face. Before executing a long-horizon task, how can an agent be sure its information is complete?

Never mind a company’s knowledge base — even when I, as an individual, hand a task to an agent, I may find partway through that it has drifted off track from what I expected. Because the agent was never aligned with me. I assumed something was self-evident, but it actually didn’t know; it didn’t know that it didn’t know, and it wouldn’t come ask me — so the job gets done poorly.

This also explains why so-called “loop engineering” doesn’t work. The term was hot for a while, championed by Peter Steinberger and Boris Cherny. The whole loop process has no intake of external information; it just lets the agent repeat itself over and over. The missing information stays missing, so the results naturally never improve.

The essence of Matt skills is that it makes the agent grill me on all kinds of details — the famous grill-me. The Matt skills bundle even has a newer skill called to-questionnaire, which organizes the current list of questions into a questionnaire to send to colleagues or others. It not only lets the user fill in the information, but also lets the user invite team members to fill it in together.

So if the model doesn’t know what it doesn’t know — it doesn’t even know — how can it possibly ask? Matt’s approach is: have the model write a plan first, then take every decision point in the plan and ask about it. Once asked, it naturally becomes clear what information is missing; the user and the model just reconcile. Based on the supplementary information, the decisions are revised, and the resulting decisions are naturally of higher quality.

Of course, some people find this annoying and feel too many questions are being asked — asking about every decision point is naturally over-saturating. But after several rounds of product-level optimization, Matt skills now asks three or four questions together per round, and most of them can just take the defaults. We only need to intervene lightly, and the experience has improved. This process of being questioned happens to also be our process of reviewing the agent’s plan. That’s actually quite nice — it means I get a look at the agent’s decisions along the way, which gives me far more control than groping entirely in the dark, and I also get to review and sort out my own requirements.

I like Matt skills far more than Superpowers. Superpowers aims to introduce some of humanity’s good engineering practices to agents. But first of all, I don’t believe any single good engineering practice can solve every problem — engineering requires flexibility. And as the model’s own general intelligence improves — after it has learned a large body of engineering techniques and mastered deep engineering thinking — it knows on its own which engineering practice fits which kind of requirement. It has flexibility, more flexible and intelligent than any rule.

But Matt skills is different: it solves the information problem. Models will always lack information, so that is something I have to provide. Although Matt skills also contains engineering content, it mainly revolves around information.

Moreover, what Matt skills embodies is a setting of the model’s personality — wanting the model to ask the user. But for many tasks, we don’t need the model to do that. With a simple question, for instance, I ask one sentence and the model should just deliver the result — no need to ask every time. Sometimes the user wants it to ask; sometimes the user doesn’t.

When calling model APIs, this happens a lot: you ask the model who it is, and the model can’t answer with its own name. Model vendors could of course train models to state their identity correctly. But none of the companies do this now.

There is a rumor that doing so would degrade the model’s general intelligence, but I haven’t found any paper or research to support that. I think there are also product-level reasons: if a model strengthens its identity information, then when I — as a downstream developer — call the model API to serve my customers, and a customer asks who it is, it will still be inclined to say “I’m a model from vendor X” rather than the product name I specified. That definitely won’t do. So identity and personality information like this is unsuitable for training into the model; it belongs in instructions — telling the model in the system prompt: “You are so-and-so.” That better fits the market needs of the model API itself as a product.

Likewise, grilling every decision, as a personality trait, is certainly unsuitable for internalizing into the model either. Model APIs also power companion apps and customer-service apps, for example — scenarios where a grilling personality is inappropriate. So “grilling” is better suited as an external add-on instruction. A friend reported that GPT 6 Astra loves to grill: when it senses a problem while working, it stops to ask the human — which is problematic in many situations. Once this trait becomes the model’s personality baseline, it becomes incompatible with many application scenarios. Therefore, I believe a love of grilling will not be the native personality of future models; Matt skills, as an external add-on, will need to exist for a long time. That is where it differs from Superpowers.

So, after using Matt skills to fully align all the information, can we then run tasks that take a dozen hours? After all, we have provided all the information before execution begins.

I think difficulties remain. If the task would take even an agent a dozen hours, then the task plan will also be extremely long, and it is hard for a human — with limited attention and working memory — to review, decide on, and respond clearly to a plan that long. Making the plan and providing the information itself constrains the task to a certain scale. The scale of information a human can process will constrain the scale of the task. Cases like the mathematical proofs and code migrations mentioned earlier — where the information lies almost entirely within the project itself — are of course exceptions.

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

Contents