Tao of Tokens

a meditation on balance, for the age of agentic coding

The model that thinks the longest
does not always find the truth first.

Speed is not shallow.
Depth is not slow.

Somewhere between haste and hesitation
the answer waits,
indifferent to how it was reached.

The Weight of Thinking

A mind that checks itself
may find the error,
or invent one.

The answer was already correct.
Thinking further did not improve it.
It only delayed it,
and sometimes lost it.

A model built for the hardest reasoning and the longest agentic work became known among practitioners for a specific complaint: it overthought simple requests. Its own maker later advised builders to stop writing "double-check your work" instructions for the next version, because those instructions now made it argue with itself. A rival shipped a model that matched its output using less than half the tokens, less than half the time, a third of the cost. This isn’t an isolated flaw. Researchers studying reasoning models keep finding the same shape: past a certain length, more thinking is more likely to talk a model out of a right answer than into one. The tool for going deeper and the tool for going too far are the same tool.

The Small Hand

The smallest hand can still thread the needle.
The largest hand does not need to.

On the phone in your palm, a small mind wakes,
answers plainly, forgets nothing it wasn’t asked,
and costs nothing to consult.

You may hold a small model in your pocket right now, or even in the palm of your hand, awake, capable, and idle. Phones ship with one built in: a few billion parameters, running entirely on the device, free per use, fast, private. The newest on-device models go further still, storing far more than they ever load: twenty billion parameters kept in reserve, only a sliver of them, one or two billion, waking for any single request. Even the largest open models built today learned this trick at their own scale, activating a small fraction of themselves per answer and leaving the rest dormant. Bigness, at every size, is mostly held in reserve. The part that’s awake is almost always smaller than the part that exists.

The Wrong Question

"How large a mind must I summon?"
is not the question.

The moment already carries its own answer.
Most of the asking is louder than the need.

Treat every request as if it deserves the largest model available, and you will overpay for most of them. Systems built to route each request to the cheapest model that can actually handle it, escalating only when the small one fails, have matched top-tier performance while cutting cost by more than half, and in some measured cases, by nearly two orders of magnitude. This is not a trick for the frugal. It is an accurate description of the work itself: most of what gets asked of a model is not hard. A quick lookup. A short rewrite. A format cleaned up. Reserving the largest mind for the fraction of requests that actually need it isn’t restraint. It’s aim.

The Hungry Ghost

It counted what it consumed
and called the counting devotion.
It ranked itself by appetite
and called the ranking progress.

The bill came due.
The work was still not finished.

There was a season when some workplaces turned raw usage into a competition, ranking people by tokens burned, treating a large budget as a badge of seriousness, one leader saying outright that everyone should maximize their consumption. The pattern broke the way these things tend to: people gamed the count with meaningless tasks just to keep their numbers up, while the actual work didn’t improve to match. Measured independently, heavy usage tracked with far more code written and then rewritten, not less, teams paying many times more for a fraction of the gain. Before long, the same organizations quietly took the leaderboards down. The measure had become the target, and the target had stopped measuring anything real.

The Spectrum

From a name too small to hold a single thought
to a name too large to say out loud,
every size has already been built.

The question was never whether it exists.
The question is whether you reached for it on purpose.

The range available today spans further than most people realize. At one end, models small enough to run from a sliver of storage, or a browser tab, answering instantly and privately, built to do one job and no other. At the far end, open models built from trillions of parameters, only a small fraction of which switch on for any single answer, priced for tasks that genuinely require depth: proofs no one has checked, work that runs for hours unsupervised, code no one has reviewed yet. Between them sits everything else, priced across a range wider than most people would guess. The spectrum wasn’t built by accident. It exists because no single size was ever going to be correct for every question.

This is not a teaching in favor of small.
The small fails the hard question
as surely as the large wastes the easy one.

Balance is not the midpoint.
Balance is the match:
the mind sized to the moment,
chosen on purpose, not by habit,
and stopped exactly when the answer is done.