Tokens per Outcome
· 3 min read
You wake up. Not because your alarm rang, but because your alarm agent spent around 300 tokens deciding it was the optimal moment to wake you up, using a tiny model tied to your sleep tracker.
You reach for your phone.
Your inbox is already summarized. One agent read everything, filtered spam, pulled a few threads, and wrote a short summary and hands you your to-do items. That’s easily 3,000 to 6,000 tokens gone.
You walk to the kitchen. While the coffee is brewing, another agent is “optimizing your day.” Calendar reorganized after calling previous agents, priorities updated, and a few recommendations generated. Call it another 1,000 tokens.
You sit down at your desk, and 5k+ LOC have already been written, reviewed, and merged by agents talking to each other all night. Dashboards show how many tokens were spent, not how much work was done.
You didn’t do anything yet… but tokens are already being burned.
By the time you actually start working, your system has already burned through 30,000 to 50,000 tokens.
The more tokens we spend, the happier and more productive we become.
What this leads to
We’re heading into a world where everything can burn tokens. Agents calling agents, models calling models, systems doing more and more thinking for us. It feels satisfying and productive. But it’s not.
Because in this world, tokens are not free. They are compute. They are cost. They are your margins.
I saw this tweet the other day by Perplexity CEO that says
“One thing that Perplexity has is, every revenue we make, unlike certain other [wrapper] companies, every revenue Perplexity makes has positive gross margins.”
“Because we route through multiple different models, we’re very efficient in terms of how we spend on the tokens… we have all this advantage with RAG and orchestration and search. We don’t actually need to blow up the [context].”
“Every single penny we make, we make profits on that. But the overall company is still yet to be profitable.”
If the real insight is
“we’re efficient in how we spend tokens”
This is the whole game. The companies that win are more efficient in how they spend tokens. Some systems just call a model and hope for the best, others treat tokens like a resource. The best AI companies will be the ones that waste the least, not the ones that generate the most.
We used to measure latency, uptime, performance.
Now there’s a new one: tokens per outcome.
Which is the fewest tokens it takes to get the user what they actually came for.
The real cost isn’t the model
Let’s take something simple. A user asks, “summarize this document.” Most systems will probably throw the entire document at the model. Let’s say that’s around 20,000 tokens. Then they use a large model by default and let it generate a long answer, maybe another 1,000 tokens. That’s roughly 21,000 tokens for one request.
But you could build the same feature differently. Instead of throwing everything, you first find the parts that actually matter. Maybe that’s 500 tokens. Then you route it to a smaller model if the task is simple, and you force a small response, maybe 120 tokens. Now you’re at around 620 tokens total.
Same outcome for the user. Completely different system behind the scene.
Now look at cost. With a cheap model mix, you might pay about $0.01 per 1,000 tokens. With stronger models, maybe $0.03 to $0.06 per 1,000.
- Naive: $0.21 → $1.26 per request
- Optimized: $0.006 → $0.037 per request
We are talking about 30 to 100x difference.
Then you scale it. At 10,000 requests a day, the naive system is burning somewhere between $2,100 and $12,600 daily. That’s $63K to $378K a month. The optimized system sits between $60 and $370 a day, or $1.8K to $11K a month.
The uncomfortable part is why this happens. Most products send too much context, default to the biggest model, and generate more text than anyone needs. Not because the problem requires it. Because it’s easier than designing the system properly.
And that’s the shift. For years we obsessed over compute, storage, and networking. Now, we’re no longer optimizing compute. We’re optimizing thinking.
And the real signal is how much thinking it took to get the answer.