This year a few companies thought it would be a good idea to setup a leaderboard for which employees spend the most AI tokens.
It radically backfired!
Right now companies are using "token spend" to determine which employees are the most productive. Spend more AI tokens, means you are producing more right?
Wrong!
And because it's so highly wrong it's springing up a whole new business model in AI, that people are making $5,000, $10,000, even $25,000 to fix this wrong.
But first, let me make sure we're on the same page about what a token even is. Because these aren't your usual Chuck-e-Cheese tokens...
Tokens, Real Quick
A token is what the AI coder nerds have decided to call a chunk of text the AI reads or writes.
We can't use "words" because some words are longer than others, so they use "tokens" instead. It's a unit of measurement.
The word "and" might be one token. While the word "artificial" would likely be 3 tokens. This sentence is around fourteen tokens.
Every time you talk to ChatGPT, Codex, Manus, or Claude, you're spending tokens. Whether you're just asking it to read something or output something.
The model reads your prompt (input tokens) and writes its answer (output tokens). The meter is always running.
In a chat app you feel this as your usage limit. You know that annoying "you've hit your limit, come back in 3 hours" message?
That's you running out of tokens.
Or I might actually say, that's you likely misusing your tokens, based on the topic of this newsletter issue.
But that's not really the big crisis I'm pointing to today...
When you build an AGENT instead of using the chat app, you feel token spend more because you pay per token usage. An agent runs on the API, not the chat interface, and the API charges you per token.
Good news, tokens are quite cheap.
They're measured in the millions and 1mil tokens might cost you $0.55 to $25.00 depending on which model you used.
When you have agents running 24/7 automating your content, or customer support, or a team of AI coders, the tokens add up.
For example, I spend a little more than 300 million tokens per month.
Like I said, they add up quick!
And that adding up quick part is the whole reason this new job exists.
Token Usage Is The New Productivity Metric
Here's where it gets a little crazy.
Because AI makes one person able to do the work of five, managers needed a new way to measure who's actually using it. So they landed on the laziest possible metric: token count.
More tokens used = more work done. At least that's the theory.
So now you've got companies building internal leaderboards that rank employees by token usage. Sales teams are gamifying it, with Slack channels celebrating the "top burner" of the week.
You see the problem already, don't you?
Token count measures CONSUMPTION, not OUTPUT. It's like ranking your sales team by how much gas they put in the company car.
And the second you make something a leaderboard, people game it. Any employee who wants to look busy just flips their AI into MAX mode, the most expensive model on the highest reasoning setting, and spams it with sloppy prompts all day.
Boom. Top of the leaderboard.
Meanwhile the sharp employee who knows exactly which model to use for which task, who gets more done with a fraction of the tokens, looks "lazy" on the board.
So the metric is broken. But the company can't see that, because nobody in management is nerdy enough to understand how this stuff gets billed or optimized.
They just know OUTPUT!!!
And THAT is the opening. You've got a broken metric burning real money in a room full of people who don't understand it.
That's your new $5,000 opportunity. Because while $5,000 might be expensive, it's a lot cheaper than wasting an extra $2,000/mo for an entire year.
How I Learned This The Hard Way
The first time token cost ever bit me, I was building the customer support agent (Build Notes #009 if you missed it).
That agent doesn't live in my normal Claude subscription. It lives outside the chat app and runs 24/7, answering support emails while I sleep. To do that, it has to reach my Claude account through the API. Which means I'm paying per token now, not a flat monthly fee.
Fine. Tokens are cheap, remember?
Then one morning I woke up, checked the dashboard, and an agent I'd built had spent $600 overnight.
Six hundred dollars. While I was asleep.
It got stuck in a loop. Kept re-running the same prompt over and over, burning tokens on every pass, with no cap to stop it. (This is exactly why I beat the "max pass count" drum so hard in #012.)
That stung. But it also taught me the lesson that's now worth thousands of dollars: most of what an agent does - doesn't need the smartest, most expensive model.
So when I rebuilt the support agent, I got surgical about it.
Reading an incoming email? That's a simple job. I don't need Opus 4.8 (the Claude genius model) for that. I use Haiku 4.5, the fast, cheap model. It's a fraction of a fraction of the cost. Literally like 100x less.
Writing the reply? Also doesn't need Opus. Writing a clear, friendly support reply is a writing task, and Sonnet 4.6 is fantastic at writing.
Opus only clocks in for when I need to troubleshoot an issue or build a new feature into my agent. That's when I bring in the big guns, for a short burst, so it doesn't cost too much.
Same agent, same quality of support, a fraction of the bill. THAT is token efficiency optimization, and it's a real skill with real dollar value.
What This Job Actually Is
At its simplest, token efficiency optimization is just knowing when to upgrade and when to downgrade your model.
Don't drive the Ferrari to check the mailbox. Don't bring a kitchen knife to chop down a tree. You match the model to the task.
Inside your chat app, that skill keeps YOU from blowing through your usage limits by lunchtime.
Optimizing token spend for one person is small money tho.
The real money is optimizing a whole COMPANY.
Picture a business with 20 employees all using AI. Not one of them knows when to upgrade or downgrade, so they're all running the biggest, most expensive model on every task, all day, because nobody ever told them not to.
This is 90%+ of the market right now!
Now picture that same company telling their employees to start building agents on the API, it's going to get out of hand FAST.
You're going to be the person who walks in and stops the bleeding.
This is just like when I used to sell ad management. They called me in because their ads were spending too much. They didn't know how to optimize their campaigns, I did, same thing today with tokens.
The 3 Mistakes Bleeding The Money
After doing this on my own stack and looking at others', it almost always comes down to three mistakes.
1. Never switching models.
This is the big one. Most people pick a model on day one and never touch the dropdown again. Usually the most powerful one, because "best = best, right?" They have no idea they're paying 10x for tasks a cheaper model would nail. I bet YOU do this 😁
2. The wrong amount of context.
This one's a balancing act. Give the AI too little context and it has to work harder, guess, and loop to fill the gaps, which burns tokens. Give it too much (dumping your entire 80-page doc in when it needed 6 pages from chapter 4) and it has to READ all of that. Reading is still token usage. Every word it reads, you pay for. The skill is learning to feed it exactly what it needs and nothing extra.
3. Matching the right model to the right job.
This is the master skill, the one from my support agent. Cheap model for the simple stuff (reading, sorting, summarizing). Mid model for writing. Premium model only for real reasoning and judgment. Most operations could cut their AI bill in half just by routing tasks to the right engine.
None of this is technical. There's no coding involved, it's judgment, which is exactly why a smart consultant can learn it fast and sell it high.
Why This Is About To Explode
Here's the part that turns this from a neat tip into a real career. The whole industry is quietly shifting you from flat-fee subscriptions to pay-per-usage.
Think about it. That $20/month plan you love?
It was ALWAYS subsidized.
These companies were eating the real cost to grab market share and get you hooked. That can't last forever, especially as they march toward going public and suddenly have to show shareholders an actual profit.
We've already seen the tremors. Claude pushing OpenClaw users off the old deal. The threatened June 15th usage fees. These aren't one-offs, they're the future arriving a little early.
Flat fee was the on-ramp, and pay-per-token is where this is headed. Once the meter is running on every employee at every company, "we have nobody who understands our AI spend" stops being a shrug and turns into a five-figure line item somebody has to fix.
The game is moving from what is AI, to how do we do this with AI while still maintaining a profit margin?
If my little AI newsletter company already has to think this hard about token cost, then every company scaling AI across multiple departments is going to hit the same wall.
The big ones especially, the kind that won't blink at paying $5,000 to $10,000 for someone to come cut their costs.
The Math That Sells The Whole Thing
Let me show you the kind of waste hiding in a normal company. This is the demonstration you'll eventually walk a client through.
Say a company bought a suite of SEO and social media agents. Blog drafts, keyword research, social posts, repurposing, the usual. Nobody ever checked what model it runs on. Out of the box, it's pointed at GPT 5.5, the frontier model.
MAJOR overkill for 90% of what it does.
Let's price one employee's daily agent workload at about 2 million input tokens and 500,000 output tokens (agents read a LOT, so they're input-heavy).
Run it across ~260 working days a year.
Here's the same exact work on three different engines, at today's published API rates (as of June 22, 2026):
| Model | Price (in / out per million) | Cost per workday | Cost per employee/year |
|---|---|---|---|
| GPT 5.5 (frontier, the default) | $5 / $30 | $25.00 | ~$6,500 |
| Claude Sonnet 4.6 (right-sized for most of this) | $3 / $15 | $13.50 | ~$3,510 |
| DeepSeek V4 Flash (for the cheap read/sort steps) | $0.14 / $0.28 | $0.42 | ~$109 |
I'll bet dollars to donuts 7 out of 10 of you DIDN'T even know what Sonnet 4.6 was before this email. And you for sure know your local dental office doesn't!!!
They definitely don't know what DeepSeek v4 is or that it's 65x less expensive than the latest ChatGPT models.
We are in a "TARGET RICH ENVIRONMENT!"
You walk in, spend a week documenting which tasks should run on which models, train the team, and you just saved them fifty grand a year. They paid you five. Who got the better deal? They did, and they'll tell their friends about it.
That's the whole offer, and the whole business.
Continue with Build Notes+
This is the free preview. The complete step-by-step guide and supporting files are available to Build Notes+ members.
