Back to All
Ai/ml
Blog

AI coding vs. AI tokens: A better way to think about ROI

Listen
AI coding vs. AI tokens: A better way to think about ROI
5:28

Dr. Ryan Ries here. Last week I received a request for commentary about AI coding and AI tokens. I thought it was a pretty interesting topic. If you’re thinking about AI token costs, this is definitely a Matrix you should read.

But before I dive in, one quick announcement. After speaking with customers, so many of you are encountering the same issue with AI. Tons of use cases, limited resources, and uncertainty on what has the biggest impact. We’re tackling that problem tomorrow – make sure to join me on my live session by registering here.

$140K per month on AI tokens

OpenAI published data showing its median researcher now spends about $12K/month on AI tokens. Its top 10% of researchers spend roughly $140K/month. Cathie Wood of ARK Invest shared this on X and commented that it could be a leading indicator of massive productivity gains.


MatrixQuote


My opinion: Token spend is not a productivity metric. This isn’t the first time someone has talked about a cost metric as a productivity metric. But confusing those two is how a company burns through a year’s coding budget in no time.

I see this pattern happen a lot. A model gets dropped into a loop, starts rewriting the same function nine different ways, and racks up quite the bill while producing... garbage. I actually talked about my experience with this frustration loop in a Matrix awhile back.

Sure, more tokens spent can mean more work is happening. But it can also mean more churn, more dead ends, and more code that somebody now has to review and throw away. The dollar figure alone cannot tell you which one you’re looking at.

More code and more experiments don't automatically equal more value

OpenAI says its researchers are writing more code and running more experiments with the work in parallel. That’s plausible. Frontier labs push their tools harder than anyone, and their researchers know exactly how to keep five agents fed and working. But writing more code and running more experiments is an activity, not an outcome.

If you’re a product company, the metric that matters is much simpler.

Did you ship a feature customers use?

Did you get a new product to market faster than you did last cycle?

Did the thing you built solve a problem somebody was willing to pay for?

None of that shows up in token count.

Five agents, one person: how are team structures actually changing

Agents are interesting because everyone assumes LLMs are giving access to knowledge, but that isn’t really accurate since you have to be able to say whether or not that knowledge was correct.

AI hallucinations haven’t gone away, and an agent that’s wrong with total conviction is more dangerous than one who hesitates. Trusting agent output without that judgment layer is how bad information ends up sent to a customer.

What this means for how teams are structured is an open question. Engineers who get the most out of agents right now tend to be the ones who already knew how to do this work without one. They can tell good output from confident nonsense.

The near-term opportunity with team structures is building teams where experienced people direct more surface area, not assuming a smaller team automatically produces the same output on autopilot.

A better way to think about ROI

If a company drops $140K a month on AI tokens, the big question is what came out on the other side. To be honest, the $140K is above any number that I have ever seen reported.

I would imagine the reason behind that number being that high is because they are using the more advanced, expensive models. For most use cases, this just isn’t necessary, and that added cost has very little value.

For example, it’s great to use Fable to check some output, but I would never use that for every task. You need to be smart about model selection and farm your tasks out appropriately.

Spending $140K on tokens is great if you are outputting a working pilot, a feature in production, or a process that used to take 30 days and now takes 30 minutes. Now THAT justifies spend.

My very honest thoughts

I don't think we’re heading toward a slowdown on AI adoption, and I wouldn’t want us to. What I do think is coming is a much sharper filter for what counts as real value, and that filter is very overdue.

Set a business outcome before you set a token budget.

Route routine, well-scoped tasks to smaller, cheaper models.

Save the frontier models for the work that actually needs them.

Build the habit of asking what shipped.

We’re helping customers build this framework every day, pairing AWS Bedrock’s model flexibility with real cost governance so token spend maps to business outcomes.

If you need a partner with experience building these systems on AWS, we’re here to help. Reach out to our team here.

Until next time,
Ryan

Now, time for this week’s AI-generated image and the prompt I used to create it.

Create a picture of a muppet standing atop a towering, still-growing mountain of glowing golden token coins, using a magnifying glass to search for one tiny gemstone buried somewhere in the pile. Below him, more coins keep pouring in from a giant faucet in the sky. Far off in the distance, on a small pedestal just out of reach, sits a single trophy labeled "SHIPPED."

MM

 

Ryan Ries avatar

3 minutes read