Token tangle

tokens are EXPENSIVE, they made boy-v-girl internet into a horror movie, AI financial trouble, and why is everyone talking about "taste"??

Good day, and welcome to another edition of Endnotes! It appears to be very hot and smokey right now for a lot of you, especially Lindsey Graham. OK! I don't have any housekeeping this week, so let's get right to it. Here we have:

Feel free to forward this e-mail or send a link to your friends!


The unknowable cost of unending tokens

[permalink]

The AI treats are ending now, as foretold, and everyone is realizing that AI is very expensive. With Copilot switching to token-based billing at the beginning of June, a lot of CFOs are saying, "wait, we are spending how much??" The list of large enterprises confirmed as pumping the brakes on their AI spend: Uber, Accenture, Amazon, Walmart, Citi, etc.

It's fun to watch them scramble, but the point I want to make here (that is probably starting to sink in as these companies grapple with what they've done) is that you can't control AI costs, not really. For the kind of open-ended, multi-step business problems where enterprises are employing AI because they have major business value (data analysis, documentation search, software development), it's structurally impossible because of the way these systems have been developed.

* * *

Early on in the AI boom, the idea was that models would steadily get bigger and better until they would be able to "one-shot" any problem. The expectation was that you would be able to put a certain number of tokens (basically, words) into a prompt, and the model would generate a certain number of tokens in response. Your problem is solved, then at the end of the month, you get a bill for the tokens you put in and the tokens the model generated.

However, it turns out that scaling models doesn't work. Specifically, it doesn't solve core problems with accuracy. No matter how big the models, they still make stuff up, because they aren't designed to find information or "think." They are designed to simulate
intelligence by generating text, and they do that probabilistically—that is, there is an element of randomness in every response that's impossible to control.

This is the original sin of large language models, and people have been trying for years to fix it.

They have tried many things. Retrieval-augmented generation (RAG), for example, is a technique for ensuring accuracy by first programmatically retrieving relevant documentation and feeding that into the LLM along with the user's prompt. Another technique feeds the LLM's output as a query to other models that "vote" on whether it's correct or not. You have model-context protocol (MCP) servers that give LLMs the "tools" they can then call to retrieve certain information. You have "agents" (essentially, more LLM queries) that can be spawned to retrieve relevant data or call their own tools. And then you can put the whole thing in a loop or a "harness" that directs the process to keep going until it meets some condition.

* * *

As you may have noticed, all these varied techniques have one thing in common: more tokens. Every solution for the LLM accuracy problem boils down to "throw more tokens at it until you get the right answer."

Obviously, this works eventually, in a "billion monkeys typing Shakespeare" kind of way, especially if you have a very specific outcome in mind that you can test, like code. But it means that while you more often get the "right" answer, the cost of doing so scales unpredictably, because all this scaffolding means that you never know how many tokens it's going to take to complete a certain task. It might be cheap, like your LLM returning the correct string of characters to call a function that returns a specific answer. Or it might be very expensive, like your harness grabbing several megabytes of context and feeding them into a loop that runs for half an hour and then returns something unusable.

On top of this, the model providers themselves are adding scaffolding behind the scenes that you can't see or control: model routers, "reasoning" models, deterministic harnesses, etc., all intended to produce a more accurate output (or control their own costs, they are running a business too), but often while burning through orders of magnitude more tokens, and with little visibility as to why.

* * *

From a business perspective, this is a problem, because while you can budget a certain dollar amount to spend on tokens, there is no way to know if it's the right dollar amount to accomplish the desired number of tasks because it is basically unknowable how much a particular task is going to cost.

It's like if a business had an expense budget for entertaining clients, but there was no way of knowing beforehand how much a single dinner might cost. You walk into a restaurant with your client, and at the end of the meal, it might cost $200 or it might cost $20,000, you can't know. How do you budget for that?

Obviously, businesses can still control costs by putting hard limits on usage and letting teams decide where best to use the tokens, or whether it's worth it at all. Enterprises are also looking at ways to route the simpler tasks to cheaper models, use cheaper Chinese models, or use open-weight models. And I do think that, in the long term, good engineers will be able to build solutions that are cost-conscious and efficient.

But the era of solving problems by throwing indiscriminate piles of heavily-subsidized tokens at them has left a legacy of sloppy, wasteful systems with little visibility or ability to control costs. Culturally, it's a mess. I suspect it will take a long time to untangle.

[back to top]


MOVIE NIGHT: "Obsession"

[permalink]

"Obsession" is a cultural phenomenon. It's the highest-grossing movie ever made for less than $1 million. Two weeks after its release in Spain, we went to a 6:30 showing and the theater was packed. I can't remember the last time I was in a movie theater that full. And it was mostly young people.

The buzz around the movie is that it was directed by a YouTuber, Curry Barker, which explains a few things about its theme and its success.

I don't think Barker's popularity on YouTube specifically drove the audience numbers. That's a Bad Idea, his YouTube channel, has 1.5 million follows, and while this is an incredible accomplishment, it's not enough to move the needle on an internationally-released film.

His real juice is that if you make stuff for YouTube, you are constantly looking over your shoulder at the algorithm. You are AB testing, you are looking at metrics, you are looking at how your audience responds to different themes, and probably a million other things. More than perhaps anyone else, YouTubers have their finger on the pulse of what drives engagement on social media.

And what drives engagement is: relationship conflict.

"Obsession" is a distillation of the boy-internet vs. girl-internet viral content that is basically everywhere. You see it on Reddit, Instagram, TikTok, X. If you're on boy-internet, it's stories of false rape allegations, humiliating rejection, height-shaming, friend-zoning, cheating wives, and crazy exes. If you're on girl-internet, it's stories of actual rape, getting roofied, clueless boyfriends, cheating husbands, and asshole fuckboys. It's catnip for a generation that was raised on screens and more physically isolated than ever before—people that live much of their social lives vicariously through the experiences of anonymous strangers.

"Obsession" is, of course, a great film, but I think the reason it's a hit is because— intentionally or not—Barker made a movie where both boy-internet and girl-internet horrors coexist in parallel. It's a perfect yin-yang of the social media obsession with gender dynamics. The girl gets roofied and manipulated by the boy; the boy gets friend-zoned but then makes it worse by getting himself in an "I-can-fix-her" loop with a maniac. It hits all the kids right in the "let's talk about it" spot.

Like any good horror movie, it's an expression of the anxieties of young people. In the 80s, that was (among other things) guilt over premarital sex. Today, as the two sexes circle each other warily, it is: How can you have a relationship? How do you talk to a girl? When can you trust a boy?

On the one hand, it's a grim reflection of the way social media algorithms have affected how people think about relationships. On social media, the "stickiest" content (the content that gets the most attention) is negative. That is, you scroll past the happy relationship content, but you stop and read all the comments on the posts about cheating, lying partners. And because you paused on that content, the algorithm gives you more and more of it. Soon, relationships seem like a battlefield where every victory is Pyrrhic. And now they've made a horror movie about it.

On the other hand, it's nice to see so many young people leaving their homes and crowding into movie theaters, to be together in physical space and enjoy a great film. Once you finally see the villain, it's less scary, and once tropes are turned into a horror movie where everyone can scream and laugh together, maybe boys and girls can walk out of the theater into the summer night feeling a little better about each other.

[back to top]


Two videos on AI's looming financial trouble

[permalink]

I try not to spend too much time on the financial side of AI, because I don't think there's any mystery at this point. It's obvious what is happening (asset bubble) and what is going to happen (market correction). The only question is when.

But I do try to keep an eye on it, and in the spirit of that, here are two great (and short!) videos that sum everything up.

First, Ed Zitron, who covers this at great length in his newsletter and podcast. He appeared on CNBC recently and gave a brief and to-the-point summary of what he's been saying for years now.

If you like what he has to say, there's a lot more where that came from, so do check out his newsletter and podcast.

Second, Garbage Day posted a very good historical review of the dotcom bubble, comparing and contrasting it with the AI bubble. Again, concise, clear, correct.

I am pretty sure this whole thing will end in tears, but as I've said in the past, they can probably keep it going much longer than seems possible, so the timing of the comeuppance is going to be hard to gauge.

[back to top]


Bad taste

[permlink]

pretty bad. prettyyyyyy prettyyyyyyyyyy prettayyyyyy bad.

I have a little theory about what makes LLM writing bad. Take, for example, the famous "it's not x, it's y" construction. People will argue these are tell-tale signs of LLM-extruded text:

"It's not a motorcycle; it's an affordable mobility solution."

"You're not just tired; you're exhausted."

"It's not surrender; it's a strategic retreat."

You get the idea. So what exactly makes this sound like synthetic text? It's a totally normal rhetorical construction, which is probably why LLMs use it so much: It's very common in English-language writing. In the right hands, it's quite elegant.

But a real human prose stylist uses those formulations deliberately when they need a dramatic pause, or a pivot point. They understand the rhythm and emotional heft of the text because they understand the text. Large language models don't understand the text; they simply generate it.

So while an LLM's use of the "it's not x; it's y" construction is grammatically and rhetorically "correct," it often makes no emotional sense. It's tone-deaf. It's uncanny valley stuff, because a stochastic text generator can produce form, but not meaning, and that's what makes it sound weird.

We see something similar with AI-generated music, graphic design, code, and images. There's a boilerplate vernacular that it reproduces that can be convincing at first, but without any real understanding, it's inert. Dead eyes. Chorus hooks that go nowhere. A button that does nothing. Everything a purple gradient.

The current LinkedIn AI consultant buzzword that is supposed to address this is "taste." "Taste is the moat" is the cliche—that is, as a Business Guy, once you've figured out the hard part of building the body of the thing you want to sell, you just gotta sprinkle some "taste" on it to fend off competitors and start printing money.

Here we have some of the least-creative people to ever live failing to grasp the difficulty of creative endeavors. It's like they think the hard part of (for example) shooting a "Seinfeld" episode was scheduling the catering, renting the building, and gathering a studio audience. The set, lighting, sound, direction, acting, and writing? Just wave a wand and say "taste."

What they will find out soon enough is that "taste" is a euphemism for what's actually missing from AI-generated content: understanding, thought, and judgment. And for that, you need people.

[back to top]


[permalink]

[back to top]

Subscribe to Endnotes

Sign up for free to get every issue in your e-mail.
jamie@example.com
Subscribe