The Twenty Minute VC cover
Technology & the Future

Are OpenAI & Anthropic Overvalued? How Token Costs Will Fall 10X & Usage Will Explode 100X

The Twenty Minute VC

Hosted by Unknown

1h 29m episode
10 min read
5 key ideas
Listen to original episode

Token costs are collapsing 10x while usage explodes 100x — but the real question is whether today's AI giants survive the math.

In Brief

Token costs are collapsing 10x while usage explodes 100x — but the real question is whether today's AI giants survive the math.

Key Ideas

1.

Customized Models Drive Trillion-Token Scale

Fireworks processes 40 trillion tokens per day — mostly from customized, not off-the-shelf, models.

2.

Specialization Replaces Generalization for Success

Jensen Huang: 'There's no specialized general company — every company exists on a special belief.'

3.

Rapid Scaling Creates Bankruptcy Risk

Scaling to bankruptcy is a real failure mode that didn't exist in the SaaS era.

4.

Companies Will Own Intelligence Stacks

Every company will own its own intelligence stack — just as every company owns its own software.

5.

Supply Chains Constrain Usage Explosion

10x token cost reduction in 3 years; 100x usage explosion to follow — but supply chains hold us back until then.

Why does it matter? Because the trillion-dollar bet on frontier model dominance is aimed at the wrong layer

Lynn Quo's Fireworks processes 40 trillion tokens per day — mostly from customized models, not off-the-shelf ones. That single data point signals something is badly wrong with the standard narrative about where AI value accumulates. This episode is the clearest articulation yet of why specialized intelligence — built on private data, tuned to specific use cases — structurally beats general intelligence built on public internet.

• Jensen Huang told Quo directly: "There's no specialized general company — every company is built on a special belief of doing things, otherwise there's no reason they should exist" • The majority of the world's data is private enterprise data that frontier models never see — and never will • Product-market fit no longer guarantees a durable AI business: companies can now "scale into bankruptcy" • Open-source models have crossed a quality threshold that puts a permanent ceiling on closed frontier API penetration

Jensen Huang accidentally proved that AGI — one model ruling everything — is a logical contradiction

AGI as defined — one model solving everything for everyone — is a logical contradiction. Jensen Huang handed Quo the proof.

After his GTC keynote, they had an unscripted conversation that ended up being recorded without Quo realizing it. Jensen said: "Ling, you're right. There's no specialized general company — every company is built on a special belief of doing things, otherwise there's no reason they should exist."

Quo's extension lands like a grenade on the AGI thesis: "Every single company is doing something unique that justifies their existence. And this something unique is deeply baked into their product design, is deeply baked into their software design and system building — and that is not learnable or assured by another company sitting outside."

If every company exists because it encodes something no outside lab can absorb, then a frontier model trained on public data cannot replace the intelligence generated from a company's own operations. The AGI thesis assumes convergence. Competitive markets guarantee divergence.

Her prediction: "I really believe the future will be — it may be scary but I think that's true — it will be millions of specialized models, one per application, per use case."

The world's most valuable training data is locked inside enterprises — and frontier models will never see it

Every GPT, every Claude, every Gemini trains on what humans happened to publish online. That's a vanishingly small fraction of actual human knowledge. "Public internet is very small corpus of data compared with world's data," Quo says. "Majority of world's data is actually private, locked inside applications, locked inside enterprise — it will never get shared with anyone else because this is the company's proprietary IP."

The training data race the press obsesses over — who scraped more, who has better synthetic data — is a competition over the smallest available pool. Decades of customer interactions, internal workflows, proprietary reasoning, enterprise decision-making: invisible to every frontier model ever trained. "Majority of data is not being activated to derive any intelligence. That's where we believe in — to activate that data — and we believe the future of the frontier of intelligence is actually private intelligence, specialized intelligence."

Fireworks was built on this asymmetry from the start. Forty trillion tokens per day, most of them flowing through customized models. The companies winning at this layer aren't serving generic intelligence at scale — they're activating data that no frontier lab will ever touch.

Token costs will fall 10x in three years — and that compression will drive a 100x explosion in usage

Costs haven't fallen yet, and Quo doesn't bother pretending. "It hasn't so far because of supply chain constraint." Transistors, chips, power, cooling, manufacturing capacity — every physical layer of Jensen's five-tier AI cake is constrained by industries never designed for 100x scaling. She doesn't expect material relief for 12 to 18 more months.

But the economics are inexorable: "We are living in a free economy. Whenever there's shortage, price is high; price will invite a lot of people coming to solve the problem, and it will invite competition, competition will bring down the cost."

Her number is concrete: "10x cost reduction in the next three years, and this 10x cost reduction will drive a 100x usage." The mechanism is partly behavioral. "The moment you don't think about it as a problem for you — if it's a utility, you just use it."

Anyone building AI products calibrated to today's inference pricing is designing for a world with a three-year shelf life. The right assumption: 10x cheaper, 100x volume. Plan accordingly, or get caught flat-footed when supply chains catch up.

Companies with product-market fit are now 'scaling into bankruptcy' — a failure mode that didn't exist in the SaaS era

Scaling fast in AI can now kill a company that's working. "During SaaS time, product-market fit and a durable business almost are equivalent to each other. Now product-market fit and durable business are two separate concepts."

The failure mode is direct: "We have great companies — they have product-market fit, customers want to pay them and they really value their product — but they cannot scale because once they scale, they could scale into bankruptcy."

Incumbents face the sharper version. A company that won the last decade and accumulated tens of millions of users ships an AI feature and simultaneously reaches every customer. "Their CFO looks at their cost forecasting — there's no way you can justify this." The same traffic that made them powerful makes AI features economically undeployable at rollout.

The exit route is control: open weights models tuned to specific workloads, inference costs you own, unit economics you can actually model before volume arrives. PMF without a path to inference cost control is a countdown timer, not a moat.

Open-source models don't just match closed frontier APIs on most enterprise tasks — they're now structurally superior for anyone who needs control

Both categories crossed the quality threshold — that's Quo's starting point, not her conclusion. "Both open model and closed model quality significantly improved over the past two years — to the point both of these two streams cross the threshold; it can solve so many problems." But the implications diverge sharply.

Closed APIs: ongoing cost, no weight access, no meaningful customization, model provider's judgment baked into every output and permanently misaligned with your company's specific design principles. Open weights: zero acquisition cost, full tuning rights, inference you own. "Open model crossed a threshold — it's so much easier to tune. Being able to steer a model is part of the model intelligence." Tune against your own eval with a small amount of unique company data, and "often the end result of hill climbing is to solve your unique problem with your data — you are better than a general purpose model."

Fireworks processes the proof internally — open models for recruiting, finance, debugging, agent workflows — and the logic scales to any enterprise that reaches meaningful AI traffic. The ceiling for closed frontier vendors isn't quality anymore. It's the structural inability to give customers control over their own intelligence stack.

Fireworks ran massive post-training RL for Cursor across five data centers using scattered GPUs — no hyperscaler cluster required

Ten thousand to one hundred thousand chips all interconnected through InfiniBand: that's what frontier-grade reinforcement learning post-training was assumed to require. Fireworks built a different architecture while training Cursor's recent models.

"We decouple these two — the trainer that is tweaking the weights and the RL rollout. In the past, in large hyperscalers, they run that all together." Instead: "We designed a fully distributed system and we run across five, six data center regions globally, tapping into scattered GPUs — and they are able to run massive jobs."

The hard engineering problem is weight synchronization across regions: stale model weights mean stale rewards, which degrade training quality. "The latency of delay of sending these weights over is going to dictate how fresh the rewards are — if it's too stale, then you are too off." Fireworks solved it well enough that "numerically it's still sound" — without the hyperscaler capex.

The capital moat around frontier model post-training is smaller than advertised. Distributed systems engineering can substitute for a monolithic GPU cluster, if you're willing to build it.

Three new hardware SKUs per year from a single vendor has made every AI infrastructure DCF built on 6-year depreciation structurally wrong

Hardware depreciation assumptions underpin gross margin calculations, infrastructure valuations, and build-vs-buy decisions across the entire AI stack. Those assumptions were calibrated for a world that no longer exists.

"In the past it's six years — solid six years. Hardware release is usually three years. And now within a year from one vendor alone, we have three SKUs." The compounding problem: newer frontier models run best on newer hardware. "Model depreciation is also very fast — every week we are launching a new model. And the new model likes the newest hardware."

Two years in, the math turns brutal. "After two years, which model runs on two-year-old hardware? It'll be the two-year-old model. Are those models still valuable? The hardware will last for six years still, but with this pace of model development, it's questionable." At three SKUs per year, a two-year-old machine is already six hardware generations behind best-in-class — not because it failed, but because the cadence of release has made it irrelevant for the workloads that matter.

Anyone doing DCF analysis on AI infrastructure companies using 6-year cycles is producing materially wrong valuations. The true useful life of front-line AI hardware in the current environment is closer to two to three years.

Jensen Huang replies to emails in one minute — and that single habit defines velocity, not micromanagement

"I sent him an email — he will reply in one minute. I just don't understand how he's like constantly in details."

Quo spent years admiring the sheer throughput. Four years running her own company clarified the theory: "Leadership is just judgment. It's not privilege. It's judgment." Judgment requires current context. "In a fast-moving space, if you do not know what's happening, what works, what doesn't work, what are the gaps — you make the wrong call. It's guaranteed there is information loss in transition layer after layers."

The conventional playbook — delegate, stay strategic, trust the reporting chain — was written for industries that move at quarterly cadences. In AI, context has a half-life measured in days, not quarters. Waiting for information to cascade up through organizational layers and back down as a decision isn't leadership; it's latency. Jensen operated this way before the current wave. Quo has built Fireworks the same way. The evidence: 40 trillion tokens per day from a 200-person company.

There's a two-year window before token costs collapse — and whoever owns the specialized intelligence layer when it closes, wins

Supply chain constraints will hold token costs elevated for another 12 to 18 months. That window is simultaneously a moat-building period. High inference costs are forcing every enterprise to choose: which tasks genuinely need AI, which models to tune, which data to activate, which vendors to trust with their intelligence stack. By the time costs fall 10x and usage explodes 100x, companies that built proprietary specialized intelligence during the expensive years will be years ahead of anyone who waited for affordable general alternatives. The constraint looks like a problem. It's actually a filter. When it lifts, the race is already over.


Topics: AI inference, specialized intelligence, open source models, token economics, frontier models, AI infrastructure, Fireworks AI, enterprise AI, model training, RL post-training, data centers, GPU supply chain, AGI, sovereign AI, startup scaling

Frequently Asked Questions

What are the key takeaways from this AI cost analysis?
Token costs are collapsing 10x while usage will explode 100x, fundamentally threatening current AI giant valuations. The analysis predicts 10x token cost reduction within three years, followed by 100x usage growth. However, this creates a novel "scaling to bankruptcy" failure mode—companies could burn capital faster than revenue if growth outpaces cost decreases. Supply chain limitations will delay the full impact. The analysis also projects that every company will eventually own its own intelligence stack, just as each owns software, moving away from reliance on centralized providers like OpenAI and Anthropic.
Are OpenAI and Anthropic overvalued?
Yes, the analysis suggests current valuations are vulnerable to fundamental threats from collapsing token costs. If token prices fall 10x in three years as predicted, revenues could crater if usage growth doesn't match that pace. Unlike traditional SaaS with predictable unit economics, AI companies face "scaling to bankruptcy"—a failure mode where aggressive growth burns capital faster than token cost decreases recoup it. Fireworks processes 40 trillion tokens daily from customized models, demonstrating market fragmentation already undercutting centralized players' pricing power and market dominance assumptions embedded in their valuations.
What is 'scaling to bankruptcy' in AI?
"Scaling to bankruptcy is a real failure mode that didn't exist in the SaaS era." This describes a unique AI vulnerability where explosive usage growth becomes financially destructive. As companies pursue market share, if usage grows faster than token costs fall, they burn capital unsustainably. Traditional SaaS achieved profitable scale; AI infrastructure demands perpetual capital investment while pricing collapses. If 100x usage expansion occurs before 10x token cost reduction completes, rapid growth becomes fatal. The 3-year window for token cost decline creates timing risk—companies scaling aggressively now could consume capital faster than decreasing costs can generate proportional revenue gains.
Why will every company build their own AI stack?
The analysis argues that as token costs collapse, competitive advantage shifts from using centralized APIs to owning customized models. Jensen Huang's principle—"There's no specialized general company — every company exists on a special belief"—supports this shift. Fireworks already processes 40 trillion daily tokens from customized, not off-the-shelf, models, proving differentiation requires ownership. As costs fall 10x and usage explodes 100x, generic models become commoditized. Companies will own intelligence stacks "just as every company owns its own software," building specialized models aligned with their unique needs rather than relying on general-purpose providers, fragmenting the AI market away from centralized dominance.

Read the full summary of Are OpenAI & Anthropic Overvalued? How Token Costs Will Fall 10X & Usage Will Explode 100X on InShort