All-In Podcast cover
Technology & the Future

Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology

All-In Podcast

Hosted by Unknown

23 min episode
6 min read
5 key ideas
Listen to original episode

The brain is 10 billion times more efficient than today's best AI chips — and the entire industry hits a hard energy wall in three years.

In Brief

The brain is 10 billion times more efficient than today's best AI chips — and the entire industry hits a hard energy wall in three years.

Key Ideas

1.

Google AI dominates U.S. power consumption

Google's AI alone devours 12 GW — nearly a third of all U.S. data center power.

2.

Energy costs dominate AI inference economics

50% of every inference token's cost is pure energy, not compute or capex.

3.

Biological brains vastly outpace AI hardware

The brain is 10 billion times more efficient than today's best AI chips.

4.

Dynamical computing prototype proves hardware viability

The first physical dynamical computer exists — built in 5 months, chip back from fab.

5.

Cost reduction unleashes unprecedented market expansion

1000x cheaper inference doesn't shrink the market — it creates the largest one in history.

Why does it matter? Because AI runs out of energy in roughly three years — and the chip that proves the fix came back from the fab this summer.

Naveen Rao, co-founder of Unconventional AI, arrives with a number that reframes every other AI conversation: one company, Google, already pulls 12 gigawatts for AI services alone — nearly a third of all U.S. data center power. Demand grows exponentially; the grid grows linearly. In roughly three years, by Rao's math, the entire industry hits a hard energy wall.

• Google's AI services consume 12 gigawatts against ~40 gigawatts for all U.S. data centers • Energy already accounts for 50% of every inference token's cost — not hardware, not capex • The brain runs on 20 watts; today's AI hardware is 10 billion times less efficient, and that gap is engineering, not physics • The world's first physical dynamical computer, built in five months, generates images at ~500 nanowatts versus GPU-scale milliwatts

The global AI industry runs out of energy in roughly three years — and half of every inference dollar is already pure electricity

Three-point-two quadrillion tokens per month, just from Google. At roughly 10 joules per token, that's 12 gigawatts from a single company — against a total U.S. data center footprint of about 40 gigawatts. "We're going to run out of energy pretty fast in like 3 years or so," Rao estimates.

Capital won't fix this. New transmission capacity takes a decade to permit. Hyperscalers have already adapted: the procurement order has completely flipped. It used to go floor space → networking → GPUs. Now: "Today it's about energy. First you think about energy. I get the energy contract and then I have to figure out how to fill it."

The cost structure tells the same story. "About 50% of the cost of serving a token — so every time you try something on ChatGPT, 50% of that cost is energy." Whoever extracts the most intelligence from a fixed watt wins the margin war in the next cycle.

Von Neumann architecture hasn't changed since 1945 — and the energy debt from that design is coming due all at once

A GPU shuttles nearly 30 trillion bits in and out of memory per second. The human cortex — running 13 to 14 billion neurons — moves about 16 billion bits. The machine burns orders of magnitude more data movement to accomplish comparable reasoning, and that constant shuttling is exactly where the energy goes.

Rao traces the root cause back to ENIAC: "The operation of that computer in 1940, 1945 is actually very similar to how they operate today." Memory on one side, compute on the other, bits moving between them at enormous scale. The design was optimized purely for speed — "it doesn't contemplate energy efficiency." Moore's Law shrinking transistors absorbed this overhead for 80 years. That runway is gone.

The brain's efficiency isn't a curiosity — it's proof. A squirrel brain runs on 8 milliwatts and nails a thousand branch-jumps out of a thousand. "You could run over 100 squirrel brains on your phone." Mammalian brains operate within one or two orders of magnitude of the thermodynamic limit. Today's AI hardware is 10 billion times away from that same limit. The gap is engineering, not physics.

The first physical dynamical computer is back from the fab — built in five months, generating images at ~500 nanowatts

Not a paper. Not a simulation. A physical chip, back in the lab, with working image outputs.

Unconventional AI started in January with no team. The design taped out on June 1st. Five months from start to working silicon. The core innovation eliminates the Von Neumann data-movement problem at the root: "Each individual computing element is a memory." No memory interface. No bits shuttling between separated blocks. Coupled oscillators compute through phase dynamics — the same way synchronized metronomes find a common beat through the physics of the system itself, without any central controller pushing instructions around.

Power: roughly 500 nanowatts per image. A standard GPU runs at milliwatts — a thousand times a nanowatt. The efficiency gap spans multiple orders of magnitude. The two-year product roadmap is a full data-center rack, tokens in and tokens out through a network cable, with completely different guts inside.

Strip connections from a dynamical network and performance improves — the architecture gets stronger at scale, not weaker

Every deep learning architect knows the N-squared scaling problem: double the elements and connectivity quadruples, making each incremental expansion more expensive than the last. Dynamical systems invert this. Remove connections between oscillators and the system doesn't degrade — it becomes more trainable.

"You can actually not only throw away some of the connections, but you can actually get better behavior out of the whole system," Rao explains. "It's one of these rare things where you get something that's more efficient, that's actually more scalable and even gives you more performance." Most architectural efficiency gains in computing erode as scale grows. This one widens precisely when it matters most.

A 1,000x drop in inference cost doesn't compress the AI market — Jevons Paradox turns it into the largest one in history

Rao ends with Jevons Paradox: cut the cost of an asset by 1,000x and consumption rises by more than 1,000x. The trillion-dollar AI market estimate isn't a ceiling to optimize margins within — it's a floor. "If you make something 1,000th the price, you'll consume more than 1/1,000th of it." His conclusion: "I think this will create the largest market that humanity's ever seen."

The right frame for this technology isn't cost reduction on existing workloads. It's which entirely new categories — ambient computing, autonomous robotics, real-time medicine — cross the economic threshold when inference approaches free. The trillion-dollar market is the starting point, not the destination.


Topics: AI hardware, energy efficiency, chip design, dynamical systems, Von Neumann architecture, brain-inspired computing, AI infrastructure, Moore's Law, data centers, Jevons Paradox, oscillator computing, compute-in-memory

Frequently Asked Questions

What is the AI energy crisis that Naveen Rao discusses?
"Google's AI alone devours 12 GW — nearly a third of all U.S. data center power," and critically, "50% of every inference token's cost is pure energy, not compute or capex." These statistics reveal an unsustainable trajectory. The entire industry hits a hard energy wall in three years, making current AI infrastructure economically and physically unsustainable. This energy consumption crisis demands fundamental architectural innovations to enable continued AI advancement without triggering industry collapse or regulatory intervention.
How much more efficient is the human brain compared to current AI chips?
"The brain is 10 billion times more efficient than today's best AI chips," representing a fundamental mismatch between biological and digital computing architectures. This staggering efficiency gap explains why current AI systems face severe energy constraints and why solving the efficiency challenge requires moving beyond traditional transistor-based computing. The brain achieves superior performance using biochemical processes consuming minimal power—a model that should fundamentally inspire next-generation computing design principles and architecture innovations.
What is 4D computing and how does it solve the energy problem?
"The first physical dynamical computer exists — built in 5 months, chip back from fab." This breakthrough represents a radical departure from conventional computing approaches. 4D computing leverages physical dynamical systems to achieve computation through fundamentally different mechanisms than silicon. By bridging the 10-billion-fold efficiency gap between current chips and biological brains, this technology offers a pathway beyond the three-year energy wall. The rapid development timeline suggests practical deployment could soon make conventional AI infrastructure architecturally obsolete.
What are the market implications of achieving 1000x cheaper AI inference?
"1000x cheaper inference doesn't shrink the market — it creates the largest one in history." This counterintuitive principle reveals how dramatic cost reductions unlock exponential demand for AI applications. When inference costs approach zero, previously economically unfeasible use cases become viable across billions of devices. This market expansion extends far beyond optimizing current applications; it fundamentally transforms how AI integrates into society and commerce, catalyzing a technological revolution rivaling the personal computer or internet's economic impact.

Read the full summary of Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology on InShort