English
Based on "The AI Model Built for What LLMs Can't Do" from Every
Watch the original video
The AI That Doesn't Guess: Unlocking Verifiable Intelligence with Energy-Based Models
In the rapidly evolving world of artificial intelligence, Large Language Models (LLMs) have taken center stage. Their ability to generate human-like text, translate languages, and answer complex questions has captivated the public imagination and drawn immense investment. Yet, for all their impressive capabilities, LLMs operate on a fundamental principle that makes them inherently unsuitable for certain critical tasks: they guess.
This "guessing game" is precisely what Logical Intelligence, a foundational AI company, aims to transcend with a different kind of model: Energy-Based Models (EBMs). Led by founder and CEO, Yee, Logical Intelligence is building AI designed for correctness, verifiability, and tasks where the stakes are too high for hallucination.
"LLMs are naturally non-autoregressive," Yee explains, righting a common misconception. "There are no sequences of tokens, and that's what makes it fundamentally different." This seemingly subtle distinction underpins a radical departure from the LLM paradigm, promising a future where AI can be trusted with our most critical systems.
The Peril of the Guessing Game
Imagine an AI driving your car. Now imagine that AI, an LLM, is prone to hallucinating 20% of the time, potentially taking you to the wrong place. While some might find that "kind of interesting" for a joyride, the stakes escalate dramatically when considering a plane. "How about the plane?" Yee challenges. "You take a plane from yourself to New York and someone says, 'You know, like 20% of the time it might just like the next word not going to match and it's going to go down.' How would you feel about it?"
The answer is obvious: terrifying. Our lives depend on deterministic, verifiable systems. Current planes are run by such systems, and while AI's integration into every facet of life seems unavoidable, Yee argues that the current LLM approach falls short for mission-critical applications like code generation, chip design, or autonomous vehicles.
The core issue lies in the LLM's architecture. It's an autoregressive model, meaning it predicts the next token in a sequence based on the preceding ones. This process, while powerful for language generation, is essentially a "guessing game" based on probabilities derived from vast training data.
"You're like a black box," Yee explains, describing the LLM's internal workings. "You don't have access to what's inside until it's all processed." To mitigate this, companies often attach external verifiers (like machine-verifiable proof languages such as Lean 4) to check the LLM's output. However, this is costly. "Even if you attach external verifier, even you fine-tune this LLM specifically for the task you're trying to create, you're still not solving the problems of tokens being expensive. It takes compute for you to play a guessing game."
A Bird's-Eye View: How EBMs Navigate Reality
To grasp the fundamental difference, consider navigating a complex environment. An LLM, in Yee's analogy, is like trying to navigate San Francisco with "tunnel vision." You choose one direction at a time, unable to see other options.
"You cannot choose... one direction at a time," Yee illustrates. "And sometimes you take the wrong turns just because you hallucinate... like there might be a hole in the road and you're just going to fall. And you might see this hole, but you cannot turn back because you're autoregressive LLM." This leads to wasted compute, endless searching, and potentially never reaching the destination.
An EBM, by contrast, has a "bird's-eye view all the time." If it sees a hole, it can choose a different route. This isn't just an analogy; it speaks to the EBM's non-autoregressive, token-free nature. There's no guessing game of what the "next word" or "next action" might be. Instead, EBMs can "oversee all possible scenarios." Their architecture allows for "self-alignment" as they process information, making them inspectable. "As it's performing, you can open it anytime during the training and you could see what's happening in there. So you cannot do this with LLMs."
This inspectability, coupled with the ability to attach external verifiers, provides "verification sort of on both sides, inside and outside."
The Physics of AI: Minimizing Energy
At the heart of EBMs is the principle of energy minimization, a concept borrowed directly from physics. "EBM just simply means energy-based model," Yee clarifies. "What is energy? Energy-based... it comes from physics. It's a very popular term when you're trying to minimize the energy."
Think of it this way: everything in nature seeks a state of lowest energy. When you're tired after a long day, you don't dance around; you collapse onto the couch to minimize your energy. This "most comfortable configuration" corresponds to the lowest potential energy.
EBMs apply this universal principle to AI. They construct an "energy landscape" – a metaphorical map of all possible states based on observed data. The highest points on this landscape represent less probable scenarios (e.g., you dancing when tired), while the lowest points represent the most probable and "comfortable" states (e.g., you on the couch).
"Essentially all of this picture can be mapped into something we call energy landscape when the... it's going to look like a map," Yee elaborates. "So, you're going to have highest points, you're going to have lowest points. The highest points we can associate less probable scenarios... the lowest point is going to be you on the couch."
During training, the EBM observes data (your behavior, your tiredness, etc.) and learns to shape this energy landscape. It then uses algorithms to navigate this landscape, always seeking the lowest energy states, which correspond to the most probable and optimal outcomes. Crucially, this process doesn't involve tokens or language; it's a direct mapping of data to a structured landscape.
Beyond Pattern Recognition: Understanding and Latent Variables
Another critical distinction between EBMs and LLMs lies in their approach to data. LLMs are exceptional at identifying patterns in vast datasets. But do they "understand" in a human sense? Yee argues no. "LLMs don't understand the data. It's just you feed a lot of data into it and it's sort of like, 'Hey, I got it. Like, okay, I know what's the most probable scenario here and here we are.'"
EBMs, particularly those with latent variables, go a step further. They not only identify patterns but also "try to understand the pattern." This "understanding" translates into basic knowledge and rules about the world. For instance, an EBM observing your apartment might infer rules like, "there's a kitchen for cooking, there's a bathroom, there's a sofa."
This knowledge is stored in "latent variables," which Yee describes as "knowledge storage" – a dataset about the rules of your data, also in the form of an energy landscape. This allows the EBM to understand its environment and adapt to changes. If you get a new couch, it still knows what to do with it because it understands the underlying rules of furniture and human behavior, not just specific examples.
This capability is particularly powerful for data analysis, where searching for patterns and rules is paramount. "It's something where language is not going to be helpful to you," Yee asserts. "If you try to attach the rules about your data and those data is like numbers and some relationships and functions to like American English and words in American English... you're losing a lot of information." EBMs work directly with the data, understanding its inherent structure without the translation layer of language.
Furthermore, EBMs are adept at working with sparse or incomplete data, a feature that evolved from traditional EBMs through diffusion models. By injecting noise and changing navigation strategies, they can reconstruct energy landscapes even when data is limited.
Building Trust: From Code to Autonomous Systems
The implications of EBMs for building trustworthy AI are profound. Consider "vibe coding," where developers use LLMs to generate code, then spend significant time debugging and fixing the patchwork of solutions. Logical Intelligence dreams of "generating formally verified code and automate the coding entirely," moving from vibe coding to "vibe code specifications."
EBMs, with their internal and external verifiers, can ensure logical compatibility and compliance. They can generate mathematical proofs to certify code, telling engineers, "Hey, this part of your code is not compatible by logic. This is potentially how you fix it."
Beyond mere compilation, EBMs address the critical question: "Is this code actually doing what you want it to be doing?" While AI cannot read a human's mind, EBMs can be constrained to follow a precise set of human-defined rules. "LLM can misbehave based because you cannot constrain it. It just hallucinates. And EBM can be constrained. You can come up with a set of constraints and EBM just forced to follow it." This determinism is vital for safety-critical applications like autopilot systems, where misbehavior could have catastrophic consequences.
Navigating the AI Ecosystem
Despite the compelling advantages of EBMs, the current AI investment landscape remains heavily skewed towards LLMs. "LLMs historically the first form of AI which gave us aha effect," Yee acknowledges. The massive investments already poured into LLM infrastructure – from data centers to hardware – create a powerful inertia. "It becomes like a one giant thing which is impossible to break."
However, Logical Intelligence isn't seeking to replace LLMs entirely but rather to complement them. "We are very much compatible with LLMs," Yee explains. "You could put LLM on top of us. EBMs compatible with transformers, transformers can, you know, work with any LLMs."
This synergistic approach offers a path forward. LLMs can handle language-related tasks, while EBMs can take on spatial reasoning, engineering, data analysis, and other non-language-dependent tasks requiring precision and verifiability. This could make LLM deployments cheaper and more reliable by outsourcing suitable tasks to EBMs. "If somebody comes to big tech LLM and say, 'Hey, can you try to do my taxes?' LLM not going to solve this. But if it's attached to EBM, we can take care of that."
While some might perceive a plateau in LLM progress, Yee believes that the apparent advancements are often within the existing paradigm, not a fundamental shift. The real breakthroughs, she suggests, will come from exploring alternative architectures like EBMs that address core limitations in correctness and understanding.
As AI permeates every aspect of our lives, the demand for systems that are not only intelligent but also reliable, verifiable, and safe will only grow. By moving beyond the "guessing game" and embracing principles of energy minimization and deep understanding, Energy-Based Models offer a powerful pathway to building the next generation of trustworthy artificial intelligence.
Based on "Are Human Drivers Finally Obsolete? | Freakonomics Radio" from Freakonomics Radio Network
Watch the original video
The Road Ahead: How Driverless Cars Went From Sci-Fi Dream to Reality
By PJ Vogt, Host of Search Engine
Imagine a future where the concept of a "driver" refers not to a person, but to a machine. A world where the most dangerous activity many of us routinely engage in—driving a car—is outsourced to an unblinking, unyielding, and un-distracted artificial intelligence. This isn't a distant sci-fi fantasy; it's a reality unfolding in cities across the globe right now.
My journey into the fascinating, complex, and often contentious world of driverless cars began not with a grand vision, but with a very personal, very painful experience. After a bench-pressing injury led to a hernia and subsequent surgery, I found myself with limited mobility, visiting a friend in San Francisco. It was then that I took my first ride in a Waymo robo-taxi. The experience was transformative. One moment, I was pressing a button on my phone; the next, a car pulled up, completely empty, its steering wheel turning itself as if by an invisible hand.
"The first time it feels like the first time you're in an airplane," I tried to explain to someone recently. "By the third time, it feels like you're in an elevator." This immediate normalization of a profoundly futuristic technology made me realize a lot was about to change, and I was confused why more people weren't talking about it. This led me down a rabbit hole, culminating in a two-part series for my podcast, Search Engine, exploring the history, technology, and human impact of this revolution.
The Echoes of the Past: When New Technology Met Old Fears
To truly understand the seismic shift driverless cars represent, it helps to look back. Picture yourself almost 200 years ago, pre-dawn on a Monday. A hard rapping at your window wakes you—it’s the "knocker-upper," a job that existed for another century, tapping on windows with a long stick because alarm clocks hadn't been invented. Outside, the "lamplighter" is still making his rounds, extinguishing the gas street lamps he lit the night before. And you? You're a "driver," a person holding the reins of a horse, taking passengers where they need to go.
The knocker-upper is now your smartphone alarm. The lamplighter is the electric streetlight. But the driver? That job, and the routine human task of driving, has persisted. This story is about whether that's about to change, and how the word "driver" might soon come to mean a machine, just like "dishwasher" or "computer" once referred to people.
As Alex Davies, author of the excellent book Driven: The Race to Create the Autonomous Car, points out, human driving is fraught with limitations. "I can't always pay attention to everything," he admits. "I get tired." Even with the best intentions, we are prone to distraction, fatigue, and road rage. Driving is, for most of us, the riskiest behavior we routinely engage in. Nationally, car accidents cause about 1 in 100 deaths, comparable to guns or opioids.
It’s this inherent human fallibility that forms the core pitch for the driverless car. They don't get drunk, tired, or distracted. They never text or feel road rage. And they are no longer a distant dream. Robo-taxis are already providing millions of rides in ten American cities, and even more widely in China.
The dream of a self-driving vehicle is almost as old as the automobile itself. When cars first arrived in the 1800s, replacing horses, society lost a crucial element: sentience. A horse wouldn't simply run off a cliff if you let go of the reins. Early automobiles, powerful and non-sentient, were met with passionate resistance. People feared for jobs—horse breeders, farriers, feed suppliers, carriage manufacturers, and especially "teamsters" (the truckers of their day). They also feared for safety. "Red flag laws" proposed requiring a person to walk in front of a car waving a giant red flag to warn people. One Pennsylvania law, thankfully vetoed, would have required drivers encountering livestock to stop, disassemble their car, and hide the parts behind bushes!
Directionally, those "crazy anti-car activists" were right. Cars did initially wipe out many jobs (though they created more), and they were incredibly unsafe. Cities like Detroit, which initially embraced cars without regulation, saw astonishing death rates, many of them children, in the early 1900s. It took decades for society to adapt, inventing laws, licenses, driver's ed, highways, seatbelts, and airbags to make driving less deadly. Yet, the safety problem remains.
The Desert Race: Where Robots Learned to Think
The quest to make cars more "sentient" like the horses they replaced continued for decades, fueled by early, often primitive, ideas like radio-controlled cars or magnets under roads. But the true breakthrough wouldn't arrive until the turn of the millennium, courtesy of a little-known military agency: DARPA.
DARPA, the Pentagon's research arm, has a mission to keep American technology a generation ahead, having funded everything from GPS to the early internet. In 2002, DARPA director Tony Tether, a former door-to-door salesman with a flair for the dramatic, decided to accelerate driverless car development with a contest: the DARPA Grand Challenge. His original idea—race self-driving cars down the Las Vegas Strip—was quickly deemed "insane" due to the gridlock it would cause. So, the desert outside Las Vegas became the proving ground.
The goal was military: to develop autonomous vehicles that could navigate roads potentially filled with explosive devices, saving American soldiers' lives. The prize: $1 million. Rules were open: any vehicle type, but no attacking other competitors. Many entrants, surprisingly, came from the world of BattleBots.
This race was not just about the vehicles; it was about the engineers who would go on to lead the billion-dollar companies creating today's driverless cars. Their differing philosophies—from methodical academic rigor to entrepreneurial risk-taking—would shape the future.
One such engineer was Chris Urmson, then a PhD student at Carnegie Mellon University. He joined the "Red Team" to build "Sandstorm," a bright red Humvee bristling with futuristic sensors. For Urmson, the challenge was teaching a computer the basics of operating a Humvee, from steering to braking.
Then there was Anthony Levandowski, a charming, gangly Berkeley grad student and hustler. Lacking the resources of academic powerhouses, Levandowski opted for a radical approach: the race's only self-driving motorcycle, "Ghost Rider." It had almost no chance of winning, but it was guaranteed to get attention.
The 2004 Grand Challenge was, in a word, "an utter hysterical disaster." Ghost Rider immediately toppled over because Levandowski forgot to flip a stabilization switch. Other vehicles fared no better, driving onto berms, flipping, or inexplicably U-turning back to the start. Even Sandstorm got stuck, its tires melting from the effort. It was a disaster, but it had flushed out and jump-started a nascent community of roboticists.
Watching this debacle from the sidelines was Sebastian Thrun, a legendary German-born roboticist and AI expert who had taught at Carnegie Mellon before moving to Stanford. Thrun saw a fundamental error in everyone's approach. "All the teams treated this like a hardware problem," he observed. "I looked at this and said, 'Well, wait a minute... that's actually really a software problem.'"
Thrun, who had lost a friend to a car accident as a teenager, envisioned something far grander than military applications. He saw the potential to save "a million lives a year" worldwide if driverless cars became ubiquitous. This personal mission would drive his subsequent work.
Eighteen months later, in October 2005, DARPA doubled the prize to $2 million for the second Grand Challenge. Chris Urmson returned with the Carnegie Mellon team. Anthony Levandowski's motorcycle still didn't work, getting knocked out in qualifiers. But Thrun, with his Stanford team, brought a modest blue Volkswagen SUV named "Stanley."
Stanley was Thrun's software-first vision brought to life. He focused on artificial intelligence, primitive by today's standards, but revolutionary for 2005. Stanley's "eyes" looked far ahead, recording what it saw as it drove down dirt roads. It could then "train itself," learning to distinguish good driving surfaces (grass) from bad ones (mud), detecting patterns and generalizing its knowledge 30 times a second.
The 2005 race was a resounding success. Multiple vehicles finished the 132-mile circular maze. Stanley, the unassuming blue SUV, passed the bulkier Carnegie Mellon entries, conquering a treacherous mountain pass, and crossed the finish line first. Thrun, dressed like a race car driver, celebrated a victory not just for his team, but for the entire "community" of roboticists. It was a made-for-TV moment, years before the cutthroat competition for driverless cars would begin.
Google's Secret Project: Teaching a Machine to Nudge
The success of the DARPA Grand Challenge caught the eye of another spectator that day: Larry Page, co-founder of Google, who attended in a baseball hat and sunglasses as a disguise. Page, who had wanted to do his grad school thesis on autonomous vehicles before being steered toward search engines, was captivated.
He soon hired Sebastian Thrun, along with Anthony Levandowski, initially for Google Street View, modifying Stanley's roof-mounted cameras to photograph American streets. But Page's true dream was the driverless car. In 2009, he approached Thrun with a mission: "Sebastian, I think you should build a self-driving car that can drive anywhere in the world."
Thrun's immediate reaction was a resounding "No." He believed taking the technology from the empty desert and putting it on a bustling San Francisco street would kill someone. Page persisted, day after day. Finally, Page challenged him: "Give me the technical reason why it can't be done." Thrun went home, and to his profound realization, he couldn't think of one. "Experts are usually experts of the past, not the future," he reflected. "If you ask an expert about innovation, something crazy new, they're the least likely person to say, 'Yes, it can be done.'"
Thus began Project Chauffeur in 2009, Google's secret self-driving car initiative. Led by Thrun, with Chris Urmson running day-to-day operations, Anthony Levandowski on hardware, and Dmitri Dolgov on planning, the small team of 11 engineers reported directly to Larry Page. They had two challenges: safely log 100,000 miles on public roads, and complete the "Larry 1K"—ten tricky 100-mile routes across California without a single human takeover.
To get started, they licensed Stanford's code and Levandowski bought eight Toyota Priuses, retrofitting them with radar, cameras, and a spinning 360° lidar system. These cars, initially given cool names, quickly became "Prius 27" and so on, as the fleet expanded.
Don Burnett, a researcher who had joined the team after losing a friend in a car accident, worked on "nudging behavior." Just as a human driver instinctively nudges left when a large truck passes on the right, Burnett's job was to teach a computer this subtle, often unconscious, behavior. "You're trying to encode the behavior that you would use as a driver under kind of partially good perception," he explained. "It's a really tricky problem."
Early testing was done in secret, in the Shoreline Amphitheatre parking lot or an empty airplane runway near Google's offices. But in spring 2009, they took a Prius onto the Central Expressway. Immediately, a critical flaw emerged: the car was "swerving wildly," like a "drunken sailor." The small oscillations that went unnoticed on a runway became a major problem on a public road.
The team adopted rigorous safety precautions: two-person teams, a safety driver ready to take over, and a partner watching a monitor, calling out discrepancies between what sensors saw and what was actually on the road. This was the painstaking process of teaching a car to drive: logging errors, troubleshooting, updating code.
Burnett became obsessed with the question, "Why do humans drive the way they drive?" He found there were no simple answers. The machine learning approach revealed that context profoundly affects perceived comfort. For example, a lateral acceleration of 2 meters per second squared might be comfortable on a highway on-ramp, but doing the same in a cul-de-sac U-turn would feel "incredibly uncomfortable," like "Mario Kart." The limit for a cul-de-sac, it turns out, is around 0.75 m/s^2. "It's almost three times less than you would be willing to tolerate as you accelerate onto a highway," Burnett noted. These subtle, contextual nuances were critical to making the experience comfortable for passengers.
By 2010, the team was on a roll. They started "knocking out" the Larry 1K routes, each presenting unique challenges, from the Bay Bridge to Lombard Street. The challenge was set up like a video game: try, fail, learn, repeat, until a route was completed without human intervention. They expected it to take two years; they did it in just over one. They celebrated each completed route with a bottle of $13.99 Korbel champagne, signing their names on it.
By fall 2010, the Larry 1K was complete. The team celebrated, throwing each other in the pool at Sebastian Thrun's house. They had pulled off a miracle, safely navigating tricky California roads with a driverless car, albeit with human supervision and extensive coding. But as the champagne bubbles settled, a new question arose: "Okay, and now what?"
The Crossroads of a Revolution
The success brought with it new challenges. What exactly was the product they were building? Thrun recalled internal debates: should they have bought Tesla (worth $2 billion at the time)? Was this an assistive technology, a feature to aid human drivers, or a disruptive replacement, waiting until cars could fully drive themselves, perhaps as robo-taxis? Thrun eventually gravitated towards the latter.
But with their secret out, competition would arrive. The team itself would begin to schism, and one member, believing the pace too slow, would take matters into his own hands in an extreme way—a story for another time.
The journey of the driverless car is a testament to human ingenuity, born from personal tragedy, fueled by visionary engineers, and refined through relentless trial and error. It's a story of how technology, once confined to desert races and secret labs, is now quietly, inexorably, reshaping our daily lives, transforming our cities, and forcing us to confront a future where the driver, as we know it, may finally become obsolete.
Based on "Jensen Huang – Will Nvidia’s moat persist?" from Dwarkesh Patel
Watch the original video
The Maestro of Moats: Jensen Huang's Bold Defense of Nvidia's AI Dominance
In the exhilarating, often dizzying world of artificial intelligence, one company has emerged as the undisputed kingmaker: Nvidia. Its GPUs power the most sophisticated AI models, driving unprecedented technological leaps and a staggering market valuation. Yet, as the industry matures and competition intensifies, a critical question looms: can Nvidia's formidable "moat" truly persist?
In a candid conversation with Dwarkesh Patel, Nvidia CEO Jensen Huang offered a robust defense of his company's enduring advantage, revealing a multi-layered strategy that spans from fundamental physics to global supply chain orchestration and a unique philosophy of strategic investment.
From Electrons to Tokens: Nvidia's Core Identity
A common, perhaps "naive," interpretation of Nvidia's business suggests vulnerability. Observers note that Nvidia designs software (GDS2 files) and relies on external partners—TSMC for logic, SK Hynix, Micron, and Samsung for memory, and Taiwanese ODMs for assembly. If AI is poised to commoditize software, could Nvidia, at its core a software company, be next?
Jensen Huang dismisses this notion with a foundational principle: "In the end, something has to transform electrons to tokens." This isn't just a technical process; it's an "incredible journey" demanding immense "artistry, engineering, science, and invention." He posits that making one "token" more valuable than another—the very essence of AI's progression—is inherently difficult to commoditize.
Nvidia's role, as Huang sees it, is to be the indispensable orchestrator of this transformation. "The input is electrons, the output is tokens. In the middle is Nvidia," he states. Their job is to do "as much as necessary and as little as possible" to enable this transformation at "incredible capabilities." The "as little as possible" part is key: Nvidia partners extensively, building the largest ecosystem of upstream and downstream partners, from chip manufacturers and memory providers to cloud companies, application developers, and model makers. This five-layer cake of AI, according to Huang, is where Nvidia embeds itself, focusing on the "insanely hard" parts that cannot be commoditized.
The Software Paradox: More Tools, Not Less
The fear that AI will commoditize enterprise software is also, in Huang's view, misplaced. He argues the opposite will happen: "I think the number of agents is going to grow exponentially, and the number of tool users is going to grow exponentially." Tools like Excel, PowerPoint, Cadence, and Synopsys won't disappear; they'll become even more widely used, albeit by AI agents rather than solely human engineers.
"Today we're limited by the number of engineers. Tomorrow, those engineers are going to be supported by a bunch of agents," Huang predicts. This means an explosion in the "instances of all these tools," leading to a "skyrocket" in software company revenues. The only current limitation, he notes, is that "the agents aren't good enough at using their tools yet." Once they are, or once companies build agents specifically for their tools, the software market will expand dramatically.
The Supply Chain: A Moat Forged in Trust and Scale
Beyond technological prowess, Nvidia's moat extends deep into the global supply chain. Recent reports of Nvidia's colossal purchase commitments—potentially $250 billion—spark questions: is this simply locking up scarce components? Huang confirms these "enormous commitments," both explicit and implicit, are a significant factor.
He describes a unique, almost evangelical role he plays in inspiring upstream partners. "I said to the CEOs, 'Let me tell you how big this industry is going to be, let me explain to you why, let me reason through it with you, and let me show you what I see.'" This process of "informing, inspiring, and aligning" convinces suppliers to make massive investments for Nvidia that they might not for others. Why? Because "they know that I have the capacity to buy their supply and sell it through my downstream."
Nvidia's GTC conference, Huang explains, serves as a physical manifestation of this ecosystem, bringing together the entire "360 degrees, the entire universe of AI." It's where upstream and downstream partners connect, see the latest AI advances, and meet "AI natives" – the startups building the future. This visibility reinforces their confidence in Nvidia's demand, allowing the company to "build for a future" where its scale could reach "a trillion dollars."
Conquering Bottlenecks: A Two-to-Three-Year Problem
The sheer scale of Nvidia's growth raises legitimate concerns about the supply chain's ability to keep pace. With Nvidia being the largest customer on TSMC's advanced nodes (N3, N2) and AI projected to consume a dominant share of future capacity, how can they continue to double output year after year?
Huang frames this as a "good condition": instantaneous demand exceeding supply signifies a booming industry. He acknowledges temporary bottlenecks, even humorously suggesting inviting "plumbers" to GTC as a critical component. However, he asserts that hardware-related bottlenecks are inherently short-lived.
"If we're too far apart, if one particular component is too far away, the industry swarms it," he explains. He cites CoWoS packaging technology as an example: initially a bottleneck, it was "swarmed" with investment and scaled rapidly. Now, TSMC is aligning CoWoS and future packaging with logic and memory demand.
Nvidia actively "prefetch[es] the bottlenecks years in advance," investing in new technologies (like silicon photonics with Lumentum and Coherent) and licensing patents to keep the ecosystem open. His confidence in scaling is striking: "None of that is impossible to scale quickly. All of that is easy to do within two or three years. You just need a demand signal. Once you can build one, you can build ten, and once you can build ten, you can build a million."
The real long-term concern, in Huang's view, isn't chip or packaging capacity, but rather fundamental energy policy. "You can't create an industry without energy," he warns, highlighting the need for robust energy infrastructure to support the reindustrialization efforts and the construction of "AI factories."
Moreover, capacity increases are only one part of the equation. Nvidia constantly drives "computing efficiency by 10x, 20x, and in the case of Hopper to Blackwell, 30x to 50x." This combination of increased supply and dramatically improved efficiency mitigates the impact of any temporary bottlenecks.
Beyond TPUs: The Power of Accelerated Computing
The rise of custom ASICs like Google's TPUs, which power major models like Claude and Gemini, presents a direct competitive challenge. Critics argue that AI primarily involves predictable matrix multiplies, for which specialized TPUs are optimally designed, potentially outperforming flexible GPUs.
Huang strongly refutes this, emphasizing Nvidia's focus on "accelerated computing, not a tensor processing unit." He stresses that while matrix multiplies are important, they are not the only part of AI. The invention of new attention mechanisms, hybrid architectures, and fused diffusion/autoregressive techniques demands a "generally programmable" architecture.
"The ability to invent new algorithms is really what makes AI advance so quickly," Huang states. Moore's Law, with its 25% annual gains, isn't enough; 10x or 100x leaps require fundamental algorithmic changes, which are only possible with programmable systems like CUDA. He points to the 50x energy efficiency jump from Hopper to Blackwell, achieved through "new models, like MoEs, that are parallelized, disaggregated, and distributed across a computing system." Nvidia's "extreme co-design company" approach, affecting processors, systems, fabric, libraries, and algorithms simultaneously, is what enables these leaps.
CUDA's Enduring Reign: Ecosystem, Install Base, and TCO
Even with hyperscalers like OpenAI developing custom kernels with tools like Triton, Huang argues that CUDA remains invaluable. He highlights three key advantages:
- Richness and Programmability: CUDA is a vast ecosystem, supporting every framework. Nvidia actively contributes to tools like Triton, ensuring their backend leverages Nvidia's technology. This provides a trusted, robust foundation for developers.
- Install Base: "The single most important thing you want is an install base," Huang says. With "several hundred million GPUs out there now," developers building on CUDA know their software will run "everywhere," from cloud servers to robots. This ubiquity is "incredibly valuable."
- Versatility and Reach: Nvidia systems are in "every cloud" (Google, Amazon, Azure, OCI) and available on-prem. This unparalleled versatility means AI companies aren't locked into a single provider, making Nvidia the default choice for those seeking maximum reach.
Huang also asserts that Nvidia offers the "best performance per TCO in the world, bar none." He challenges competitors like TPU and Trainium to demonstrate superior cost-effectiveness using benchmarks like InferenceMAX and MLPerf, claiming "nobody wants to show up." He also points to Nvidia's "highest tokens per watt architecture," which maximizes revenue generation for data centers. This combination of superior TCO, energy efficiency, and market reach creates a powerful "flywheel" effect.
A Surprising Admission: Nvidia's Strategic Investments
Perhaps the most surprising revelation came when discussing why major AI companies like Anthropic have diversified their compute away from Nvidia, partnering with Google/Broadcom for TPUs. Huang conceded a past "mistake": "I didn't deeply internalize how difficult it would be to build a foundation AI lab like OpenAI and Anthropic, and the fact that they needed huge investments from the supplier themselves."
He explains that Google and AWS were able to make multi-billion dollar upfront investments in Anthropic, securing their compute business, something Nvidia was not "in a position to do" at the time. "I didn't realize we needed to," he admitted, thinking they could "just go raise from VCs, for God's sakes."
"But I'm not going to make that same mistake again," Huang declared. Nvidia is now "delighted to invest in OpenAI" and Anthropic, recognizing that "the world needs them to exist." This strategic shift, from purely selling hardware to also investing in key AI labs, underscores Nvidia's commitment to nurturing the entire AI ecosystem.
Despite having vast cash reserves, Nvidia resists becoming a cloud provider itself. This adheres to their "as much as needed, as little as possible" philosophy. Building the core computing platform is a unique contribution, but cloud services are a crowded field where others can step in. Instead, Nvidia invests in "neoclouds" like CoreWeave, Nscale, and Nebius, ensuring the ecosystem thrives.
Finally, Huang reveals a core principle: "Don't pick winners." Drawing from Nvidia's own early struggles in the graphics industry, where their initial architecture was "precisely wrong," he emphasizes humility. "If you would have taken those 60 graphics companies and asked yourself which one was going to make it, Nvidia would be at the top of that list not to make it," he muses. Therefore, Nvidia supports all foundation model companies, viewing it as "imperative to our business."
The Allocation Myth: No Highest Bidder, Just Orders
Addressing persistent rumors about how Nvidia allocates its scarce GPUs, Huang firmly refutes the idea of "fracturing the market" or favoring the "highest bidder." He outlines a straightforward process: forecasting, placing purchase orders, and "first in, first out."
"If you don't place a PO, all the talking in the world won't make a difference," he states. While minor adjustments might occur based on a customer's data center readiness to maximize throughput, the core principle is simple. He even debunks the widely reported anecdote of Larry Ellison and Elon Musk "begging for GPUs" at dinner, confirming the dinner happened but "at no time did they beg for GPUs. They just had to place an order."
Why not simply sell to the highest bidder? "Because it's a bad business practice," Huang asserts. "You set your price and then people decide to buy it or not."
A Moat Built on Vision and Execution
Jensen Huang's narrative paints a picture of a company whose dominance is not accidental but deeply strategic, built on a unique blend of technological innovation, ecosystem cultivation, and prescient market understanding. From the fundamental transformation of "electrons to tokens" to the orchestration of a global supply chain, Nvidia's moat is multifaceted.
It's a moat reinforced by an unwavering commitment to programmability, a vast install base, and an ecosystem that fosters continuous innovation. And while past missteps are acknowledged, Nvidia's willingness to adapt its investment strategy shows a dynamic, forward-looking approach. In the relentless race for AI supremacy, Jensen Huang believes Nvidia's position, earned through decades of dedicated effort, is not just persistent, but increasingly indispensable.
Based on "Trump’s Risky Strategy to Blockade Iran’s Blockade" from New York Times Podcasts
Watch the original video
Strait of Hormuz: Trump's High-Stakes Gambit in the Blockade of Blockades
The waters of the Strait of Hormuz, a critical choke point for global oil supplies, have become the stage for a perilous game of "blockade chicken." In a bold and risky maneuver, the United States has deployed a formidable naval presence, enforcing what it calls a blockade of Iran, designed to bring an end to a simmering conflict on American terms. But this isn't just any blockade; it's a direct counter to Iran's own control over the vital shipping channel, creating a complex and potentially explosive situation that experts warn could reshape the global economy for years to come.
At its core, the American strategy is stark: strangle Iran's economy by cutting off its oil and gas revenues, forcing Tehran back to the negotiating table. "We're blockading the ports in Iran where they get oil and gas shipments," explains White House correspondent David Sanger. "Without oil and gas money, Iran has no economy." This aggressive approach, however, carries immense dangers and uncertain outcomes, prompting a critical look at its strategy, the risks it poses, and whether it can truly work.
The Genesis of a Risky Strategy
The decision to impose a naval blockade emerged directly from a diplomatic impasse. Following Vice President JD Vance's recent, fruitless negotiations in Pakistan—which saw him return empty-handed—the Trump administration found itself confronting a "pretty messy situation." A ceasefire was in place, but Iran continued to exercise control over who transited the Strait of Hormuz, a vital artery for global commerce.
For 47 years, since the Islamic revolutionary government came to power in 1979, Iran had largely allowed free passage through the strait. But in the recent conflict, they had begun to assert control, stopping traffic and even imposing "tolls" on passing ships, some reportedly as high as $2 million. This was a dynamic the US administration found "intolerable." The US Navy's objective became clear: reverse the situation and ensure that it was the US, not Iran, controlling traffic through the strait. While this sounds straightforward on paper, given the might of the US Navy, its execution promises to be anything but.
An Act of War: Defining the Blockade
Militarily, a naval blockade is a grave measure. "By definition, a blockade is an act of war," states military correspondent Eric Schmidt. It involves one nation using its military power to block the transit of ships from other countries, either through the threat of force or the actual boarding and seizing of vessels.
The US deployment is substantial: approximately 10,000 sailors aboard more than a dozen warships, ranging from aircraft carriers to destroyers and Marine-carrying vessels, are now positioned outside the Strait of Hormuz. Their mission is to intercept ships either leaving the Persian Gulf or attempting to enter it.
The primary target of this economic strangulation is Iran's oil exports. Interestingly, Iran had managed to maintain its own oil shipments through the strait even while exerting control and threatening other vessels. This allowed Tehran to continue funding its war efforts, particularly the Islamic Revolutionary Guard Corps (IRGC), which is almost entirely dependent on oil revenues. The US blockade, therefore, aims to cut off this crucial lifeline.
The ultimate goals of this audacious move are two-fold: to reclaim control of the strait from Iranian influence and, critically, to force Iran back to the negotiating table. President Trump aims to leverage control over the strait to compel Iran to concede on key issues, such as its nuclear stockpiles and uranium enrichment program.
The Looming Dangers: Blowback and Boomerangs
Such a high-stakes gamble comes with equally high risks. The panel of experts identified several potential "blowbacks" that could escalate the conflict dramatically.
- Iranian Retaliation: The most immediate concern is that Iran's IRGC could lash out, attacking US Navy ships and triggering a major escalation of hostilities.
- China's Wrath: A significant portion—an estimated 90%—of Iran's oil exports are destined for China, often carried on Chinese-flagged vessels. White House correspondent David Sanger highlights the danger: "A simple, efficient way to really piss off your superpower adversary is to start to systematically deprive them of oil." This could severely complicate an upcoming meeting between President Trump and Beijing, intended to focus on trade and security, and further strain relations already tense with reports of China considering arming Iran.
- Regional Energy Infrastructure Attacks: Energy reporter Rebecca Elliot warns of a "third big category of risk": Iran restarting attacks on energy infrastructure across the Persian Gulf. Such actions carry "long-term risks for the global energy system and the global economy." The International Energy Agency has already estimated that over 80 energy sites in the region have been damaged, and restoring pre-war production levels could take up to two years.
Early Days: A "Quarantine" in Action
In its initial 48 hours, the US blockade has presented a complex picture. Eric Schmidt clarifies that what President Trump terms a "blockade" is "really more of a quarantine." The US Navy is strategically positioned outside the strait, using drones and open-source intelligence to monitor ships bound for Iranian ports. Suspect vessels are contacted via radio, with the threat of boarding parties—Marines or Navy Seals—to inspect cargo if they refuse to comply.
The deterrent effect appears to be working, at least partially. Central Command reported no "interdictions" or "violations" in the first 24 hours. Crucially, six vessels departing Iranian ports reportedly turned around after being contacted by US Navy ships. "It looks like for now he has stopped the Iranians from shipping oil out and thus fueling their economy and their government," Sanger notes, calling it a "halfway" success.
However, a major question mark remains: Can commerce from other Gulf states—like the UAE—now pass safely through the strait? The situation is, as Sanger puts it, "a blockade by the US of a blockade by Iran," making it incredibly messy. The global economy hinges on whether shippers will gain the confidence to navigate this precarious route.
Global Ripples and a Strait Transformed
The trepidation among shipping companies and captains is "very high," according to Rebecca Elliot. Insurance costs for vessels attempting to transit the strait have soared. Many in the industry are now asking not if the strait will reopen, but if it will ever return to its pre-war state of free passage and minimal risk. The consensus among experts is grim: "Probably not."
David Sanger agrees, noting that Iran has "discovered this kind of superpower they have" in their ability to control the strait with minimal military effort. The genie, it seems, cannot easily be put back in the bottle.
Proposals for the strait's future include President Trump's brief, intriguing suggestion of a joint business venture with Iran's Ayatollah Khamenei, or more seriously, an international consortium. Such a body, involving Iran, Oman, the US (providing security), China, India, and European nations, would monitor traffic, potentially charge tolls, and divide revenues. However, this would demand a level of international cooperation and power-sharing that the current US administration has not typically embraced.
The long-term implications for global energy infrastructure are profound. If the Strait of Hormuz no longer guarantees free, unimpeded commerce, the world will need alternatives. These include:
- Alternative Routes: Building more pipelines from Gulf countries to world markets, bypassing the strait. While some options exist in Saudi Arabia and the UAE, these are complicated by the need to traverse other nations' territories.
- Diversified Oil Sources: Increased demand for oil from other regions that don't require transit through the strait.
- Accelerated Energy Transition: Higher oil prices inherently make alternative energy sources—nuclear, solar, batteries—more economically attractive, potentially accelerating the global shift away from fossil fuels.
"The entire world of energy and energy infrastructure could end up changing because of this war and to a degree because of this blockade," says Elliot, "and that could end up long outlasting the conflict itself."
A Contest of Endurance
Ultimately, this blockade is a test of wills. "We are now entering a contest over which country, the United States or Iran, can endure the pain of this blockade over the next few months," Sanger posits.
Iran is betting that President Trump will be forced to back down. Rising gas prices, particularly as midterm elections approach, pose a significant political problem for the Republican party and Trump's domestic agenda. Iranian leadership has openly taunted Trump about gas prices, with one negotiator reportedly stating, "You could be nostalgic for five or six dollar a gallon gas."
The US, conversely, believes economic pressure will succeed where military action failed. By cutting off Iran's main revenue source and specifically targeting the IRGC, Washington hopes to compel Tehran to "cry uncle" and accept American terms.
The outcome remains highly uncertain. "If anybody tells you they know which side is not going to be able to withstand the pain first, I'm not sure I'd believe them yet," Sanger concludes.
The Costs of Sustained Pressure
Even if successful, the blockade comes at a steep price. The Pentagon asserts its ability to sustain the operation indefinitely, but this drains resources from other critical areas. Ships and munitions urgently needed in the Indo-Pacific to counter China or address North Korea are being diverted. Interceptors and other armaments are being pulled from European Command, impacting support for Ukraine.
The operation alone requires some 10,000 US Navy, Marine Corps, and other service personnel, a number that could swell to over 50,000 if the region remains a primary focus and hostilities resume.
While the current full-on war has largely ceased, the underlying tensions and strategic gaps between the US and Iran—over nuclear programs, regional conflicts, and more—remain enormous. Both sides face immense pressure to avoid a return to widespread military conflict, as Trump's base fragmented during the earlier fighting, Congress was frustrated, and allies were conspicuously absent. Iran, too, is desperate for the war to end due to its devastating economic impact.
The question now is not whether the war ends, but "on whose terms it ends." As France and Britain begin to develop their own plan to reopen the Strait of Hormuz—a plan that may exclude the United States and only commence after the war's conclusion—the enduring legacy of this high-stakes blockade may be a permanently altered geopolitical and energy landscape.