Nvidia’s 2026 strategy, as described by Jensen Huang in this Stratechery interview after GTC, is to expand the meaning of accelerated computing from chips and training clusters into full-stack AI factories. Huang’s core thesis is that agents will not live solely inside chat interfaces or model APIs; they will increasingly operate the software humans already use, from SQL databases to desktop applications and EDA tools. That pushes Nvidia beyond selling GPUs into accelerating legacy software, designing CPUs, integrating networking and storage, and helping customers construct entire data-center-scale systems. He quantified the scale of the new buildout in unusually concrete terms: a 1 GW AI factory could cost $50-60 billion, of which roughly $15-17 billion is just land, power, and shell. His point is that customers will not risk tens of billions of dollars unless Nvidia can de-risk throughput, utilization, and system integration across cooling, power, networking, and compute. That framing is particularly relevant to the data-center and infrastructure buildout around AI, because Huang treats power as the fundamental economic constraint and argues system-level co-design is the only way to maximize intelligence produced per watt and per dollar.
On the technical side, Huang drew a line between the first wave of generative AI and the current wave of useful AI. He said reasoning, reflection, retrieval, and search improved enough over the last year to ground models, reduce hallucinations, and enable paid applications—especially coding agents. His description of coding as a separate modality was notable: code must be generated in coherent chunks and validated by execution, not merely by token probability. He also argued that transformers alone are not enough for future workloads, citing quadratic attention costs, KV-cache bloat, and the need for new architectures for continuous motion, geometric symmetry, and physical constraints. He named Nemotron 3’s transformer-plus-SSM hybrid and cuEquivariance as examples. At the infrastructure layer, he defended Nvidia’s CPU effort as complementary to GPUs rather than a reversal: the goal is to prevent expensive accelerators from sitting idle. Vera, he said, emphasizes very high single-thread performance and 3x higher bandwidth-per-CPU than prior designs to support agent tool use over NVLink. The Groq acquisition fits the same logic: Nvidia wants finer-grained heterogeneous inference, including disaggregating pieces of decode attention, to push latency-sensitive coding and enterprise agent workloads beyond what a general GPU-only setup can economically deliver.
The interview also highlighted constraints and geopolitics. Huang said the bottleneck in 2026 is not one thing but everything at once—power, fabs, supply chain, and site readiness—though he sounded confident in Nvidia’s planning across hundreds of suppliers. More striking was his China argument: keeping an American AI stack present in China is, in his view, strategically essential because Chinese labs and open-source communities are too important to ignore. He explicitly praised DeepSeek, Kimi, and Qwen as technically meaningful contributors and warned that exclusion could let rival ecosystems harden across chips, platforms, models, and applications. That makes the conversation highly relevant not just as an Nvidia profile, but as a window into how the dominant AI infrastructure vendor thinks about power markets, data-center economics, hardware/software co-design, and U.S.-China competition.