On Aug. 24, Nvidia published a blog post about its new Vera Rubin NVL72 system that led not with a speed record but with a ratio: up to 30 times more throughput per megawatt than the generation it replaces. The next day, OpenAI and Broadcom published the first real benchmark numbers for Jalapeno, the inference chip they have been building together since October 2025, and led with the same kind of number -- 1.5 to 1.9 times more AI work per watt than Nvidia's own current chips. In between, Nvidia's Groq 3 LPX -- the chip built from the $20 billion licensing deal that absorbed Groq's engineering team and LPU architecture last December -- entered full production promising 4x faster responsiveness than "the nearest alternative." None of the three companies used to talk this way. A year ago, the AI chip industry's headline metric was tokens per second, or FLOPS, or how many GPUs fit in a rack. This week, three separate hardware announcements from two of the industry's biggest players all reached for a different unit: the watt.
The shift is not just a marketing choice. It tracks a real, well-documented constraint: the AI industry now has more chips than it has electricity to run them. Microsoft CEO Satya Nadella said as much months before any of this week's announcements, in a November 2025 interview on the Bg2 Pod: "The biggest issue we are now having is not a compute glut, but it's power -- it's sort of the ability to get the builds done fast enough close to power. So, if you can't do that, you may actually have a bunch of chips sitting in inventory that I can't plug in. In fact, that is my problem today." If GPUs are stacking up in warehouses waiting for substations, the fastest chip in the world is worth nothing until it has somewhere to plug in -- which is exactly the fact this week's hardware roadmap visibly reoriented around.
The pivot, in short
- What changed
- 3 major hardware announcements, all led with performance-per-watt
- Who
- Nvidia (x2), OpenAI/Broadcom
- Biggest claimed multiple
- 30x throughput/MW
- Independently verified?
- Not yet
- Real bottleneck, per Microsoft's CEO
- Grid power, not chip supply
Tokens per second was never a bad metric on its own -- it measures the thing a user actually experiences, how fast an answer arrives. It became a bad primary metric once the industry started measuring success in gigawatts rather than GPUs. A cluster's maximum tokens-per-second figure assumes unlimited power to run it at full tilt; a real data center's actual output is capped by the power it can draw, not by how fast any individual chip can go. Once that cap became the binding one -- once, as Nadella put it, chips started sitting in warehouses instead of racks -- the question that mattered stopped being how fast one chip is and became how much useful work the whole fleet does inside the megawatts a company actually has. Tokens per watt is that second question, restated as a single number.
Nvidia's own numbers, published on its corporate blog: the Vera Rubin NVL72 system delivers up to 30 times higher throughput per megawatt and 35 times lower token costs than the GB300 NVL72 rack it replaces, measured on agentic coding sessions across five open models -- Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro. That is a multiple built on top of another multiple: Nvidia says the GB300 generation being used as the baseline already delivered up to 15x better throughput per megawatt than the older Hopper architecture, measured on a different model (DeepSeek V4 Pro alone) and a different workload (SemiAnalysis's AgentX benchmark, an earlier version of the one used for Vera Rubin). The company's own post flags the new figure as unfinished business: the results are "currently pending SemiAnalysis review" -- not yet confirmed by the outside benchmarking group whose workload produced them.
The same day, Nvidia announced that Groq 3 LPX -- its dedicated inference chip, built from the Groq engineering team and LPU architecture it absorbed in a reported $20 billion deal last December -- has entered full production. (Groq the company is now, in effect, an early customer of a chip built from technology it sold to Nvidia less than a year ago.) In one demonstration, paired with Vera Rubin NVL72, it produced a record 3,400 output tokens per second running the open Gemma 4 31B model with a 100,000-token context window, which Nvidia says is 4x faster response than "the nearest alternative" it tested against. Nebius will be the first cloud to offer the chip, through its Token Factory platform, and Groq itself -- now an independent inference-cloud company after the licensing deal took its architecture and most of its staff -- says it plans to be among the platform's earliest adopters.
The next morning, OpenAI and Broadcom published the first public benchmark numbers for Jalapeno, the inference-only chip the two companies have been co-designing since OpenAI first disclosed the partnership in October 2025. Tested on InferenceX -- a public, third-party benchmark suite from the semiconductor analysis firm SemiAnalysis -- across three open models (GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T), Jalapeno delivered 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end-to-end latency than the Nvidia Blackwell systems (GB200 and GB300) it was benchmarked against; on the most latency-sensitive, interactive workloads, OpenAI reported a gap as wide as 2.1x to 4.1x. "Jalapeno can serve more AI work per unit of power, while also returning responses more quickly," said Richard Ho, OpenAI's head of hardware. The chip is not shipping yet -- OpenAI plans to deploy it inside its own infrastructure in "very small volumes" by the end of 2026, with any broader rollout pushed to 2027 -- but unlike Nvidia's Vera Rubin figures, Jalapeno's numbers were measured against a benchmark a third party can actually rerun, rather than a workload run and reported by the vendor itself pending outside review.
Three efficiency claims, three different baselines
| Nvidia Vera Rubin NVL72 vs. GB300 NVL72 | OpenAI/Broadcom Jalapeno vs. Nvidia GB200/GB300 | Positron Atlas vs. Nvidia H100 | |
|---|---|---|---|
| Claimed efficiency gain | Up to 30x throughput/MW | 1.5x-1.9x AI work per watt | ~3x compute per watt |
| Baseline chip's age | Prior generation (2025) | Current generation (2025) | Three-plus generations back |
| Measured by | Nvidia's own AgentX run, pending outside review | Public InferenceX benchmark (SemiAnalysis) | Positron's own reported figures |
| Available today? | Yes, Aug. 2026 | No -- small volumes end of 2026 | Yes, shipping to customers |
Line the three claims up and the multiples stop looking comparable. A bigger number is not automatically a bigger advance -- it can just mean an older baseline. Positron's roughly 3x-per-watt claim against Nvidia's three-plus-year-old H100 is a real, shipping, independently ordered product, but H100 was never built for the efficiency era; comparing to it is comparing against the chip the rest of the industry has already spent two full generations improving past. Jalapeno's more modest 1.5x-1.9x is measured against Nvidia's current-generation Blackwell -- the machine actually running most of today's production inference -- on a public benchmark a third party can rerun. Nvidia's own 30x is the biggest number of the three and the hardest to check: it is Nvidia's own workload, on Nvidia's own hardware, awaiting a review from the outside firm whose benchmark produced it. None of that makes any of the three claims false. It does mean a reader comparing "30x" to "3x" to "1.9x" at face value is comparing three different questions, not three answers to the same one.
Money followed the same pitch. Three inference-focused chip companies closed funding rounds in 2026 pitching efficiency as their core differentiator, not raw speed: SambaNova closed the first $1 billion of a Series F at an $11 billion valuation on July 8, with JPMorgan Chase signing on to run its SN40 and SN50 systems in-house; Positron AI raised $230 million in February at a valuation just above $1 billion, on the strength of Atlas, a chip fabricated by Intel that it says delivers roughly 3x the compute-per-watt of an H100 in an air-cooled rack; and OLIX Computing, a London startup, raised $312 million in August at a $3.3 billion valuation to build an optical chip -- light, not electricity, carrying the data -- that it says cuts power draw well below a comparable GPU setup while skipping the HBM memory that has been the industry's tightest supply chokepoint all year.
The three approaches are not interchangeable, either. SambaNova's SN40/SN50 systems and Positron's Atlas chip are both conventional silicon, competing with Nvidia on architecture and interconnect design within the same physics as a GPU. OLIX is betting on a different substrate entirely -- Optical Tensor Processing Units that move data as light rather than electrical current, which the company says generates less waste heat and sidesteps the memory-bandwidth ceiling that has made HBM the industry's tightest supply chokepoint. That is a much larger bet: DX-1, OLIX's first commercial chip, does not ship until the second half of 2027, a full generation behind Positron's already-shipping Atlas and roughly contemporaneous with whatever Nvidia calls the chip after Vera Rubin. Photonics has promised to unseat conventional silicon before and has not yet done so at scale; a $3.3 billion valuation on a chip that doesn't exist yet is a bet that the promise holds this time.
Money chasing power-efficient inference chips in 2026
None of this spending or engineering effort touches the actual bottleneck Nadella described. The Energy Department's own Lawrence Berkeley National Laboratory tracks how long it actually takes a proposed power project to move from an interconnection queue request to commercial operation nationally, and its 2026 edition -- covering projects that reached commercial operation through the end of 2025 -- put the median at 5+ yrs (Median US grid interconnection wait, request to power-on (LBL, 2026)) That is the wait before a data center's own power supply is even guaranteed, separate from the time it takes to build the data center itself.
The scale of the backlog is not small. As of the end of 2025, roughly 8,200 projects representing more than 1,300 gigawatts of generation capacity -- plus another 749 gigawatts of storage -- were sitting in US interconnection queues, according to the same Berkeley Lab report. Most of that capacity will never actually get built: historically, the large majority of queued projects are withdrawn rather than completed, which is its own signal that the queue functions less like a line and more like a filter. A gigawatt of capacity that clears the queue is a gigawatt an AI data center can actually draw on. A gigawatt still waiting is exactly the chips-in-a-warehouse problem Nadella described, just measured on the supply side instead of the demand side.
The equipment side is no better. Large power transformers -- the unglamorous boxes that step grid voltage down to something a data center can actually use -- are now running lead times of up to four years in the most constrained parts of the US market, pv magazine USA reported in May, citing severe supply constraints across the sector; switchgear backlogs run into 2028 in many channels. None of Nvidia's, OpenAI's, or any chip startup's tokens-per-watt number changes either of those two clocks. A chip that does 30 times more work per megawatt still needs the megawatt, and the megawatt is still queued behind a multi-year wait that has nothing to do with silicon.
The two clocks a faster chip doesn't reset
The same underlying power constraint was already forcing a fight over who pays for it: a ratepayer dispute that moved through Congress, five states and the White House this summer -- a bipartisan House bill that cleared committee 52-0, a New York executive order that sidestepped the state's own legislature, Ohio electricity rates up 175% since 2005. That fight was about cost allocation: who absorbs the price when a data center strains a local grid. This week's hardware announcements are the other half of the same constraint. Even a buyer willing to pay any price for electricity still waits behind the same multi-year queue for the physical infrastructure -- a transformer, a substation, an interconnection agreement -- and no chip announcement moves that queue.
The inference-chip funding is not isolated to the three companies named above. AI chip and semiconductor startups have closed roughly $4.16 billion in disclosed venture funding across 2026, and inference-specific accelerators account for close to half of that capital and roughly half of the sector's deals -- a reversal from the training-chip gold rush of 2023 and 2024, when investment chased whoever could pack the most FLOPS onto a die, not the most tokens per watt.
There is a wrinkle in calling any of this week's benchmarking "independent." SemiAnalysis, the firm behind both AgentX (the workload Nvidia used for its Vera Rubin numbers) and InferenceX (the benchmark OpenAI used for Jalapeno), is a semiconductor analysis outfit that vendors routinely cite in their own marketing precisely because it is not the vendor itself -- but it is not a neutral academic body either, and it built its reputation partly by becoming the benchmark every major chipmaker wants to be measured on. That doesn't make its numbers wrong. It does mean "pending SemiAnalysis review" is a real, checkable milestone, not a rubber stamp -- but it is one benchmarking firm's methodology standing in for independent verification, not a chorus of separate labs.
- Vera Rubin NVL72 delivers up to 30x more throughput per megawatt than GB300 NVL72
- Jalapeno delivers 1.5x-1.9x more AI work per watt than Nvidia GB200/GB300
- Positron Atlas delivers roughly 3x the compute-per-watt of an Nvidia H100
- OLIX's optical chip cuts power draw well below a comparable GPU setup
- The real bottleneck is grid power, not chip efficiency
None of this is dishonest, exactly -- vendor benchmarks are still benchmarks, run on real hardware. But it isn't settled, either. Here is the strongest case that the efficiency framing matters less than this week's press cycle suggested:
You may actually have a bunch of chips sitting in inventory that I can't plug in. In fact, that is my problem today. -- Satya Nadella, Microsoft CEO, Bg2 Pod interview, November 2025
The industry didn't stop competing on speed this week -- it admitted, all at once, that speed was never the constraint that mattered. Every chip named here, from Nvidia's shipping hardware to OLIX's chip that won't exist until 2027, is racing to do more work inside a power budget that isn't growing as fast as demand for it. That race produces real engineering and, in Jalapeno's case, a number a third party can actually check. It does not produce more electricity, and it does not move a transformer order that was placed four years before delivery. The vendors making the tokens-per-watt pitch this week are not wrong that watts are now the scarce resource. They are the wrong entities to ask how long it takes to get more of them.
- Nvidia, OpenAI/Broadcom, and Groq's own former chip all led this week's hardware news with watts, not speed.
- Vera Rubin NVL72 claims up to 30x more throughput per megawatt; Jalapeno claims 1.5x-1.9x more work per watt.
- Three inference-chip startups -- OLIX, Positron, SambaNova -- raised over $1.5 billion in 2026 on the same pitch.
- These efficiency multiples share no common baseline, and Nvidia's own headline figure is still unreviewed.
- Caveat: Microsoft's CEO says the real bottleneck isn't chip efficiency -- it's grid power, now a 4-to-5-year wait.