OpenAI's Jalapeño Chip Beat Nvidia. Now What?

OpenAI's Jalapeno chip beat Nvidia Blackwell on efficiency and latency benchmarks. See what frontier labs owning their own silicon does to your API pricing.

Scott Armbruster
12 min read
OpenAI's Jalapeño Chip Beat Nvidia. Now What?

On August 25, OpenAI published Jalapeño’s first benchmark results, and the headline number is real: its first custom inference chip delivered 1.5x to 1.9x more AI work per watt than the best commercially available systems at peak throughput, with 1.7x to 3.6x lower end-to-end latency. CNBC covered it the next day. The comparison system was Nvidia Blackwell.

Four days earlier, Broadcom was reportedly seeking up to $100 billion in debt to fund custom AI silicon for Anthropic.

Same week. Same chipmaker. Two different frontier labs both walking away from buying Nvidia GPUs off the shelf. That’s the story, and almost every take I’ve read so far has been about Nvidia’s stock price. Wrong altitude. The thing worth your attention is what happens to your renewal leverage when the company selling you tokens also owns the factory floor.

Quick Verdict

QuestionThe Answer
What happened?OpenAI published first benchmark results for Jalapeño, its custom inference chip, on August 25, 2026.
How much faster?1.5x to 1.9x more work per watt at peak throughput, 1.7x to 3.6x lower end-to-end latency. Interactive workloads showed 2.1x to 4.1x higher performance.
Tested on what?GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, on SemiAnalysis’s public InferenceX benchmark.
Who built it?OpenAI designed it. Broadcom did silicon implementation and networking. Celestica did board, rack, and system integration.
Can it train models?No. Inference only.
Can you buy one?No. It runs inside OpenAI’s own infrastructure, initial deployment by end of 2026, ramping through 2027.
Is the comparison fair?Partly. SemiAnalysis calls it “somewhat incomplete and unfair” because Jalapeño uses HBM4 memory and Blackwell doesn’t.
What’s the honest matchup?Nvidia’s Vera Rubin platform, which also uses HBM4. SemiAnalysis already ran it: total cost per token comes out essentially tied.
Does anything change this quarter?No. Nothing in your stack moves in 2026.
What actually changes for you?Your provider’s cost structure, its switching cost, and eventually the shape of your negotiation.

What is OpenAI’s Jalapeño chip?

Jalapeño is OpenAI’s first custom-designed AI accelerator, built for inference only, meaning it runs trained models but cannot train them. OpenAI designed the architecture, Broadcom implemented the silicon and networking, and Celestica handles system integration. It deploys inside OpenAI’s own data centers starting late 2026 and will not be sold to anyone else.

That last clause is the one that matters commercially, and I’ll come back to it.

The Numbers, With the Asterisk Attached

Let’s be precise about what got measured, because the gap between the headline and the fine print is where the useful analysis lives.

OpenAI ran three models through InferenceX, a public benchmark run by SemiAnalysis. Per The Decoder’s reporting, Jalapeño hit roughly 1,400 tokens per second per user on GPT-OSS 120B and over 700 tokens per second on DeepSeek R1 670B at a single concurrent request. SemiAnalysis CEO Dylan Patel’s summary: “Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin.”

Now the asterisk, which SemiAnalysis put in its own writeup rather than burying: comparing Jalapeño to Blackwell is “somewhat incomplete and unfair.” Jalapeño ships with HBM4 memory. Blackwell does not. Nvidia’s Vera Rubin platform also uses HBM4, which makes it the fairer matchup, and SemiAnalysis has already run that comparison against Rubin’s own published July 2026 results. On total cost per token, the two platforms come out essentially tied.

There’s a second wrinkle, and it points in Jalapeño’s favor. Jalapeño’s results were measured using single-token prediction with no speculative decoding, the stricter and less flattering way to benchmark. Rubin’s published numbers use multi-token prediction, a software optimization that inflates throughput. Even compared on those uneven terms, Jalapeño’s throughput surpasses Rubin’s. Both caveats are real, and the net of them is closer to “Jalapeño holds its own” than “this is unproven.”

Here’s my read on where that leaves things:

ClaimHow much I’d trust it
Jalapeño is a competitive first-generation inference chipHigh. Three models, public benchmark, independent analyst confirmation.
Jalapeño beats Blackwell on perf-per-wattHigh, with the memory-generation caveat stated.
Jalapeño beats Nvidia’s next platformContested, not unproven. SemiAnalysis’s own Rubin comparison shows a TCO/token tie, with Jalapeño ahead on the stricter throughput terms.
Nvidia is in troubleNo. One inference-only chip that isn’t for sale doesn’t move that needle.
OpenAI’s inference costs are about to dropDirectionally yes, timing unknown, and the savings are OpenAI’s to keep.

The design timeline is the part that impressed me most. Design work started mid-2024. Fabrication began in November 2025. Roughly sixteen months from design start to fabrication, on a first attempt, at a company that had never taped out a chip. That’s a compressed schedule by any standard, and OpenAI has said its own models helped compress it.

The Broadcom Detail Nobody Circled

Read the partner list on Jalapeño again: Broadcom for silicon implementation and networking.

Now read the $100 billion debt story from August 20: Broadcom arranging financing so investors can buy custom AI chips and lease them to Anthropic.

Broadcom is the common denominator. It is co-designing the silicon that OpenAI will use to serve you tokens, and it is co-designing plus partly financing the silicon that Anthropic will use to serve you tokens. Two labs that compete for your budget on every dimension that shows up in a bake-off are converging on the same supplier one layer down.

I wrote in the Broadcom piece that the vendor decision, the infrastructure decision, and the credit decision were collapsing into a single decision. Jalapeño is the same collapse from the other side. OpenAI’s original October 2025 agreement with Broadcom covered 10 gigawatts of OpenAI-designed accelerators, with the full buildout running through 2029. Jalapeño is the first chip in that program, and OpenAI has been explicit that it’s a multi-generation platform rather than a one-off.

So when you evaluate “multi-vendor AI strategy” in 2027, understand what you’re actually diversifying. Two model providers, two APIs, two contracts, one chip partner. That’s thinner insulation than it looks like on the slide.

Why This Shows Up in Your Bill

Custom silicon changes three things about a model provider’s economics, and each one eventually reaches your invoice.

It lowers marginal cost. If OpenAI serves inference on hardware it designed, at better perf-per-watt, on chips it isn’t paying Nvidia’s margin for, its cost per token falls. That’s the whole point of the exercise.

It raises fixed cost. A 10-gigawatt accelerator program is an enormous capital commitment with a long payback period. Fixed obligations don’t flex when demand wobbles.

It raises switching cost for the provider itself. OpenAI can’t easily walk away from a platform it spent years and billions designing. Neither can Anthropic. That commitment gets defended.

Put those together and you get an outcome that a lot of buyers won’t expect. The last two years trained everyone to assume per-token prices only fall, and mostly they have. But cheaper hardware does not automatically mean cheaper API pricing. It means better margins for whoever owns the hardware, and the split between margin and price cut is a business decision made in a room you’re not in.

We’ve already seen how that decision goes. When OpenAI shipped GPT-5.5 in April, its serving cost dropped roughly 35x while list price doubled. Efficiency gains and customer pricing are separate levers. Custom silicon adds torque to the first one and changes nothing structural about the second.

How should you respond to frontier labs building their own chips?

Five moves. A technical lead and whoever owns the vendor relationship can work through all of them in about a week, and none require a budget request.

  1. Stop treating hardware news as irrelevant to procurement. Add one line to your vendor review: what silicon does this provider run on, and do they own it? It takes five minutes and it’s now a real variable.
  2. Model your 2027 spend at flat pricing, not declining pricing. If your budget assumes token costs keep dropping 30% a year, build the scenario where they hold steady and see what breaks first.
  3. Ask for price-protection language on any contract past 12 months. A cap on increases costs nothing to request. Providers with heavy fixed infrastructure commitments will resist harder next year than they do today, which makes now the better time to ask.
  4. Keep your prompt and eval layer portable. Not your vendor list, your actual plumbing. Model-agnostic workflows are the only version of optionality that survives contact with a renewal conversation.
  5. Route production traffic through a second provider at least once a quarter. A tested alternate path is leverage. An untested one on a slide is a bluff, and vendors can tell the difference.

Moves 1 through 3 are this month. Moves 4 and 5 are habits.

What Would Have to Go Wrong

I’m not forecasting a bad outcome here. Anti-hype cuts both ways, and the doom read on custom silicon is as lazy as the “Nvidia is finished” read.

For this to actually hurt an AI buyer, a specific chain has to connect:

Failure pointWhat it looks likeHow you’d notice
Volume ramp slipsJalapeño stays in low-volume production past 2027Quiet timeline language in OpenAI infrastructure posts
Nvidia answersVera Rubin closes the perf-per-watt gapBenchmark parity reporting, no more efficiency claims
Savings stay internalCosts drop, list prices don’tYour renewal quote, unchanged, with a capability story attached
Supplier concentration bitesOne chip partner has a problem affecting multiple labsSimultaneous capacity issues at providers you thought were independent

Row three is the likeliest one and also the least dramatic. The realistic outcome of this whole shift isn’t a price spike. It’s the disappearance of the automatic annual discount that a lot of AI budgets are quietly built on.

My Read

Three things I think are true.

The engineering result is more impressive than the competitive result. A first-generation chip from a company that had never designed one, competitive with Nvidia’s shipping platform on perf-per-watt, in roughly sixteen months from design start to fabrication. That is a genuinely hard thing to do, and Richard Ho, OpenAI’s hardware lead, wasn’t overselling when he told TechCrunch the results show “a very, very significant performance advance over state of the art.” What it doesn’t do is dent Nvidia’s business, because Jalapeño isn’t for sale to anyone.

“Not for sale” is the most strategically loaded fact in the announcement. Google has run TPUs internally for years and it hasn’t made Gemini cheaper than the competition in any consistent way. It has made Google’s inference margins structurally better. OpenAI is copying that model. When your provider’s cost advantage is locked inside its own data centers, you don’t get to buy the advantage, you get to buy access to it at whatever price they set. That’s a different relationship than the one most buyers think they’re in.

The supplier layer is consolidating while the vendor layer looks like it’s diversifying. This is the part I’d want on a board slide. From where a buyer sits, 2026 looks like more choice than ever: more labs, more models, more gateways, more price competition. One layer down, the same names keep appearing. Broadcom in OpenAI’s chip. Broadcom in Anthropic’s chip and its financing. TSMC fabricating both. If you built a vendor risk model around “we use two providers, we’re covered,” this is the year to redraw it with the infrastructure layer included.

Here’s what I’d tell a business owner who reads this and concludes it’s above their pay grade. It isn’t, and you don’t need to know what HBM4 is. You need one sentence: the companies selling you AI are becoming the companies that make the hardware, which makes them harder to replace and less pressured to discount. Everything else here is detail for somebody else to track.

The Bottom Line

OpenAI’s Jalapeño posted 1.5x to 1.9x better perf-per-watt and up to 3.6x lower latency than Nvidia Blackwell on a public benchmark, days after Broadcom went to credit markets for as much as $100 billion to fund Anthropic’s custom silicon. Both results point at the same shift: frontier labs are moving from renting compute to owning it.

The benchmark caveat is legitimate for the Blackwell comparison. Blackwell lacks the HBM4 memory Jalapeño has. But the fairer fight, Jalapeño against Nvidia’s HBM4-equipped Vera Rubin, has already been run: SemiAnalysis puts the two essentially tied on cost per token, with Jalapeño ahead on the stricter throughput terms. Take the “beats Nvidia” framing with a smaller discount than the Blackwell number alone suggests.

What survives the caveat is the direction. Every serious lab now has a silicon program, those programs carry large fixed commitments, and fixed commitments show up as pricing floors long before they show up as pricing cuts. Nothing in your AI stack changes in 2026 because of this. Your negotiating position in 2028 changes quite a bit.

Your Next Step: This week, open your largest AI vendor contract and check two things: whether there’s a cap on price increases, and how much notice you’re owed before pricing changes. If either is missing, write them down as asks for your next renewal. Then spend twenty minutes rebuilding next year’s AI budget with token prices held flat instead of declining. If that scenario is merely annoying, you’re fine. If it breaks the plan, you just found the work worth doing before anyone’s silicon actually ships.


Related Reading:

TAGS

OpenAI Jalapeno chipAI inference chip 2026Nvidia Blackwell alternativeAI vendor lock-incustom AI silicon enterprise

SHARE THIS ARTICLE

What is this worth in your business?

The free Build Audit is 30 minutes. You leave with a ranked list of the automations worth doing in your business, whether or not we build them.