The Gist Post logo

Friday, October 9, 2026

AboutContact
The Gist Post logoThe Gist Post logo

The Gist Post publishes clear guides, practical explainers, and honest reviews across technology, programming, business, finance, investing, and everyday life.

Categories

  • Technology
  • Business & Finance
  • Gaming & Entertainment
  • Health & Fitness
  • Travel & Hospitality
  • Education & Learning
  • Lifestyle
  • Marketing & SEO
  • Productivity & Work
  • Programming & Software
All categories →

Company

  • About
  • Contact
  • Privacy policy
  • Affiliate disclosure
  • DMCA policy

© 2026 The Gist Post. All rights reserved.

Some links on this site are affiliate links. See our disclosure.

Home/Technology

OpenAI's Jalapeño Chip: What the Hot Chips Reveal Actually Told Us

TechnologyTech News & Trends
By The Gist Post·August 2, 2026·8 min read

OpenAI unveiled its first custom AI chip, Jalapeño, at Hot Chips 2026: a 700W inference ASIC built with Broadcom that OpenAI says beats NVIDIA's best on efficiency. The benchmarks are OpenAI's own, and the chip is for internal use only. Here's the full picture.

Engineer examining a custom AI processor representing OpenAI's Jalapeño chip
Engineer examining a custom AI processor representing OpenAI's Jalapeño chip

On this page

  • Key takeaways
  • What Jalapeño is (and isn't)
  • The specs: a reticle-sized inference engine
  • The design bet: one balanced chip, not a fleet of specialists
  • The benchmarks: OpenAI's numbers, labeled as such
  • The snub: AMD hosts, not NVIDIA Vera
  • What it means beyond OpenAI
  • Practical next steps
  • The bottom line
  • Sources

The most consequential chip announcement of 2026 didn't come from NVIDIA, AMD, or Intel. It came from a software company. At Hot Chips 2026 in August, OpenAI presented the full architecture and first published benchmarks for Jalapeño, its inaugural custom AI accelerator, built with Broadcom and designed to run ChatGPT and Codex inference on OpenAI's own silicon.

This matters for reasons that go beyond one chip. When the world's largest buyer of AI compute starts designing its own processors, the economics of the entire AI infrastructure market shift. Here's what OpenAI actually disclosed, what it claimed, and what to make of both.

Key takeaways

  • Jalapeño is OpenAI's first custom AI chip: an inference-only ASIC built with Broadcom on TSMC's 3nm process, taped out in November 2025, with limited deployment planned for late 2026.
  • The chip delivers 13.4 petaflops of MXFP4 compute with 216 GB of HBM4 at a 700W TDP, scaling to 128 chips per rack and 2,048 chips across 16 racks.
  • OpenAI's own benchmarks claim 1.5–1.9x more AI operations per watt and 1.7–3.6x lower latency than NVIDIA's GB200/GB300. These are OpenAI's numbers, not independently verified.
  • It is for internal use only. You cannot buy it. Its purpose is cutting OpenAI's inference costs and NVIDIA dependence.
  • OpenAI pairs Jalapeño with AMD EPYC "Turin" CPUs, explicitly passing over NVIDIA's Vera on maturity grounds.

What Jalapeño is (and isn't)

Jalapeño is an inference accelerator, full stop. It is designed to run already-trained models, serving ChatGPT conversations, Codex coding sessions, and agentic workloads, as efficiently as possible. OpenAI continues to rely on NVIDIA and other merchant GPU suppliers for training new models, which is a different and far more demanding job. This split is deliberate: inference is where OpenAI burns the most compute every day, so it's where custom silicon pays back fastest.

The partnership structure is notable. Broadcom handled silicon implementation; Celestica handled system design. This sits inside a broader strategic collaboration OpenAI and Broadcom announced in 2025 to deploy up to 10 gigawatts of OpenAI-designed accelerators, with deployment planned from the second half of 2026 through 2029. The Wall Street Journal reported in October 2026 that Broadcom was arranging more than $50 billion in financing for the OpenAI chip program, with Apollo and Blackstone among the lenders approached. That number tells you the scale of OpenAI's ambition better than any spec sheet.

The chip's internal program reportedly goes by "Nexus," with pepper-themed generation names: Jalapeño is generation one, Serrano is generation two (already approaching tapeout, according to OpenAI), and a third generation is in planning.

The specs: a reticle-sized inference engine

What Hot Chips added beyond the original name reveal was the full technical picture. Jalapeño is built on TSMC's N3 process (N3P/N3E variants were discussed), making it a cutting-edge 3nm part. Each chip delivers 13.4 petaflops of MXFP4 compute, packs 216 GB of HBM4 memory with up to 15.4 TB/s of memory bandwidth, and operates at a 700W thermal design power, with OpenAI reporting actual consumption dipping below 550W during sustained workloads.

For context on that power figure: NVIDIA's GB200 runs above 1,200W. If OpenAI's efficiency claims hold, Jalapeño does competitive inference work at roughly half the power draw, which is the entire economic argument in one number. Data centers are increasingly power-constrained, so watts per token is the metric that decides purchasing.

At the system level, the design scales aggressively. Server designs integrate 128 accelerators per rack (roughly 1.7 exaflops of 4-bit compute and 27.5 TB of HBM4 per rack), and up to sixteen racks can connect into a single domain of 2,048 chips. Networking comes from Broadcom's Tomahawk 6 switches at 600 Gb/s per chip for the scale-up domain. The rack topology uses evocative internal names: "Katsu" CPU trays, "Vindaloo" accelerator trays, and "Chana" switch trays, with a two-rack layout drawing roughly 160 kW split between host and accelerator trays.

The design bet: one balanced chip, not a fleet of specialists

The most interesting disclosure wasn't a number but a philosophy. The industry has been drifting toward disaggregated inference: separate hardware for the "prefill" phase (processing your prompt) and the "decode" phase (generating tokens one by one), since the two phases stress hardware differently. OpenAI is betting against that trend.

Keep reading

  • The Coolest AI Gadgets of 2026: The Wearables Actually Worth Your Attention
  • Deepfake Scams in 2026
  • Canada's New National AI Council: What It Means for Jobs and Business

Jalapeño is a single "balanced" chip designed to handle prompt processing, token generation, and the low-latency draft models used in speculative decoding, all on one die, with power-gated regions that shut off unused sections. The OpenAI engineers' summary, as relayed by analysts covering the talk: "dark silicon is cheaper than idle accelerators." In other words, they'd rather waste a little chip area than waste a lot of expensive accelerator time sitting idle waiting for the other phase to finish.

There's a second philosophical point worth noting. OpenAI says it optimizes for end-user experience ("time to last token") and energy per request ("tokens per joule") rather than total cost of ownership, because unlike a merchant chip vendor, its customer is the end user chatting with ChatGPT. That's a genuinely different objective function from NVIDIA's, and it explains why a software company's chip looks different from a chip company's chip.

One claim to hold loosely: the "nine months from RTL to tapeout" timeline, which has been widely repeated as evidence of AI-accelerated chip design. Engineers discussing the disclosure noted that nine months from RTL freeze to tapeout is typical-to-unimpressive for a large 3nm chip; the impressive version would be nine months from concept to tapeout, and OpenAI's hardware chief has taped out with Broadcom before, so this wasn't a cold start. The timeline is real but less remarkable than the headlines suggest.

OpenAI's Jalapeño Chip: What the Hot Chips Reveal Actually Told Us: The design bet: one balanced chip, not a fleet of specialists

The benchmarks: OpenAI's numbers, labeled as such

Now the part that needs the clearest labeling. Every performance figure below comes from OpenAI's Hot Chips presentation. No independent verification has been published.

Against NVIDIA's GB200 and GB300 on inference workloads, OpenAI claims Jalapeño delivers 1.5 to 1.9 times higher throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency. The test workloads were GPT-OSS 120B and DeepSeek R1 670B. The sharpest comparison OpenAI showed was a 1-trillion-parameter Kimi K2.5 run: Jalapeño at 700W versus the GB300 at 1,400W, with 1.5x higher peak mixed tokens per second per kilowatt and 3.4x lower end-to-end latency.

These are strong numbers if true, and they're plausible: an ASIC designed for exactly one company's inference stack should beat a general-purpose GPU on that stack's metrics. But "should" isn't "did," and OpenAI chose the workloads, the baselines, and the metrics. Treat the figures as OpenAI's opening bid in a negotiation with reality, not as settled fact. Independent testing of Jalapeño silicon, if it ever happens publicly given the chip isn't for sale, is the only thing that settles it.

OpenAI's Jalapeño Chip: What the Hot Chips Reveal Actually Told Us: The benchmarks: OpenAI's numbers, labeled as such

The snub: AMD hosts, not NVIDIA Vera

One of the most revealing details had nothing to do with Jalapeño's own specs. Each host tray in the Jalapeño rack carries two AMD EPYC "Turin" CPUs and 1.5 TB of DRAM. NVIDIA's Vera CPU, the shiny new agent-optimized processor, was evaluated and passed over.

OpenAI's VP and Head of Hardware, Richard Ho, was candid about why: Vera standalone is "a little bit behind on that maturity level" for OpenAI's current scale, making Turin the pragmatic, de-risked choice. That's a meaningful signal about Vera's readiness, and a meaningful endorsement of AMD's data center CPU position: when the most demanding buyer in the industry needs host silicon today, it buys AMD.

It also underscores that the chip war isn't winner-take-all. OpenAI will happily buy NVIDIA GPUs for training, AMD CPUs for hosting, Broadcom networking for scale-out, and its own ASICs for inference, all in the same rack. The future is heterogeneous, and every vendor's "complete platform" pitch runs into customers who mix and match.

What it means beyond OpenAI

Jalapeño confirms Broadcom's position as the leading merchant partner for custom AI silicon, a role it already holds with Google and Meta. It puts direct pressure on inference margins that NVIDIA's forecasts may not have fully absorbed. And it accelerates a trend worth watching from Canada: the buyers becoming the builders. Canada's National AI Council will be navigating a compute landscape where the biggest labs increasingly run on silicon you can't buy, which has implications for sovereign AI capacity. For developers, the nearer-term story is what all this inference capacity enables, including the AI coding agents reshaping software work and the skills the Canadian AI job market rewards.

Practical next steps

  • If you buy AI inference at scale: Jalapeño itself isn't for sale, but its existence is leverage. Ask vendors how their inference pricing responds to custom-silicon competition.
  • If you build on OpenAI's APIs: expect inference costs to keep falling as Jalapeño deploys through 2027; price that into long-term product planning.
  • If you evaluate AI hardware claims: apply the same standard everywhere. OpenAI's 1.5–1.9x is OpenAI's number, exactly as NVIDIA's 1.8x is NVIDIA's number.
  • If you follow the money: watch the Broadcom financing talks and the Gen 2 (Serrano) tapeout as the two signals of whether this program scales or stalls.

The bottom line

OpenAI's Jalapeño is the clearest proof yet that the AI chip war has a new kind of combatant: the customer. A 3nm, 700W inference ASIC with 216 GB of HBM4, built with Broadcom, benchmarked (by its maker) at up to 1.9x the efficiency of NVIDIA's best, and deployed only inside OpenAI's own data centers. The benchmarks are OpenAI's own, the chip is internal-use only, and the nine-month timeline is less miraculous than advertised. But the strategic message needs no verification: at hyperscale, renting someone else's silicon forever is a tax, and OpenAI just started building its own mint.

Sources

  • https://nand-research.com/openais-jalapeno-inference-accelerator/
  • https://aiweekly.co/alerts/openai-pairs-jalapeo-asic-with-amd-turin-skips-nvidia-vera
  • https://github.com/hczhu/stock-research/blob/HEAD/memos/2026-08-28-semi-doped-openai-jalapeno-hot-chips-teardown.md
  • https://autonainews.com/openais-jalapeno-chip-claims-inference-efficiency-over-nvidia/
  • https://world-today-journal.com/openai-unveils-jalapeno-a-new-ai-chip-to-rival-nvidias-dominance/
  • https://investinglive.com/stocks/wsj-broadcom-seeks-over-50-billion-for-openai-chips-as-oracle-spacex-chase-ai-debt-deals/

About the author

TG

The Gist Post

Clear guides, practical explainers, and honest reviews across technology, programming, business, finance, investing, and everyday life.

Published August 2, 2026

On this page

  • Key takeaways
  • What Jalapeño is (and isn't)
  • The specs: a reticle-sized inference engine
  • The design bet: one balanced chip, not a fleet of specialists
  • The benchmarks: OpenAI's numbers, labeled as such
  • The snub: AMD hosts, not NVIDIA Vera
  • What it means beyond OpenAI
  • Practical next steps
  • The bottom line
  • Sources

Related

Server racks and a circuit board close-up representing the 2026 AI chip war

Technology

The AI Chip War in 2026: NVIDIA, AMD, and Intel Battle for the Data Center

Close-up of a modern processor representing NVIDIA's Vera CPU

Technology

NVIDIA Vera CPU Explained: The Chip Built for the Age of AI Agents

Quick answers

Frequently asked questions

01

What is OpenAI's Jalapeño chip?

Jalapeño is OpenAI's first custom AI accelerator, revealed at Hot Chips 2026 in August. It's an inference-only ASIC built with Broadcom on TSMC's 3nm process, designed to run models like ChatGPT and Codex inside OpenAI's own data centers. It taped out in November 2025 and limited production deployment is planned for late 2026.

02

Is the Jalapeño chip available to buy?

No. Jalapeño is for OpenAI's internal use only. It is not a merchant product and won't appear on any price list. Its significance is strategic: it shows OpenAI building supply independence from NVIDIA for inference compute.

03

How fast is the Jalapeño chip?

On OpenAI's own benchmarks presented at Hot Chips 2026, Jalapeño delivered 1.5 to 1.9 times more AI operations per watt and 1.7 to 3.6 times lower end-to-end latency than NVIDIA's GB200 and GB300 on inference workloads. These are OpenAI's figures and have not been independently verified.

04

Who built the Jalapeño chip with OpenAI?

Broadcom handled silicon implementation and Celestica handled system design, under a strategic collaboration announced in 2025 to deploy up to 10 gigawatts of OpenAI-designed accelerators. The chip is fabricated by TSMC on its N3 process.

05

Why did OpenAI pair Jalapeño with AMD CPUs instead of NVIDIA Vera?

Each Jalapeño host tray carries two AMD EPYC "Turin" CPUs with 1.5 TB of DRAM. OpenAI's hardware chief Richard Ho said NVIDIA's Vera was "a little bit behind on that maturity level" for OpenAI's scale, calling the Turin choice a pragmatic de-risking decision.

06

What's next after Jalapeño?

OpenAI says a second-generation chip, reportedly codenamed Serrano, is already approaching tapeout, with a third generation in planning. The internal program reportedly uses pepper names, with Jalapeño as generation one.

Newsletter

Get the week's gist.

One short email every Sunday: the most useful guides we published that week, plus one thing worth knowing. Free forever, no spam, unsubscribe anytime.

Subscribe

Launching soon. Check back after our first issues ship.

Keep exploring

Related posts

Server racks and a circuit board close-up representing the 2026 AI chip war

Technology

The AI Chip War in 2026: NVIDIA, AMD, and Intel Battle for the Data Center

Close-up of a modern processor representing NVIDIA's Vera CPU

Technology

NVIDIA Vera CPU Explained: The Chip Built for the Age of AI Agents

Modern desktop computer setup representing Apple's M6 Mac mini and M5 Ultra Mac Studio

Technology

Apple M6 and M5 Ultra Explained: 2nm, Quad-Die, and a Big Bet on Local AI

Close-up of a microchip on a circuit board, representing China's domestic AI chips

Technology

China's Domestic AI Chips Just Served 62 Trillion Tokens

From across the spot

People also read

  • Starlink in Canada in 2026: What It Costs, Why Ontario Dumped It, and What's Next
  • On-Device AI in 2026: Your Phone Is the New Data Centre
  • Every Streaming Service That Raised Prices in Canada in 2026, and What It Costs Now
  • How to Spot AI-Powered Phishing in 2026
  • The Best VPNs for Canada in 2026, Compared in Canadian Dollars
  • 1Password vs Bitwarden in 2026: Canada's Own Password Manager Just Got Pricier, Should You Switch?