OpenAI publishes first benchmarks of Jalapeño, its first custom inference chip
On Aug 25, 2026 OpenAI published the first measured results for Jalapeño, its first in-house AI inference chip. On SemiAnalysis's public InferenceX benchmark, serving GPT-OSS 120B, DeepSeek R1 and Kimi K2.5, OpenAI says Jalapeño did 1.5-1.9x more work per watt at peak throughput and had 1.7-3.6x lower end-to-end latency than the Nvidia GB200/GB300 systems it was compared with. OpenAI says AI helped take the chip from design to tapeout in nine months, and it plans to start deploying Jalapeño in its own data centers by the end of 2026.
Key facts
- OpenAI's first custom inference chip; first measured results published Aug 25, 2026 (Engineering blog), with a companion essay 'The full stack behind abundant intelligence'
- InferenceX (SemiAnalysis) results, normalized by rated chip power: 1.5-1.9x more peak throughput per watt, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher performance on highly interactive workloads vs. the comparison systems (OpenAI)
- GPT-OSS 120B vs GB200: ~1.9x peak mixed tokens/s per kW (85,448 vs 44,960), ~1.7x lower end-to-end latency (1.03 s vs 1.80 s), 1,459 vs 535 tokens/s per user at minimum TBT
- DeepSeek R1 670B (MXFP4) vs GB300: ~1.7x per kW, ~3.6x lower latency; Kimi K2.5 1T (MXFP4) vs GB300: ~1.5x per kW, ~3.4x lower latency
- Rated at 700 W package TDP (GB200 1,200 W, GB300 1,400 W); measured sustained power stayed at or below 550 W on the tested workloads
- AI-assisted design: design to tapeout in nine months; AI also optimized arithmetic circuits. Codex with GPT-Astra brought three open-weight models not in the original plan to high performance within two months; AI-written kernels for selected GPT-OSS attention/MoE blocks ran 1.5-1.8x faster than human-expert kernels
- Deployment in OpenAI's compute infrastructure planned 'by the end of the year'; Gen 2 'deep in development', Gen 3 'taking shape'; OpenAI will keep deploying Nvidia and other accelerators
- Inference-only design: OpenAI describes it purely as an inference chip; press (e.g. Quasa) stresses that it cannot be used to train models
What happened
OpenAI published the first measured performance numbers for Jalapeño, its first custom inference chip. It tested the chip on SemiAnalysis's public InferenceX benchmark, which measures the whole process of serving a request, using three public models: GPT-OSS 120B (against an Nvidia GB200 system), DeepSeek R1 670B and Kimi K2.5 1T (against GB300). OpenAI normalized results by each accelerator's rated chip power. Jalapeño is rated at 700 W, but OpenAI says it stayed at or below 550 W in practice. On that basis Jalapeño sat on the Pareto frontier of throughput per watt against latency for all three models. OpenAI says the lead grew further on its own frontier models in internal tests. Those internal results were not published.
OpenAI says the design keeps model state such as the KV cache local and uses a large network domain, so a whole request stays in one system. It also says AI shortened the design-to-tapeout cycle to nine months. Codex with GPT-Astra ported three extra open-weight models in two months, and for some GPT-OSS blocks AI-generated kernels beat human-written ones by 1.5-1.8x. OpenAI says the chip will enable "ultra-fast-mode inference at efficiencies previously available only in fast mode".
Why it matters
This is the first published evidence that OpenAI's own silicon works, and it puts OpenAI next to Google (TPU), Amazon (Trainium/Inferentia) and Meta (MTIA) as a lab with first-party accelerators. The comparison is OpenAI's own and uses rated rather than measured power for the Nvidia systems. Critics (e.g. MLQ) note the limits of that comparison and that Jalapeño is an inference-only part. The results support OpenAI's push for faster serving tiers (Ultrafast, launched at DevDay on Sept 29) and lower cost to serve.
Changelog
- 2026-09-30: created (official-blog audit)
Related posts (1)
- tae kim original ↗ tae kim @firstadopter · x · 2026-08-25
Cited as a source by: 2026-08-25-openai-jalapeno-first-results
Related events
- Altman says OpenAI will "definitely" build its own humanoid robots ★★★
- OpenAI adds a $500/month Pro 500 plan and the Ultrafast speed tier, and halves the $200 Pro allowance ★★★
Sources (7)
- officialOpenAI: Jalapeño's first results show industry-leading speed and efficiency in AI inference
- officialOpenAI: The full stack behind abundant intelligence
- officialOpenAI (Sarah Friar): The Work Now Within Reach (Sept 8, cites Jalapeño results)
- pressTechCrunch: OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show
- pressMLQ: Jalapeño shows strong inference gains, but its Nvidia comparison has limits
- pressQuasa: OpenAI Jalapeño leads InferenceX, but cannot train models
- discussionTae Kim on X: quoting OpenAI's Jalapeño results
id: 2026-08-25-openai-jalapeno-first-results · updated 2026-09-30 · open in the interactive timeline