Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. OpenAI publishes first benchmarks of Jalapeño, its first…

OpenAI publishes first benchmarks of Jalapeño, its first custom inference chip

★★★★after cutoffhardware-computeOpenAIconfidence: high

On Aug 25, 2026 OpenAI published the first measured results for Jalapeño, its first in-house AI inference chip. On SemiAnalysis's public InferenceX benchmark, serving GPT-OSS 120B, DeepSeek R1 and Kimi K2.5, OpenAI says Jalapeño did 1.5-1.9x more work per watt at peak throughput and had 1.7-3.6x lower end-to-end latency than the Nvidia GB200/GB300 systems it was compared with. OpenAI says AI helped take the chip from design to tapeout in nine months, and it plans to start deploying Jalapeño in its own data centers by the end of 2026.

Key facts

What happened

OpenAI published the first measured performance numbers for Jalapeño, its first custom inference chip. It tested the chip on SemiAnalysis's public InferenceX benchmark, which measures the whole process of serving a request, using three public models: GPT-OSS 120B (against an Nvidia GB200 system), DeepSeek R1 670B and Kimi K2.5 1T (against GB300). OpenAI normalized results by each accelerator's rated chip power. Jalapeño is rated at 700 W, but OpenAI says it stayed at or below 550 W in practice. On that basis Jalapeño sat on the Pareto frontier of throughput per watt against latency for all three models. OpenAI says the lead grew further on its own frontier models in internal tests. Those internal results were not published.

OpenAI says the design keeps model state such as the KV cache local and uses a large network domain, so a whole request stays in one system. It also says AI shortened the design-to-tapeout cycle to nine months. Codex with GPT-Astra ported three extra open-weight models in two months, and for some GPT-OSS blocks AI-generated kernels beat human-written ones by 1.5-1.8x. OpenAI says the chip will enable "ultra-fast-mode inference at efficiencies previously available only in fast mode".

Why it matters

This is the first published evidence that OpenAI's own silicon works, and it puts OpenAI next to Google (TPU), Amazon (Trainium/Inferentia) and Meta (MTIA) as a lab with first-party accelerators. The comparison is OpenAI's own and uses rated rather than measured power for the Nvidia systems. Critics (e.g. MLQ) note the limits of that comparison and that Jalapeño is an inference-only part. The results support OpenAI's push for faster serving tiers (Ultrafast, launched at DevDay on Sept 29) and lower cost to serve.

Changelog

  • 2026-09-30: created (official-blog audit)

Related posts (1)

Related events

  1. Altman says OpenAI will "definitely" build its own humanoid robots ★★★
  2. OpenAI adds a $500/month Pro 500 plan and the Ultrafast speed tier, and halves the $200 Pro allowance ★★★

Sources (7)

id: 2026-08-25-openai-jalapeno-first-results · updated 2026-09-30 · open in the interactive timeline