tae kim
tae kim @firstadopter · x · 2026-08-25 · ★★★ · archived
Cited as a source by: 2026-08-25-openai-jalapeno-first-results
Summary
Archived text
OpenAI: "Today, we shared the first measured performance results from Jalapeño, OpenAI’s first custom inference chip. On InferenceX, a public benchmark using GPT‑OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in the comparison. It also performed strongly on DeepSeek R1 and Kimi K2, showing that its gains extend across model families."
"This is Jevons paradox: greater efficiency makes more uses worthwhile, expanding consumption and creating new economic activity through more work completed, better decisions, more products launched, and more revenue generated."
"AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule."
"We plan to begin deploying Jalapeño within OpenAI’s compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will build on what we learn and further advance both efficiency and speed."
"As we prepare for deployment, we are continuing production qualification, maturing the software, preparing to operate Jalapeño at scale, and validating performance across more models."
Media: https://pbs.twimg.com/media/HQk4lXBa4AA0-bG.jpg?name=orig https://pbs.twimg.com/media/HQk5L4DakAAHWTw.jpg?name=orig https://pbs.twimg.com/media/HQk5L4Ja4AAJt8t.jpg?name=orig
views 298641 · likes 528 · reposts 63 · replies 34 (at fetch time)
Archived 2026-09-30 via fxtwitter (unofficial).
Archived text
OpenAI: "Today, we shared the first measured performance results from Jalapeño, OpenAI’s first custom inference chip. On InferenceX, a public benchmark using GPT‑OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in the comparison. It also performed strongly on DeepSeek R1 and Kimi K2, showing that its gains extend across model families."
"This is Jevons paradox: greater efficiency makes more uses worthwhile, expanding consumption and creating new economic activity through more work completed, better decisions, more products launched, and more revenue generated."
"AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule."
"We plan to begin deploying Jalapeño within OpenAI’s compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will build on what we learn and further advance both efficiency and speed."
"As we prepare for deployment, we are continuing production qualification, maturing the software, preparing to operate Jalapeño at scale, and validating performance across more models."
Media: https://pbs.twimg.com/media/HQk4lXBa4AA0-bG.jpg?name=orig https://pbs.twimg.com/media/HQk5L4DakAAHWTw.jpg?name=orig https://pbs.twimg.com/media/HQk5L4Ja4AAJt8t.jpg?name=orig
views 298641 · likes 528 · reposts 63 · replies 34 (at fetch time)
Archived 2026-09-30 via fxtwitter (unofficial).
Related events
All posts · id: x-firstadopter-2092266927377039838