OpenAI Jalapeño Inference Chip just announced at Hot Chips

OpenAI Jalapeño chip delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance.

In June, OpenAI unveiled the chip program in partnership with Broadcom, built from a blank slate exclusively for LLM inference. Design work began in the middle of 2024, going from initial team hiring to manufacturing tape-out in ~16 months, an extremely fast ASIC development cycle.

Semianalysis thinks the comparison to Blackwell is somewhat incomplete and unfair. Jalapeño is really competing against chips like Rubin that also use HBM4. Vera Rubin systems are starting to ship to customers now. OpenAI has to scale beyond engineering samples of Jalapeño.

Vera Rubin NVL72 delivers 5.4x the perf/MW of GB200 NVL72

Screenshot
Screenshot

Screenshot

Screenshot

Screenshot