
What OpenAI actually announced
I build AI agents for a living, so I read chip launches through one lens: does this change what the models on top can do, or what they cost to run? Jalapeño is squarely the second kind, and it is a big one.
On June 24, 2026, OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom in-house chip. The name is a joke about heat and speed; the substance is that OpenAI is no longer renting all of its compute from someone else. It designed its own silicon, TechCrunch reported, and Bloomberg noted it had already received the first physical samples and was testing them.

The scale is the headline most people missed. This chip is one piece of a partnership to deploy 10 GW of OpenAI-designed accelerators over the next few years, a build-out OpenAI frames as part of its wider "full stack" plan to own the whole thing: models, products, serving software, chips, memory, and networking. Jalapeño is where that plan stops being a slide and becomes working silicon.
What Jalapeño is, and what it isn't
A few distinctions matter here, because "OpenAI made a chip" gets flattened into "OpenAI is fighting Nvidia" and that misses the point.
Jalapeño is an ASIC, an application-specific chip. Where an Nvidia GPU is a flexible general-purpose processor that can train models, run graphics, and serve inference, an ASIC does one narrow job extremely well and nothing else. Jalapeño's one job is inference, the serving side where a trained model reads your prompt and generates an answer. It does not train models. OpenAI still uses Nvidia and other partners for that.
And you cannot buy it. Jalapeño is not a product; it runs inside OpenAI's own infrastructure to serve ChatGPT, Codex, the API, and agentic products. For everyone else, the chip shows up indirectly, as faster and cheaper OpenAI models, not as hardware in your rack. If you want the specialised-inference-hardware-you-can-actually-rent version of this story, that market already exists in platforms like Baseten.
One more thing worth being fair about: this was fast. OpenAI says it went from design to tapeout in nine months, which is fast for custom silicon and tells you how much of the industry's talent and money is now pointed at inference cost.
The benchmark numbers
On August 25, 2026, OpenAI published Jalapeño's first results, and the numbers are the reason this launch got talked about.

Tested across three public models, GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, OpenAI reported that Jalapeño delivered:
- 1.5 to 1.9 times more AI work per watt at peak throughput
- 1.7 to 3.6 times lower end-to-end latency
- 2.1 to 4.1 times higher performance on highly interactive workloads
The comparison chip in OpenAI's own charts is the Nvidia GB200, rated at 1,200 watts against Jalapeño's 700. The point OpenAI keeps hammering is that it hit both higher throughput and lower latency from one design, where existing hardware usually forces a tradeoff between the two.

The independent read comes from SemiAnalysis, whose InferenceX benchmark OpenAI used. Their verdict was blunt in the headline, "Better Than Nvidia Blackwell," citing Jalapeño hitting nearly 700 tokens per second per user and more than 9 times the next best chip at a matched interactivity target. Worth holding one thing in mind, which I get to below: SemiAnalysis owns the benchmark OpenAI chose to run.
The most interesting part: the chip designed itself
Here is the detail I would put on the front page, and OpenAI almost buried it. The reason nine-month silicon was even possible is that OpenAI used its own models to design the chip.
Earlier models helped bring the chip up; newer ones are optimising and programming it. Using Codex with GPT-Astra, the team brought three open-weight models that were not in the original plan to high performance in two months. For selected attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than the human-expert versions.

That is a loop: models help design the silicon, the silicon runs the models faster, the faster models help design the next silicon. Gen 2 is "deep in development" and Gen 3 is "taking shape." Whatever you think of the benchmarks, this feedback loop is the part that compounds, and it is the same reason AI coding assistants keep getting a bigger say in how real products get built.
What the community made of it
The reaction was not a clean victory lap, and the skepticism is fair enough to include.
The sharpest critique on Hacker News was that this is not an apples-to-apples fight:
"OpenAI's chip is targeting inference (FP8, FP4), while most of the chips it is being compared to are general purpose."
A purpose-built inference ASIC should beat a general-purpose GPU at inference per watt. That is what specialising buys you, so the win is real but also somewhat expected. Others pointed at who wrote the scorecard:
"This semi-analysis article reads a lot more like an OpenAI press release than a real analysis."
On the investing side of Reddit, the framing was less about tokens and more about margins, with r/NVDA_Stock treating custom silicon as a slow-burn threat to Nvidia's pricing power rather than an overnight one. My own read lands in the middle: the numbers are probably directionally true and real for OpenAI's workloads, and also exactly the numbers you would expect a specialised chip, benchmarked on its author's own test, to produce.
Why a chip announcement matters if you're not OpenAI
You are not going to install Jalapeño. So why care? Because inference cost is the quiet number under every AI product you use or build.

Every AI support reply, every drafted email, every AI agent that reads an order and answers a customer is an inference call that costs real money to serve. When OpenAI drives that cost down, it improves its own margins first, but the effect flows up the stack over time, in cheaper API pricing and models that do more per dollar. That is the tailwind that keeps turning AI support from a demo into a line item that obviously pays for itself, the same math behind how much AI can actually save in customer support.
It is also a reminder of where the real leverage sits for most teams. You do not compete by owning silicon; you compete by putting the models to work on your actual problems. The chip is the factory floor. The AI agent handling your tickets is the employee, and the employee is the part you get to hire.
Try eesel
Jalapeño is infrastructure, the layer under the models. eesel is the layer on top, the part you actually put to work. It is an AI teammate platform: you hire ready-to-work teammates for specific jobs, and the AI helpdesk teammate joins your existing Zendesk, Freshdesk, or Gorgias queue, learns from your past tickets and help center, and starts drafting or sending replies.
The differentiator I would point to is that you can simulate a teammate against your historical tickets before it ever answers a real customer, so you see how it will perform first instead of flipping a switch and hoping. Pricing is usage-based at 40 cents per ticket with no per-seat fees, and you can start free. Cheaper silicon under the hood is nice; a teammate you can actually try today is the part that changes your week.
Frequently Asked Questions
What is OpenAI's Jalapeño chip?
Is the Jalapeño chip better than Nvidia?
Can I buy or use the Jalapeño chip?
Why does OpenAI's Jalapeño chip matter for businesses?
What does the Jalapeño chip mean for AI customer service?

Article by
Alicia Kirana Utomo
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.








