OpenAI Built Its Own AI Chip. Watch out, Nvidia!
OpenAI's first custom AI chip beat Nvidia GB200 and GB300 systems in its opening inference tests. The results are early, but the threat to Nvidia is already real.
OpenAI built its first AI chip, called it Jalapeño, tested it against Nvidia's GB200 and GB300 systems, and published the sort of opening numbers that should make Jensen Huang pay attention.
Across three large AI models, OpenAI says Jalapeño completed 1.5 to 1.9 times more work for the same amount of power, while returning the full response 1.7 to 3.6 times faster than the Nvidia systems in the comparison. In the most interactive part of the test, where every small delay becomes noticeable to the person waiting, the advantage reached 4.1 times.
Sam Altman reduced the announcement to six words: "we made a chip and it is fast." Fair enough, because this one does not need much decorating.
The meme going around X shows Nvidia's CEO watching one of his biggest customers build its own hardware, and while the exact customer ranking is difficult to prove, the business point is obvious. OpenAI still needs Nvidia, still plans to buy millions of its GPUs and will probably remain an enormous customer for years, yet it now has a working chip for the repeated inference tasks that keep ChatGPT, Codex and the API running all day, which means a part of Nvidia's future revenue has quietly become OpenAI's engineering problem.
That is close enough to an Nvidia killer for us.
Quick answer
OpenAI has built a credible Nvidia killer for AI inference, the part of AI that runs a trained model whenever ChatGPT answers, an API processes a request or an agent moves to its next step.
In OpenAI's first published test, Jalapeño beat Nvidia GB200 and GB300 systems across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, delivering 1.5 to 1.9 times more work per watt at peak throughput, cutting full-response latency by 1.7 to 3.6 times and reaching up to 4.1 times higher performance in highly interactive workloads.
The numbers come from OpenAI, using the public InferenceX benchmark, so they deserve attention without being treated as the final word. Production deployment still has to show whether Jalapeño remains this fast and efficient when thousands of chips are running continuously, software is being updated, racks fail and real customers behave less neatly than a benchmark.
The threat to Nvidia is already serious because OpenAI and Broadcom are planning 10 gigawatts of OpenAI-designed accelerators, with the first systems due inside OpenAI by the end of 2026. As soon as Jalapeño becomes the cheaper default for a large share of the predictable work OpenAI currently pays other companies to handle, it has done enough damage.
The first numbers are genuinely bad for Nvidia
OpenAI tested Jalapeño on GPT-OSS 120B, DeepSeek R1 670B and the one-trillion-parameter Kimi K2.5, then compared the results with Nvidia GB200 and GB300 systems through InferenceX, an open benchmark maintained by semiconductor research firm SemiAnalysis.
Here is the useful part, without the chip-launch vocabulary:
| AI model | Nvidia system | Jalapeño peak work per watt | Full-response latency |
|---|---|---|---|
| GPT-OSS 120B | GB200 | 1.9 times higher | 1.03 seconds vs 1.80 seconds |
| DeepSeek R1 670B | GB300 | 1.7 times higher | 1.65 seconds vs 5.99 seconds |
| Kimi K2.5 1T | GB300 | 1.5 times higher | 1.56 seconds vs 5.31 seconds |
The DeepSeek result is the one we kept looking at. Jalapeño completed the full response in 1.65 seconds while the GB300 system took 5.99 seconds, and the minimum delay between generated tokens was 4.1 times lower, which is a large difference for an agent completing one step after another while somebody waits.
OpenAI also says that, at the fastest token rate previously achieved by the Nvidia system, Jalapeño handled more than 100 times as much throughput per kilowatt. That does not mean the chip is simply "100 times faster than Nvidia," because the figure comes from a matched operating point where both systems are held to the same token delay, although the less dramatic version remains impressive enough: Jalapeño was faster and more power efficient across every public model in the test.
Power is part of the story. Jalapeño carries a published rating of 700 watts, compared with 1,200 watts for GB200 and 1,400 watts for GB300, while OpenAI says its measured sustained consumption stayed at or below 550 watts during these workloads. The company still used the higher 700-watt figure in its calculations, which makes the comparison less flattering to its own chip than using the measured number would have been.

Why inference is the bill OpenAI wants to reduce
Training creates a model, inference runs it, so every ChatGPT reply, API request, coding task and agent decision adds another small amount to a bill that never stops arriving.
Training receives the giant supercomputer photographs because it needs enormous bursts of processing power, but inference follows the product everywhere. One question costs very little on its own, then hundreds of millions of users arrive, developers build products on the API, agents start taking dozens of steps for one request and those tiny costs become a permanent infrastructure problem.
We have already seen what happens when companies encourage heavy AI use before checking how quickly the token bill can grow. OpenAI sits on the expensive side of that equation because it has to buy accelerators, power them, connect them, cool them and keep the whole system fast enough that a conversation still feels like a conversation.
The price of each model call also changes which model developers can afford to use repeatedly. In our Grok 4.6, Claude Opus 5 and GPT-5.6 Sol test, the cheapest model completed the full set for much less, although higher token use reduced part of the advantage advertised by its rate card. OpenAI cannot control how many tokens every task needs, but it can try to reduce what the hardware underneath those tokens costs.
OpenAI also knows its own workload better than an outside supplier ever could. It sees which models people use, which parts of a response make the hardware wait, how long the model keeps information in memory and where the same expensive patterns repeat millions of times, then it can design a processor around those habits instead of buying the flexibility of a general AI accelerator for every request.
Jalapeño was built around that advantage, with the chip, memory, networking and software designed together so the model can keep more information close to the processing work. Less movement means less time waiting for data and less electricity spent carrying it around the system, which sounds boring until the saving is multiplied across a product the size of ChatGPT.
This is a 10-gigawatt programme
A fast laboratory chip can produce a lovely announcement and still vanish before it reaches a data centre, although the scale behind Jalapeño makes that outcome much harder to dismiss.
OpenAI and Broadcom announced their plan in October 2025, with 10 gigawatts of OpenAI-designed accelerators scheduled for deployment from the second half of 2026 through 2029. OpenAI designed the architecture around its own models and serving systems, Broadcom handled silicon implementation and networking, and Celestica is helping turn the design into boards, racks and production hardware.
The first chip moved from the beginning of development to manufacturing tape-out in nine months, helped by OpenAI's own models, and engineering samples are already running machine-learning workloads at the intended production speed and power. OpenAI plans to begin deployment by the end of 2026, while generation two is deep in development and generation three is already taking shape.
Ten gigawatts is too large to dismiss as a clever experiment. It is an infrastructure programme capable of redirecting a meaningful part of OpenAI's future hardware spending, and it gives the company enough volume to justify years of custom-chip development even if Jalapeño only handles the workloads it knows best.
Nvidia is still receiving an enormous order
The awkward part of this story is that OpenAI is building an Nvidia rival while placing some of the largest Nvidia orders ever announced.
In September 2025, the companies announced plans for at least 10 gigawatts of Nvidia systems, representing millions of GPUs, with Nvidia intending to invest as much as $100 billion as the capacity is deployed. The first gigawatt is expected in the second half of 2026 on Nvidia's Vera Rubin platform.
The relationship expanded again in August, when Nvidia announced that it would become the exclusive AI-compute provider for the PORTS-Pike campus in Ohio, where OpenAI will be the customer for eight gigawatts of capacity, with the first phases expected from 2028. OpenAI has also agreed to deploy six gigawatts of AMD GPUs, so its infrastructure plan now looks less like loyalty to one supplier and more like a company grabbing every serious source of compute it can find.
That does not make Jalapeño harmless, it shows how large OpenAI's demand has become. Nvidia can remain one of its most important suppliers while losing the assumption that every future ChatGPT answer must run on Nvidia hardware, and even a modest percentage moved onto OpenAI's own chips represents a lot of equipment that no longer needs to be bought at Nvidia's margins.
The benchmark is public, the measurements are OpenAI's
InferenceX is more useful than a private graph with no method behind it. SemiAnalysis describes the benchmark as open, reproducible and designed to compare the complete process of serving an AI request, with public recipes, logs and raw data that other people can inspect.
OpenAI still produced these measurements, chose the systems and configurations, then published the results while Jalapeño was preparing for deployment. We have three public models and useful data, rather than months of evidence from a large production fleet, and several costs that matter in the real world remain unknown, including complete rack cost, manufacturing yield, maintenance, uptime and performance across a wider collection of private workloads.
The comparison also uses GB200 and GB300 systems, while Nvidia's newer Vera Rubin platform is outside the test. Jalapeño is aimed at inference, where OpenAI's knowledge of its own products gives it the clearest advantage, so these results do not tell us that the chip can replace the enormous Nvidia clusters OpenAI still plans to use for training.
None of those limits erase the result. On three public models, using a public inference test, OpenAI's first chip beat two Nvidia systems on speed and power efficiency, which is exactly what the company built it to do, and the next useful evidence will arrive when the first large deployment starts working for real customers.
Broadcom may collect a lot of Nvidia's missing money
OpenAI designed Jalapeño, but Broadcom supplied the silicon implementation and Tomahawk networking needed to turn that design into hardware that can be produced at gigawatt scale, while Celestica contributed the board, rack and system expertise.
This gives OpenAI more control without asking it to become a semiconductor manufacturer overnight, and it moves part of the money that might have gone to Nvidia towards another group of suppliers. Broadcom is well placed for this change because the largest AI companies increasingly want processors shaped around their own services, yet those companies still need somebody who knows how to turn an architecture into working silicon and connect huge numbers of chips together.
Nvidia's defence has always included much more than the processor. CUDA, networking, software, developer support and complete systems make Nvidia difficult to remove once a company has built around it, and OpenAI is now creating its own answer to more of that stack, with Broadcom filling the parts it does not want to own itself.
What Jalapeño could mean for ChatGPT
Nobody opening ChatGPT tomorrow should expect a Jalapeño button, a new plan or an immediate price cut, because OpenAI has announced none of those things and the first systems are only due to begin deployment by the end of 2026.
If the published gains survive production, ChatGPT could begin and finish answers faster, coding and research agents could complete long chains of steps with less waiting, and OpenAI could serve more requests from the same amount of power. Lower costs would also give the company room to reduce API prices or increase usage limits, although it could just as easily keep the saving, buy more capacity and improve its own margin.
Speed may be the first benefit people notice. Saving two seconds on one answer is pleasant, while saving the same two seconds across 40 linked decisions inside an agent can change whether the product feels useful or spends most of its time making the user stare at a progress message.
TGK take on this
We think Nvidia's difficulty starts with a very ordinary business decision. OpenAI looks at the inference bill every month, knows exactly which expensive tasks repeat, has enough demand to justify its own hardware and can now point to a first chip that completed those tasks faster while using less power.
Jalapeño can be worse than Nvidia at a hundred general-purpose jobs and still be a financial success for OpenAI, because it only needs to be better at the work OpenAI performs all day. That narrow advantage is more dangerous than a broad promise, especially when the customer behind it has committed to 10 gigawatts and already has two more generations in development.
Nvidia will continue selling an absurd amount of hardware, its software advantage remains difficult to copy and Vera Rubin may produce a much tougher comparison, although the warning has already arrived: Nvidia's largest buyers have reached a size where its prices become an invitation to design around it, and OpenAI has moved beyond threatening to do that.
Calling Jalapeño an Nvidia killer is fair because the phrase describes the job, not the whole company. This chip was built to kill OpenAI's dependence on Nvidia for a large part of inference, and its first numbers suggest it can.
Frequently asked questions
Can another company buy an OpenAI Jalapeño chip?
OpenAI has announced Jalapeño for deployment inside its own compute infrastructure and has not announced it as a chip that other companies can order. A business using OpenAI products may eventually benefit from the hardware indirectly, but there is no public Jalapeño catalogue, price or sales date.
Why does OpenAI describe the plan in gigawatts instead of chip numbers?
The number of chips can change as each generation becomes faster or more efficient, while the electrical capacity of a data-centre programme gives a more stable picture of its size. Ten gigawatts describes planned power capacity, not ten gigawatts continuously consumed by one chip design and not a confirmed chip count.
Who actually builds Jalapeño?
OpenAI designed the accelerator architecture, Broadcom handles silicon implementation and networking, and Celestica contributes board, rack and system work. The public announcement does not identify every company in the manufacturing chain, so describing Jalapeño as entirely made by OpenAI would give the wrong picture.
Why did OpenAI publish results for public models instead of only its own models?
Public models make the comparison easier for outsiders to inspect and reproduce through InferenceX. OpenAI says Jalapeño's advantage became larger in its internal frontier-model testing, although those private results are less useful to readers who cannot examine the models or rerun the same work.
What evidence would make this result more convincing?
Independent reruns would help, but production evidence matters more: thousands of chips working for months, complete rack costs, uptime, manufacturing yield, software stability and performance across many workloads. OpenAI's current numbers are a strong opening test, while the deployment beginning at the end of 2026 will show whether the chip is also a dependable product.
Could Nvidia lower its prices to protect inference business?
Nvidia can change pricing, bundle more of its software and networking, or use newer hardware such as Vera Rubin to improve the comparison. The important change is that a customer with its own credible chip gains leverage, even if it continues buying Nvidia systems, because it no longer has to accept the same answer for every workload.
Primary sources
- OpenAI: Jalapeño's first performance results
- OpenAI and Broadcom unveil the Jalapeño inference chip
- OpenAI and Broadcom's 10-gigawatt deployment plan
- OpenAI and Nvidia's 10-gigawatt systems partnership
- Nvidia: PORTS-Pike campus and OpenAI's eight-gigawatt capacity
- OpenAI and AMD's six-gigawatt GPU agreement
- SemiAnalysis: How the InferenceX benchmark works