Jalapeño is OpenAI's inference-cost bet, not a Nvidia replacement

Jalapeño is OpenAI's first custom AI inference chip, built with Broadcom in nine months and positioned as the company's answer to a $20.92 billion 2025 operating loss. The chip is being framed as a step toward infrastructure independence. In the same funding cycle, OpenAI took $30 billion from Nvidia, $50 billion from Amazon, signed AMD MI450 agreements, and counted on Cerebras, all while Microsoft shipped the Maia 200 accelerator that already runs OpenAI's GPT-5.2 models in Azure. That is not consolidation; it is vendor expansion. Jalapeño is a unit-economics bet, not a replacement for the multi-vendor stack the company is simultaneously deepening.

OpenAI's 2025 financials make the urgency legible. The source attributes the figures to financial documents posted by Ed Zitron, whom the source identifies as an AI critic and public-relations specialist, and reports $13.07 billion in 2025 revenue against $34 billion in operational expenses, producing a $20.92 billion operating loss. R&D alone accounted for $19.18 billion, roughly 56 percent of total spending, and the source reports OpenAI paid Microsoft more than $10.59 billion for R&D and compute in 2025. With a 2026 IPO that the source characterizes as heavily anticipated, the question of inference cost control has moved from engineering optimization to investor narrative. Greg Brockman's statement included in Broadcom's release, that designing more of the stack lets OpenAI "serve more intelligence with greater efficiency," is a unit-economics argument dressed as an architectural one.

What Jalapeño actually is, in the terms the source provides, is an Application-Specific Integrated Circuit, or ASIC, designed for LLM inference rather than general GPU workloads. Broadcom is contributing core silicon implementation and Tomahawk networking silicon; Celestica is handling board, rack, and system integration. The OpenAI-Broadcom partnership was publicly announced only in October 2025, and the source says the chip moved from early schematics to fabrication readiness within nine months. The companies attribute the pace to a software-hardware co-development process that used OpenAI's own models to accelerate parts of the design. The source does not specify which parts of the design pipeline those models touched or how that compares to conventional EDA toolchains.

The first production test is GPT-5.3-Codex-Spark, running in a test environment on at least one prior-generation workload. OpenAI says rollout across active data centers is planned by the end of the year. None of the source's reporting includes benchmark numbers, dollar-per-token comparisons, power-efficiency figures, or competitive performance data against Nvidia's Vera Rubin platform, AMD's MI450, or Microsoft's Maia 200. The source explicitly flags performance versus competitors, cost, and manufacturing viability as outstanding questions. That gap matters because the central claim, that vertical integration will reduce inference unit cost, depends on numbers the announcement does not provide.

The "vertical integration" framing also runs against the deal flow the source describes. In February 2026, Nvidia finalized a $30 billion direct investment as part of a $110 billion funding round, securing commitments for 10 gigawatts of computing systems, including 3 gigawatts of dedicated inference capacity and 2 gigawatts of training capacity on the next-generation Vera Rubin platform. The source reports, citing sources close to the companies, that Nvidia will remain central to OpenAI, particularly on the model training and development side. In the same round, Amazon committed $50 billion alongside an agreement for OpenAI to consume approximately two gigawatts of AWS Trainium capacity over eight years. AMD Instinct MI450 and Cerebras, which the source notes executed its IPO in May 2026, round out a portfolio that now spans at least four silicon families in addition to Jalapeño. That is vendor diversification, not consolidation.

Google's Tensor Processing Units have matured for years, Amazon's Trainium has a similar runway, Microsoft has shipped Maia 100 and the Maia 200 on TSMC's 3-nanometer process that already runs OpenAI's GPT-5.2 models, and Meta's MTIA line spans the 300, 400, 450, and 500 series. The source frames Jalapeño as OpenAI's way to "match and offset" that infrastructure lead. A nine-month ASIC project, supported by Broadcom's networking stack and Celestica's integration, does not yet close the per-chip maturity gap those programs accumulated over multiple generations and years of in-house iteration, and the source does not show that it does.

The deployment target is gigawatt-scale data centers with Microsoft and other partners beginning in 2026. Broadcom's share price has responded, up 18 percent year-over-year in the first part of 2026 and roughly 7x since the end of 2022, according to the source's CNBC citation. The chip's first public test will be whether dollar-per-token data arrives in 2026, or whether the next concrete numbers come from IPO disclosures and the gigawatt-scale rollouts the source describes, not from independent benchmark comparisons.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe