返回新聞
產品展示AI Understanding 簡報

OpenAI 在 2026 年有限部署之前報告 Jalapeo 晶片基準

Firstpost 報告稱,在該公司運行的測試中,OpenAI 的客製化 Jalapeo 推理晶片比基於 Nvidia Blackwell 的系統提供了更高的速度和能效,計劃在 2026 年底進行有限部署,並在 2027 年進行更廣泛的部署。

6 min readRead the linked source
Source-provided image accompanying OpenAI reports Jalapeño chip benchmarks ahead of limited 2026 deployment
來源參考來源記錄
出版商
firstpost.com
來源連結
firstpost.comhttps://www.firstpost.com/tech/openai-reveals-jalapeno-ai-chip-benchmark-results-plans-wider-deployment-in-2027-14040784.html
來源類型
連結來源-主要來源狀態尚未確定。
還引用了

故事最後修訂

背景60 秒內了解這一點

從這裡開始

關鍵術語

記憶體(代理記憶體)
AI 代理程式跨步驟或會話使用儲存的上下文來提高連續性。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
推理
經過訓練的模型產生預測或輸出的運行時階段。
測試一下自己AI 模型解釋測驗

自發布以來發生了什麼變化

  1. 首次發表
  2. OpenAI’s primary-source update materially advances the existing Jalapeño chip story by publishing first-party benchmark results. It reports higher performance per watt and lower latency across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, describes AI-assisted chip design and programming, and sets a planned deployment target of the end of 2026 while acknowledging that qualification and scale validation remain unfinished.
  3. This TechCrunch report materially advances the existing Jalapeño update with the first reported InferenceX benchmark results and an expected deployment schedule. TechCrunch says OpenAI reported higher tokens per user and throughput per kilowatt than an Nvidia Blackwell system, with very small-volume deployment expected by the end of 2026 and broader deployment in 2027. The benchmark claims and timeline are not independently confirmed in the supplied source.
  4. This is a material first-party update to the continuing Jalapeño inference-chip event already represented by the canonical entry. OpenAI now publishes its methodology, model-specific benchmark figures, power ratings, architecture description, AI-assisted development claims and planned end-of-year deployment. The new results remain OpenAI-reported; the source does not provide independent validation or confirm production availability.
  5. This source materially advances the existing Jalapeño chip update by adding OpenAI’s broader full-stack strategy, its stated infrastructure portfolio, the claim that future chip generations are underway, Project Camellia’s facility commitments, and a separate claim that GPT-5.6 Sol used 54% fewer output tokens than another leading model on a coding-agent index.
  6. OpenAI’s new primary-source update provides first detailed measured Jalapeño results across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, including claimed gains in throughput per watt, end-to-end latency and selected AI-generated kernels, while setting out planned infrastructure deployment by the end of 2026.
  7. This materially advances the continuing Jalapeño-chip event already covered by the canonical update. The Verge adds reported briefing details, the InferenceX comparison, model-specific performance ranges, and OpenAI’s small-volume deployment target for late 2026 with a 2027 ramp. The benchmark claims and deployment plans remain attributed to OpenAI and are not independently confirmed in the source material.
  8. Firstpost materially advances the existing Jalapeño report with OpenAI’s first reported benchmark results across three named language models, claimed efficiency and latency ranges, technical details about prefill, KV-cache placement and system communication, and a stated timeline of limited deployment by the end of 2026 followed by broader rollout in 2027. The figures remain OpenAI’s claims and are not independently confirmed in the source.

發生了什麼事

Firstpost reports that OpenAI presented its first public results for Jalapeño, a custom AI chip developed with Broadcom. OpenAI says the chip outperformed Nvidia Blackwell-based comparison systems on selected language-model workloads, including tests involving GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. The company expects limited internal deployment by the end of 2026 and a broader rollout in 2027.

Firstpost reports that OpenAI presented Jalapeño in greater detail at the Hot Chips conference and released what it described as the chip’s first public results. The chip was developed with Broadcom as a custom accelerator for AI , the stage in which a trained model generates responses. The report describes Jalapeño as part of OpenAI’s longer-term effort to build specialized infrastructure around its own models and products.

According to Firstpost, OpenAI tested three models on Jalapeño: GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. The tests used SemiAnalysis’ InferenceX and compared Jalapeño with systems based on Nvidia’s Blackwell architecture. Firstpost attributes the results to OpenAI and does not report an independent reproduction, audit or competing assessment of the methodology.

Firstpost says OpenAI reported that Jalapeño processed between 1.5 and 1.9 times more work per watt than the comparison systems while reducing overall latency by between 1.7 and 3.6 times. For workloads requiring frequent interaction, OpenAI reported performance improvements ranging from 2.1 to 4.1 times. These figures describe the company’s selected test results, not a general finding about every model, workload or deployment environment.

The report also describes a comparison at an operating point associated with the previous best time between generated tokens. At that setting, OpenAI reported between 8.6 and 104.3 times more work per watt, depending on the model. Firstpost does not provide enough information in the supplied text to determine how that operating point was selected, how the comparison systems were configured, or how representative the result is of normal production use.

Firstpost reports that OpenAI expects to deploy Jalapeño in limited volumes within its own infrastructure by the end of 2026, followed by a more meaningful rollout in 2027. The company also says work on its next two chip generations is already under way. The report does not establish that the chip is currently available to outside customers or that the planned deployment schedule is firm.

來源詳情: firstpost.com ↗

為什麼這很重要

The results suggest OpenAI is pursuing greater control over the hardware used to serve its models, particularly as multi-step AI agents increase demand. Higher performance per watt could affect operating costs and system design, but the reported figures are OpenAI’s claims from tests and have not been independently confirmed in the source.

The immediate significance is strategic as well as technical. Firstpost reports that OpenAI wants a multigenerational hardware platform in which future models, products, chips and memory systems can be designed together. If that plan succeeds, OpenAI could have more influence over how its models are served and optimized instead of relying entirely on general-purpose accelerators supplied by other companies.

efficiency matters because serving a model involves repeated computation for every response. Firstpost reports that OpenAI designed Jalapeño for workloads in which latency can accumulate across many sequential steps, such as AI agents that call models repeatedly while using tools, checking results and deciding what to do next. Faster individual operations could therefore affect the responsiveness of longer-running agent tasks, although the source provides no independent measurements of complete agent workflows.

The reported energy-efficiency gains could also matter for the cost and physical scale of AI services. More work per watt may reduce electricity demand for a given workload or allow a system to provide more within a fixed power budget. Those implications remain conditional: the source does not provide purchase prices, total operating costs, manufacturing yields, cooling requirements, utilization rates or results from a production-scale facility.

Jalapeño’s reported design choices target known bottlenecks. Firstpost says the architecture addresses the prefill stage, memory bandwidth during token generation and communication between processing units. It also reports that the system keeps model state, including the KV cache used during response generation, closer to where it is needed and combines computing, memory and networking resources. The source does not independently verify whether those choices deliver the claimed benefits outside the reported tests.

The story also illustrates the limits of headline comparisons. Nvidia’s Blackwell systems are the reference point in the reported tests, but Firstpost notes that competing hardware is likely to advance before Jalapeño reaches broader deployment. The result is therefore best understood as an early company-reported comparison rather than a settled ranking of AI infrastructure. OpenAI also expects to continue using Nvidia accelerators and hardware from other partners, so the chip is not described as an immediate replacement for its existing suppliers.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The key test will be whether Jalapeño can be produced and deployed at scale, and how it compares with newer Nvidia and rival accelerators available when that happens. Important unknowns include production volume, cost, reliability, deployment sites, customer access and whether the reported gains persist across OpenAI’s broader workload mix.

The first practical milestone is whether OpenAI begins the limited internal deployment it described for the end of 2026. Monitoring that step should include how many systems are installed, which models or services use them, whether the deployment is experimental or production-facing, and whether OpenAI reports operational results beyond the figures. None of those details is established by Firstpost’s report.

The 2027 rollout will show whether Jalapeño is a functioning platform rather than a one-generation engineering project. Key evidence would include manufacturing scale, availability of the required memory and networking components, system reliability, utilization and cost per unit of useful . The source gives no production volume, supplier breakdown beyond Broadcom’s development role, pricing or service-level information.

Future comparisons will need to account for the hardware available at the time of deployment. Firstpost explicitly notes that newer Nvidia systems and other rival processors may be on the market by 2027. A meaningful assessment should therefore compare the same models, response-quality requirements, batch sizes, latency targets, power assumptions and software stack across contemporary systems rather than rely on today’s Blackwell baseline.

It is also important to watch whether the reported gains generalize beyond the three models named in the article. GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T represent specific model and workload choices, but the source does not say how Jalapeño performs across OpenAI’s proprietary models, multimodal systems, smaller models, long-context requests or agent tasks involving external tools. Those unknowns limit what can be inferred about broad product impact.

Finally, OpenAI’s continued use of Nvidia and other suppliers will reveal how the company balances custom and commercial hardware. A hybrid strategy could let OpenAI use Jalapeño for workloads where its architecture is advantageous while retaining outside accelerators for flexibility or capacity. The report does not say whether Jalapeño will ever be sold externally, licensed, or used only inside OpenAI’s own infrastructure.

相關指引和測驗

人工智慧模型解釋人工智慧代理人工智慧培訓變形金剛測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器

更新和更正

當正在發生的事件發生重大變化時,這個典型的故事就會被更新。它的 URL 和原始發布日期永遠不會改變。

  • Firstpost materially advances the existing Jalapeño report with OpenAI’s first reported benchmark results across three named language models, claimed efficiency and latency ranges, technical details about prefill, KV-cache placement and system communication, and a stated timeline of limited deployment by the end of 2026 followed by broader rollout in 2027. The figures remain OpenAI’s claims and are not independently confirmed in the source.
  • This materially advances the continuing Jalapeño-chip event already covered by the canonical update. The Verge adds reported briefing details, the InferenceX comparison, model-specific performance ranges, and OpenAI’s small-volume deployment target for late 2026 with a 2027 ramp. The benchmark claims and deployment plans remain attributed to OpenAI and are not independently confirmed in the source material.
  • OpenAI’s new primary-source update provides first detailed measured Jalapeño results across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, including claimed gains in throughput per watt, end-to-end latency and selected AI-generated kernels, while setting out planned infrastructure deployment by the end of 2026.
  • This source materially advances the existing Jalapeño chip update by adding OpenAI’s broader full-stack strategy, its stated infrastructure portfolio, the claim that future chip generations are underway, Project Camellia’s facility commitments, and a separate claim that GPT-5.6 Sol used 54% fewer output tokens than another leading model on a coding-agent index.
  • This is a material first-party update to the continuing Jalapeño inference-chip event already represented by the canonical entry. OpenAI now publishes its methodology, model-specific benchmark figures, power ratings, architecture description, AI-assisted development claims and planned end-of-year deployment. The new results remain OpenAI-reported; the source does not provide independent validation or confirm production availability.
  • This TechCrunch report materially advances the existing Jalapeño update with the first reported InferenceX benchmark results and an expected deployment schedule. TechCrunch says OpenAI reported higher tokens per user and throughput per kilowatt than an Nvidia Blackwell system, with very small-volume deployment expected by the end of 2026 and broader deployment in 2027. The benchmark claims and timeline are not independently confirmed in the supplied source.
  • OpenAI’s primary-source update materially advances the existing Jalapeño chip story by publishing first-party benchmark results. It reports higher performance per watt and lower latency across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, describes AI-assisted chip design and programming, and sets a planned deployment target of the end of 2026 while acknowledging that qualification and scale validation remain unfinished.
查看公開更正日誌
覺得有用嗎?