Google 的 Gemini 4 Argon 发布增加了新的基准测试结果,在 19 项测试中的 13 项中击败了领先竞争对手,并通过 Fairwind 计划向经过审查的网络安全团队进行了有限的推广,介绍性 API 定价为每百万输入代币 2 美元。
发生了什么
Google 于 9 月 30 日发布了 Gemini 4 Argon,报告称其在 19 个内部基准测试中的 13 个上优于 GPT-6 Astra 和 Claude Opus 5.5。该模型最初仅通过 Google Fairwind 计划向部分网络安全合作伙伴开放,并计划稍后扩展到付费 API 客户和 Google AI Ultra 订阅者。 Argon 的高设定在 9 月 30 日的 Arena 文本排行榜上名列前茅(4,942 票),并在人工分析的高推理指数上获得 53 分,与 GPT-6 Astra 相当。它还在 DeepSWE v1.1 软件工程基准测试中取得了 77.9% 的成绩。 API 费率从每百万输入代币 2 美元和每百万输出代币 10 美元开始,在介绍期后升至 4 美元和 20 美元。
9月30日,Google发布了Gemini 4 Argon,称其为迄今为止最强大的型号。该公司的内部比较显示,Argon 在 19 项基准测试中的 13 项中击败了 GPT-6 Astra 和 Claude Opus 5.5,其中包括基于 4,942 票在 Arena 文本排行榜上名列前茅。
Access is currently restricted to vetted cybersecurity teams via the Fairwind Program, a trusted‑defender initiative. Google 表示稍后将向付费 API 客户和 Google AI Ultra 订阅者开放该模型,但没有给出推出日期。
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Models Explained Quiz
Which component of an AI application is the machine-learning model itself?