返回新聞
產品展示AI Understanding 簡報

Anthropic 推出 Claude Sonnet 5.5,這是一款更快、更便宜的型號,效能接近 Opus 5.5

Anthropic 宣布推出 Claude Sonnet 5.5,這是一種新的中端 LLM,運行速度提高了 30%,每項任務的成本降低了 30%,同時在多個基準測試中提供接近旗艦 Opus 5.5 的性能。

4 min readRead the primary source
Source-provided image accompanying Anthropic unveils Claude Sonnet 5.5, a faster, cheaper model that nears Opus 5.5 performance
主要來源文件來源記錄
出版商
anthropic.com
來源連結
anthropic.comhttps://www.anthropic.com/claude-sonnet-5-5
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

大語言模型(LLM)
在海量文本語料庫上訓練來產生和分析文本的語言模型。
蒸餾
將知識從大型教師模式壓縮到較小的學生模式。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己AI 模型解釋測驗

發生了什麼事

Anthropic introduced Claude Sonnet 5.5, the second model in its Claude 5.5 family. The company says the new model runs more than 30 % faster than its predecessor, Claude Sonnet 5, and typically costs up to 30 % less per task because it needs fewer tokens. In internal testing, Sonnet 5.5 scored 70.6 % on the Terminal‑Bench 4.0 coding and only two points behind Opus 5.5 on the GDPval‑AA occupational benchmark. It also became the first Sonnet model to beat Pokémon Red using only screenshots, indicating improved long‑horizon reasoning and image understanding. Pricing remains $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads, but the reduced token usage translates into lower effective cost. The model is now available on AWS, Google Cloud, and Azure, and can be accessed via the Claude Platform using the identifier “claude‑sonnet‑5‑5.”

Anthropic’s announcement details Claude Sonnet 5.5 as the latest addition to the Claude 5.5 family, positioned as a faster, lower‑cost alternative to Claude Opus 5.5. The model improves on Sonnet 5 across multiple dimensions: it generates outputs more than 30 % faster, requires fewer tokens for comparable tasks, and delivers higher scores on coding‑focused evaluations such as Terminal‑Bench 4.0 (70.6 % vs. 10.3 % for Sonnet 5).

Performance testing shows Sonnet 5.5 trailing Opus 5.5 by only two points on the GDPval‑AA occupational , indicating near‑parity on real‑world work across 44 occupations. The model also excels in long‑horizon tasks and image understanding, becoming the first Sonnet variant to succeed at playing Pokémon Red using only visual inputs.

Pricing remains unchanged from Sonnet 5 ($2 / M input tokens, $10 / M output tokens, $0.20 / M cache reads), but the reduced token consumption translates into up to a 30 % cost reduction per task. The model is available across major cloud providers and can be accessed via the Claude Platform using the identifier “claude‑sonnet‑5‑5.”

Safety and alignment enhancements include cyber‑security safeguards comparable to Opus 5.5, new safety classifiers to mitigate attacks, and unchanged biology safeguards. Anthropic’s automated behavioral audit reports that Sonnet 5.5 matches or exceeds Sonnet 5 on most alignment metrics, with only marginal differences from Opus 5.5 in sandbox‑escape behavior.

來源詳情: anthropic.com ↗

為什麼這很重要

Claude Sonnet 5.5 represents a meaningful shift in the AI market by offering near‑flagship capabilities at a lower price point, narrowing the gap between mid‑tier and top‑tier offerings. For enterprises and developers, the speed and cost improvements mean faster iteration on routine coding, document creation, and design tasks, potentially lowering total cost of ownership for AI‑augmented workflows. The model’s alignment upgrades—including cyber‑security safeguards and new safety classifiers—address growing concerns about misuse and model extraction, setting a higher baseline for responsible deployment in commercial settings. By positioning Sonnet 5.5 as a cost‑effective complement to Opus 5.5, Anthropic expands the range of use cases that can be economically justified, from internal tooling to customer‑facing applications, and pressures competitors to improve pricing and safety features.

The launch narrows the performance‑cost gap between mid‑tier and flagship LLMs, making advanced AI capabilities more accessible to a broader set of developers and enterprises. Faster generation and lower token usage can accelerate development cycles for code‑heavy or document‑intensive workflows, reducing both time and compute expenses.

Anthropic’s emphasis on safety—cybersecurity safeguards, ‑attack classifiers, and a detailed alignment audit—addresses industry‑wide concerns about model misuse and data extraction. By embedding these controls in a mid‑tier model, Anthropic raises the baseline for responsible AI deployment across the market.

The model’s availability on all major cloud platforms simplifies integration for existing cloud‑native pipelines, potentially driving higher adoption rates and encouraging competition among cloud providers to offer optimized pricing or specialized services for Claude models.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

Future updates will reveal how Sonnet 5.5 performs in real‑world deployments, especially in high‑volume environments where its speed advantage matters most. Watch for adoption metrics on the Claude Platform and any pricing adjustments that could further compress the cost gap with Opus 5.5. Anthropic’s rollout of the Cyber Verification Program and Life Sciences Verification Program will indicate how the company balances expanded capabilities with tiered safety controls. Finally, monitor competitor responses—particularly OpenAI and other LLM providers—to see whether similar mid‑tier models with comparable cost‑performance ratios emerge.

Adoption metrics: Track usage statistics on the Claude Platform to gauge how quickly developers shift from Sonnet 5 or other mid‑tier models to Sonnet 5.5.

Pricing dynamics: Observe whether Anthropic adjusts token pricing or introduces volume discounts that could further lower the effective cost per task.

Safety program rollouts: Monitor enrollment and outcomes of the Cyber Verification Program and Life Sciences Verification Program, which will reveal how the new safeguards are applied in practice.

Competitive response: Watch for announcements from OpenAI, Google, and other LLM providers that may introduce comparable mid‑tier models with similar speed and cost advantages.

相關指引和測驗

人工智慧模型解釋AI 倫理Prompt Engineering測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?