返回新闻
产品展示AI Understanding 简报

Sakana AI 推出 Fugu Ultra v2 和成本更低的 Fugu Max

据报道,Sakana AI 推出了两款编排产品,声称 Fugu Ultra v2 超过了选定的前沿模型基准分数,而 Fugu Max 则降低了使用成本。

4 min readRead the linked source
Source-provided image accompanying Sakana AI unveils Fugu Ultra v2 and lower-cost Fugu Max
来源参考来源记录
出版商
finance.biggo.com
来源链接
finance.biggo.comhttps://finance.biggo.com/news/b97559a4-53ff-450c-afaa-38c98658fc76
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

从这里开始

关键术语

API(应用程序编程接口)
一种软件系统向另一个系统发送请求并接收响应的结构化方式。
基准测试
用于测量和比较模型性能的标准化测试或数据集。
推理
经过训练的模型生成预测或输出的运行时阶段。
测试一下自己AI 模型解释测验

发生了什么

According to BigGo Finance, Sakana AI launched Fugu Ultra v2 and Fugu Max, two versions of its system for routing tasks among specialized AI models. Sakana claims Fugu Ultra v2 exceeded GPT-6 Astra and Claude Fable 5.1 on selected benchmarks, while Fugu Max emphasizes lower cost and speed. Both are reportedly available through an OpenAI-compatible API with metered pricing and subscriptions.

BigGo Finance reports that Tokyo-based Sakana AI introduced Fugu Ultra v2, a flagship version of its Sakana Fugu orchestration system, and Fugu Max, a lower-cost variant. The system dynamically routes tasks among multiple specialized models through a single OpenAI-compatible API, so users reportedly do not need to manage the internal model selection. The report attributes the product descriptions and performance claims to Sakana AI.

Sakana claims Fugu Ultra v2 scored 74.3 on DeepSWE, compared with 74.1 for GPT-6 Astra and 67.4 for Claude Fable 5.1. It also reports a 48.3 score on Chartography, while noting that comparison scores for the other two models were not publicly disclosed. Sakana says Fugu Ultra v2 has an August 28, 2026 training-data cutoff and does not use GPT-6 Astra or Claude Fable 5.1 as underlying collaborators. These figures and claims were not independently confirmed in the source.

The report says Fugu Max incorporates NVIDIA’s open-weight Nemotron models and routes tasks to smaller models when appropriate. It claims Fugu Max reduces per-output usage costs by 40–60% compared with leading competitors and achieved the strongest position on a cost-versus-score comparison using Terminal Bench 2.1. Both products are described as available through API metered billing or subscriptions. Reported API prices are $5 input and $30 output per million tokens for Fugu Ultra v2, and $2 input and $6 output for Fugu Max. Subscription tiers are listed at $20, $100, and $200 per month.

来源详情: finance.biggo.com ↗

为什么这很重要

The reported launch is notable because it presents model orchestration—not simply scaling one large model—as a route to competitive performance. If Sakana’s results and pricing are independently validated, businesses could have another way to deploy AI agents while reducing dependence on individual frontier-model providers. The practical significance remains uncertain because the report does not establish the products’ availability, reliability, or performance outside the cited tests.

The reported products illustrate a potentially important shift in AI-system design. A well-coordinated collection of specialized models could let a provider match some capabilities of larger proprietary systems while controlling costs and avoiding dependence on a single closed model. That would matter most for companies running continuous agent workloads, where token costs and routing efficiency can materially affect deployment economics.

The evidence is currently limited. BigGo Finance is the named reporting outlet, and the results are presented primarily as Sakana’s claims. The source does not provide independent test results, methodology details, model versions beyond the names given, error analysis, latency measurements, or evidence that the systems perform consistently across ordinary business tasks. The comparison is also incomplete because Chartography scores for the cited comparison models were reportedly unavailable.

If the prices and performance claims hold, Fugu Max could give enterprises another option for high-volume agent processing, while Fugu Ultra v2 could appeal to users prioritizing quality and reasoning. The source does not establish general availability, eligibility requirements, service-level commitments, data-retention terms, privacy protections, rate limits, regional access, or whether the subscription plans include API usage. Those unknowns limit the immediate purchasing significance of the announcement.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下来看什么

Independent evaluations of the claims, documentation of the underlying model lineup, and confirmation of API access will determine how meaningful the launch is. Buyers should also examine rate limits, data handling, geographic availability, service reliability, and whether the reported cost savings hold for their workloads.

The first priority is independent replication of the DeepSWE, Chartography, and Terminal Bench 2.1 results, including consistent prompts, model configurations, routing policies, and full cost calculations. Without that information, the reported advantages cannot be compared reliably with the named frontier models.

Potential users should verify whether the API is accessible to the public or only to approved customers, and should confirm token accounting, context limits, latency, uptime, data-use policies, and pricing in current documentation. The source reports prices but does not independently confirm them.

It will also be important to see whether Sakana publishes the underlying model roster and explains how routing decisions are made. Those details would clarify reproducibility, vendor dependence, security exposure, and whether the system’s performance comes from broadly available open models or from undisclosed components.

相关指南和测验

人工智能模型解释人工智能代理ChatGPT 与大语言模型测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?