What happened
According to BigGo Finance, Sakana AI launched Fugu Ultra v2 and Fugu Max, two versions of its system for routing tasks among specialized AI models. Sakana claims Fugu Ultra v2 exceeded GPT-6 Astra and Claude Fable 5.1 on selected benchmarks, while Fugu Max emphasizes lower cost and speed. Both are reportedly available through an OpenAI-compatible API with metered pricing and subscriptions.
BigGo Finance reports that Tokyo-based Sakana AI introduced Fugu Ultra v2, a flagship version of its Sakana Fugu orchestration system, and Fugu Max, a lower-cost variant. The system dynamically routes tasks among multiple specialized models through a single OpenAI-compatible API, so users reportedly do not need to manage the internal model selection. The report attributes the product descriptions and performance claims to Sakana AI.
Sakana claims Fugu Ultra v2 scored 74.3 on DeepSWE, compared with 74.1 for GPT-6 Astra and 67.4 for Claude Fable 5.1. It also reports a 48.3 score on Chartography, while noting that comparison scores for the other two models were not publicly disclosed. Sakana says Fugu Ultra v2 has an August 28, 2026 training-data cutoff and does not use GPT-6 Astra or Claude Fable 5.1 as underlying collaborators. These figures and claims were not independently confirmed in the source.
The report says Fugu Max incorporates NVIDIA’s open-weight Nemotron models and routes tasks to smaller models when appropriate. It claims Fugu Max reduces per-output usage costs by 40–60% compared with leading competitors and achieved the strongest position on a cost-versus-score comparison using Terminal Bench 2.1. Both products are described as available through API metered billing or subscriptions. Reported API prices are $5 input and $30 output per million tokens for Fugu Ultra v2, and $2 input and $6 output for Fugu Max. Subscription tiers are listed at $20, $100, and $200 per month.
Source details: finance.biggo.com ↗
Why it matters
The reported launch is notable because it presents model orchestration—not simply scaling one large model—as a route to competitive performance. If Sakana’s results and pricing are independently validated, businesses could have another way to deploy AI agents while reducing dependence on individual frontier-model providers. The practical significance remains uncertain because the report does not establish the products’ availability, reliability, or performance outside the cited tests.
The reported products illustrate a potentially important shift in AI-system design. A well-coordinated collection of specialized models could let a provider match some capabilities of larger proprietary systems while controlling inference costs and avoiding dependence on a single closed model. That would matter most for companies running continuous agent workloads, where token costs and routing efficiency can materially affect deployment economics.
The evidence is currently limited. BigGo Finance is the named reporting outlet, and the benchmark results are presented primarily as Sakana’s claims. The source does not provide independent test results, methodology details, model versions beyond the names given, error analysis, latency measurements, or evidence that the systems perform consistently across ordinary business tasks. The comparison is also incomplete because Chartography scores for the cited comparison models were reportedly unavailable.
If the prices and performance claims hold, Fugu Max could give enterprises another option for high-volume agent processing, while Fugu Ultra v2 could appeal to users prioritizing quality and reasoning. The source does not establish general availability, eligibility requirements, service-level commitments, data-retention terms, privacy protections, rate limits, regional access, or whether the subscription plans include API usage. Those unknowns limit the immediate purchasing significance of the announcement.
What to watch next
Independent evaluations of the benchmark claims, documentation of the underlying model lineup, and confirmation of API access will determine how meaningful the launch is. Buyers should also examine rate limits, data handling, geographic availability, service reliability, and whether the reported cost savings hold for their workloads.
The first priority is independent replication of the DeepSWE, Chartography, and Terminal Bench 2.1 results, including consistent prompts, model configurations, routing policies, and full cost calculations. Without that information, the reported advantages cannot be compared reliably with the named frontier models.
Potential users should verify whether the API is accessible to the public or only to approved customers, and should confirm token accounting, context limits, latency, uptime, data-use policies, and pricing in current documentation. The source reports prices but does not independently confirm them.
It will also be important to see whether Sakana publishes the underlying model roster and explains how routing decisions are made. Those details would clarify reproducibility, vendor dependence, security exposure, and whether the system’s performance comes from broadly available open models or from undisclosed components.