What happened
OpenAI released GPT‑6.1 Sol, an upgraded version of GPT‑6 Sol that OpenAI says matches Astra’s capabilities in coding, computer‑use, and professional workflows while costing roughly one‑fifth of Astra’s standard token rates. At the same time, the company rolled out a premium Ultrafast tier that promises up to 300 tokens per second—about 8× faster in Codex and 6× faster via the API—at a six‑fold price premium over standard processing.
OpenAI’s VentureBeat report states that GPT‑6.1 Sol is priced at $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens—exactly one‑fifth the uncached rates of GPT‑6 Astra and half the rates of the older GPT‑5.6 Sol. The cached‑input price is also cut by 50% versus GPT‑6 Sol.
OpenAI claims the model narrows the performance gap with Astra across several internal benchmarks: DeepSWE v1.1 (software‑engineering tasks), GDP.pdf (complex professional documents), AutomationBench (end‑to‑end business tool workflows), and OSWorld 2.0 (offline computer‑use tasks). In each case, OpenAI reports that GPT‑6.1 Sol either matches or exceeds Astra’s performance at roughly one‑fifth the cost.
The new Ultrafast tier is described as a premium lane that can reach up to 300 generated tokens per second. OpenAI says this translates to up to 8× faster generation in Codex and up to 6× faster via the API. The tier costs six times the standard token price, which the article derives as $12 input and $60 output per million tokens for GPT‑6.1 Sol (assuming the 6× multiplier applies uniformly).
Availability: GPT‑6.1 Sol is already accessible through the API (model name gpt‑6.1‑sol) and for Plus, Pro, Business, Enterprise, and Edu customers in ChatGPT Work and Codex, but not in the regular Chat interface. Ultrafast is live for GPT‑6 Astra and will be added for GPT‑6.1 Sol in the coming days, with access controlled via a service_tier=ultrafast parameter.
Source details: venturebeat.com ↗
Why it matters
The launch reshapes the cost‑performance calculus for enterprise developers. By offering a model that approaches frontier capability at a fraction of the price, OpenAI enables more budget‑conscious autonomous agents to run at scale. The Ultrafast tier, meanwhile, gives organizations a high‑latency option for human‑in‑the‑loop scenarios such as interactive coding, customer support, and financial analysis, where response time directly impacts business value. Together, the two offerings force architects to balance three variables—intelligence, cost, and latency—rather than simply picking the most capable model.
Cost reduction: The one‑fifth token price relative to Astra makes high‑capability models affordable for large‑scale autonomous agents, lowering the barrier for enterprises to adopt AI‑driven automation across sales, support, finance, and HR workflows.
Latency trade‑off: The Ultrafast tier creates a clear option for latency‑sensitive applications, where the premium cost can be justified by the business value of faster responses. This differentiates OpenAI’s offering from competitors that typically bundle speed and cost together.
Strategic positioning: By separating cost‑effective and high‑speed tiers, OpenAI nudges developers to optimize deployments based on specific workload characteristics rather than defaulting to the most expensive, highest‑capability model.
Market impact: The announced token‑per‑second rates place OpenAI’s Ultrafast tier above many standard offerings but still behind specialized speed‑focused models like Mercury 2 and Celeris‑1, indicating a focus on balanced capability rather than raw throughput.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
crm_get_transaction(id='4092').In AI, what are a model's "parameters"?
What to watch next
Future pricing details for the Ultrafast tier, especially on GPT‑6.1 Sol, and independent verification of the speed and quality claims. Adoption patterns among enterprise customers will indicate whether the premium latency tier gains traction beyond niche use cases. Additionally, any policy or safety updates related to the accelerated models could affect broader deployment strategies.
Independent verification: Third‑party benchmarks of GPT‑6.1 Sol’s performance and the Ultrafast tier’s token‑per‑second claims are needed to confirm OpenAI’s internal results.
Pricing clarity: OpenAI has not published exact Ultrafast token rates for GPT‑6.1 Sol; the article’s figures are derived from the announced 6× multiplier. Future official pricing tables will clarify actual costs.
Adoption signals: Monitoring enterprise usage patterns—especially in sectors that value low latency—will reveal whether the Ultrafast tier gains traction beyond experimental use cases.
Safety and policy updates: Any future safety or policy changes affecting high‑speed could influence the availability or pricing of the Ultrafast tier.