Back to News
InnovationAI Understanding briefing

Nous Research launches Hermes Index to benchmark agentic AI models and costs

Nous Research unveiled the Hermes Index, a new benchmark that ranks 14 frontier AI models on task performance and per‑task cost within its Hermes Agent framework, and announced a $90 million funding round.

4 min readRead the linked source
Source-provided image accompanying Nous Research launches Hermes Index to benchmark agentic AI models and costs
Source referenceSource recorded
Publisher
cryptobriefing.com
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Key terms

Benchmark
A standardized test or dataset used to measure and compare model performance.

What happened

Nous Research released the Hermes Index on Oct 6 2026, ranking 14 AI models on performance and cost using its Hermes Agent framework and the new Hermes Bench suite. Claude Opus 5.5 topped the list, while DeepSeek V4.1 Flash was the cheapest. The lab raised $90 million in a funding round led by Nvidia and Microsoft’s M12 on Oct 7 2026.

On Oct 6 2026, open‑source AI lab Nous Research launched the Hermes Index, an evaluation platform that measures both task performance and per‑task cost for frontier AI models operating inside the lab’s Hermes Agent framework. The index aggregates results from four test suites, including the newly introduced Hermes Bench—a collection of 150 tasks across 25 categories—alongside established benchmarks such as Terminal‑Bench 4.0 and SkillsBench.

The first edition evaluated 14 models. Claude Opus 5.5 achieved the highest overall score of 63.31 with an average cost of $4.99 per task. GPT‑6 Astra placed second with a score of 56.25 but a higher cost of $11.61 per task. Claude Sonnet 5.5 ranked third (score 53.14, $2.82 per task). At the low‑cost end, DeepSeek V4.1 Flash scored 36.91 at $0.259 per task, and Ling 3.0 Flash scored 21.56 at $0.054 per task.

One day after the index went live, Nous Research announced a $90 million financing round, bringing total capital raised to roughly $160 million and valuing the company at $1.5 billion. Investors included Nvidia and Microsoft’s venture arm M12. The funding is earmarked for scaling enterprise deployments of the Hermes technology, which will remain open‑source under an MIT license.

Source details: cryptobriefing.com ↗

Why it matters

The Hermes Index provides a rare combined view of how well agentic AI models complete tasks and how much they cost, giving enterprises concrete data for large‑scale deployments where per‑task expense can dominate budgets. By publishing cost‑aware rankings, Nous Research pushes the industry toward more transparent pricing and performance trade‑offs, potentially influencing model selection, pricing strategies, and future design.

Enterprises that run thousands of AI‑driven agent tasks daily need clear data on both effectiveness and cost. The Hermes Index directly addresses this need, allowing decision‑makers to compare models not just on scores but on the financial impact of large‑scale use.

By publishing cost‑per‑task figures, the index may pressure model providers to improve pricing transparency and efficiency, potentially reshaping competitive dynamics in the agentic AI market.

The open‑source nature of Hermes Agent and the index could encourage broader community participation, fostering a shared standard for evaluating agentic AI performance and economics.

The simultaneous funding round signals strong investor confidence in commercializing cost‑aware AI evaluation tools, suggesting that similar benchmarking services may emerge as a new segment of the AI ecosystem.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

What to watch next

Future Hermes Index updates, adoption of Hermes Agent by enterprise customers, and competitive responses from model providers seeking to improve cost efficiency or challenge the ’s methodology.

Updates to the Hermes Index, including additional models, expanded task suites, and refined cost metrics.

Enterprise adoption rates of Hermes Agent and any reported cost savings or performance gains from real‑world deployments.

Responses from AI model vendors, such as new pricing tiers, performance optimizations, or challenges to the methodology.

Potential collaborations or integrations between Nous Research and cloud providers or AI platforms that could broaden the index’s reach.

Related guides & quizzes

Found this useful?