ΟΔΗΓΟΣ Εφαρμογών

Gross Margins and Inference Costs in AI Businesses

An AI business's gross margin is its revenue minus the direct cost of delivering the product, divided by revenue.

  • 4 min read
  • Τελευταία ενημέρωση
Σε αυτήν τη σελίδα4 min read
  1. Επισκόπηση
  2. Βαθιά κατάδυση
  3. Στρατηγικός αντίκτυπος
  4. The Future of Gross Margins and Inference Costs in AI Businesses
  5. Υλοποίηση σε πραγματικό κόσμο
  6. Κίνδυνοι & προστατευτικά κιγκλιδώματα
  7. Οδικός Χάρτης Εφαρμογής
  8. Συνεχίστε την εξερεύνηση
  9. Συχνές ερωτήσεις

Επισκόπηση

For AI software that cost includes inference, the compute spent every time a model answers a request. Inference grows with usage instead of staying close to zero for each extra user, so AI products often have lower and less predictable margins than classic SaaS. That affects how they set prices, how they raise money and whether they survive.

Βαθιά κατάδυση

Gross margin is revenue minus cost of revenue, divided by revenue. In classic SaaS, cost of revenue is mostly hosting, customer support and third-party software. Serving one more user costs almost nothing, so mature SaaS companies commonly report gross margins of 70 to 80 percent or higher. AI products add a cost that behaves more like a raw material. Every prompt uses GPU time. The company pays for it either per token to an API provider such as OpenAI, Anthropic or Google, or through GPUs it owns or rents. A heavy user can cost many times more than a light user on the same subscription. A widely read 2020 Andreessen Horowitz essay, "The New Business of AI", argued that many AI companies looked like a mix of software and services, with gross margins pulled down by cloud compute and humans in the loop. Human work still counts: labeling data, reviewing outputs and handling escalations often sit in cost of revenue alongside moderation, retrieval infrastructure such as vector databases, and evaluation runs. Companies protect margins with a familiar set of levers: - **Model routing** sends easy requests to small, cheap models and keeps large models for hard ones. - **Caching** reuses work. Providers can cache repeated prompt prefixes, and a product can store answers to identical queries. - **Smaller or distilled models**, sometimes fine-tuned for one task, can replace a general model for high-volume work. - **Leaner requests** help too: tighter prompts, limits on output length, and batch processing for jobs that can wait. - **Pricing** is the last lever. Usage-based tiers, credits or caps tie revenue to cost. A common misconception is that falling token prices will fix margins on their own. Prices for a given level of capability have dropped sharply. But products often use up the savings by moving to newer models, longer contexts and agents that make many calls per task. Margin comes from engineering and pricing choices, not from some fixed property of AI.

Στρατηγικός αντίκτυπος

Δημιουργήστε επιλογές

Ο σχεδιασμός σε επίπεδο εφαρμογής καθορίζει εάν η τεχνητή νοημοσύνη βελτιώνει τα πραγματικά αποτελέσματα.

Ομάδα και ροή εργασίας

Η καλή ενσωμάτωση ροής εργασιών δημιουργεί κέρδη παραγωγικότητας που μπορούν να εμπιστευτούν οι χρήστες.

Κίνδυνος και ασφάλεια

Οι καλές περιπτώσεις χρήσης μειώνουν την κόπωση λόγω αλλαγής και τον κίνδυνο εφαρμογής.

The Future of Gross Margins and Inference Costs in AI Businesses

Inference costs per unit of capability have been falling. Hardware, serving software and smaller models keep improving, and competition among model providers holds prices down. That trend should help, but it will not automatically raise margins, because longer contexts, reasoning models and multi-step agents use up much of the saving. More companies are likely to treat cost engineering as a core discipline and build in routing, caching and task-specific models from the start. Many will likely move from unlimited flat pricing to plans that account for usage. A business that depends entirely on one upstream provider will keep facing margin risk from that supplier's pricing decisions.

Υλοποίηση σε πραγματικό κόσμο

A customer-support chatbot startup sends simple FAQ questions to a small, cheap model and passes only complex tickets to a frontier model. This lowers its average cost per conversation without hurting answers on the hard cases.

A coding assistant sold at a flat monthly price per seat finds that a small group of power users accounts for a large share of its inference bill. It adds usage caps and a premium tier so revenue follows cost.

A contract-summarization tool puts its long system prompt and standard clause library at the start of every request. Provider prompt caching then bills those repeated tokens at a discounted rate.

A legal research company sees that one type of query dominates its traffic. It fine-tunes a smaller open-weight model for that task and stops paying a general-purpose API for it.

Κίνδυνοι & προστατευτικά κιγκλιδώματα

  • Η αυτοματοποίηση μιας διαλυμένης διαδικασίας μπορεί να ενισχύσει τα υπάρχοντα προβλήματα.

  • Οι ομάδες μπορεί να αυτοματοποιήσουν υπερβολικά και να αφαιρέσουν την απαραίτητη ανθρώπινη κρίση.

  • Η ποιότητα μπορεί να αλλάξει αν τα αποτελέσματα δεν αξιολογούνται συνεχώς.

Οδικός Χάρτης Εφαρμογής

  1. Χαρτογραφήστε την τρέχουσα ροή εργασίας και εντοπίστε το βήμα της υψηλότερης τριβής.

  2. Καθορίστε ανθρώπινα σημεία ελέγχου πριν από την πλήρη αυτοματοποίηση.

  3. Εκπαιδεύστε τους χρήστες σε προτροπές, διαδρομές κλιμάκωσης και πρότυπα ποιότητας.

  4. Παρακολουθήστε τα αποτελέσματα σε επίπεδο εργασίας για να επιβεβαιώσετε τη σταθερή αξία.

Συνεχίστε την εξερεύνηση

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Gross Margins and Inference Costs in AI Businesses quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Έναρξη κουίζ

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Συχνές ερωτήσεις

What is Gross Margins and Inference Costs in AI Businesses?

An AI business's gross margin is its revenue minus the direct cost of delivering the product, divided by revenue. For AI software that cost includes inference, the compute spent every time a model answers a request. Inference grows with usage instead of staying close to zero for each extra user, so AI products often have lower and less predictable margins than classic SaaS. That affects how they set prices, how they raise money and whether they survive.

According to the guide, why does serving an extra user usually cost more for an AI product than for classic SaaS?

In classic SaaS, an extra user adds almost no cost. In AI products, every request uses compute, paid per token or through GPUs, so cost rises with usage and behaves more like a raw material.

What does model routing do?

Routing matches each request to the cheapest model that can handle it well. That lowers average cost per request. Storing outputs for repeated questions is a different lever: caching.

What did the 2020 Andreessen Horowitz essay "The New Business of AI" argue?

The essay argued that cloud compute costs and human involvement, such as labeling and review, gave many AI companies lower gross margins than typical software businesses.

Why don't falling token prices automatically fix AI gross margins?

Cheaper tokens help. But when products adopt larger models, longer contexts or multi-call agents, total spend per task can stay flat or even rise.

To get the most out of provider prompt caching, how should you structure a prompt?

Caching works on a stable, repeated prefix. Keeping unchanging material at the start and variable content at the end means more of each request matches the cached prefix.