Learn
AI Tutor
News
Tools
Jobs
Glossary
Certification
Quizzes
Mission
Support
English
Search
⌘K
Submit AI Tool
Donate
English
Search
⌘K
Learn
AI Guides & Foundations
AI Tutor
Ask AI anything, free
News
Latest AI Developments
Tools
Top AI Directory
Jobs
AI Hiring Board
Glossary
AI Terms Dictionary
Certification
Get Your AI Certificate
Quizzes
Interactive AI Assessments
Mission
Why We Exist
Support
Help and Contact
Submit AI Tool
Donate
English
Home
/
Glossary
/
Inference-Time Compute
AI Glossary Term
What is Inference-Time Compute?
Definition
The amount of processing power consumed while producing each response.
Related terms
Test-Time Compute
Additional inference computation used during response generation to improve quality or reasoning.
Usage-Based Billing
Pricing where costs scale with API calls, tokens, inference time, or consumed compute.
Distilled Model
A smaller model trained to imitate a larger model's behavior while using less compute at inference.
Compute
The processing resources required to train and run models, often measured in FLOPS or GPU hours.
RAG (Retrieval-Augmented Generation)
A method that retrieves external knowledge and feeds it into generation at inference time.
Alignment Tax
The extra cost in time, compute, or product velocity required to make systems safer and more controllable.
Learn more in our free guides
AI Inference
Computer Vision
Sentiment Analysis
Real-Time Voice Agents
Browse the full AI Glossary
Explore all free AI guides