Skip to content
Learn
News
Tools
Jobs
Mission
Search
⌘K
English
Donate
Menu
×
Start here
Start learning
Learn
Go to Learn
Courses
Guides
AI Tutor
Prompt library
Quizzes
Certification
Glossary
News
Go to News
Latest AI news
Topic trackers
Research & datasets
Blog
Editorial standards
Corrections
Tools
Go to Tools
Tool directory
Best-of lists
Compare tools
Cost calculator
Prompt Refiner
Submit AI Tool
Jobs
Go to Jobs
Browse AI jobs
Post a job
Sponsor AIU
Mission
Go to Mission
Mission
Impact & stats
Authors
Help centre
What's new
Support
Donate
Home
/
Glossary
/
Inference-Time Compute
AI Glossary Term
What does “
Inference-Time Compute
” mean?
Definition
The amount of processing power consumed while producing each response.
Related terms
Test-Time Compute
Additional inference computation used during response generation to improve quality or reasoning.
Usage-Based Billing
Pricing where costs scale with API calls, tokens, inference time, or consumed compute.
Distilled Model
A smaller model trained to imitate a larger model's behavior while using less compute at inference.
Compute
The processing resources required to train and run models, often measured in FLOPS or GPU hours.
RAG (Retrieval-Augmented Generation)
A method that retrieves external knowledge and feeds it into generation at inference time.
Alignment Tax
The extra cost in time, compute, or product velocity required to make systems safer and more controllable.
Learn more in our free guides
AI Inference
Computer Vision
Sentiment Analysis
Real-Time Voice Agents
See also
Inference
Instruction Tuning
In-Context Learning
Intent Classification
Hyperparameter
Jailbreak
Human-in-the-Loop
Knowledge Cutoff
Previous
Inference
Next
Instruction Tuning
Browse the full AI Glossary
Explore all free AI guides