Skip to content
Learn
News
Tools
Jobs
Mission
Search
⌘K
English
Donate
Menu
×
Start here
Start learning
Learn
Go to Learn
Courses
Guides
AI Tutor
Prompt library
Quizzes
Certification
Glossary
News
Go to News
Latest AI news
Topic trackers
Research & datasets
Blog
Editorial standards
Corrections
Tools
Go to Tools
Tool directory
Best-of lists
Compare tools
Cost calculator
Prompt Refiner
Submit AI Tool
Jobs
Go to Jobs
Browse AI jobs
Post a job
Sponsor AIU
Mission
Go to Mission
Mission
Impact & stats
Authors
Help centre
What's new
Support
Donate
Home
/
Glossary
/
Usage-Based Billing
AI Glossary Term
What does “
Usage-Based Billing
” mean?
Definition
Pricing where costs scale with API calls, tokens, inference time, or consumed compute.
Related terms
Inference-Time Compute
The amount of processing power consumed while producing each response.
Test-Time Compute
Additional inference computation used during response generation to improve quality or reasoning.
Speculative Decoding
An inference acceleration method where a small draft model proposes tokens that a larger model verifies in parallel.
Sliding Window Attention
An attention pattern where each token attends only to a fixed-size window of nearby tokens to reduce compute.
Inference
The runtime phase where a trained model generates predictions or outputs.
RAG (Retrieval-Augmented Generation)
A method that retrieves external knowledge and feeds it into generation at inference time.
Learn more in our free guides
Entropy-Based Sampling
Energy-Based Models
Score-Based Generative Models
WaveGlow Flow-Based Vocoder
See also
Trust Calibration
Zero Data Retention
Structured Output
KV Cache
Shadow Deployment
MCP (Model Context Protocol)
Safety Case
Agentic Loop
Previous
Trust Calibration
Next
Zero Data Retention
Browse the full AI Glossary
Explore all free AI guides