Learn
AI Tutor
News
Tools
Jobs
Glossary
Certification
Quizzes
Mission
Support
English
Search
⌘K
Submit AI Tool
Donate
English
Search
⌘K
Learn
AI Guides & Foundations
AI Tutor
Ask AI anything, free
News
Latest AI Developments
Tools
Top AI Directory
Jobs
AI Hiring Board
Glossary
AI Terms Dictionary
Certification
Get Your AI Certificate
Quizzes
Interactive AI Assessments
Mission
Why We Exist
Support
Help and Contact
Submit AI Tool
Donate
English
Home
/
Glossary
/
Inference Endpoint
AI Glossary Term
What is Inference Endpoint?
Definition
A deployed API interface that receives model requests and returns predictions in production.
Related terms
API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Inference
The runtime phase where a trained model generates predictions or outputs.
Usage-Based Billing
Pricing where costs scale with API calls, tokens, inference time, or consumed compute.
Inference-Time Compute
The amount of processing power consumed while producing each response.
Decision Tree
A model that makes predictions through a sequence of if-then feature splits.
Ensemble
Combining predictions from multiple models to improve robustness or accuracy.
Learn more in our free guides
AI Inference
AI Inference Optimization
LLM Inference Routing and Load Balancing
Seldon Core and Inference Graphs
Browse the full AI Glossary
Explore all free AI guides