Skip to content
Learn
News
Tools
Jobs
Mission
Search
⌘K
English
Donate
Menu
×
Start here
Start learning
Learn
Go to Learn
Courses
Guides
AI Tutor
Prompt library
Quizzes
Certification
Glossary
News
Go to News
Latest AI news
Topic trackers
Research & datasets
Blog
Editorial standards
Corrections
Tools
Go to Tools
Tool directory
Best-of lists
Compare tools
Cost calculator
Prompt Refiner
Submit AI Tool
Jobs
Go to Jobs
Browse AI jobs
Post a job
Sponsor AIU
Mission
Go to Mission
Mission
Impact & stats
Authors
Help centre
What's new
Support
Donate
Home
/
Glossary
/
Model Quantization
AI Glossary Term
What does “
Model Quantization
” mean?
Definition
Reducing numeric precision of model weights to decrease memory and inference cost.
Related terms
Quantization
Converting model weights to lower precision formats such as 8-bit or 4-bit.
QLoRA
A fine-tuning technique that combines 4-bit weight quantization with LoRA adapters to reduce memory needs.
Flash Attention
An optimized attention algorithm that reduces memory use and speeds up transformer training and inference.
Memory (Agent Memory)
Stored context an AI agent uses across steps or sessions to improve continuity.
AI Agent
A software system that can observe, reason, and take actions to achieve a goal, often using tools and memory.
AI Safety
A field focused on reducing harmful behavior, failures, and misuse risks in AI systems.
Learn more in our free guides
AI Models Explained
Model Collapse
Model Lifecycle
Model Context Protocol
See also
Model Drift
Multimodal Model
Model Card
Named Entity Recognition (NER)
Mixture of Experts (MoE)
Natural Language Processing (NLP)
Neural Network
Machine Learning (ML)
Previous
Model Drift
Next
Multimodal Model
Browse the full AI Glossary
Explore all free AI guides