Skip to content
Learn
News
Tools
Jobs
Mission
Search
⌘K
English
Donate
Menu
×
Start here
Start learning
Learn
Go to Learn
Courses
Guides
AI Tutor
Prompt library
Quizzes
Certification
Glossary
News
Go to News
Latest AI news
Topic trackers
Research & datasets
Blog
Editorial standards
Corrections
Tools
Go to Tools
Tool directory
Best-of lists
Compare tools
Cost calculator
Prompt Refiner
Submit AI Tool
Jobs
Go to Jobs
Browse AI jobs
Post a job
Sponsor AIU
Mission
Go to Mission
Mission
Impact & stats
Authors
Help centre
What's new
Support
Donate
Home
/
Glossary
/
Reinforcement Learning
AI Glossary Term
What does “
Reinforcement Learning
” mean?
Definition
Training by reward signals where an agent learns actions that maximize long-term return.
Related terms
Reinforcement Learning from Human Feedback (RLHF)
A training method that uses human preference signals to shape model behavior.
Reward Model
A model that scores outputs based on preference signals, often used in RLHF pipelines.
AI Agent
A software system that can observe, reason, and take actions to achieve a goal, often using tools and memory.
Constitutional AI
A training and behavior-shaping approach where model outputs are guided by a fixed set of written principles.
Agentic Loop
An iterative cycle where an AI agent observes, plans, acts, and reflects until it completes a goal or hits a stop condition.
DPO (Direct Preference Optimization)
A training method that fine-tunes models directly on preference pairs without needing a separate reward model.
Learn more in our free guides
Deep Learning
Reinforcement Learning
Machine Learning Basics
Supervised Learning
See also
Red Teaming
Recommendation System
Retrieval
Recall
RAG (Retrieval-Augmented Generation)
Robustness
Quantization
Safety Filter
Previous
Red Teaming
Next
Reinforcement Learning from Human Feedback (RLHF)
Browse the full AI Glossary
Explore all free AI guides