सीखें
एआई ट्यूटर
समाचार
टूल्स
नौकरियां
शब्दावली
प्रमाणपत्र
क्विज़
मिशन
सहायता
English
Search
⌘K
टूल भेजें
दान
English
Search
⌘K
सीखें
AI Guides & Foundations
एआई ट्यूटर
Ask AI anything, free
समाचार
Latest AI Developments
टूल्स
Top AI Directory
नौकरियां
AI Hiring Board
शब्दावली
AI Terms Dictionary
प्रमाणपत्र
Get Your AI Certificate
क्विज़
Interactive AI Assessments
मिशन
Why We Exist
सहायता
Help and Contact
टूल भेजें
दान
English
Home
/
Glossary
/
Reinforcement Learning from Human Feedback (RLHF)
AI Glossary Term
What is Reinforcement Learning from Human Feedback (RLHF)?
Definition
A training method that uses human preference signals to shape model behavior.
Related terms
Reinforcement Learning
Training by reward signals where an agent learns actions that maximize long-term return.
Reward Model
A model that scores outputs based on preference signals, often used in RLHF pipelines.
DPO (Direct Preference Optimization)
A training method that fine-tunes models directly on preference pairs without needing a separate reward model.
Continual Learning
Training approaches that let a model keep learning from new data without forgetting prior knowledge.
Few-Shot Learning
Learning or adapting behavior from only a small number of examples.
Model Drift
Performance degradation over time as real-world conditions diverge from training assumptions.
Learn more in our free guides
Deep Learning
Reinforcement Learning
Machine Learning Basics
Supervised Learning
Browse the full AI Glossary
Explore all free AI guides