เรียนรู้
ติวเตอร์ AI
ข่าว
เครื่องมือ
งาน
อภิธาน
ใบรับรอง
ควิซ
ภารกิจ
ช่วยเหลือ
English
Search
⌘K
ส่งเครื่องมือ
บริจาค
English
Search
⌘K
เรียนรู้
AI Guides & Foundations
ติวเตอร์ AI
Ask AI anything, free
ข่าว
Latest AI Developments
เครื่องมือ
Top AI Directory
งาน
AI Hiring Board
อภิธาน
AI Terms Dictionary
ใบรับรอง
Get Your AI Certificate
ควิซ
Interactive AI Assessments
ภารกิจ
Why We Exist
ช่วยเหลือ
Help and Contact
ส่งเครื่องมือ
บริจาค
English
Home
/
Glossary
/
Reward Model
AI Glossary Term
What is Reward Model?
Definition
A model that scores outputs based on preference signals, often used in RLHF pipelines.
Related terms
Reinforcement Learning from Human Feedback (RLHF)
A training method that uses human preference signals to shape model behavior.
Reinforcement Learning
Training by reward signals where an agent learns actions that maximize long-term return.
Ground Truth
Trusted reference labels used to train or evaluate model outputs.
DPO (Direct Preference Optimization)
A training method that fine-tunes models directly on preference pairs without needing a separate reward model.
Mid-training
An intermediate training phase between pretraining and post-training, often used for capability or domain adjustments.
AI Agent
A software system that can observe, reason, and take actions to achieve a goal, often using tools and memory.
Learn more in our free guides
AI Models Explained
Model Collapse
Model Lifecycle
Model Context Protocol
Browse the full AI Glossary
Explore all free AI guides