Skip to content
Learn
News
Tools
Jobs
Mission
Search
⌘K
English
Donate
Menu
×
Start here
Start learning
Learn
Go to Learn
Courses
Guides
AI Tutor
Prompt library
Quizzes
Certification
Glossary
News
Go to News
Latest AI news
Topic trackers
Research & datasets
Blog
Editorial standards
Corrections
Tools
Go to Tools
Tool directory
Best-of lists
Compare tools
Cost calculator
Prompt Refiner
Submit AI Tool
Jobs
Go to Jobs
Browse AI jobs
Post a job
Sponsor AIU
Mission
Go to Mission
Mission
Impact & stats
Authors
Help centre
What's new
Support
Donate
Home
/
Glossary
/
Vision-Language Model (VLM)
AI Glossary Term
What does “
Vision-Language Model (VLM)
” mean?
Definition
A multimodal model that jointly processes visual and textual information.
Related terms
Artificial Intelligence (AI)
The broad field of building systems that perform tasks requiring pattern recognition, reasoning, language, or decision-making.
CLIP
A multimodal model architecture that learns shared representations between text and images.
Computer Vision
The branch of AI that extracts meaning from images and video.
Context Window
The maximum amount of input tokens a language model can process at once.
Hallucination
When a model generates fluent but false or unsupported information.
Large Language Model (LLM)
A language model trained on massive text corpora to generate and analyze text.
Learn more in our free guides
AI Models Explained
Computer Vision
Model Collapse
Model Lifecycle
See also
Vector Database
Weak Supervision
Validation Set
Weight
Unsupervised Learning
Word Embedding
Training Loss
XAI (Explainable AI)
Previous
Vector Database
Next
Weak Supervision
Browse the full AI Glossary
Explore all free AI guides