Series Path · 7 parts
Vision-Language Models
Understanding, fine-tuning, and applying vision-language models — from BLIP-2 and LLaVA to LoRA fine-tuning and vision-language-action systems for robotics.
Curriculum — 7 parts
- 01
Vision-Language Models: Bridging Visual and Textual Understanding
Vision-Language Models (VLMs) represent one of the most exciting frontiers in artificial intelligence, combining computer vision and natural language…
8 min read - 02
Fine-tuning Vision-Language Models: A Comprehensive Guide
Vision-Language Models (VLMs) represent a significant advancement in artificial intelligence, combining computer vision and natural language processing to…
24 min read - 03
LoRA for Vision-Language Models: A Comprehensive Guide
Low-Rank Adaptation (LoRA) has emerged as a revolutionary technique for efficient fine-tuning of large language models, and its application to Vision-Language…
39 min read - 04
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Paper: arXiv:2301.12597 (v3, 15 Jun 2023) Authors: Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi — Salesforce Research Code:…
20 min read - 05
Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment
A section-by-section walkthrough of arXiv:2312.12148v1
39 min read - 06
Visual Instruction Tuning (LLaVA)
Paper: Visual Instruction Tuning Authors: Haotian Liu¹, Chunyuan Li², Qingyang Wu³, Yong Jae Lee¹ ¹University of Wisconsin–Madison, ²Microsoft Research,…
28 min read - 07
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
A Comprehensive Survey — IEEE Access 2025
22 min read