Series Path · 4 parts
Vision-Language Models
Understanding, fine-tuning, and applying VLMs with LoRA.
~91 min total·4 articles
Start from Part 1Curriculum — 4 parts
- 01
Vision-Language Models: Bridging Visual and Textual Understanding
8 min read - 02
Fine-tuning Vision-Language Models: A Comprehensive Guide
24 min read - 03
LoRA for Vision-Language Models: A Comprehensive Guide
39 min read - 04
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
20 min read