Type to search across all posts
Keyboard Shortcuts
Press ? to close

Series Path · 7 parts

Vision-Language Models

Understanding, fine-tuning, and applying vision-language models — from BLIP-2 and LLaVA to LoRA fine-tuning and vision-language-action systems for robotics.

~180 min total·7 articles
Start from Part 1

Curriculum — 7 parts

  1. 01

    Vision-Language Models: Bridging Visual and Textual Understanding

    Vision-Language Models (VLMs) represent one of the most exciting frontiers in artificial intelligence, combining computer vision and natural language…

    8 min read
  2. 02

    Fine-tuning Vision-Language Models: A Comprehensive Guide

    Vision-Language Models (VLMs) represent a significant advancement in artificial intelligence, combining computer vision and natural language processing to…

    24 min read
  3. 03

    LoRA for Vision-Language Models: A Comprehensive Guide

    Low-Rank Adaptation (LoRA) has emerged as a revolutionary technique for efficient fine-tuning of large language models, and its application to Vision-Language…

    39 min read
  4. 04

    BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

    Paper: arXiv:2301.12597 (v3, 15 Jun 2023) Authors: Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi — Salesforce Research Code:…

    20 min read
  5. 05

    Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment

    A section-by-section walkthrough of arXiv:2312.12148v1

    39 min read
  6. 06

    Visual Instruction Tuning (LLaVA)

    Paper: Visual Instruction Tuning Authors: Haotian Liu¹, Chunyuan Li², Qingyang Wu³, Yong Jae Lee¹ ¹University of Wisconsin–Madison, ²Microsoft Research,…

    28 min read
  7. 07

    Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications

    A Comprehensive Survey — IEEE Access 2025

    22 min read