
CAP6412 Advanced Computer Vision - Spring 2024
Dormant Last read · last published · next check
Read 3 days ago and current, but nothing has been published for 2 years.
Latest videos
Saves to your Watch queue, to pick up on another day or another device.


Lecture 14 - Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Lecture 4 - Visual-Language Models Introduction Part-I: CoCA, PALI

Lecture 3 - CLIP

Lecture 2 - Transformers Introduction

Lecture 1 - Introduction

Lecture 5 - Visual-Language Models Introduction Part-II: FLAMINGO, FLAVA, PAINTER, BLIP-2

Lecture 6 - Visual-Language Models Introduction Part-III: Image-Bind, Language-Bind, LLaVA

Lecture 7 - Visual-Language Models Introduction Part-IV: Video ChatGPT, PG-Video LLaVA

Lecture 13 - MERLOT RESERVE: Neural Script Knowledge through Vision and Language and Sound

Lecture 12 - MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks

Lecture 10-BLIP:Bootstrapping Language-Image Pretraining for Unified VL Understanding and Generation

Lecture 11 - BLIP-2 : Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMs

Lecture 9 - HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

