RSS Amplifier

Video feed

CAP6412 Advanced Computer Vision - Spring 2024

youtube.comSource feed ↗15 videos

Dormant Last read · last published · next check
Read 3 days ago and current, but nothing has been published for 2 years.

Written by

Latest videos

Saves to your Watch queue, to pick up on another day or another device.

Lecture 15 - Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Play

Lecture 14 - Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Play

Lecture 4 - Visual-Language Models Introduction Part-I: CoCA, PALI

Play

Lecture 3 - CLIP

Play

Lecture 2 - Transformers Introduction

Play

Lecture 1 - Introduction

Play

Lecture 5 - Visual-Language Models Introduction Part-II: FLAMINGO, FLAVA, PAINTER, BLIP-2

Play

Lecture 6 - Visual-Language Models Introduction Part-III: Image-Bind, Language-Bind, LLaVA

Play

Lecture 7 - Visual-Language Models Introduction Part-IV: Video ChatGPT, PG-Video LLaVA

Play

Lecture 13 - MERLOT RESERVE: Neural Script Knowledge through Vision and Language and Sound

Play

Lecture 12 - MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks

Play

Lecture 10-BLIP:Bootstrapping Language-Image Pretraining for Unified VL Understanding and Generation

Play

Lecture 11 - BLIP-2 : Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMs

Play

Lecture 9 - HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

Play

Lecture 8 - FILIP: Fine-grained Interactive Language-Image Pre-Training

Play