RSS Amplifier

Topic · bootstrapping language-image

bootstrapping language-image

The 15 most recent episodes and tracks on this topic.

Saves to your Watch queue, to pick up on another day or another device.

Pick anything below and it plays in the bar at the foot of the window — and keeps playing while you go on browsing the directory.

  1. Lecture 15 - Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionCAP6412 Advanced Computer Vision - Spring 2024Notes
  2. Lecture 14 - Shikra: Unleashing Multimodal LLM's Referential Dialogue MagicCAP6412 Advanced Computer Vision - Spring 2024Notes
  3. Lecture 4 - Visual-Language Models Introduction Part-I: CoCA, PALICAP6412 Advanced Computer Vision - Spring 2024Notes
  4. Lecture 3 - CLIPCAP6412 Advanced Computer Vision - Spring 2024Notes
  5. Lecture 2 - Transformers IntroductionCAP6412 Advanced Computer Vision - Spring 2024Notes
  6. Lecture 1 - IntroductionCAP6412 Advanced Computer Vision - Spring 2024Notes
  7. Lecture 5 - Visual-Language Models Introduction Part-II: FLAMINGO, FLAVA, PAINTER, BLIP-2CAP6412 Advanced Computer Vision - Spring 2024Notes
  8. Lecture 6 - Visual-Language Models Introduction Part-III: Image-Bind, Language-Bind, LLaVACAP6412 Advanced Computer Vision - Spring 2024Notes
  9. Lecture 7 - Visual-Language Models Introduction Part-IV: Video ChatGPT, PG-Video LLaVACAP6412 Advanced Computer Vision - Spring 2024Notes
  10. Lecture 13 - MERLOT RESERVE: Neural Script Knowledge through Vision and Language and SoundCAP6412 Advanced Computer Vision - Spring 2024Notes
  11. Lecture 12 - MaMMUT: A Simple Architecture for Joint Learning for MultiModal TasksCAP6412 Advanced Computer Vision - Spring 2024Notes
  12. Lecture 10-BLIP:Bootstrapping Language-Image Pretraining for Unified VL Understanding and GenerationCAP6412 Advanced Computer Vision - Spring 2024Notes
  13. Lecture 11 - BLIP-2 : Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMsCAP6412 Advanced Computer Vision - Spring 2024Notes
  14. Lecture 9 - HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware AttentionCAP6412 Advanced Computer Vision - Spring 2024Notes
  15. Lecture 8 - FILIP: Fine-grained Interactive Language-Image Pre-TrainingCAP6412 Advanced Computer Vision - Spring 2024Notes