RSS Amplifier

hardeep's life updates, essays and more · Feb 8, 2026

AI video generation Blog #1

0
Sign in to vote or save

Hardeep Gambhir · hardeep's life updates, essays and more

After hosting the AFM festival in Mumbai, I got very interested in how AI video generation actually works. It is a field of interest of the world and technology that continuously interests me. This year I ventured out to write these blogs every month about what I’m learning from the internet about the space that’s constantly emerging and growing.

This is my MIT challenge like how Scott Young did it. These essays are apart from my daily update blogs.

My year of learning AI image generation and video generation started in January 2024 when I was trying to get a job at Perplexity. I loved their design department and fell in love with how the team was approaching its branding. It’s true that branding has nothing to do with how you get your users, but it surely attracts the right talent by showcasing what your company stands for—and that’s critical.

After a tweet I made went viral, I generated a couple of images through Midjourney using the Perplexity SREFs. The stickers we made for The Residency, the previous company I was working at, were extracted using this method and gained significant admiration from the community.

Then I left my obsession with image-gen and design for a bit to understand virality through Instagram reels. I found myself back in the realm of video generation when I stumbled across an AI video-gen hackathon in Paris.

That gave birth to the Bangalore AI Film Hackathon.

And then the Mumbai AI Film Festival, which garnered 25 million views and landed 80% of participants jobs at companies like Netflix India and Eros Now.

Since then, I’ve rabbitholed as much as I can about AI video generation:

  • How the public is using it

  • How technically, the models are working

  • How videos and images are being generated (Shoutout 3Blue1Brown)

  • How the economy and film production studios are reacting to it and adapting by chatting with these studios

  • Who the biggest players in the space are

    • Aggregators

      • ComfyUI

      • Invideo

      • Higgsfield

      • Morphic

      • Krea

      • Leonardo AI

    • Foundational models

      • Luma Labs

      • Sora 2

      • Google Veo 3.1

      • Kling

      • Black Forest Labs

      • Stability AI

At one point, I realized I had collected, understood, and processed much information. So I went on to write this blog.

Here’s something I’ve noticed: there’s a scarcity of people who understand both taste AND understand how generation works. Who are the people responsible for the creative direction for these models? There can’t possibly be too many.

This intersection—between aesthetic judgment and technical knowledge—is where the real magic happens. And it’s surprisingly rare.

One of the most counterintuitive things I learned: adding random noise surprisingly helps generate clearer images than not adding random noise.

The process works like this:

  1. Text → Text vector → Remove noise → Picture

Think of it as working backwards from chaos. The model learns to recognize patterns in pure static, then gradually refines them based on your text prompt encoded as a vector through CLIP (Contrastive Language-Image Pre-training).

The transformer architecture treats images and text fundamentally differently:

  • Pictures are processed as columns

  • Text is processed as rows

This is why cross-attention mechanisms are so powerful—they use both text and image input to put information into the generation process. The model isn’t just following your prompt; it’s understanding the relationship between language and visual concepts.

As one researcher put it: “Our ability to steer the diffusion process is still very limited.”

We can pass in prompts through a CLIP text encoder to generate an encoding vector, but the model’s interpretation of that vector remains somewhat of a black box. This is why:

  • Reference images are critical

  • LoRAs (Low-Rank Adaptations) are essential for style consistency

  • A/B testing across different models is necessary

  1. Write prompts naturally by imagination FIRST, then get AI to refine them—not before

  2. Always use reference images when possible

  3. Do A/B testing across different AI generation models

  4. For specific camera angles → try to have a LoRA that captures that style

LoRAs are essentially fine-tuned weights that teach the model a specific style or concept. They use “trigger words” to activate—specific phrases that tell the model to apply that learned style.

For example, if you need consistent camera angles across your generations, training or finding a LoRA with those specific cinematography styles will give you far more control than prompting alone.

While diving deep into the technical side, I also became fascinated by the business models emerging in this space. Midjourney stands out as the leanest high-growth AI company in the world and a standout example of product-market fit in consumer AI.

Building these models is expensive. Consider just the data collection:

Midjourney’s founder described their training data collection as “a big scrape of the internet.” But scraping is costly:

  • Web scraping services charge around $3.33/hour

  • Scraping just one week’s worth of internet photos (roughly 20 billion) at 10 milliseconds per photo = 55K hours

  • At $3.33/hour, that’s around $185K just to collect one week’s worth of photos

  • This doesn’t include proxy costs to prevent IP blocking or server costs for running the collection process

Once collected, data needs cleansing and the models need training:

  • Stable Diffusion was trained using 256 Nvidia A100 GPUs on AWS

  • Total: 150K GPU hours at a cost of $600K

  • As one industry observer noted: “Without a doubt, there has never been a service before where a regular person is using this much compute”

Midjourney’s customers likely fall into two main categories:

  1. Advertisers seeking rapid iteration on visual concepts

  2. Artists using it for inspiration and creation

Before generative AI, artists relied on Pinterest, Dribbble, or stock photo sites for inspiration. These give you pieces, but only generative AI can help artists combine those pieces during the inspiration phase.

Artist adoption varies—some are suing against AI art, others embracing it—but the value proposition is clear for early adopters.

Runway focuses on professional and enterprise use, while Midjourney remains more targeted toward individuals. Interestingly, some users combine both: using Midjourney for image generation and Runway for video, together creating complete movie trailers.

Shutterstock has its own AI generator, but Midjourney has a unique advantage: it’s much harder to find a specific Midjourney generation compared to a Shutterstock image. However, Midjourney’s lack of off-platform access could be a disadvantage compared to Shutterstock’s web-based generator.

Here’s the profound shift: “To create incredibly lifelike and beautiful images and video, you no longer need a camera. You don’t need to know how to draw or paint or use animation software. All you need is language.”

This democratization is real, but it’s not as simple as just typing anything. The best results come from understanding:

  • How the models interpret language

  • How to structure prompts for maximum control

  • When to use reference images vs. pure text

  • How to leverage LoRAs for consistency

  • Which models excel at which tasks

I’m continuing to learn about:

  • How video-gen and image-gen have advanced since 2023

  • The roadmap for where these technologies are heading

  • New techniques and tools emerging in the space

I have an extremely weird skillset now—somewhere between creative director, technical researcher, and community builder. The intersection of taste and technical understanding in AI generation is still largely unexplored territory.

I want to keep learning more about AI video generation while hosting more of these AI Film festivals across the world. We’re doing one in Los Angeles, Tokyo, Paris, Mumbai, San Francisco this year.

If you’re building in this space or experimenting with AI video/image generation, I’d love to connect. The field is moving fast, and the people who understand both the aesthetic and technical sides will shape where this technology goes next.

No posts

Read the original on hardeepgambhir.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.