Google’s new AI video model, Gemini Omni, is being billed as “Nano Banana - but for video”.
First released last August, Google Nano Banana represented a step change in being able to make fine-grained modifications to AI-generated images using natural language without mucking up the elements of the image you didn’t want to change.
Subsequent Nano Banana updates improved the model’s handling of text, the number of reference inputs it could handle and its ability to draw on Gemini’s ‘world knowledge’ and ‘reasoning’.
Gemini Omni promises to match those capabilities for video, or - in Google's words - to let you “create anything from any input and edit naturally using conversational language”.
There have been some impressive videos posted on X but those posts rarely tell you how many attempts (and therefore credits) it took, or show you the comedy outtakes.
I decided to pit Gemini Omni against Aleph 2.0 - an update to Runway’s flagship video editing model, released two days after Omni with the strapline ‘Get the video you need. From the video you have.’
I also decided to benchmark Omni against Seedance 2.0, which isn’t positioned as a dedicated editing model, but is capable of edits and is currently topping most AI video leaderboards.
I only allowed the models one shot at each challenge and didn’t include any additional visual references - this was purely a test of natural language editing.
Challenge 1
Editing uploaded videos via Gemini Omni isn’t currently supported in the UK, so I had to fly to the USA use a VPN for this first challenge.
I uploaded 10 seconds (Omni’s max duration) of the most boring video I could find on my camera roll (a video I’d taken of the back of our old car to determine which of the brake lights wasn’t working) to all three models with the prompt: “as the driver gets into the car, the boot springs open by itself and out springs a lion”.
Gemini Omni followed the prompt and handled the physics of the boot opening well, although it made the driver door visible (and therefore in slightly the wrong place). The lion could be more realistic and the colours lost some of their vibrancy, although the model’s handling of light and shadow is impressive.
Runway Aleph 2.0 had a bit of a nightmare, depicting an open passenger door and a resolutely shut boot, forcing the very-obviously-AI-generated lion to enter from right of shot and sit behind the car.
Seedance 2.0 followed the prompt to the letter and gave us the most convincing lion. It did make a couple of small changes to my movements, but kept the vibrancy of the colours. The car dipping slightly as the lion gets out is a nice touch, as is the lion’s growl.
Winner: Seedance 2.0
Challenge 2
For the second challenge, I gave all three models a short video I’d generated (using Seedance 2.0) from a photo of my desk. It shows a small furry red creature (I’d prompted ‘monster’, but it looks more like a bird to me) emerging from my desktop wallpaper, jumping out of my monitor and onto my iPad keyboard before diving headfirst into a mug of coffee.
I uploaded the video to all three models with the prompt: “make the monster blue, the mouse run away from him and the mug bottomless” (maybe a bit mean when the mug in the original video is full of coffee).
Gemini Omni struggled with this one, turning my character into the cookie monster (#trainingdatabias), putting a hole in the side, rather than the bottom, of my mug and altering the action in various unrequested ways. The mouse did at least run away.
Runway Aleph 2.0 managed to turn the monster blue, but the mouse stayed put and the lamp started inexplicably dispensing coffee into a different (but definitely not bottomless) mug.
Seedance 2.0 turned the monster blue and the mouse very literally (and rather charmingly) ran away. It didn’t freestyle any edits. My only note would be that the coffee splash indicates the mug still has a bottom.
Winner: Seedance 2.0 (by a mouse’s whisker)
Challenge 3
For my third and final challenge, I gave the models a video of the lifeless puppet the kids had left on the bottom of the banisters this morning with the brief “the bird puppet reads today’s weather forecast for Bath direct to camera”.
Gemini Omni did a pretty bang-up job on this one, maintaining the camera motion and rotating the puppet to maintain a direct-to-camera address. The voice sounds like it was recorded in situ and the lip syncing is decent. There have also been some sunny spells today (along with some showers).
Runway Aleph 2.0 refused this edit for some reason, claiming “the aspect ratio is not supported”, despite it being a standard 16:9 video. Probably something to do with the frame rate or resolution 🤷♂️.
I forgot to change the duration from the previous challenge for Seedance 2.0, so it got an extra 5 seconds in which to freestyle a weather report. It did a pretty decent job although only fulfilled the direct-to-camera element of the brief by ditching the original arc shot.
Winner: Gemini Omni
Conclusions
Whilst none of the edited videos were flawless, editing videos using natural language is, as predicted, becoming increasingly viable. Multimodal models are able to use their language smarts to parse some pretty informal briefs. That said, prompt adherence remains mixed, with ambiguously worded prompts a human would understand often misinterpreted and unrequested changes introduced.
Based on these challenges, Gemini Omni tends to do the best job of preserving camera movement but can wobble off when it comes to prompt adherence. My wider testing suggests it’s also very cautious when it comes to editing videos of people, which will limit its use in production scenarios.
Seedance 2.0 did a creditable job on all three challenges, with only minor divergences from the brief. ByteDance might want to more actively market Seedance 2.0’s natural language editing capabilities.
I didn’t get great results from Runway Aleph 2.0 in any of these challenges, although I suspect that may be my insistence on providing text with no additional references (the Edit Studio on Runway’s website forces you to generate image references to guide Aleph’s edit). I’ve had more success with Aleph on straightforward object removal / background change scenarios, which Runway provides dedicated ‘apps’ for.
Of course, as always with AI, this is only going to get better from here…
Related posts:

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.