Trois Sonneries
I entered a 48-hour video contest run by Nous Research and Black Forest Labs. As part of the collaboration, Hermes users got exclusive free access to the image-and-video model FLUX 3 Preview. Users could compete for rewards by uploading their work on Twitter. The video model’s standout feature was that it generated video and audio together and derived sounds from physical events.
I joined the contest on the second day. The deadline was at 7 PM PT on August 1st, 2026. I submitted my project just before the deadline, which was at dawn on August 2nd in Europe.
This was my first time working with generative video, as opposed to video from code. My entry, Trois Sonneries, was a minute-and-a-half-long trailer for a French policier that didn’t exist.
Claude Opus 5 wrote the script and the prompts. GLM-5.2 in a Hermes Agent harness actually prompted the video model. Opus 5 and I refined the short starting from an initial concept by Kimi K3.
The Opus 5 session I stuck with started by offering a strategic read of the contest: the judges would pick those clips that were good marketing material by demonstrating what FLUX 3 could do and previous models couldn’t. I let AI produce the video concept. I borrowed Gwern Branwen’s “brainstorming” technique again: generate a number of items, critique each, revise based on the critique, rate the final version of the item 1 to 5 stars, and pick the best at the end. Each concept required a title, visual inspirations, an art style with a unique twist (like Gwern’s comics), and a plot with a shot list. I wanted the model to name genres and movements as inspirations, not particular films, because in a test FLUX had refused to generate a legally distinct reptilian kaiju.
Soon I had a catalog of 120 concepts: 80 by Claude Opus 5 instances, 20 by GLM-5.2, and 20 by Kimi K3. Out of those, 30 had been picked as the best by their respective models.
I considered that other contestants would be prompting AI models for concepts and wanted to avoid collisions. In order to avoid them and learn more about the models’ attractors, I asked Opus 5 to analyze the recurring themes.
Opus 5 found that four of the six sets (Opus, Kimi) included a film about a foley artist. Five sessions conjured an interpreter booth (Opus, GLM); two Opus instances titled the short Booth Four. Two Opus sessions also titled their concept The Quiet Car.
AI allegory?
Foley, like teletext below, is a potential AI allegory. The case is unclear when the writing model initially searched the Internet for information about the contest and found hype for the video model’s ability to create sound synchronized with visuals.Other convergences:
- A silent film with modern sound: 4/6 (Opus, Kimi)
- A nature documentary about mundane objects: 3/6 (Opus, Kimi)
- A forge, bell foundry, or letterpress: 4/6 (Opus, GLM)
- A degrading weather broadcast: 5/6 (Opus, GLM, Kimi)
- Instrument tuning (Opus, GLM)
- A robot arm learning to grip an egg: 2/6 (Opus)
Every model generated a deliberate audio-video desync and rejected it for the same reason: a judge would read it as a mistake.
Opus 5 thought the deepest mode-collapse finding wasn’t the particular themes but the structure. Every concept went, “establish a rule, break the rule once, stop”. None reversed the rule and kept going.
Opus noted Kimi thought in terms of broadcast formats (teletext, home shopping, aerobics tapes) rather than subjects. This made its concepts distinctive. Kimi reasoned about judge psychology unprompted when picking the winning concepts (“Nous and BFL people will have watched forty neon-dreamscape submissions by hour twelve, and a film with genuine format discipline reads as a director, not a prompt.”)
The Opus session that analyzed and selected the concepts deemed GLM “the sentimental one”. Its picks were about departure and loss. Opus concluded GLM was miscalibrated because it never rated a concept under three stars.
Opus called itself and its three sibling sessions “near-clones of each other”. One Opus session surprised me and broke the pattern on its own. It saw a mention of my “usual methods” in the opening prompt and before I could explain them used search to find a review of the seed-word experiment from Unslop that accidentally wasn’t part of my Unslop claude.ai project. This session understood the method and drew random seeds from a system wordlist for half of the concept.
Opus suggested that the easiest concept to realize on a deadline was Kimi’s Teletext Evening. It envisioned a concerning teletext transmission about disappearances with concrete poetry and structural film as stated inspirations. Opus thought that if we did it first, we could guarantee a finished entry without faces or continuity to worry about; then we could move on to Weld (Opus, a forge in the dark lit only by the sparks) or Drolleries (Opus, an animation of illuminated manuscript marginalia—the result of the model searching my chat history). Weld risked being cliché and Drolleries seemed too difficult to make. Kimi’s concept, on the other hand, felt like underdeveloped analog horror. I wondered if we could develop it into something that cut between silent teletext and people and sound and perhaps ruptured the boundary.
Kimi K3’s concept.
- TELETEXT EVENING. Inspirations: teletext-era broadcast; concrete poetry; structural film. Style: pure mosaic-text pages on black, period palette, page-flip chirps over CRT hum; twist: one pixel-block in the page border flickers all film and resolves, on the final page, as a tiny window with a room in it. Plot: an evening’s pages cycle — index, news, weather, TV guide; one recurring headline keeps revising itself toward a local disappearance; the 9pm listing is the viewer’s name; final page: “PLEASE TURN OFF YOUR SET.” Shots: S1 index and news; S2 weather map mosaic, headline mutating; S3 TV guide reveal; S4 the warning page, border-window showing a hand reaching for a remote. Critique: zero humans means zero continuity risk and maximum text-adherence flex, but gimmick risk is real. Revision: narrative told through headline revision (concrete-poetry method); sound design carries emotion — the page chirps slow to a heartbeat. Rating: ★★★★★
A brainstorming round yielded a concept called 888 after the UK teletext subtitle page: a 1980s analog television broadcast with teletext closed captions. It was a domestic scene that combined normal and creepy sounds (the kettle boiling, a staircase creaking out of frame). It would show off the model’s ability to generate synchronized visuals, text, and sound. The teletext caption would announce sounds from off-screen.
I asked Opus for a shot script and generated the first video.

The alt text method.
Here and below, the images’ alt text derives from descriptions generated by Gemma 4 31B. Here is the Gemma prompt for standalone descriptions:
Please create a detailed, factual description of this picture in 3 to 8 sentences. Write 20 candidate descriptions, rate each 1 to 5 stars, pick the best, and explain why.
For the comparisons, I used this prompt and edited the text more heavily.
Please describe this image in terms of how it differs from the previous. Write 10 candidate descriptions, rate each 1 to 5 stars, pick the best, and explain why.
FLUX was easy to work with through the Hermes Agent. The characters changed from shot to shot, since (as GLM in the harness explained) the shots were generated independently and concatenated by the agent using FFmpeg. Character appearance needed better specification. I never turned to reference images on this project.
Some sounds came out better than others. It was not clear whether the video model could dub a door opening upstairs or a staircase creaking out of sight. The telephone ringing was, however, perfect. My test audience said the clip was “a bit scary and sad”.
With the lessons from the previous round, I wanted to change the concept while reusing the idea of pre-digital footage and subtitling. It didn’t have to be teletext. The short could be the yellow DVD subtitles of a movie release or something else. We’d drop the horror element entirely to be less sad and would have a different setup with telephones ringing, since that had worked well. I wanted to zero in on a particular genre: a 90-second film trailer. It could be a foreign-language film with, like, an office, a police department, or a spy agency—colorful but not garish, like The Pink Panther and 1960s-to-1970s French cinema.
Opus brainstormed, focusing on the perfect ring sound, and came up with Trois Sonneries (Three Rings), a trailer where every cut was a telephone ring in a police headquarters.

After a shot script and the render, I thought there was definitely something to this concept and setting. “Look at those characters”, I told Opus. They looked authentic.
However, the trailer felt claustrophobic and overpowered by the telephones. The test audience noted two problems. The corridor was repetitive and the collection of phones was “straight out of the uncanny valley.” This (trailer for a nonexistent) film about phones ringing had too many phones ringing.
On the agent side, I asked GLM to take a screenshot of the video every five seconds and create an ImageMagick montage of the screenshots with a timestamp for each. GLM’s container didn’t have ImageMagick, so it wrote a script in Python with Pillow. The script is preserved.
I proposed we reimagine the trailer to include shots outside the office. France, so a cafe? Also, a shot inside a car (visibly faked with rear projection) and a car-chase shot (obligatory). A big factor making our little trailer uncanny was that there wasn’t an actual story for the whole imaginary film. I proposed to Opus that we invent one, lighthearted but not deliberately silly. To convey it, we’d add old-school English narration in scenes where people weren’t speaking French. It’d maximize the “we are watching a trailer for a foreign film” effect as opposed to “we are assumed to understand French”.
To Opus, the test looked better than expected: the period mise-en-scène was convincing and the subtitle rendering was clean. Opus attributed the previous short’s claustrophobia to ring-on-every-cut. It brainstormed a revised concept that preserved the three rings.
Opus 5’s full revised concept.
LE MONTE-EN-L’AIR ★★★★★, keeping the title Trois Sonneries. A cat burglar telephones the Brigade before every job, lets it ring three times, and hangs up. By the time anyone traces the call, a safe is empty and the café has cleared. Marchand has learned to count; he hasn’t yet learned to answer.
It wins on three grounds. It demotes the telephone from subject to device, which is the correction your footage demands. Every shot you asked for follows from the plot rather than being bolted on — the thief calls from cafés, the Brigade drives to them, the drive becomes a chase. And it’s a game between two professionals with no bodies and no menace, which is light without being silly.
The revision that makes it producible: the thief is gloved hands and a white cuff, never a face, not once. That reads as style — the elegant unseen adversary is period-correct — and it happens to mean the character who appears most often is the one character who physically cannot drift.
Ending: Marchand telephones him back. Three rings, then a fourth. Deluc, from the doorway: That's four.

Opus wished me luck with the render, and luck manifested. This was a huge improvement. It looked so much more like a real film. We continued iterating.
Opus included hard and soft QA criteria in the production notes (the document framing the script to pass to the agent). One hard QA criterion was that the thief was fully visible in any shot. This happened at the start. In fact, he was so fully visible that I initially thought it must be scripted misdirection.
The rear-projection scene worked, although with the steering wheel on the wrong side. I only got the steering wheel right later, at the last moment. GLM and I debugged this just before submission time. You couldn’t say the steering wheel was on the left (from the driver’s perspective); you had to use the camera’s right.
The chase scene looked the fakest, like it was shot with a modern camera and old vehicles. It had generated-looking vegetables scattering unconvincingly when hit by a car. A dialing scene looked wrong: an overly repetitive hand motion with too many digits. I suggested specifying the number, but instead cut it per Opus’s advice. The vegetables were cliché; we tried chairs, but the stack of chairs collapsed without contact with the cars. We cut it.
There was an audio problem: one scene cut to another with similar-but-different inconsistent music. It sounded like an editing error. I proposed combining two solutions: (1) to alternate between shots with and without music (for example, music in the first, none in the second, etc.) and (2) to fade out the music before the shot’s end and fade it in during the next.
The thief collecting small jewelry didn’t work. The jewelry reappeared in its box after being obscured. Opus replaced it with a necklace being lifted and put into a bag. This was better but still janky. The bag didn’t expand to accommodate the necklace, and the thief’s hand was clumsy. A retry was worse: the necklace clipped through the mannequin. Opus suggested an elegant lift with the glove, no bag, which worked.
Wall-mounted phones at the cafe didn’t work and were replaced with regular ones. The ringing phones caused visual problems in general. They dialed themselves and vibrated. They didn’t look right until I asked the GLM agent to ground them in real French period telephones.
Version 3 generated twice.

Version 3B kept the script but asked the agent to add two seconds of runtime to every shot with dialogue as a crude fix for dialogue clipping. Opus suggested choreography as a better fix in the next script. The speaker would have a physical action after the line.
Dissatisfied with the imprecision of version 3, I found the FLUX 3 prompting guide) in the Hermes Agent source code. I gave it and the agent-written skill file to Opus. I made a mistake not looking up the guide earlier. Opus seriously updated on these:
A quoted line becomes speech only if a speaker is visible on camera. Without one it tends to render as burned-in text instead.
Opus thought every narration line had been a coin-flip between voice-over and a caption because the narrator wasn’t on-screen. A reasonable precaution, but I didn’t see it actually happen.
Start from the user’s own wording. The prompt is rewritten by a reasoning harness before anything is generated, so restyling it yourself just stacks a second rewrite on top and gives their intent another chance to drift.
More importantly, Opus realized its prompts had been the wrong shape. The emphasis in caps did nothing. The giant prohibition block at the end got rewritten along with everything else. Opus thought the sheer volume of the block gave GLM more surface to reinterpret. Instead, constraints should be stated once, in prose, in the shot they applied to.
For the final version, assembled near the deadline, I’d stopped iterating on the script and retried individual shots. The agent had been giving me unassembled shots on request. I put them together with FFmpeg on my machine.

The results were announced on August 11th. My entry didn’t win. I attribute it to not matching the vibe and also to attempting too much continuity with the characters. The winning entries were all compilations of disconnected or loosely connected clips featuring different characters. This covers choosing what to work on.
There were also notable flaws in execution besides the characters. The opening scene was botched. I couldn’t get the thief to handle the phone in a way that made sense within the allotted time. The plot still felt confusing, cryptic, and off, as if under the simple surface it was an intentionally esoteric riff on the French detective film.