I had $300 API credits to burn and nothing else to show for it.
That is only slightly unfair. I also got a folder full of tiny watercolor kitchens, cursed market maps, inconsistent character sheets, and a better sense of which image model I should bother when I need artifacts for this site.
I have been experimenting with image generation models for blog images, visual headers, and the odd little fake artifact that makes a post feel less like a wall of text. The problem is that most model comparisons are too polite. Ask for a cozy cafe or a dramatic astronaut and everyone looks competent. You get six pretty pictures and learn almost nothing.
So I made the prompts irritating.
The benchmark had 10 prompts. Each one was built to catch a different failure mode: exact text, countable objects, maps, negative constraints, character consistency, editable whitespace, low light, and "please keep this looking like a normal phone photo instead of an ad campaign."
This was not a lab benchmark. This was closer to a junk drawer test. If I am going to use these models for my own site, I care less about leaderboard purity and more about what breaks when the prompt asks for something specific.
I ran it through nanoghibli, my local image/video generation repo. The useful bit is that the benchmark is now standardized enough to reuse: prompts live in a prompt library, models go through a catalog, outputs are cached by model, and the leaderboard is generated from the same scoring pass each time. The next time a new image model shows up, I should be able to add it as another model entry and rerun the same prompt set instead of rebuilding the comparison from vibes and screenshots.
The models
The comparison was between the models I would realistically reach for when making visual artifacts for this site. I was not trying to crown the best image model in the abstract, because that is a fake job. I wanted to know which one I should trust with a post header, a fake map, a tiny sign, or a character sketch I might reuse later.
| Model | Why it was here |
|---|---|
| ChatGPT Images 2.0 | The practical default. I use it a lot, and it is usually the one I trust when the image needs to obey text, layout, or editing intent. |
| Nano Banana Pro | The premium Gemini image option in this run (gemini-3-pro-image-preview). This is the one I expected to be most competitive with ChatGPT Images 2.0 on instruction following. |
| Nano Banana Flash | The faster Gemini image model (gemini-3.1-flash-image-preview). I was curious how much quality I would lose when optimizing for iteration speed. |
| Imagen 4 Ultra | The expensive Imagen option. In theory, this should be the strongest Imagen result when quality matters more than cost. |
| Imagen 4 Standard | The middle Imagen option, mostly interesting because middle tiers are often where pricing pages get psychologically annoying. |
| Imagen 4 Fast | The cheap/fast Imagen option. I did not expect it to win, but I wanted to know when "good enough" was actually good enough. |
The setup was deliberately boring: same 10 prompts, same scoring pass, 60 total images. The point of putting this inside nanoghibli was to make the comparison repeatable. When I want to test the next model, I can plug it into the catalog, run the same prompt library, and get a leaderboard plus raw outputs without inventing a new ritual.
| Model | Average score |
|---|---|
| ChatGPT Images 2.0 | 9.6 |
| Nano Banana Pro | 8.6 |
| Nano Banana Flash | 8.3 |
| Imagen 4 Ultra | 8.1 |
| Imagen 4 Standard | 7.3 |
| Imagen 4 Fast | 7.0 |
The lazy read is "ChatGPT Images 2.0 won." The more useful read is that the gap showed up when the image had to obey something measurable: text, layout, counts, routes, consistency. On pure mood, all of these can look good. Once the prompt asks for a tiny sign with exact wording and no invented symbols, the difference becomes less polite.
The comparisons
The previous version of this post had composite grids. They looked nice until you tried to actually read anything. So this version uses expandable, full-width comparison strips. Open a prompt, scroll sideways, and click any image for the raw file.
Kitchen sign: exact text on a tiny object
Text rendering is still the easiest way to embarrass an image model. A pretty kitchen does not help if the sign is corrupted.






Market map: readable labels and a route that makes sense
Models love the idea of a map. They are less excited about the accountability of a map.






Character sheet: one subject, three views, same outfit
This is where "almost the same person" becomes a real problem. A site can reuse a character language only if the model can remember what it just drew.






Phone photo: keep the amateur snap feeling
Sometimes the awkwardness is the evidence. Over-polished output can be worse than a technically messier image.






Low light: warm room without orange mush
This is less about correctness and more about taste. Some outputs make the room warm. Some make it radioactive.






Train platform: countable objects in specific places
A prompt for spatial obedience. The scene can look charming and still fail if the lanterns, cat, or platform details drift.






Poster layout: useful whitespace for later editing
Pretty posters are easy. Posters that leave usable edit space without filling every corner are more useful.






No extra animals: negative constraints
Models love adding bonus details. This prompt checks whether "do not add more animals" survives contact with style.






Rainy bicycle: mood plus object discipline
This is closer to a normal blog-artifact prompt: mood matters, but the bicycle and rain still need to hold together.






Tiny workshop: small tools and readable clutter
A clutter prompt without permission to become visual soup. Good for seeing who can keep small details organized.






What I am taking from this
The grid is still the benchmark, but the grid has to let the images breathe. Tiny contact sheets are good for a gut check and bad for decisions. When text, labels, routes, and character details matter, the comparison needs to be large enough that the failure is visible without zooming like a detective.
For future runs I would keep the rules simple:
- Use the exact same prompt across every model.
- Add new models as catalog entries, not one-off screenshots.
- Include traps: text, counts, spatial relationships, negative constraints, and editability.
- Keep raw outputs, not only the winners.
- Score the images, but also write down the dumb failures.
- Put the outputs side by side before trusting your memory.
- Avoid vibe prompts unless vibe is the actual product.
The annoying answer is that I will probably keep using more than one model.
ChatGPT Images 2.0 won this run for the kind of site artifacts I care about. Nano Banana Pro was strong enough that I would keep it in the mix, especially when I want API-driven iteration. Nano Banana Flash is the one I would use when speed matters and the prompt is not too fussy. Imagen 4 Ultra had some strong moments, but not enough to make it the obvious default. Imagen 4 Fast is useful when price and speed matter more than obedience. Imagen 4 Standard is the awkward one: sometimes good, sometimes not different enough from the cheaper or better option to justify another decision.
Anyway, that is what $300 of credits got me: a spreadsheet, a pile of generated artifacts, and a better sense of which model to bother when I need a fake market map that does not lie to me.