Why one image gen model is never enough


Hello Reader

Everyone has a favorite image model.

Ask a creative director or marketer which one they use and they'll tell you Midjourney or Gemini’s Nano Banana. If you’re like me you rely on ChatGPT’s Image 2.

Once you’ve found the image gen model you like, you’re likely to stick with it for everything — concept boards, product shots, social graphics, webinar banners.

Which is the entirely wrong approach.

AI design expert (and my sister!) Toyah Perry does it differently.

She runs Design Method, a creative AI studio in Melbourne, and builds client campaigns across a range of models instead of one.

One handles concepts, another editing and a third product placement. She has a different model if she's looking for speed and another if she needs to move elements around the canvas.

People always ask what platform they should use if they can only use one. It turns out that if you want truly effective image gen outputs, you'll need to use them all, but for completely different reasons.

Pick the Right Model for the Job

There are countless image gen models available but these five are Toyah’s favorites. She shared why she likes each and which job she assigns them depending on the project.

1. Midjourney: concept and direction

When a brief calls for something unexpected, a distinctive look, an unusual composition, this is where Toyah starts.

The latest model, which was released in June is showing a big improvement in outputs. Standard jobs render four to five times faster than the previous version, HD now generates at 2K by default, and the model holds on to small prompt details it used to drop.

Toyah relies on Midjourney for artistic expression and experimentation. It’s great for creating mood boards and campaign concepts, but it’s the weaker choice for editing an existing asset.

2. Nano Banana Pro: editing and consistency

Once your artistic direction is locked, Google's Nano Banana is your best choice. Hand it a product shot and a brand reference and it keeps the product, the model's face and the color story intact across every variation, up to five character references and fourteen reference images in a single generation.

Toyah recently built a full set of client assets including gradient backgrounds, hero images, product placements and marketing assets.

She says it’s much harder to spot if a Nano Banana output is AI generated. The image outputs appear more natural (compared to ChatGPT), which matters for agencies whose clients expect creative that doesn't look AI-generated.

3. ChatGPT (GPT Image 2): speed

A speaker graphic for tomorrow's webinar, a newsletter illustration, an internal slide, this is the tool for anything with a short shelf life where an hour of refinement won't change the outcome.

The current model renders dense text at roughly 99% character accuracy across a range of languages and you can spin up edits on each iteration in less than a minute.

I use Image 2 almost exclusively and I’ve become increasingly impressed with the accuracy of the layouts, its headshot rendering and color choices. It’s great at creating infographics, hero banners and images I can use to support my slide materials.

I needed images for my recent rebrand launch and instead of creating them by hand in Canva I asked ChatGPT to create these for me. It required a little finessing but what would have taken a lot of fiddling in Canva was done in less than five minutes.

4. Reve: in image editing

Most image models regenerate the whole picture when you ask for a change. Reve builds the scene as a layout first, with each object placed and described on its own, so you can move, resize or swap a single element without touching the rest of the image.

If your images require inline editing, this is your model.

5. Recraft: editorial and fashion

Recraft describes its own output as art-directed, atmospheric, and stylistically alive.

“The outputs are really evocative," says Toyah.

It’s also great for native SVG generation for logos and icons that open cleanly in Illustrator or Figma.

AI still needs creative direction

Regardless of how good your image gen model is, getting the right output is going to take some effort.

“There’s no easy button here,” says Toyah.

A recent client brief she tackled is a great example.

Her task was to take a real photo of the brand's model, and place her in a series of elaborate scenes that matched the brand’s aesthetic.

The model's features and body had to stay exactly the same (a tricky proposition with image gen models that make slight variations to the inputs you give it).

“On my first pass, Nano Banana, shifted the model’s body proportions by a fraction and the art director caught it immediately,” says Toyah.

Toyah sourced reference images, tested several versions and reworked the composition until the setting looked natural and the model’s position and proportions matched. Then came the detail edits, minor fixes to each element in Photoshop. All in, the one image took six hours from start to finish.

When I told her that seemed like an awfully long time, she pointed out that compared to a full photoshoot which would include sourcing a location, hiring a photographer, a lighting person, a hairstylist and makeup artist, it was relatively quick.

"Plus it would cost thousands of dollars," she says.

These image gen models compress the production time that comes after the creative direction is set, but the direction itself still requires time and effort to get right.

What makes these models produce better output

The model matters less than what you feed it. Building a detailed prompt takes time, but it’s the key to getting the output you want..

To streamline this part of the process Toyah built two Claude skills that work in concert for image generation.

The first is her pages long brand skill, which includes everything about Design Method's visual identity: her exact color palette, typefaces, the layouts she favors, lighting and mood. It also covers her tone of voice backed by real examples of her own copy.

The second skill is her prompt writer.

Using the brand skill as a foundation, it takes whatever Toyah gives it (even something as loose as "gradient, orange and lavender, blobby shapes”) and expands it into a full prompt covering the subject, scene, photography style, lens, lighting and mood, and then produces a finely detailed prompt.

Here’s an example of a prompt from Toyah’s prompt creator and the final output from Nano Banana:

The prompt

The output

This isn’t the kind of image you’re going to get by asking for “a birds eye view of a white sneaker in a basket," but if you have your brand skill already loaded, you can spin up these images in minutes.

Your prompt strategy in six moves

Building a skill in Claude is a great way to speed up your process but it’s still helpful to understand the best practices for effective prompting.

Here are Toyah’s six key moves:

  1. Be specific about the subject. A generic prompt produces a generic result. "A dog in a living room" gives you the average of everything the model has seen. "A chocolate French bulldog in a plush round bed, chewing a pink duck toy" gives you something nobody else is generating.
  2. Name the style and mood. Tell the model what visual world the image lives in. Cinematic photography with dramatic lighting produces something very different from a flat product shot.
  3. Direct the lighting. Soft natural light, dramatic side light and a studio spotlight all produce different levels of contrast, realism and emotion.
  4. Don't try to perfect the first prompt. If the image is close but not exactly what you want, use the edit function rather than rewriting from scratch. You’ll speed up the process and burn fewer tokens.
  5. When quality starts to drift, open a new session. Every additional edit inside one conversation degrades the image a little. Download the version you like, start a new chat and upload it as the starting point.
  6. Put exact text in quotation marks. Any word you want rendered in the image, a label, a sign, a headline, goes in quotes in your prompt.

Stitching it all together

If you’re producing images in bulk it might be time to move on from single image gen platforms to node based work.

Figma Weave, Flora and Kive all allow you to build a single workflow and pull in your image gen models of choice. You can branch, remix and compare using the same canvas. These platforms also allow you to work across your team, so you can manage feedback and version control.

This newsletter is already too long to cover these platforms in detail but if you're interested in learning more, hit reply and let me know, and I'll schedule it for a future edition.

If you want to see this system built in real time, Toyah and I are walking through the whole thing, concept to finished asset, in our Creative with Claude Masterclass on July 23rd. We'd love to see you there.

More info 👇

Copyright and AI Generated Images

If you’re generating images for campaigns or clients it’s a good idea to stay up to date on copyright law.

In the US, an image generated entirely from a text prompt doesn't qualify for copyright.

Copyright law requires a human author, and the U.S. Copyright Office doesn't consider a prompt alone to meet that bar. That means if you create an image without editing it, it belongs in the public domain where anyone can use it, even if you paid for the generation.

You can avoid this issue if you apply editing to the image.

Collaging multiple generations together, painting over sections or editing heavily in Photoshop or Canva means the result can qualify for copyright.

The same logic applies to compilations: arranging a set of AI images into a layout, a comic book or a full campaign can earn copyright on the arrangement itself even when the individual images don't qualify on their own.

You also need to be aware that an image gen model can produce an output that closely resembles an existing copyrighted character or artwork, and using or selling that image can still expose you to an infringement claim, regardless of whether you hold copyright on it yourself.

The training data behind these models is also still working its way through the courts, with several lawsuits over fair use and unauthorized use of copyrighted material yet to be resolved.

None of this is legal advice, and the rules here are still being tested case by case. If copyright ownership matters for a specific client asset, project or campaign, then it’s best to consult your legal team before you send it out into the world.

The details: links and pricing

Midjourney: midjourney.com. Accessed via Discord or the web app. Plans run from Basic at $10/month (3.3 hours of fast generation) to Mega at $120/month (60 hours, private generation mode). Annual billing saves 20%. Agencies over $1M in annual revenue need the Pro tier or above per Midjourney's terms.

Nano Banana Pro: gemini.google.com. Free with a limited daily quota inside the Gemini app, then you're bumped to the standard (non Pro) model. Google AI Plus runs $4.99/month for a higher quota, Google AI Pro is $19.99/month. Volume users can go pay per use through Google AI Studio.

ChatGPT (GPT Image 2): chatgpt.com. Free and Go ($8/month) get Instant mode only. Plus ($20/month) turns on Thinking mode, the layout planning and multi-turn editing that makes the text accuracy numbers above possible. Business plans start around $20 to $25 per seat per month and exclude your data from training by default.

Reve: app.reve.com. Free tier with a daily energy allowance that resets each day. Lite runs $7.99/month, Pro is $19.99/month with 100 times the free energy plus 250 video credits a month.

Recraft: recraft.ai. Free tier available for testing but images are public and carry no commercial license. Basic starts at $12/month for 1,000 credits with full commercial rights. Pro and Teams plans scale by credit volume from there.

Your Image Gen Reference Guide

Here’s a quick guide you can use to decide which image gen model is the right one for the task.

Use Midjourney when you need:

  • Creative concepts
  • Mood boards
  • Campaign exploration
  • High-end visual direction

Use Nano Banana Pro when you need:

  • Image editing
  • Product consistency
  • Campaign variations
  • Brand asset refinement

Use ChatGPT (GPT Image 2) when you need:

  • Presentation graphics
  • Newsletter illustrations
  • Social graphics
  • Fast B2B marketing assets

Use Reve when you need:

  • Midjourney-style concepts with more control
  • Element-level edits without a full regenerate
  • A layout you can adjust piece by piece

Use Recraft when you need:

  • Fashion-forward, art-directed stills
  • Native vector logos and icon sets
  • Files that open cleanly in Illustrator or Figma

The Masterclass is next week. Toyah and I are spending 90 minutes building a full campaign system live, from Claude configuration to finished creative assets.

Newsletter subscribers get $75 off with code NEWSLETTER. That code expires this Friday.

If you've been meaning to fix how your team uses Claude, this is the session that gets it done.

Register for the Masterclass

PS. You’ve also got one more chance to sign up for our free lightning lesson this Thursday. Sign up here.

Hope to see you there!

Did some one forward you this email? You can subscribe here.

2120 Contra Costa Blvd #1059 , Pleasant Hill, CA 94523
Unsubscribe · Preferences

AI at Work

AI at Work is a weekly newsletter on how marketing teams redesign workflows, roles, and systems with AI. Real examples, practical frameworks, and repeatable processes operators can use immediately. Join thousands of successful marketing leaders by subscribing below!

Read more from AI at Work
AI Agents

Hello Reader You’ve been meaning to set up AI correctly. Then a campaign needs approving, a client asks for changes and three meetings appear on your calendar. The setup gets pushed to next week while you continue dragging work through the same collection of chats, files and tools. You may already have several useful pieces. A Project containing your brand documents or a prompt that produces a decent brief. Maybe you even have an agent with a name and job title. But every task still begins...

AI Creative Campaigns

Hello Reader, A recent post in r/marketing collected dozens of AI-generated posters from different companies. Same blocky type, painted illustrations and crowded layout, even when the companies had nothing in common. At the start of this week the post had 932 upvotes and 250 comments. People had opinions. Some defended the posters as an effective way for small organizations to make something presentable. Others said they’d started scrolling past the style because every poster looked the same....

AI Agents and Pipeline

Hello Reader “Sorry we couldn’t find time to chat.” That sounds like a reasonable email to send a prospect who failed to schedule a meeting. Except this prospect had never received a meeting link. The email came from Zapier and landed in the inbox of a qualified enterprise buyer. The company’s automation blamed the buyer for failing to complete an action they’d never been given the chance to take. Angela Ferrante, who led enterprise marketing at Zapier (and is now head of marketing at...