Every Nano Banana prompt guide on the internet points back to the same place, which is Google’s own documentation. That documentation has a problem: Google publishes the prompt formula in three separate places, with the same five elements listed in three different orders. The DeepMind guide leads with style. The Google Cloud guide leads with subject. The Gemini blog post leads with subject too, but then adds a sixth element the other two leave out.

That inconsistency is not a footnote. The Cloud guide also tells you to list details “from most important to least important, as earlier details have more influence on the final result.” So the order is doing real work, and three official orders cannot all be the best one. This guide covers the five elements that stay stable across all three sources, how to decide your own order, and the four specific mistakes that ruin more Nano Banana prompts than anything else. If you are still working out where the model lives, start with how to use Nano Banana first.

The Key Takeaways

  • Five elements, three official orders: subject, action, location, composition and style are stable across every Google source. The sequence is not.
  • Lead with whatever is non-negotiable: earlier details carry more weight, so the first thing you write should be the thing you refuse to compromise on.
  • Say what you want, never what you do not: Google’s own best practice is positive framing. A prompt full of prohibitions leaves the model to invent everything you failed to specify.
  • Quote your text and name the font: the official example is a headline “URBAN EXPLORER” rendered in bold, white, sans-serif, not a request for a poster with a title on it.
  • “Attach 14 images to Pro” is wrong: Nano Banana Pro takes 6 object images plus 5 character images. Fourteen is the cap on the cheapest model, and separately an output figure.

What Is a Nano Banana Prompt?

De l'éditeur

Tous les modèles d'IA dans une seule app

Fello AI réunit GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 et plus dans une seule app native pour Mac et iPhone.

Téléchargez maintenant !

A Nano Banana prompt is the plain language description you give Google’s Gemini image models to generate or edit a picture. It is not code and it has no required syntax. What separates a good one from a bad one is coverage: a strong prompt specifies the subject, what it is doing, where it is, how the shot is framed, and what it looks like. A weak prompt names a subject and leaves the model to guess the other four, which it does by defaulting to the blandest possible answer.

That last part explains most disappointment with these models. When people say the output looks generic, what usually happened is that they wrote one element and the model filled in four. The fix is rarely a longer prompt. It is a more complete one.

The Five Element Nano Banana Prompt Formula

Here is what each element controls, and what the difference looks like in practice.

ElementWhat it controlsWeakStrong
SubjectWho or what the image is ofa womana woman in her sixties with cropped grey hair and reading glasses
ActionWhat is happening, and the implied storystandinghalfway through pulling a tray of bread out of the oven
LocationThe setting and everything it impliesin a kitchenin a narrow tiled bakery kitchen, flour dust in the air
CompositionFraming, angle, depth, aspect ratioa photowaist up, slightly low angle, shallow depth of field
StyleMedium, lighting, colour, eranice lookingdocumentary photography, hard overhead fluorescent light, muted colour

Google’s ultimate prompting guide for Nano Banana writes this as subject, then action, then location, then composition, then style. The DeepMind image prompt guide opens with style instead and calls location “setting”. The Gemini prompting tips post from November 2025 lists six, adding editing instructions as a separate slot.

So Which Order Do You Use?

Yours. The elements are the stable part and the order is the variable, so treat the sequence as a decision rather than a rule you are failing to remember correctly.

The practical version is this: lead with the element you refuse to compromise on. If you need a specific person to look like themselves, subject goes first. If you are producing a set of images that all have to match visually, style goes first, which is likely why the DeepMind guide is written that way. If you are illustrating a moment, action goes first. Everything else follows in descending order of how much you care.

Why Your Nano Banana Prompt Failed

These four account for most bad results, and all four are directly addressed by Google’s own best practices. They are also the part that no prompt library will teach you, because a library gives you prompts that already work rather than showing you why yours did not.

Mistake 1: You Stacked Competing Styles

The most common bad prompt is not too short. It is a pile of quality words that contradict each other.

Weak: “a beautiful cinematic photorealistic oil painting of a woman in a city, anime style, 8k, hyperrealistic, dramatic lighting, trending”

Strong: “an oil painting of a woman waiting at a tram stop in a rainy city, visible brush texture, muted greens and greys, three quarter view”

An oil painting is not photorealistic, and neither of those is anime. The model cannot satisfy all three, so it picks, and what it picks changes run to run.

We ran the weak prompt twice. One run abandoned the oil painting entirely and returned a glossy anime illustration of a neon street. The other kept the brushwork, the rain and the tram, then sat an anime face on top of it, and threw in a destination board reading “LINE 4” that nothing in the prompt asked for. Same prompt, same model, two unrelated results, neither of them what anyone would have asked for on purpose.

Weak prompt result, a glossy anime illustration of a woman in a dark coat on a wet neon lit city street at dusk, showing none of the oil painting texture the prompt asked for.
Weak prompt, first run. Anime won outright and the oil painting vanished.
Second run of the same weak prompt, an oil painted rainy tram stop with visible brushwork but the woman's face rendered in anime proportions, plus an invented destination board reading LINE 4.
Weak prompt, second run. Same words, an entirely different failure.
Strong prompt result, an oil painting of a woman under an umbrella at a rainy tram stop, visible brush texture and muted greens and greys, a green tram approaching over wet cobbles.
Strong prompt. One style, asked for once, and repeatable across runs.

That is the real cost of stacking styles. You are not blending them, you are running a lottery, and a prompt you cannot repeat is worth very little. The focused version returned the same kind of picture every time. Words like 8k and trending contributed nothing to any of the runs.

Mistake 2: You Left the Lighting Blank

Google’s guide has an entire creative director section built around designing your lighting, choosing your camera and lens, and defining the colour grading. Most people write none of it, and lighting is the single element that does the most work.

Weak: “a man drinking coffee in a kitchen”

Strong: “a man drinking coffee at a kitchen counter, low winter sun coming through the window behind him, long shadows across the worktop, warm highlights and cool shadows”

Same subject, same action, same location. The second one has a time of day, a light source with a direction, and a colour relationship. That is the whole difference between a stock photo and a photograph.

Weak prompt result, a man drinking coffee in a bright kitchen under flat even daylight with no directional light and no visible shadows.
Weak prompt. Flat, even, sourceless light.
Strong prompt result, the same kitchen with low sun through the window, long shadows thrown across the cabinets and floor, warm highlights against cooler shadows.
Strong prompt. One named light source, and the shadows it throws.

Mistake 3: You Described What You Did Not Want

Google states the rule plainly in its best practices: use positive framing, and describe what you want rather than what you do not. This is the mistake people carry over from older image generators that had a separate negative prompt field.

Weak: “a product shot of a ceramic mug, no text, no watermark, not cartoonish, no people”

Strong: “a product shot of a plain white ceramic mug on a light oak surface, soft diffused studio light from the left, clean seamless background”

We ran both. The interesting part is that the negatives were obeyed: no text, no watermark, no people. What came back instead was a styled lifestyle scene, a speckled stoneware mug on weathered wood surrounded by ferns, folded linen and a teapot. Perfectly nice, and useless as a product shot.

That is how negative instructions actually fail. They do not summon the thing you forbade. They spend your whole prompt on what the image must not contain and leave the model to invent everything it must, so it fills the frame with props. The fix is not a firmer prohibition, it is naming the surface, the light and the background so there is no space left to fill.

Weak prompt result, a speckled stoneware mug on weathered wood surrounded by ferns, folded linen and a teapot, an attractive lifestyle scene rather than the clean product shot requested.
Weak prompt. Every prohibition obeyed, and still not a product shot.
Strong prompt result, a plain white ceramic mug on a light oak surface against a clean seamless background, lit softly from the left.
Strong prompt. The surface and background named, so nothing else crept in.

Mistake 4: You Asked for Text Without Quoting It

Legible text is the thing these models got good at, and it is also where vague prompts fail most visibly.

Weak: “a travel poster for Lisbon with the city name at the top”

Strong: “a travel poster of Lisbon rooftops at sunset, the headline ‘LISBON’ rendered in bold, white, sans-serif type across the top third”

Google’s own example follows exactly this shape, quoting the words and then naming weight, colour and typeface. Quote the string, specify the type, and say where on the canvas it goes. Anything you leave out is a coin flip.

Testing this produced something we have not seen written down anywhere, so treat it as the practical rule underneath the official one. Nano Banana adds text you did not ask for. The vague poster prompt came back carrying an invented strapline, “DISCOVER THE HEART OF PORTUGAL, SUNSHINE, FADO, HISTORY, TRAMS”, in a typeface nobody chose. A product shot of a water bottle, with no mention of text at all, arrived with “HYDRATE // 32oz” printed down the side. Even the careful Pro version of the poster added “PORTUGAL” and a line reading “THE CITY OF GOLDEN LIGHT”.

So the model has a strong prior that posters and packaging carry words, and it will fill that slot whether you asked or not. Specifying your text is not only how you get the words you want. It is how you stop the words you did not want.

Weak prompt result, a busy vintage Lisbon travel poster with an ornate serif title and an invented strapline reading Discover the Heart of Portugal, Sunshine, Fado, History, Trams.
Weak prompt. A strapline nobody asked for, in a typeface nobody chose.
Strong prompt result, a Lisbon travel poster carrying the single headline LISBON in bold white sans-serif across the top third above terracotta rooftops at sunset.
Strong prompt. The quoted string, the named typeface, the stated position.

Six Worked Nano Banana Prompts

Each of these covers all five elements, and each one is annotated so you can see the formula doing the work rather than take it on faith. Copy one, then swap the subject for yours and leave the structure alone. Portraits are deliberately missing here because our 14 copy and paste portrait prompts already cover them in depth.

1. A Poster With Legible Text

“A travel poster of Lisbon rooftops at golden hour, terracotta tiles descending toward the river, viewed from a high balcony, mid century screen print style with flat colour separations and visible registration, the headline ‘LISBON’ in bold cream sans-serif across the top third.”

Subject: Lisbon rooftops. Action: descending toward the river. Location: from a high balcony. Composition: top third reserved for type. Style: mid century screen print. Run this one on Pro, since text is where Pro separates itself.

Mid century screen print poster of Lisbon rooftops at golden hour seen from a wrought iron balcony, headline LISBON with the words Portugal and The City of Golden Light added by the model.
Run on Nano Banana Pro. Note it still added Portugal and a strapline of its own.

2. A Product Shot

“A matte black insulated water bottle standing upright on a wet slate surface, condensation beading down the side, shot square on at product height, soft diffused studio light from the left with one subtle rim highlight on the right edge, clean seamless mid grey background.”

Subject: the bottle, described physically. Action: condensation beading. Location: wet slate. Composition: square on at product height. Style: the two light sources. Notice there is no instruction to avoid clutter, only a description of the surface that leaves no room for any.

A matte black insulated water bottle on wet slate, condensation beading down the side, soft light from the left with a rim highlight on the right edge, and HYDRATE 32oz printed on it by the model.
No text was requested. HYDRATE // 32oz arrived anyway.

3. An Implied Story

“A golden retriever mid shake after a swim, water flying off in a ring around him, shallow lake edge behind, low evening sun catching the spray, shot side on at dog height.”

Action leads here instead of subject, because the moment is the point. Note how short it is. Thirty words with a real moment in them beat sixty words of set dressing, and “mid shake” is doing more work than every adjective you could stack around a dog standing still.

A golden retriever mid shake at a lake edge at sunset, water flying off in a backlit ring around him, shot side on at dog height.
Thirty two words, one real moment.

4. A Diagram That Has to Be Right

“A clean instructional diagram showing the water cycle, four labelled stages arranged clockwise, arrows indicating direction of flow, labels reading ‘EVAPORATION’, ‘CONDENSATION’, ‘PRECIPITATION’ and ‘COLLECTION’ in a plain bold sans-serif, flat vector illustration on a pale background, muted blues and greens.”

Every label is quoted and the arrangement is specified. Diagrams are the use case where vague framing fails hardest, because a diagram that is nearly right is worse than no diagram.

A flat vector water cycle diagram on a pale background, four stages arranged clockwise and labelled Evaporation, Condensation, Precipitation and Collection, joined by curved arrows.
Four quoted labels, four correct spellings, arranged as asked.

5. A Set That Has to Match

“Three square icons in one consistent style for a recipe app, a chef’s knife, a mixing bowl and a stovetop kettle, each centred on its own tile, thick uniform line weight, two colour palette of deep navy on cream, flat with no gradients or shadows.”

Style goes first because consistency across the set outranks any single icon. This is the case the DeepMind ordering is built for, and it is worth generating as one prompt rather than three so the model holds the treatment steady.

Three cream square tiles on a deep navy wall carrying matching navy line icons of a chef's knife, a mixing bowl and a stovetop kettle in uniform line weight.
One prompt, three icons, one treatment held steady across all of them.

6. Editing a Photo You Already Have

“Relight this photo as late afternoon, sun low and behind the subject on the left, long soft shadows falling toward the camera, warm highlights against cooler shadows. Keep the framing, the pose and the background exactly as they are.”

Editing prompts start with a strong verb, which tells the model the primary operation before anything else. The second sentence matters as much as the first: naming what stays fixed is how you stop an edit turning into a fresh generation.

Before the edit, a white curly coated dog lying on grass in a garden of roses and lavender under flat daylight.
Before.
After the edit, the same dog in the same pose and framing relit as late afternoon, with warm low sun and longer shadows across the grass.
After. Same pose, same framing, same background, different hour.

Which Model Runs Your Nano Banana Prompt

There are four models in the Nano Banana family, and the prompt you write should account for which one is answering. This is also where the most widely repeated claim about Nano Banana turns out to be wrong.

ModelOfficial nameReference images it accepts
Nano Banana 2 LiteGemini 3.1 Flash Lite Imageup to 14 images, all objects
Nano Banana 2Gemini 3.1 Flash Image10 object, 4 character, 3 style
Nano Banana ProGemini 3 Pro Image6 object, 5 character
OriginalGemini 2.5 Flash Imagelegacy, superseded

Read that middle column again, because it inverts the usual advice. Nano Banana Pro accepts the fewest object references of the current three. The cheapest model in the family takes the most. Pro earns its name on reasoning, text rendering and resolution, not on how much you can hand it at once. A prompt built around eight product photos belongs on a Flash model.

The figure you have probably read elsewhere is that you can attach fourteen reference images to Nano Banana Pro. You cannot. Pro takes six object images and five character images, and Google’s API documentation states no combined total beyond that.

Fourteen is a real number, which is why the mistake spreads. It is the input cap for Nano Banana 2 Lite, and separately it is how many objects DeepMind says the models keep faithful in the picture they produce, alongside five characters. So fourteen is either a limit on the cheapest model or a fact about the output, depending on which page you read, and it is never a licence to hand Pro fourteen photos. Our breakdown of what changed in Nano Banana 2 keeps the specifications apart.

Keeping a Character Consistent

Character drift is the complaint that comes up more than any other, and the prompt-side fix is unglamorous: name your characters. When you attach references, define the role of each one rather than assuming the model will infer it, and give every person a distinct name you then reuse in the instruction. “Put Maya and Tomas at the same table” gives the model two anchors. “Put them at the same table” gives it none.

The second half is to change one thing at a time. These models edit conversationally. A follow up message that adjusts the light while keeping everything else fixed will hold a face together. A fresh prompt that redescribes the whole scene quietly rerolls it. If you want fully worked examples rather than principles, our copy and paste Gemini photo prompts cover portraits in detail.

Writing Nano Banana Prompts on a Mac

Prompting well is iterative, and iteration is where the browser hurts. A good prompt usually arrives on the fourth attempt, after you have changed the light, then the framing, then the style, keeping notes on which change did what. Doing that across browser tabs, with a separate tab for the model you use to write the prompt in the first place, is most of the friction.

Running Nano Banana next to the text models in one desktop app removes that. You draft the prompt, generate, compare the result against another image model on the same brief, and keep the whole thread in one window. Fello AI puts Gemini, ChatGPT, Claude and Grok behind a single interface on macOS, which is the setup this kind of iteration actually wants. There is a per generation cost to be aware of either way, and our Nano Banana pricing breakdown covers what each tier includes. Our Nano Banana desktop client guide for macOS walks through that setup step by step.

The Bottom Line

The five elements are settled and the order is not, which is why chasing a single canonical Nano Banana prompt template is a waste of time. Cover subject, action, location, composition and style, lead with the one you cannot compromise on, and you are already past most of what a prompt library will give you.

Then check the four mistakes before you blame the model. Competing styles, missing light, negative instructions and unquoted text explain most bad generations. All four are cheaper to fix than to work around. If none of that helps, the problem is usually the model rather than the prompt, and the answer is a different one in the family rather than a fifth rewrite.

FAQ

What is the Nano Banana prompt formula?

Subject, action, location, composition and style. Those five elements appear in every official Google source, though the three sources list them in different orders. Cover all five and lead with whichever one matters most for your image.

Does the order of a Nano Banana prompt matter?

Yes. Google’s Cloud guide says earlier details have more influence on the result, so put your non-negotiable element first. That is also why the three official orderings cannot all be optimal, and why you should choose the order per image rather than memorise one.

How many reference images can I attach?

It depends on the model. Nano Banana Pro takes 6 object images plus 5 character images, Nano Banana 2 takes 10 object plus 4 character plus 3 style, and Nano Banana 2 Lite takes up to 14 objects. The commonly quoted figure of fourteen reference images confuses this with an output specification.

Why is the text in my image garbled?

Usually because the words were described rather than quoted. Put the exact string in quotation marks, name the weight and typeface, and say where it sits in the frame. Nano Banana Pro is the model to reach for when the image has to carry text.

Why do I keep getting generic looking images?

Almost always missing elements rather than a short prompt. If you name a subject and leave action, location, composition and style unspecified, the model picks the safest option for all four. Adding a light source and a camera angle fixes more images than adding adjectives.