Design · · 6 min read
The sizes AI image models can actually give you, and how to art direct around them
Every model has hard ceilings on size and shape. The real numbers for Nano Banana and gpt-image-2, and what they mean for print, social and billboards.
Two numbers decide whether a generated image can do the job you had in mind: how many pixels it can produce, and what shapes it will produce them in. Both are published, both are stricter than people assume, and neither is what you get by typing "4K, high resolution" into a prompt.
Here are the documented limits for the two most widely used hosted families, read on 20 September 2026, and the art direction decisions that follow from them.
Google's models: a fixed menu, with the pixels printed on it
Google's Nano Banana quickstart lists the aspect ratios its image models accept and, usefully, the pixel dimensions each one produces at the 1K setting.
| Ratio | Dimensions at 1K | Models |
|---|---|---|
| 1:1 | 1024x1024 | All three |
| 2:3 / 3:2 | 832x1248 / 1248x832 | All three |
| 3:4 / 4:3 | 864x1184 / 1184x864 | All three |
| 4:5 / 5:4 | 896x1152 / 1152x896 | All three |
| 9:16 / 16:9 | 768x1344 / 1344x768 | All three |
| 21:9 | 1536x672 | All three |
| 1:4 / 4:1 | Ultrawide and panoramic | Nano Banana 2 and Pro only |
| 1:8 / 8:1 | Extreme panoramic | Nano Banana 2 and Pro only |
Three things fall out of that table.
The default is square, unless input images say otherwise, so every image you do not deliberately shape is a 1:1. That alone explains a lot of centred, symmetrical AI output.
A 16:9 frame at 1K is 1344x768. That is smaller than a 1080p video frame. If you are making a hero banner, the 1K tier is a comp, not an asset.
Google documents four resolution tiers rather than free-form sizes: 512px on the Lite and 2 models, 1K everywhere, and 2K and 4K on Nano Banana 2 and Pro. The quickstart prints exact dimensions for the 1K tier only, so for anything higher, measure what you actually receive rather than assuming it scaled cleanly.
OpenAI's model: free-form, inside four hard rules
gpt-image-2 takes any size you pass, as long as all four of these hold:
- the long edge is under 3840px
- both edges are multiples of 16
- the ratio of long edge to short edge is no greater than 3:1
- the total pixel count is between 655,360 and 8,294,400
OpenAI adds a warning that matters more than the ceilings: above 2560x1440 the results become more variable, and it calls that territory experimental.

Work out the biggest A4 you can ask for
This is where the rules bite. A4 at 300dpi is 2480x3508, which is 8,699,840 pixels. That is over the cap, so the request is invalid no matter how you word the prompt. Hold the A4 proportion and solve for the pixel limit instead and you land near 2422x3425, which rounds down to 2416x3408 once both edges have to be multiples of 16. That is roughly 292dpi at A4 size.
So the honest answer for print is: generate at the largest legal size, then upscale outside the model, and plan the artwork so the detail that matters is not two pixels wide in the original.
The 3:1 rule is a real constraint
A web banner at 1600x400 is exactly 4:1, which gpt-image-2 will not produce. Google's models will, and go as far as 8:1. If your deliverable is a header, a billboard or a panoramic key visual, that single line in the documentation decides which model you open.
The alternative is the one photographers have always used: generate at a legal ratio with the composition planned for a crop, and leave the space you are going to cut.
Art direction follows the frame, not the other way round
The temptation is to generate square and crop later. It produces weak images, for the same reason cropping a photo to a different aspect never quite works: the model composed for the frame it was given. Subject placement, negative space, where the light falls and how much room the background gets are all decided against the edges of the canvas.
So pick the frame first, then write the prompt for that frame. A 9:16 story needs the subject low and the air above it. A 21:9 banner needs a horizon and a place for type. A 4:5 feed post needs the subject filling more of the frame than you think, because it will be viewed at thumb size.
If layout matters, say so in the prompt. OpenAI's guide recommends calling out placement in words: logo top right, subject centred with negative space to the left. That is cheaper than fixing it in the crop.

A working set of defaults
| Deliverable | Ask for | Then |
|---|---|---|
| Feed post | 4:5, the tallest ratio most feeds show in full | Check the subject reads at 400px wide |
| Story or reel still | 9:16 | Keep the top and bottom sixths clear of anything important |
| Deck slide | 3:2 or 16:9 at the highest quality setting | Use the high quality setting if there is small type |
| Wide banner or billboard comp | 21:9, or Google's 4:1 and 8:1 for the extremes | Compose with a type zone, do not add the type in the model |
| Print piece | The largest size inside the model's pixel cap | Upscale outside the model, then check detail at 100% |
| Cutout for reuse | A square or portrait frame with generous padding | Ask for transparency explicitly, keep the file as PNG or WebP |
What none of this fixes
Resolution is not detail. A 4K generation of a face can still have soft eyelashes and invented stitching, because the model is inventing texture rather than resolving it. Enlarging it makes the invention bigger. When you need real fine detail, the reliable path is still a real photograph, or a generated base composited with real assets on top.
And image size has nothing to do with type quality. If a layout has small copy in it, the quality setting and the model choice matter far more than the pixel count. The prompt recipes in prompt recipes for gpt-image-2 cover how to ask for copy that survives.
FAQ
Can I just ask for "300dpi" in the prompt?
No. Generated images have no meaningful dpi until you assign one. What you control is the pixel count, and dpi is that number divided by the physical size you print at. Work backwards from the print size instead.
Why does my request get rejected at a size the model supposedly supports?
With gpt-image-2, almost always because an edge is not a multiple of 16, or because the total pixel count crossed the cap even though both edges look reasonable. Check the multiplication, not just the edges.
Is 4K worth using?
On Google's models it is available on Nano Banana 2 and Pro, and on gpt-image-2 anything above 2560x1440 is flagged as more variable by OpenAI itself. Use the high tiers for final frames you have already approved at a smaller size, not for exploration, where you will pay for pixels you throw away.
Which model should I pick if my formats are extreme?
Google's, on the documented ratios alone: 1:8 and 8:1 have no equivalent in a model that caps at 3:1. The wider comparison is in which AI image model for which job.
Sources
- Google, Get started with Nano Banana models, read 20 September 2026
- OpenAI, GPT Image Generation Models Prompting Guide, read 20 September 2026
More to read
- AI
AI · · 6 min read
An image model that hands you layers instead of a flat PNG
Qwen-Image-Layered splits a picture into separate RGBA layers you can move, recolour and export as a PSD. What it does, and the thing it does not do.

AI · · 7 min read
Which image models can actually set Chinese, Japanese and Korean type
One maker names Chinese, one claims dozens of languages, one says nothing about scripts at all. What the documentation supports, and how to sign off on CJK copy.