← Read

Design · · 6 min read

The sizes AI image models can actually give you, and how to art direct around them

Every model has hard ceilings on size and shape. The real numbers for Nano Banana and gpt-image-2, and what they mean for print, social and billboards.

Two numbers decide whether a generated image can do the job you had in mind: how many pixels it can produce, and what shapes it will produce them in. Both are published, both are stricter than people assume, and neither is what you get by typing "4K, high resolution" into a prompt.

Here are the documented limits for the two most widely used hosted families, read on 20 September 2026, and the art direction decisions that follow from them.

Google's models: a fixed menu, with the pixels printed on it

Google's Nano Banana quickstart lists the aspect ratios its image models accept and, usefully, the pixel dimensions each one produces at the 1K setting.

RatioDimensions at 1KModels
1:11024x1024All three
2:3 / 3:2832x1248 / 1248x832All three
3:4 / 4:3864x1184 / 1184x864All three
4:5 / 5:4896x1152 / 1152x896All three
9:16 / 16:9768x1344 / 1344x768All three
21:91536x672All three
1:4 / 4:1Ultrawide and panoramicNano Banana 2 and Pro only
1:8 / 8:1Extreme panoramicNano Banana 2 and Pro only

Three things fall out of that table.

The default is square, unless input images say otherwise, so every image you do not deliberately shape is a 1:1. That alone explains a lot of centred, symmetrical AI output.

A 16:9 frame at 1K is 1344x768. That is smaller than a 1080p video frame. If you are making a hero banner, the 1K tier is a comp, not an asset.

Google documents four resolution tiers rather than free-form sizes: 512px on the Lite and 2 models, 1K everywhere, and 2K and 4K on Nano Banana 2 and Pro. The quickstart prints exact dimensions for the 1K tier only, so for anything higher, measure what you actually receive rather than assuming it scaled cleanly.

OpenAI's model: free-form, inside four hard rules

gpt-image-2 takes any size you pass, as long as all four of these hold:

  • the long edge is under 3840px
  • both edges are multiples of 16
  • the ratio of long edge to short edge is no greater than 3:1
  • the total pixel count is between 655,360 and 8,294,400

OpenAI adds a warning that matters more than the ceilings: above 2560x1440 the results become more variable, and it calls that territory experimental.

Pitch deck slide titled Market Opportunity with concentric circle diagram and bar chart, generated with gpt-image-2 at 1536x864
OpenAI's published example of a wide deck slide, generated at 1536x864 with the high quality setting. Image: OpenAI

Work out the biggest A4 you can ask for

This is where the rules bite. A4 at 300dpi is 2480x3508, which is 8,699,840 pixels. That is over the cap, so the request is invalid no matter how you word the prompt. Hold the A4 proportion and solve for the pixel limit instead and you land near 2422x3425, which rounds down to 2416x3408 once both edges have to be multiples of 16. That is roughly 292dpi at A4 size.

So the honest answer for print is: generate at the largest legal size, then upscale outside the model, and plan the artwork so the detail that matters is not two pixels wide in the original.

The 3:1 rule is a real constraint

A web banner at 1600x400 is exactly 4:1, which gpt-image-2 will not produce. Google's models will, and go as far as 8:1. If your deliverable is a header, a billboard or a panoramic key visual, that single line in the documentation decides which model you open.

The alternative is the one photographers have always used: generate at a legal ratio with the composition planned for a crop, and leave the space you are going to cut.

Art direction follows the frame, not the other way round

The temptation is to generate square and crop later. It produces weak images, for the same reason cropping a photo to a different aspect never quite works: the model composed for the frame it was given. Subject placement, negative space, where the light falls and how much room the background gets are all decided against the edges of the canvas.

So pick the frame first, then write the prompt for that frame. A 9:16 story needs the subject low and the air above it. A 21:9 banner needs a horizon and a place for type. A 4:5 feed post needs the subject filling more of the frame than you think, because it will be viewed at thumb size.

If layout matters, say so in the prompt. OpenAI's guide recommends calling out placement in words: logo top right, subject centred with negative space to the left. That is cheaper than fixing it in the crop.

Mobile app mockup for a farmers market in an iPhone frame, generated with gpt-image-2 at 1024x1536
A vertical 1024x1536 generation from OpenAI's guide. The frame was chosen before the prompt, not after. Image: OpenAI

A working set of defaults

DeliverableAsk forThen
Feed post4:5, the tallest ratio most feeds show in fullCheck the subject reads at 400px wide
Story or reel still9:16Keep the top and bottom sixths clear of anything important
Deck slide3:2 or 16:9 at the highest quality settingUse the high quality setting if there is small type
Wide banner or billboard comp21:9, or Google's 4:1 and 8:1 for the extremesCompose with a type zone, do not add the type in the model
Print pieceThe largest size inside the model's pixel capUpscale outside the model, then check detail at 100%
Cutout for reuseA square or portrait frame with generous paddingAsk for transparency explicitly, keep the file as PNG or WebP

What none of this fixes

Resolution is not detail. A 4K generation of a face can still have soft eyelashes and invented stitching, because the model is inventing texture rather than resolving it. Enlarging it makes the invention bigger. When you need real fine detail, the reliable path is still a real photograph, or a generated base composited with real assets on top.

And image size has nothing to do with type quality. If a layout has small copy in it, the quality setting and the model choice matter far more than the pixel count. The prompt recipes in prompt recipes for gpt-image-2 cover how to ask for copy that survives.

FAQ

Can I just ask for "300dpi" in the prompt?

No. Generated images have no meaningful dpi until you assign one. What you control is the pixel count, and dpi is that number divided by the physical size you print at. Work backwards from the print size instead.

Why does my request get rejected at a size the model supposedly supports?

With gpt-image-2, almost always because an edge is not a multiple of 16, or because the total pixel count crossed the cap even though both edges look reasonable. Check the multiplication, not just the edges.

Is 4K worth using?

On Google's models it is available on Nano Banana 2 and Pro, and on gpt-image-2 anything above 2560x1440 is flagged as more variable by OpenAI itself. Use the high tiers for final frames you have already approved at a smaller size, not for exploration, where you will pay for pixels you throw away.

Which model should I pick if my formats are extreme?

Google's, on the documented ratios alone: 1:8 and 8:1 have no equivalent in a model that caps at 3:1. The wider comparison is in which AI image model for which job.

Sources

More to read