← Read

AI · · 8 min read

Z-Image runs on a 16GB card. Turbo is the variant to be careful with

Tongyi-MAI's 6B open model ships in four variants. Its own table rates Turbo's quality Very High and its diversity Low, and that trade decides which you want.

If you want one open image model sitting on your own machine, download Z-Image-Turbo to make a single finished-looking picture fast, and the plain Z-Image when you need twenty different ideas from one brief. That is not a hedge. It is the trade the Tongyi-MAI team prints in its own model table, where Turbo scores Very High on visual quality and Low on diversity, and the undistilled Z-Image scores High and Medium.

It is easy to lead with this model's leaderboard placement. The more useful number for a designer is the diversity column, because it predicts what happens on the fifth image rather than the first.

Everything below comes from the Z-Image repository, read on 22 September 2026. The showcase images are the team's own picks, not a controlled test, and nothing here was generated for this article.

What it actually is

Z-Image is a family of image models with 6B parameters, built on what the team calls a Scalable Single-Stream DiT, in which text, visual semantic tokens and image tokens are concatenated into one input stream. Six billion is small. FLUX.2 [dev] is 32B and HunyuanImage-3.0 is 80B, which is why the biggest open model needs a datacentre and this one does not.

The repository claims Turbo "fits comfortably within 16G VRAM consumer devices", and links a community project, stable-diffusion.cpp, with a guide to running Z-Image on a GPU with only 4GB of VRAM. The repository carries an Apache 2.0 licence file. Weights live on Hugging Face and ModelScope, which set their own terms, so check those before client work; the distinction between a permissive model licence and the rights to the picture is covered in what the open image model licences actually say.

The four variants, read as a design decision

VariantStepsGuidanceJobVisual qualityDiversityAvailable
Z-Image-Turbo8OffGenerationVery HighLowYes
Z-Image50OnGenerationHighMediumYes
Z-Image-Omni-Base50OnGeneration and editingMediumHighTo be released
Z-Image-Edit50OnEditingHighMediumTo be released

The quality and diversity ratings are the team's own, from the Model Zoo table in the repository. Two of the four are still marked "To be released", so a piece of writing that tells you to edit images with Z-Image-Edit is describing something you cannot download yet.

Turbo is a distilled model: the team compressed a 50-step process into 8, using a method they call Decoupled-DMD, then post-trained it with reinforcement learning. Distillation is what buys the speed, and it is also what costs the diversity. A distilled model has been taught the shortest route to a good answer, so it keeps taking it.

Collage of photorealistic generated pictures: a woman on a train holding a croissant and a coffee cup, rugby players mid-match, prayer flags in front of a snowy peak at sunset, a woman posing in a toy shop, a man walking a poodle in a show ring, a steel stadium facade, and fireworks over a river at night
The team's own photorealism showcase for Z-Image-Turbo. These are the maker's picks, not a controlled comparison. Image: Tongyi-MAI

Why "Low diversity" is the line that matters

Think about what you actually ask an image model to do in a working week. Three routes for a landing page. Six thumbnails for a client to react to. A set of twelve, each different enough to sit next to each other in a deck.

A low-diversity model will give you twelve very well made versions of one idea. Same composition instinct, same light, same face type, same warm grade. You will notice it on the contact sheet, not in any single frame, and it is exactly the failure described in why AI images read as AI: unspecified decisions get filled with an average, and a distilled model has a narrower average than the model it came from.

The undistilled Z-Image is the one the team describes as "well-suited for creative generation, fine-tuning, and downstream development", supporting "a wide range of artistic styles, effective negative prompting, and high diversity across identities, poses, compositions, and layouts". It takes 50 steps instead of 8. For a moodboard, pay the time.

The settings the repository actually recommends

For the full Z-Image model, the repository publishes recommended parameters, and they are unusually specific:

  • Resolution: 512×512 to 2048×2048, described as total pixel area at any aspect ratio, so you spend a pixel budget rather than pick from a menu
  • Guidance scale: 3.0 to 5.0
  • Steps: 28 to 50
  • Negative prompts: "Strongly recommended for better control"
  • CFG normalisation: off for general stylism, on for realism

That last one is the interesting switch. One flag moves the model between a stylised register and a photographic one, which is a decision most models make you argue for in words.

For Turbo the instructions invert. Guidance must be set to 0, because "Guidance should be 0 for the Turbo models", and the step count is 9 in the example, with a comment noting this "actually results in 8 DiT forwards". Guidance is not a dial you still have on Turbo, which is another way of saying the same thing the diversity rating says.

The total pixel area rule is worth holding next to the fixed menus other makers ship, which are laid out in the sizes AI image models can actually give you.

Text in the image, in two scripts

Bilingual text rendering, English and Chinese, is one of the three things the team claims Turbo excels at. The showcase below is their evidence for it.

Eight generated posters combining Chinese and English type: a steam train in snow, a night sky over desert trees, a black portfolio cover with a cartoon figure, a botanical moss exhibition sheet, a period film poster, a watercolour landscape exhibition poster, a close-up leaf product poster, and a blue festival poster with a tiger
The repository's bilingual text rendering showcase. Image: Tongyi-MAI

Now read the English in that image rather than looking at it. The Chinese headlines are clean. The English is not: the festival poster announces a "Sofa Montain Slumnerfest", the film credits include "ARTHUR PENHLDGON", and the portfolio cover, which is advertising the model's own bilingual rendering, sets the word "Biliingual". These are the maker's chosen showcase images, so treat that as the ceiling rather than the average.

The practical rule that falls out: headline-length Chinese is plausible, long English body copy is not, and either way you are setting the real type yourself. The sign-off process is the same as for every other model, and how to check CJK copy in a generated image does not change because this one is smaller.

The model's own example prompt is instructive in itself, and it is the maker's, not mine:

Young Chinese woman in red Hanfu, intricate embroidery. Impeccable makeup,
red floral forehead pattern. Elaborate high bun, golden phoenix headdress,
red flowers, beads. Holds round folding fan with lady, trees, bird. Neon
lightning-bolt lamp (⚡️), bright yellow glow, above extended left palm.
Soft-lit outdoor night background, silhouetted tiered pagoda (西安大雁塔),
blurred colorful distant lights.

Note the shape: no camera specs, no "8K masterpiece", no adjectives about quality. It is a list of nouns with materials and positions attached, and a named light source. That is how these prompts are built.

About the leaderboard rank

The repository states that on 8 December 2025, Z-Image-Turbo "ranked 8th overall on the Artificial Analysis Text-to-Image Leaderboard", making it the top open-source model, and reproduces the leaderboard.

Leaderboard table of open-weights image models ranked by Elo, with Z-Image Turbo at 6B parameters in first place, followed by FLUX.2 dev at 32B, HunyuanImage 3.0 at 80B and Qwen-Image at 20B, alongside columns for release date and API price
The repository's own screenshot of the Artificial Analysis leaderboard, open-weights models only, as it stood in December 2025. Image: Tongyi-MAI

The parameter column is the part to look at. In that snapshot a 6B model sits above a 32B one and an 80B one, which is the actual news here: for the work most designers do, the size of the model stopped being the thing that decides quality.

Two cautions. It is a maker reporting its own placement, and the date is December 2025, which is nine months before this article. Leaderboards move. The durable facts here are the parameter count, the VRAM figure, the licence file and the published settings, because those do not change when someone else ships a model.

Who should actually download this

  • You have a gaming GPU and want to stop paying per image. Start with Turbo, and accept that your set will look related.
  • You need options, not one picture. Use the full Z-Image at 28 to 50 steps with negative prompts.
  • You want to fine-tune on your own illustration style. The table rates Z-Image and Omni-Base as easy to fine-tune, and Turbo as not applicable. Distilled models are not the ones you train on.
  • You need editing. Wait, or use something else. Z-Image-Edit is announced, not released.

FAQ

Is Z-Image better than FLUX.2 or Qwen-Image?

Different question to ask. It is smaller and cheaper to run than either, which for most designers matters more than a ranking. If you want the cross-maker comparison on documented capabilities, which AI image model for which job covers it.

Can I really run it on 4GB of VRAM?

The repository links a community guide for exactly that, using stable-diffusion.cpp rather than the official PyTorch path. Expect it to be slow, and expect to lose features that the main pipeline provides.

Does the diversity rating mean Turbo is worse?

No. It means Turbo is narrower. For a single hero image where you already know what you want, narrow is an advantage, because the model wastes fewer attempts on interpretations you did not ask for.

What is the Prompt Enhancer?

A component the repository says gives the model reasoning, so it can "transcend surface-level descriptions and tap into underlying world knowledge". In practice that means your short prompt gets expanded before it is drawn. Useful when you are exploring, dangerous when you have written a precise brief, for the reasons set out in the rewrite trap.

Sources

More to read