← Read

AI · · 7 min read

The largest open image model is Tencent's, and its own chart shows where it loses

80 billion parameters, reasoning before it draws, and three obstacles: a datacentre-sized hardware bill, a licence that excludes three regions, and maker-run evals.

HunyuanImage-3.0 is the largest open image model anyone has published, by its maker's own description, and almost no designer will ever run it. Tencent's own model card asks for at least three 80GB GPUs for the base model and eight for the Instruct version. What you can use is the hosted version, and what you should read before you build a workflow on it is the licence, which does not apply in the European Union, the United Kingdom or South Korea.

This is a walk through what Tencent documents in its own repository, read on 21 September 2026: what the model does, what its January 2026 update added for creative work, and what its published evaluations actually show.

What it is

Tencent open sourced HunyuanImage-3.0 on 28 September 2025 with inference code, weights and a technical report. Two things make it unusual.

It is architecturally different from most image models. Instead of the diffusion transformer design nearly everyone else ships, Tencent describes "a unified autoregressive framework" that models text and images in one system. In practice that is why the Instruct version can think in words before it draws.

And it is enormous. The repository calls it "the largest open-source image generation Mixture of Experts (MoE) model to date", with 64 experts, 80 billion total parameters and 13 billion active per token. Those are Tencent's descriptions of its own model, not an independent measurement.

What the January update added, and why a designer would care

On 26 January 2026 Tencent released HunyuanImage-3.0-Instruct and a distilled variant. The feature list is the part that matters for creative work:

  • Image-to-image editing, described as adding elements, removing objects, modifying styles and replacing backgrounds "while preserving key visual elements".
  • Multi-image fusion, combining up to three reference images into one composition.
  • Prompt self-rewrite, which "automatically enhances sparse or vague prompts into professional-grade, detail-rich descriptions".
  • CoT Think, a structured reasoning pass that breaks a brief into subject, composition, lighting, colour palette and style before generating.

The rewrite is the double-edged one. On a thin brief it will invent an art direction for you, and that art direction is the model's house taste rather than yours. If you have a look you care about, write the full description yourself and leave nothing for the rewrite to fill in.

A showcase panel from Tencent's repository headed with the Chinese words for style transformation: on the left a phone selfie of a young woman in a black leather jacket shot from above, on the right the same face redrawn as a pop-art graffiti portrait in saturated colour on a yellow album sleeve, with the Chinese prompt set below
One of Tencent's own showcase panels for HunyuanImage-3.0-Instruct, a style transformation from a reference photo. A maker's showcase pick, not a controlled test. Image: Tencent

Obstacle one: you cannot run it

Tencent's model card is blunt about hardware.

ModelParametersRecommended VRAMWhat it does
HunyuanImage-3.080B total, 13B active3 x 80GB or moreText to image
HunyuanImage-3.0-Instruct80B total, 13B active8 x 80GB or moreText to image, image to image, prompt self-rewrite, reasoning
HunyuanImage-3.0-Instruct-Distil80B total, 13B active8 x 80GB or moreThe above, with 8 sampling steps recommended

Eight 80GB accelerators is a rack, not a workstation. Compare that with FLUX.2 [klein] 4B, which Black Forest Labs says fits in around 8GB of VRAM, and the difference in who can actually download these weights becomes obvious. "Open" here means the weights are published, not that they are yours to run.

What is left is Tencent's hosted route: the repository links an official site and chat product where you can try the model, and the showcase images carry a footnote saying the results come from the Instruct model available in Tencent's own apps.

Obstacle two: the licence has a map

The weights ship under the Tencent Hunyuan Community License, and the first thing it says, in capitals, is that the agreement does not apply in the European Union, the United Kingdom and South Korea. "Territory" is then defined as "the worldwide territory, excluding the territory of the European Union, United Kingdom and South Korea".

Two clauses matter for client work. The grant is "for the Territory only". And section 5 states that you must not use, reproduce, modify, distribute or display the works, "Output or results" outside the Territory, and that any such use is unlicensed.

Read plainly, that means a studio in Sydney, Singapore, Mumbai, Tokyo or Jakarta is inside the licence, a studio in Seoul or London is not, and the restriction follows the pictures, not just the weights. There is also a commercial threshold: if your products had more than 100 million monthly active users in the month before the version release date, you have to request a separate licence.

That is a summary of what the text says, not legal advice. If Hunyuan output is heading into a campaign that runs across regions, put the licence in front of whoever signs off your contracts. The same reading exercise for the other open models is in what the open image model licences actually say.

Obstacle three: the evaluations are Tencent's own, and they are interesting anyway

Most makers publish charts where they win. Tencent published one where it does not, which is worth looking at closely.

Bar chart titled HunyuanImage3.0-Instruct Win Rate GSB, with green bars for an internal R&D test set and orange for a user preference test set. Nano Banana Pro shows negative bars at minus 1.43 and minus 2.60 per cent, Seedream-4.5 shows 7.13 and 10.48 per cent, Qwen-Image-Edit-2511 shows 23.59 and 34.71 per cent
Tencent's published Good/Same/Bad win rates for the Instruct model on editing tasks. The bars against Nano Banana Pro sit below zero. Image: Tencent

The method, as described in the repository: more than 1,000 single-image and multi-image editing cases, one inference per prompt with no cherry-picking, default settings for every model, judged by more than 100 professional evaluators. On that test the Instruct model wins comfortably against Qwen-Image-Edit-2511, wins clearly against Seedream-4.5, and loses narrowly to Nano Banana Pro on both test sets.

A maker publishing its own defeat in the same chart as its wins tells you something about how the numbers were run. It does not turn the chart into an independent benchmark: the prompts, the pairings and the judging panel are all Tencent's.

The text-to-image side uses a machine metric, SSAE, built from 3,500 key points across 12 semantic categories and scored by multimodal models.

Two radar charts comparing Seedream 4.0, Nano Banana, GPT-Image and HunyuanImage 3.0 across fourteen axes including mean accuracy, global accuracy, subject nouns and attributes, scene attributes, shot, style and composition, one chart for English prompts and one for Chinese
Tencent's SSAE comparison. The four models trace almost the same shape, in both languages. Image: Tencent

Look at the shape rather than the ranking. Four models from four companies trace nearly the same outline on both the English and Chinese charts, with the visible gaps in shot, style and composition rather than in whether the right things appear in the picture. Prompt adherence has largely converged at this level. Taste has not, and no radar chart measures it.

So when would you use it?

Three honest cases. When you are working in Chinese and want a model whose documentation, prompt handbook and showcase were written in Chinese first, which is not true of the American models. When you want a reasoning pass over a vague brief and are happy for the model to propose the art direction. And when you are researching what open weights can do, in a territory the licence covers.

For everything else, the models with better documentation and smaller hardware bills are compared in which AI image model for which job, and the specific question of which models can set Chinese, Japanese and Korean type is in CJK type in AI images.

FAQ

Is HunyuanImage-3.0 open source?

The weights and inference code are public, but the Tencent Hunyuan Community License is not an OSI open source licence. It limits the territory, attaches an acceptable use policy, restricts trademark use and adds a commercial threshold above 100 million monthly active users.

Can I use the images commercially?

Inside the licence territory, subject to the agreement and its acceptable use policy. Outside it, including the EU, UK and South Korea, the licence states that use of the output is not authorised. Take the actual text to your legal advisers rather than relying on a summary.

Do I need eight GPUs to try it?

No, to try it. Tencent's repository links its own hosted site and chat product. You need that hardware only to run the weights yourself, and the recommended VRAM in the model card is at least 3 x 80GB for the base model.

What does the distilled version change for me?

Speed. It is tuned for 8 sampling steps instead of the default 50, with the same feature list as Instruct. If you are generating variations, that is the difference between waiting and working.

Sources

More to read