AI · · 6 min read
An image model that hands you layers instead of a flat PNG
Qwen-Image-Layered splits a picture into separate RGBA layers you can move, recolour and export as a PSD. What it does, and the thing it does not do.
The most annoying thing about a generated image is not the hands. It is that everything arrives welded together. Move the bottle two centimetres left and you are back to prompting, hoping the light and the label survive the trip.
Qwen-Image-Layered is Alibaba's answer to that. It takes an image and decomposes it into several RGBA layers, each one independently editable, and its reference app exports the stack as .psd, .pptx or a zip. The weights are Apache 2.0 and the code is public. Everything below comes from the Qwen-Image-Layered repository, read on 20 September 2026, and from the Qwen-Image repository's release notes. No images were generated for this article.
What it actually does
The repository's own description is short and worth reading closely: the model decomposes an image into "multiple RGBA layers", a representation Qwen describes as having "inherent editability", because "each layer can be independently manipulated without affecting other content".
That phrasing is doing real work. The claim is not that edits are cleverer. The claim is that the problem goes away, because the thing you want to change is physically separate from the things you do not. Qwen lists the operations this buys you: resizing, repositioning and recolouring, described as "high-fidelity elementary operations".
If you have spent an afternoon trying to talk a model into moving one object without redrawing the room, you already know why that matters. Every other approach to consistency is persuasion. This one is geometry.
The part nobody mentions in the demos
Here is the limit, stated plainly in the repository's own note, and it changes how you would use this.
The released weights are fine-tuned for image-to-multi-RGBA decomposition. Text-to-multi-RGBA generation is supported, but Qwen says its performance there is "limited". So this is not a model that generates a layered composition from a prompt. It is a model that takes a finished flat image and pulls it apart.
The second note is about the prompt. It is meant to describe the overall content of the input image, including elements that are partly hidden behind something else, and you can name text that sits behind a foreground object. It is explicitly "not designed to control the semantic content of individual layers". You cannot ask for the logo on its own layer. You get the decomposition the model thinks is right.
Practically, the workflow is: generate or shoot the image somewhere else, then decompose, then edit one layer, then recombine.
Variable layers, and layers inside layers
Two details raise this above a background remover.
The layer count is yours to set. The repository shows the same image decomposed into three layers or eight, and the quick start passes layers as an ordinary argument alongside resolution and steps. Three for a product on a backdrop; eight when you want the type, the highlight and each prop apart.
Decomposition is also recursive. Any layer can itself be decomposed again, which Qwen describes as enabling "infinite decomposition". That is the bit designers will actually use. Pull the foreground out, then pull the foreground apart.
Quick start defaults published in the repository
layers: 4
resolution: 640 (640 and 1024 buckets exist; 640 is recommended for this version)
num_inference_steps: 50
true_cfg_scale: 4.0It exports a PSD
The repository ships three small apps, and this is where it stops being a research artefact. src/app.py opens a web interface that decomposes an image and exports the layers "into pptx, zip, and psd files, where you can edit and move these layers flexibly". A second one edits an individual RGBA layer using Qwen-Image-Edit. A third recombines them, and the README is specific about the order: upload the layers from the bottom up.
A PSD export means the handoff is to Photoshop, not to another prompt. A PPTX export means the handoff is to whoever asked for the deck. Both are unglamorous and both are the reason this will get used.
Where it sits next to what you already do
| If you want | The usual approach | What layers change |
|---|---|---|
| Move one object | Re-prompt, or mask and inpaint | Drag the layer. Nothing else can shift |
| Recolour one element | Prompt "change only X", repeat the preserve list | Recolour that layer, alone |
| Remove an object | Inpainting, then fix what it invented underneath | Delete the layer. Qwen's showcase shows objects removed cleanly |
| A cutout on transparency | Generate with a transparent background | Different job: one subject, one alpha, no stack |
| A character across a series | Reference images, chained edits, locked constraints | Different job: see keeping a character consistent |
The transparent background route is the closer cousin, and the two solve different halves of the same problem. One gives you an asset with clean alpha from the start. This gives you a composition you can take apart after the fact.
The licence question, answered for once
Qwen-Image-Layered is Apache 2.0, stated in the repository's licence agreement and in the LICENSE file. That is the permissive end of the spectrum, and it puts this model in the same bracket as Qwen-Image itself rather than the non-commercial weights elsewhere in the open ecosystem. The distinction between a permissive model licence and the rights to the picture you made with it is a separate question, covered in what the open image model licences actually say.
What to be sceptical about
Every image in the repository's showcase was chosen by the team that built the model. They show clean decompositions: a recoloured layer, a figure swapped, text revised, an object deleted, an object resized, an object moved. Nothing published tells you how it behaves on a hair mask, a glass bottle, a motion-blurred edge or a photograph where the subject is genuinely ambiguous against its background.
Those are exactly the cases that decide whether a tool earns a place in a real project, and they are also the cases a showcase never includes. Run it on your own worst image before you plan a campaign around it.
The release timeline is worth knowing too: weights and blog on 19 December 2025, the research paper the day before, hosted demos on Hugging Face Spaces and ModelScope from 22 December 2025. This is young software. Treat the version numbers as moving.
github.com ↗Qwen-Image repository release notesThe parent repository, where the dated release notes for Qwen-Image-Layered, Qwen-Image-2512 and Qwen-Image-2.0 are kept. Source: QwenFAQ
Do I need a GPU for this?
To run the weights, yes, and the quick start assumes CUDA. Qwen also links hosted demos on Hugging Face Spaces and ModelScope, which is the sensible way to find out whether decomposition works on your kind of image before you set anything up locally.
Can I ask for a specific number of layers?
Yes. layers is a parameter, and the repository demonstrates three and eight on the same image. What you cannot do is say what goes on each one.
Will it work on a photograph I took?
Nothing in the repository restricts it to generated images, and the pipeline takes any RGBA image as input. Whether it separates your photograph cleanly is a different question, and the only honest answer is to try it on a few frames that matter to you.
What about the layered output other tools advertise?
Treat any "layered" claim as a question rather than a feature until you have read the documentation behind it: how many layers, whether the alpha is real, and whether the layers are generated from scratch or derived from a finished image. Those answers differ, and they change what the feature is worth to you. Qwen's is the one whose code, weights, licence and stated limitations you can read for yourself today, which is why this article is about that one.
Sources
- Qwen-Image-Layered repository, read 20 September 2026
- Qwen-Image repository, release notes for Qwen-Image-Layered, Qwen-Image-2512 and Qwen-Image-2.0, read 20 September 2026
More to read

AI · · 7 min read
Which image models can actually set Chinese, Japanese and Korean type
One maker names Chinese, one claims dozens of languages, one says nothing about scripts at all. What the documentation supports, and how to sign off on CJK copy.

AI · · 7 min read
Keeping a character, a product and a style consistent across a set of images
Four mechanisms for stopping drift across a series: chained edits, locked constraints, reference images and structural conditioning, with the limits each one has.