← Read

AI · · 6 min read

An image model that hands you layers instead of a flat PNG

Qwen-Image-Layered splits a picture into separate RGBA layers you can move, recolour and export as a PSD. What it does, and the thing it does not do.

The most annoying thing about a generated image is not the hands. It is that everything arrives welded together. Move the bottle two centimetres left and you are back to prompting, hoping the light and the label survive the trip.

Qwen-Image-Layered is Alibaba's answer to that. It takes an image and decomposes it into several RGBA layers, each one independently editable, and its reference app exports the stack as .psd, .pptx or a zip. The weights are Apache 2.0 and the code is public. Everything below comes from the Qwen-Image-Layered repository, read on 20 September 2026, and from the Qwen-Image repository's release notes. No images were generated for this article.

github.com ↗Qwen-Image-Layered on GitHubThe repository, with the quick start, the Gradio apps and the Apache 2.0 licence. Source: Qwen

What it actually does

The repository's own description is short and worth reading closely: the model decomposes an image into "multiple RGBA layers", a representation Qwen describes as having "inherent editability", because "each layer can be independently manipulated without affecting other content".

That phrasing is doing real work. The claim is not that edits are cleverer. The claim is that the problem goes away, because the thing you want to change is physically separate from the things you do not. Qwen lists the operations this buys you: resizing, repositioning and recolouring, described as "high-fidelity elementary operations".

If you have spent an afternoon trying to talk a model into moving one object without redrawing the room, you already know why that matters. Every other approach to consistency is persuasion. This one is geometry.

The part nobody mentions in the demos

Here is the limit, stated plainly in the repository's own note, and it changes how you would use this.

The released weights are fine-tuned for image-to-multi-RGBA decomposition. Text-to-multi-RGBA generation is supported, but Qwen says its performance there is "limited". So this is not a model that generates a layered composition from a prompt. It is a model that takes a finished flat image and pulls it apart.

The second note is about the prompt. It is meant to describe the overall content of the input image, including elements that are partly hidden behind something else, and you can name text that sits behind a foreground object. It is explicitly "not designed to control the semantic content of individual layers". You cannot ask for the logo on its own layer. You get the decomposition the model thinks is right.

Practically, the workflow is: generate or shoot the image somewhere else, then decompose, then edit one layer, then recombine.

Variable layers, and layers inside layers

Two details raise this above a background remover.

The layer count is yours to set. The repository shows the same image decomposed into three layers or eight, and the quick start passes layers as an ordinary argument alongside resolution and steps. Three for a product on a backdrop; eight when you want the type, the highlight and each prop apart.

Decomposition is also recursive. Any layer can itself be decomposed again, which Qwen describes as enabling "infinite decomposition". That is the bit designers will actually use. Pull the foreground out, then pull the foreground apart.

Quick start defaults published in the repository
layers: 4
resolution: 640 (640 and 1024 buckets exist; 640 is recommended for this version)
num_inference_steps: 50
true_cfg_scale: 4.0

It exports a PSD

The repository ships three small apps, and this is where it stops being a research artefact. src/app.py opens a web interface that decomposes an image and exports the layers "into pptx, zip, and psd files, where you can edit and move these layers flexibly". A second one edits an individual RGBA layer using Qwen-Image-Edit. A third recombines them, and the README is specific about the order: upload the layers from the bottom up.

A PSD export means the handoff is to Photoshop, not to another prompt. A PPTX export means the handoff is to whoever asked for the deck. Both are unglamorous and both are the reason this will get used.

Where it sits next to what you already do

If you wantThe usual approachWhat layers change
Move one objectRe-prompt, or mask and inpaintDrag the layer. Nothing else can shift
Recolour one elementPrompt "change only X", repeat the preserve listRecolour that layer, alone
Remove an objectInpainting, then fix what it invented underneathDelete the layer. Qwen's showcase shows objects removed cleanly
A cutout on transparencyGenerate with a transparent backgroundDifferent job: one subject, one alpha, no stack
A character across a seriesReference images, chained edits, locked constraintsDifferent job: see keeping a character consistent

The transparent background route is the closer cousin, and the two solve different halves of the same problem. One gives you an asset with clean alpha from the start. This gives you a composition you can take apart after the fact.

The licence question, answered for once

Qwen-Image-Layered is Apache 2.0, stated in the repository's licence agreement and in the LICENSE file. That is the permissive end of the spectrum, and it puts this model in the same bracket as Qwen-Image itself rather than the non-commercial weights elsewhere in the open ecosystem. The distinction between a permissive model licence and the rights to the picture you made with it is a separate question, covered in what the open image model licences actually say.

What to be sceptical about

Every image in the repository's showcase was chosen by the team that built the model. They show clean decompositions: a recoloured layer, a figure swapped, text revised, an object deleted, an object resized, an object moved. Nothing published tells you how it behaves on a hair mask, a glass bottle, a motion-blurred edge or a photograph where the subject is genuinely ambiguous against its background.

Those are exactly the cases that decide whether a tool earns a place in a real project, and they are also the cases a showcase never includes. Run it on your own worst image before you plan a campaign around it.

The release timeline is worth knowing too: weights and blog on 19 December 2025, the research paper the day before, hosted demos on Hugging Face Spaces and ModelScope from 22 December 2025. This is young software. Treat the version numbers as moving.

github.com ↗Qwen-Image repository release notesThe parent repository, where the dated release notes for Qwen-Image-Layered, Qwen-Image-2512 and Qwen-Image-2.0 are kept. Source: Qwen

FAQ

Do I need a GPU for this?

To run the weights, yes, and the quick start assumes CUDA. Qwen also links hosted demos on Hugging Face Spaces and ModelScope, which is the sensible way to find out whether decomposition works on your kind of image before you set anything up locally.

Can I ask for a specific number of layers?

Yes. layers is a parameter, and the repository demonstrates three and eight on the same image. What you cannot do is say what goes on each one.

Will it work on a photograph I took?

Nothing in the repository restricts it to generated images, and the pipeline takes any RGBA image as input. Whether it separates your photograph cleanly is a different question, and the only honest answer is to try it on a few frames that matter to you.

What about the layered output other tools advertise?

Treat any "layered" claim as a question rather than a feature until you have read the documentation behind it: how many layers, whether the alpha is real, and whether the layers are generated from scratch or derived from a finished image. Those answers differ, and they change what the feature is worth to you. Qwen's is the one whose code, weights, licence and stated limitations you can read for yourself today, which is why this article is about that one.

Sources

More to read