🔥 Trending on HN

Qwen-Image-2.1, Qwen’s image model, combines creation and editing

2 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
7B parameters(seven billion parameters)

A count of learned numerical settings in the model.

reference image(reference image)

An image shown to guide a new result.

mask(mask)

A guide that marks which area may change.

On September 20, 2026, the Qwen team announced Qwen-Image-2.1, an image-generation model from Qwen. The team says the model is open source. Its visual-generation component has 7B parameters. It combines text-to-image generation and image editing. It also supports transparent images as a native output.

What happened

The model can use a prompt to choose a normal image or a transparent one. Qwen shows it changing a subject’s expression while keeping the transparent background. It also shows text replacement inside a transparent layer. A normal photo can become a reusable transparent subject.

Qwen says the model accepts up to 10 reference images. A group portrait can combine separate portraits. A virtual outfit can combine a model, clothes, shoes, a bag, and a hat. An interior can use 10 furniture images. Local edits can be guided by circles, painted marks, or a separate mask. The release also emphasizes preserving people and products, including faces, text, textures, and shapes.

The background

Qwen introduced Qwen-Image-Layered in December 2025 as a dedicated transparent-image model. Qwen-Image-2.1 brings that ability into one creation-and-editing model. The team describes a lightweight architecture with 32 Single-Stream DiT layers and 7B parameters in the visual-generation component. It also describes mixed-granularity attention and KV-cache reuse for multi-image editing. In simpler terms, some instruction and image information can be prepared once and reused. Qwen presents this as an efficiency and memory improvement.

Why it matters

A transparent subject can be reused across designs without first removing its background. Combining generation and editing could reduce the need to switch tools. Multiple references allow a user to specify more parts of a scene. Local masks allow a smaller change instead of requesting a whole new image. Qwen lists design, content creation, e-commerce, and visual storytelling as possible uses. These are practical directions, not independent proof.

What is confirmed

The primary source is Qwen’s official announcement, which describes these features and shows examples. It includes transparent images, up to 10 references, local editing, portrait and product fidelity, typography, panoramas, infographics, and storyboards. The post received 634 points and 170 comments on Hacker News. Those numbers show community attention. They do not prove the model’s claims or performance.

What remains unknown

The announcement does not provide a neutral comparison under matching conditions. It does not give one speed or cost number that applies across computers. It also does not show a failure-rate study for faces, product text, or complex edits. The examples demonstrate possibilities, but they cannot show how often every user gets similar results. Hardware needs and practical setup also require separate testing.

What to watch next

The next useful evidence will be reproducible tests by people outside Qwen. They should check transparent output, multi-image references, local edits, and text rendering. They should report hardware, speed, cost, and failures. The key question is whether one model can preserve the requested parts while changing only the requested area. If it can, Qwen-Image-2.1 may make image work feel more like one continuous task.

Attention source: Hacker News post.

💬 Qwen Image 2.1: What HN commenters noticed

HN commenters found Qwen Image 2.1 impressive as a locally runnable, 7B-class image diffusion model. The performance figures below are anecdotal user reports or thread explanations, not an independent benchmark.

  • Some users felt local image generation was faster and more impressive than local code generation. Others stressed that visual quality is only part of the story: prompt adherence can be unreliable, and artists may judge the results differently.
  • One explanation in the thread is that images benefit from spatial structure, while code must preserve long-range relationships and keep many tokens correct. This is a commenter’s interpretation, not a measured result.
  • For local use, a user reported that stable-diffusion.cpp worked out of the box and generated a 512×512 image in about three minutes on a CPU.
  • Another runtime report for a Q8 quantized setup showed about 15.6 GB of VRAM for the model components. Actual memory needs and speed depend on the quantization and hardware.
  • A commenter praised the model’s CJK text rendering, but that was an individual impression rather than a controlled comparison.
  • The license was described in the thread as restricting commercial use unless a separate commercial license is obtained. Commenters disagreed about enforcement and legal risk, and the discussion did not resolve those questions.

mature digest at 198 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Qwen-Image-2.1, Qwen’s picture model, can create and edit images

📰 Full story: Qwen-Image-2.1, Qwen’s image model, combines creation and editing

It can make see-through pictures and change only the part you choose.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
reference image(reference image)

An image used as a guide for a new picture.

mask(mask)

An image that shows where a change should happen.

7B parameters(seven billion parameters)

A count of learned numerical settings in the model.

💡 The gist

  • Qwen-Image-2.1 can create and edit pictures in one model.
  • It can make see-through pictures and use up to 10 reference images.
  • Hacker News, a technology news site, gave the story 634 points and 170 comments. Those numbers show attention, not proof.

Qwen-Image-2.1, an image model from the Qwen team, was announced on September 20, 2026. Qwen says it is open source. The model follows a text instruction. It can create a normal picture or a picture with a see-through background. It can also edit that picture without losing the see-through background.

Why does this matter? A see-through background makes a subject easier to place over another picture. A person, product, or object can become a reusable design piece. The model can also take up to 10 reference images. Qwen shows examples with people, clothes, shoes, bags, hats, and furniture. These images can become one scene.

The user can mark a small area for editing. They can draw a circle, paint over the area, or provide a separate mask. A mask is an image that shows where the change should happen. Qwen says the model is better at keeping faces and products consistent. It also highlights product words, textures, and shapes.

Qwen says the visual-generation part uses 7B parameters. It also describes a design that reuses information from instructions and input images. The goal is better speed and lower memory use. Qwen does not give one simple speed promise for every computer.

The announcement covers panoramas, infographics, and storyboards. These examples suggest uses in design, shopping images, and visual stories. They are examples from Qwen, not independent test results.

The next step is independent testing. Tests should check transparent images, local edits, multiple references, and text. They should also report hardware, speed, cost, and failures. That will show whether this unified workflow helps beyond demonstrations.

💬 Qwen Image 2.1: An easier summary

HN commenters see Qwen Image 2.1 as a powerful image model that can run locally. Its speed and memory figures come from user reports, not a formal benchmark.

  • Some users found this 7B-class model surprisingly good at local image generation. It can still miss instructions, and artists may evaluate it differently.
  • One explanation is that images provide useful clues about shape and position, while code must stay correct across a long sequence.
  • A user reported that stable-diffusion.cpp worked on a CPU and made a 512×512 image in about three minutes. Another reported about 15.6 GB of VRAM at Q8; hardware and settings change the result.
  • Some commenters praised its CJK text rendering, but that was anecdotal. The license was also understood to require a separate license for commercial use, with legal enforcement still debated.

mature digest at 198 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Qwen-Image-2.1, Qwen’s picture AI, can make and fix pictures

📰 Full story: Qwen-Image-2.1, Qwen’s image model, combines creation and editing

It can make pictures with see-through spaces behind them.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
Qwen-Image-2.1(Qwen Image two point one)

An AI from Qwen that makes pictures.

see-through background(see-through background)

A background you can see through.

Hacker News(Hacker News)

A website where people share technology news.

Qwen-Image-2.1 is a picture-making AI from the Qwen team. It can make a picture from words. It can also fix one small part. A picture can have a see-through background. Then you can place it over another picture.

You can show it up to 10 pictures. It can use them to make one new picture. It can also change one marked spot. You can circle the spot or paint over it.

Qwen says the model can keep faces and products looking alike. It can also keep product words, shapes, and textures. These are claims from Qwen’s own announcement.

Hacker News, a technology news site, gave the story 634 points. It also had 170 comments. Those numbers show interest. They do not prove the AI works perfectly.

Qwen showed examples of pictures, posters, wide views, and story plans. Another test is still needed. People must see whether the same results happen elsewhere.

The big idea is simple. One AI may make and fix the same picture. That could make picture work easier.

💬 Qwen Image 2.1: A 5-year-old’s version

This is a computer model that makes pictures. Some HN users thought it was very impressive for something that can run on a local computer.

  • It does not always follow the picture instructions perfectly. People who know a lot about art may judge it differently.
  • Pictures have shapes and places that help the model. Code is like a long sentence where many parts must all be correct.
  • One user said a CPU made a 512×512 picture in about three minutes. Another said the Q8 version used about 15.6 GB of VRAM. Commercial work may need a separate license.

mature digest at 198 comments (revision 2). We fetched 100 comments and sampled 100 across the thread. These are HN users’ reports, not independently verified facts.

Sources