News · Medium impactBack to News

Alibaba Tongyi Lab Releases Qwen-Image-3.0: High-Density Multimodal Layouts and Precise Micro-Detail Generation for Real-World Productivity

Alibaba Tongyi Lab has released Qwen-Image-3.0, its third-generation image generation base model. Focusing on practical productivity, the model supports up to 4.5k token inputs, delivers 10px micro-text rendering, handles multi-layered UI nesting, and supports 12 languages alongside real-time web search integration.

Enhanced Generation Capacity and Layout Control

Qwen-Image-3.0 increases the prompt input limit to 4.5k tokens, enabling the comprehension and rendering of highly dense visual layouts with both horizontal expansion and vertical nesting capabilities.

  • Horizontal Layout Expansion: The model generates multi-themed infographics in a single execution. In official demonstrations, it successfully output a single nine-grid layout containing distinct educational topics such as tunnel safety comics, spatial geometry lessons, classical text analysis, and projectile motion diagrams, complete with precise texts, formulas, and charts without post-process image stitching.
  • Vertical Logical Nesting: The model handles semantic deconstruction and multi-layered interface nesting. Demonstrations show that a single prompt can direct the model to render a nested screen progression from a VSCode programming interface down to a Qwen chat box, a WeChat dialog, and finally a pour-over coffee poster, with each layer preserving its native UI aesthetics.

update2

Micro-Detail Rendering Precision

The model refines the rendering of fine text and textures, delivering micro-level visual fidelity.

  • Fine Text Precision: The model supports precise rendering of text as small as 10px, successfully displaying academic papers filled with complex LaTeX equations, and simulating natural red handwritten annotations on books or newspapers.
  • Texture Fidelity and Image Editing: For portraits and objects, the model captures micro-textures like skin pores and hair strands. In editing tasks, it repairs stained or torn classical ink paintings, restoring missing areas while maintaining the original brushstroke and ink gradient style.

Multilingual Support and World Knowledge

Qwen-Image-3.0 expands its knowledge base to support broader layout and creative scenarios.

  • Multilingual Rendering and UI Simulation: The model natively renders text in 12 languages, including Japanese, Korean, and Spanish, and simulates interactive interfaces of popular websites, live streams, and operating systems like iOS, macOS, Windows, and Android.
  • Search Integration and IP Creation: The model leverages real-time web search to construct up-to-date content, such as generating daily weather forecast graphics, and recognizes specific IP figures to combine artists like Qi Baishi and Vincent van Gogh in a virtual live stream for creative layouts.

Next step

Keep tracking Alibaba Cloud

Continue along the same topic.

Open entity record