News · High impactBack to News

Alibaba Releases Qwen3.8-Max Frontier Foundation Model

Alibaba has officially released its next-generation frontier model, Qwen3.8-Max. Built on a 2.4-trillion-parameter sparse Mixture-of-Experts (MoE) architecture with 95 billion activated parameters, it natively supports a 1-million-token context window and multimodal inputs. The new model substantially advances long-horizon autonomous coding and complex agent workflows, achieving top-tier rankings on the Arena overall and vision leaderboards. Alibaba also announced that open weights will be released the following week.

2.4-Trillion-Parameter MoE Architecture and Million-Token Context

Alibaba has officially released its next-generation frontier foundation model, Qwen3.8-Max. Built on a sparse Mixture-of-Experts (MoE) architecture, the model scales to 2.4T total parameters while activating approximately 95B parameters per token, natively supporting a context window of up to 1M tokens.

Core Specifications and Metrics

  • 2.4T

    Total parameter scale, designed with a sparse MoE expert network.

  • 95B

    Activated parameters per inference step, balancing model capacity and compute efficiency.

  • 1M

    Native context window length, supporting full codebases and long video indexing.

The model adopts a 2.4-trillion-parameter MoE architecture with a 1M-token context window, scaling capacity while sustaining efficient per-step inference costs. This architectural upgrade enables seamless processing of extensive technical documentation, academic papers, and complex code repositories, establishing the foundation for long-horizon task planning.

16-Day Autonomous Coding and Long-Horizon Execution

To validate engineering capabilities in real-world production environments, the team conducted an autonomous development experiment on the oh-my-cli project. Operating without human intervention, the model ran continuously for 16 days, completing an end-to-end cycle from requirements decomposition and architectural design to code implementation and integration testing.

Long-Horizon Autonomous Execution Loop

  1. Requirements Decomposition and Environment Initialization

    Parse engineering objectives, generate execution plans, and configure test harnesses and runtime environments.

  2. Code Generation and Multi-Step Evolution

    Execute multi-file modifications, track build statuses, and iteratively commit code patches.

  3. Automated Testing and Self-Healing

    Run integration test suites, capture runtime exceptions, and autonomously generate debugging fixes.

Across the autonomous execution run, the model generated 265 code commits, 127 Pull Requests, and 151 GitHub Issues, while building a self-improving evaluation harness. Beyond software engineering, the model ran for 125 continuous hours on research paper replication, improving complex reasoning scores through 4 rounds of iterative exploration.

The model completed 265 code commits across 16 days of unattended operation, demonstrating a transition from single-turn interaction to autonomous long-horizon engineering delivery. This long-horizon capability transforms traditional single-step assistance into autonomous software engineering delivery.

Benchmark Performance and Open-Weights Schedule

On third-party evaluation benchmarks, Qwen3.8-Max ranks #5 globally on the authoritative Arena leaderboard (with an Elo rating of ~1496), standing as the highest-ranked Chinese foundation model; it ranks #2 globally on the vision leaderboard, trailing only the Claude series. Furthermore, the model scored 86.1 on the desktop agent benchmark OSWorld-Verified and 93.0 on PaperBench, establishing leading marks across agent evaluations.

For enterprise access, Alibaba Cloud Model Studio has opened public API access for Qwen3.8-Max, priced at $2 per million input tokens and $6 per million output tokens.

Qwen3.8-Max ranks among the global top 1 tier on authoritative benchmarks, becoming the first open-weights trillion-parameter MoE frontier model. This initiative significantly lowers the barrier for researching and deploying frontier foundation models.

Core Conclusion

Next step

Keep tracking Qwen3.8-Max

Continue along the same topic.

Open entity record