Today for AI

IT之家 · 智能时代 · 10/6/2026, 7:45:52 AM

Strata Engine Lets 12GB GPUs Run 125B Qwen3.8 Model

By IT之家Original title: 12GB 显存显卡跑 125B Qwen3.8 模型:Strata 登场,单张 RTX 5070 跑出 94 词元 / 秒
68AI Score
Executive Summary

Developer Niko1221 has open-sourced the Strata inference engine, enabling quantized Qwen3.8-Flash-Next (125B parameters) to run on consumer GPUs with 12GB+ VRAM, such as the RTX 5070. The solution achieves a generation speed of 94 tokens per second, demonstrating efficient deployment of large models on low-memory hardware.

SOURCE COVERAGEOriginal coverage

科技媒体 gigazine 今天(10 月 6 日)报道,报道称开发者 Niko1221 开源推出 Strata 引擎,可以在 12GB 及以上显存的消费级显卡上,运行量化的 Qwen3.8-Flash-Next 模型(1250 亿参数)。