News · High impactBack to News

DeepSeek-V4-Flash Official API Released in Public Beta

DeepSeek has announced the public beta release of the official DeepSeek-V4-Flash API (DeepSeek-V4-Flash-0731). Retaining the same lightweight MoE architecture with 284B total and 13B active parameters, the re-post-trained model delivers significantly enhanced agent capabilities. Benchmark scores on Terminal Bench 2.1 and Toolathlon approach or exceed the performance of the early V4-Pro-Preview.

Re-Post-Training and Agent Performance

Artificial intelligence research laboratory DeepSeek has announced the public beta launch of the official API for deepseek-v4-flash, its lightweight Mixture-of-Experts (MoE) model. While retaining its original parameter scale, this release incorporates intensive re-post-training to substantially strengthen the model's inference performance on agentic tasks.

Parameter Scale and Agentic Benchmark Scores

  • 284B

    The total parameter size designed in the architecture of the updated model.

  • 13B

    The active parameter size utilized during a single forward pass in the updated model.

  • 82.7

    The score achieved on the Terminal Bench 2.1 benchmark for terminal command generation and execution.

  • 70.3

    The score achieved on the Toolathlon evaluation for multi-tool calling and integration capabilities.

The table below illustrates the benchmark scores of the updated model compared with legacy previews and other mainstream industry models:

BenchmarkDeepSeek-V4-Flash-0731DeepSeek-V4-Flash-PreviewDeepSeek-V4-Pro-PreviewGLM-5.2Opus-4.8
Terminal Bench 2.182.761.872.181.085.0
NL2Repo54.239.438.548.969.7
Cybergym76.738.752.7-83.1
DeepSWE54.47.312.846.258.0
Toolathlon-Verified70.349.755.959.976.2
Agents' Last Exam25.215.816.523.825.7
AutomationBench (Public)25.110.812.812.927.2
DSBench-FullStack68.737.041.861.871.6
DSBench-Hard59.625.831.154.571.7

The updated model has undergone intensive re-post-training, significantly boosting its agent capabilities, which now approach or exceed the early V4-Pro-Preview on Terminal Bench 2.1 and Toolathlon benchmarks. This update significantly expands the practical boundaries of lightweight MoE models in automation systems and complex workflows.

Low API Pricing and Peak Hours Adjustments

With the public beta of the updated model, DeepSeek continues to maintain its highly competitive and low-cost API pricing. Meanwhile, this version fully supports an ultra-large context window of up to 1M tokens with a maximum output limit of 384K tokens, featuring a default thinking mode that bills reasoning tokens at standard output rates.

Input Token Price Comparison (Cache Miss)

DeepSeek-V4-Pro Input Token Pricing (Cache Miss)0.435 $/1M
DeepSeek-V4-Flash Input Token Pricing (Cache Miss)0.14 $/1M

DeepSeek-V4-Flash maintains highly competitive low pricing, but official updates indicate that rates will double during Beijing Time peak hours daily in the near future. The exact effective date for this peak/off-peak policy is subject to official announcement, and developers should plan high-concurrency requests accordingly during peak times.

API Identifiers and Retirement of Legacy Aliases

Developers can access the latest version by setting the model parameter to deepseek-v4-flash in their API calls. Currently, this update applies only to the API endpoint; the official web and app interfaces remain unchanged for now, while the official release of V4-Pro is still under preparation.

All API calls must specify the concrete model identifiers, as the legacy general aliases were fully retired on July 24, 2026. Teams are advised to complete global parameter replacements in their applications as soon as possible.

Core Conclusion

Next step

Keep tracking DeepSeek-V4-Flash

Continue along the same topic.

Open entity record