DeepSeek-V4

DeepSeek's MIT-licensed fourth-generation MoE family, led by the 1.7T-parameter V4-Pro flagship.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

DeepSeek's fourth generation keeps frontier weights under plain MIT while splitting the family into two lines: V4-Pro, a 1.7 trillion parameter mixture-of-experts flagship whose 0813 build left preview in mid-August 2026, and the lighter V4-Flash aimed at cheaper high-volume serving, both preceded by preview releases. The models expose three reasoning effort levels, low, high, and max, with recommended output budgets reaching 384K tokens at the top setting, and they ship with DSpark, a speculative decoding module that accelerates generation during agentic coding and multi-turn reasoning sessions. Serving guidance targets vLLM with DSpark enabled or SGLang with the FlashInfer backend, using FP8 KV caches to hold memory down; the Pro model is sized for a four-way GB300 node with expert-parallel and data-parallel layouts. For teams that cannot host it, DeepSeek's own API mirrors the open checkpoints, but the MIT license means clouds and enterprises can serve, fine-tune, and resell the weights without restriction.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Freemium
Platform
Hybrid
Difficulty
Advanced (4/5)
License
MIT
Added
Aug 24, 2026

Related Tools

Featured

Lightweight open-weight LLM by Google available in 1B to 27B sizes.

Open SourceSelf HostedOfflineGPU 4GB+
Easy

Open-source code LLM family by IBM for enterprise code generation.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate

Open-weight LLM by Meta in 8B and 70B sizes with strong general capabilities.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
Featured

High-performance open-weight MoE LLM with 671B total parameters.

Open SourceSelf HostedOfflineGPU 24GB+
Advanced

Hybrid SSM-Transformer model by AI21 Labs combining Mamba with attention layers.

Open SourceSelf HostedOfflineGPU 12GB+
Intermediate

Open-weight code LLM trained on 2 trillion tokens of code and natural language.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
Browse all Large Language Models (LLMs) tools