Qwen3.8
Alibaba's 2026 open-weight flagship family, from a 2.4T MoE down to an Apache-licensed 27B model.
About
Alibaba's August 2026 generation spans the widest range of any open family. At the top, Qwen3.8-2.4T-A95B, the open-weight release of the hosted Qwen3.8-Max, is second in size only to Kimi K3 among open models: 2.4 trillion total parameters with 95B active, built from a hybrid stack that interleaves Gated DeltaNet linear-attention blocks with gated full attention, routing through 512 experts per MoE layer. It runs exclusively in thinking mode, handles text only, reads 262K tokens natively with extension toward 1M, and downloads from Hugging Face under the custom Qwen3.8-Max license. At the other end, Qwen3.8-27B arrived in mid-August 2026 under Apache 2.0 as a dense vision-language model that reads images and video and posts strong coding scores, with 4-bit quantized builds around 17GB that fit a single consumer GPU. Recommended serving for the big checkpoints is SGLang, vLLM, or TokenSpeed on multi-GPU clusters, with Alibaba Cloud offering hosted endpoints; the 27B's permissive terms make it the practical choice for commercial local deployment.
Should you use Qwen3.8?
Pick it when
Pick the Apache-2.0 27B when you want a current vision-language model with image, video, and strong coding on one consumer GPU at 4-bit, or the 2.4T flagship when you run a multi-GPU cluster and want open weights close to Qwen3.8-Max.
Look elsewhere when
Skip the flagship if you need image input, fast non-thinking replies, or a standard license: it is text-only, thinking-only, and under a custom license. For permissive frontier MoE weights, DeepSeek-V4 and GLM-5 are MIT.
Alternatives to Qwen3.8
- Qwen 2.5 / Qwen 3
Earlier Apache-2.0 generation with many more sizes from 0.5B and a non-thinking mode, better for small or fast deployments.
- Kimi K3
The only larger open model, 2.8T with native image input and 1M context, under a modified MIT license, where Qwen3.8's flagship is text-only.
- DeepSeek-V4
MIT-licensed 1.7T Pro flagship with three reasoning effort levels and a lighter Flash line, with no custom license to review.
- Muse Glimmer
Apache-2.0 30B with image input for a 24 GB card, trained for tool calling and error recovery rather than the 27B's coding focus.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- Qwen3.8-Max License
- Added
- Aug 24, 2026