Tencent Hunyuan Hy3
Tencent's 295B mixture-of-experts LLM activating 21B per token with a 256K context window.
About
Tencent's entry in the open MoE race, Hy3 packs 295B parameters into 192 experts and routes each token through 8 of them, so only about 21B activate per step, while a 3.8B multi-token-prediction layer accelerates decoding through speculative execution. Context stretches to 256K tokens, a reasoning_effort switch moves between deep chain-of-thought and a no_think mode for direct answers, and tool calling is built in. The summer 2026 open-sourcing, following an April preview, includes BF16 and FP8 weights on Hugging Face and ModelScope under Apache-2.0, serving recipes for vLLM and SGLang with speculative decoding enabled, TensorRT-LLM support, a fine-tuning pipeline, GRPO reinforcement learning via the verl framework, and the AngelSlim quantization toolkit. Tencent recommends eight large-memory GPUs such as H20-3e class hardware for full-precision serving, placing self-hosting firmly in cluster territory, though hosted endpoints picked it up quickly and it climbed OpenRouter usage rankings soon after launch.
Should you use Tencent Hunyuan Hy3?
Pick it when
Pick this when you have an eight-GPU cluster and want an Apache-2.0 MoE with 256K context, a switch between deep reasoning and direct answers, and a full fine-tuning and GRPO pipeline to adapt it in-house.
Look elsewhere when
Skip it if you lack eight large-memory GPUs: self-hosting is cluster work, so you would depend on third-party endpoints. gpt-oss-120b fits one 80 GB GPU with similar reasoning-effort control.
Alternatives to Tencent Hunyuan Hy3
- Qwen 2.5 / Qwen 3
Qwen 3's 235B-A22B is a similar-size Apache-2.0 MoE with thinking and non-thinking modes and a large community fine-tune base.
- gpt-oss
Apache-2.0 reasoning MoE with effort levels whose 120b fits one 80 GB GPU, far easier to host, though with fewer total parameters.
- Trinity Large
Apache-2.0 400B MoE with 256K context and a thinking variant; 13B active versus Hy3's 21B means less compute per token, but more weights to store.
- MiniMax-M1
Apache-2.0 456B hybrid-attention reasoning MoE with native 1M context, better for very long inputs but heavier at 45.9B active.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026