Trinity Large
Arcee AI's US-trained 400B sparse MoE family under Apache 2.0, with a dedicated reasoning variant.
About
Trained entirely in the United States by startup Arcee AI, the Trinity family argues that sovereign open models can compete at frontier scale. Trinity-Large packs about 400B total parameters into an unusually sparse mixture-of-experts layout, activating 4 of 256 experts for roughly 13B parameters per token, a 1.56 percent routing fraction that yields 2 to 3x the throughput of comparably sized dense models, with a 256K context window. The reasoning variant Trinity-Large-Thinking, released April 1, 2026 under Apache 2.0, adds extended chain-of-thought post-training and agentic reinforcement learning, scoring 91.9 on the PinchBench agentic benchmark, just behind the strongest proprietary systems. Smaller siblings Trinity-Nano (6B total, 1B active) and Trinity-Mini (26B total, 3B active) cover edge and mid-range deployment. Weights live on Hugging Face and run under vLLM, SGLang, llama.cpp, LM Studio, and Transformers, with managed APIs on Arcee's platform and OpenRouter; Apache terms permit unrestricted enterprise use.
Should you use Trinity Large?
Pick it when
Pick this when you want a US-trained open model under Apache-2.0 with high throughput per parameter, since only about 13B of 400B activate, plus a Thinking variant for agents and Nano and Mini siblings for smaller hardware.
Look elsewhere when
Skip it if you must self-host on one GPU: the 400B model needs a multi-GPU setup unless you use Arcee's or OpenRouter's API, while gpt-oss-120b fits one 80 GB GPU. For image or audio input or 1M context, Inkling is the US-built Apache-2.0 option.
Alternatives to Trinity Large
- Inkling
Also US-built and Apache-2.0, with text, image, and audio input and 1M context, but 41B active at 975B total costs more to serve.
- gpt-oss
Apache-2.0 US reasoning models; the 120b fits one 80 GB GPU, far easier to host but much smaller than Trinity-Large.
- Kimi K2
Trillion-parameter agentic MoE known for long tool-call chains, but non-US and under a Modified MIT license with a display clause.
- Tencent Hunyuan Hy3
Apache-2.0 295B MoE with 21B active and 256K context, a similar sparse design plus a multi-token prediction layer for faster decoding.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026