Open Source AI License Commercial Use: What You Can Ship in 2026
This guide covers open source AI license commercial use for people shipping products: which licenses on AI code and model weights let you sell, host, or embed, and which quietly do not. It is written for developers and indie makers who pull tools from GitHub and checkpoints from Hugging Face and need to know, before launch, whether the stack is actually clear to ship.
Every license cited below was checked against the project's LICENSE file or model card in early October 2026, and licenses change between versions, so recheck the release you pin. This is not legal advice.
Code and weights carry separate licenses
An AI project usually has at least two licenses. The repository license covers the source code. The model weights, often hosted separately on Hugging Face, carry their own terms, and the training data can push those terms somewhere the code author never intended.
GitHub's license badge reports the repository license and nothing else. When GitHub cannot match a file to a standard license, its API reports the license as "Other" with NOASSERTION as the SPDX ID. Treat that as an instruction to read the file: MinerU, Dify and Open WebUI all show up that way. On Hugging Face, a model card with license: other means the same thing. Llama 4, Qwen 2.5 72B, Tencent's Hunyuan video and 3D releases and Stability's newer models are all tagged other.
So the audit question is not "is this open source?" but "what are the terms on each artifact I ship or run": the code, the weights, and any auxiliary models downloaded at runtime.
Permissive code licenses: MIT, Apache-2.0, BSD
Permissive licenses are the easy case. MIT, Apache-2.0 and BSD let you use, modify, embed and sell the software in closed products, hosted or distributed, as long as you keep the copyright and license notices.
- vLLM is Apache-2.0, as are Qdrant and Unsloth.
- Ollama and llama.cpp are MIT, as are LangChain and Docling.
- PyTorch uses a BSD-style 3-clause license.
The differences show up at the edges:
| License | Patent grant | Notice duties | Extra condition |
|---|---|---|---|
| MIT | None explicit | Keep copyright and license text | None |
| Apache-2.0 | Explicit grant from contributors, ends if you sue over patents in the work | Keep the license and pass along any NOTICE file | Mark files you changed |
| BSD-3-Clause | None explicit | Keep copyright and license text | No using contributor names to endorse your product |
Apache-2.0's main advantage is the explicit patent grant in Section 3. If a dependency ships a NOTICE file, Section 4(d) requires you to carry its attribution text in your distribution, a step often skipped when vendoring.
The catch with permissive runtimes is what they load. Ollama's MIT license covers the Ollama binary, not the models you pull through it. Each model keeps its own license, and Ollama prints it for any model you have pulled:
ollama show llama3.1 --license
A permissive engine serving a non-commercial checkpoint is a non-commercial deployment.
Copyleft: GPL, AGPL, and why AGPL matters for hosted apps
Copyleft licenses allow commercial use but require derivative works to carry the same license once a trigger is met. The trigger is what separates GPL from AGPL.
GPL-3.0 triggers on distribution, which the license calls conveying. ComfyUI is GPL-3.0. Running it on your own server behind an API and selling the images does not oblige you to publish anything, because you are not distributing ComfyUI. Shipping a desktop app that bundles a modified ComfyUI is distribution, and the GPL then covers the combined work.
AGPL-3.0 closes that network gap. Section 13 says that if you modify the program, "your modified version must prominently offer all users interacting with it remotely through a computer network" access to the corresponding source. For a hosted product, users of your SaaS can ask for the source of the modified AGPL component, and if your own code is combined with it into one program, your code too.
AGPL turns up in places developers do not expect:
- Ultralytics YOLO is AGPL-3.0. The alternative is the Ultralytics Enterprise License, which the README positions as the route for business products, internal tools and production deployments, covering the company's AI models as well as its code.
- PyMuPDF, the fast PDF library inside many RAG pipelines, is AGPL-3.0, with commercial licenses sold by Artifex for proprietary applications.
- text-generation-webui, KoboldCpp, Khoj and the AUTOMATIC1111 Stable Diffusion web UI are all AGPL-3.0.
A common practical line: running an unmodified AGPL server as a separate process and talking to it over HTTP is lower risk. Importing an AGPL library into your own application, say from ultralytics import YOLO in your backend, makes that backend part of the combined work. For a closed hosted product, your options are to open-source it under AGPL, buy the commercial license, or swap the dependency. For detection, RF-DETR's core package and its Apache-designated models are Apache-2.0, though its smallest and largest detection sizes ship as Plus models under a separate PML 1.0 license. For PDFs, Docling is MIT; our PDF parsing comparison covers the quality tradeoffs.
Licenses also move. MinerU was AGPL-3.0 through the 3.0.x releases. From 3.1.0, released in April 2026, it ships under the "MinerU Open Source License": Apache-2.0, plus a separate commercial license above 100 million monthly active users or USD 20 million in monthly revenue, plus a duty to state in your interface or public docs that an online service uses MinerU. Pinning an old release and upgrading are different legal positions.
OpenRAIL and use-based restrictions
Responsible AI Licenses (RAIL) look permissive and grant commercial rights, then add a list of banned uses that must flow down to everyone downstream. Stable Diffusion XL ships under the CreativeML Open RAIL++-M License. Stable Diffusion 1.5 uses the earlier CreativeML OpenRAIL-M, and StarCoder2 uses BigCode OpenRAIL-M.
What the SDXL license actually requires:
- You may host the model as a service and sell outputs. The licensor "claims no rights in the Output You generate."
- Attachment A bans specific uses, including providing medical advice or interpreting medical results, fully automated decision making that adversely impacts someone's legal rights, and use in law enforcement, immigration or asylum processes such as predicting that an individual will commit fraud.
- The use-based restrictions "MUST be included as an enforceable provision" in any agreement governing use or distribution, and the license names software-as-a-service explicitly. In practice, your terms of service have to carry them.
- Fine-tunes may be released under different terms, but must keep at least the same use restrictions.
For a consumer image app this is paperwork. For a health, HR, lending or legal product it can be a hard stop, because the banned use is the product.
The OpenRAIL name has also been borrowed for licenses that are far from permissive. Marker 2.0, released in July 2026, moved its code from GPL to Apache-2.0, but its model weights use a "modified AI Pubs Open Rail-M" license that is free only for research, personal use, and startups under USD 5 million in funding or revenue. Surya OCR uses the same split. Datalab's Chandra sets its cap at USD 2 million and adds that the weights cannot be used competitively with Datalab's API. Read the modifications, not the acronym.
Meta's SAM License for SAM 3 is a use-based variant with no revenue cap: commercial use is allowed, but ITAR-regulated, military, nuclear, espionage and weapons uses are excluded. SAM 2 was plain Apache-2.0.
Community licenses: user caps, revenue caps and regional exclusions
The big-lab community license is a custom agreement: free commercial use up to a threshold, plus attribution, naming and sometimes territorial rules. These carry the most clauses that bite as you grow.
| Model family | License | Free-use ceiling | Conditions that catch people |
|---|---|---|---|
| Llama 4, Llama 3.1 | Llama Community License | 700M monthly active users | Display "Built with Llama"; distributed fine-tunes must start their name with "Llama"; Llama 4 multimodal rights not granted to EU-based individuals or companies |
| Qwen 2.5 72B, Qwen2.5-VL 72B | Qwen License | 100M monthly active users | Show "Built with Qwen" or "Improved using Qwen" in the docs of any model you train with it or its outputs and make available |
| HunyuanVideo, Hunyuan3D 2, HunyuanImage 3.0 | Tencent Hunyuan community licenses | 100M monthly active users (1M for Hunyuan3D 2) | No rights in the EU, UK or South Korea; outputs may not improve other AI models |
| Stable Diffusion 3.5, Stable Audio Open | Stability AI Community License | USD 1M annual revenue | Display "Powered by Stability AI"; no using outputs to build other foundation models |
| Kimi K2 | Modified MIT | None | Show "Kimi K2" in the UI above 100M monthly active users or USD 20M monthly revenue |
| LTX-2.x | LTX-2.x Community License | Entities under USD 10M annual revenue | Paid license above the threshold |
Regional exclusions. The acceptable use policies attached to the Llama 3.2 and Llama 4 licenses both state that the Section 1(a) rights for multimodal models are not granted to an individual domiciled in, or a company with its principal place of business in, the European Union. Llama 4 Scout and Maverick are natively multimodal, so an EU-headquartered startup does not get those rights. The restriction does not apply to end users of a product that incorporates the model. Tencent goes further: the Hunyuan license "DOES NOT APPLY IN THE EUROPEAN UNION, UNITED KINGDOM AND SOUTH KOREA", and Section 5(c) forbids using or displaying the works or their output outside the licensed territory. A global SaaS built on HunyuanVideo is hard to square with that text without geo-blocking.
Output restrictions. Tencent bars using outputs to improve any other AI model, and Stability bars using them to create or improve any foundational generative model. Generating synthetic training data for your own model is exactly what these clauses forbid.
Thresholds. Llama's 700 million user test uses your user count in the month before the version's release date. Stability's USD 1 million test is annual revenue, and your rights "shall terminate as of such date" when you cross it, so arrange an enterprise license before you get there.
The 2026 trend runs toward plain permissive weights, which makes per-checkpoint checking more important, not less. Qwen3 models are Apache-2.0 while the 72B Qwen 2.5 models stay on the Qwen License and the 3B Qwen 2.5 sits under a research-only license. Gemma 4 is Apache-2.0, while Gemma 3 still uses Google's Gemma terms with gated access. The original DeepSeek-V3 release shipped weights under a custom DeepSeek License; V3-0324, R1 and the V4 family are MIT. Same family name, different rules.
Non-commercial and research-only licenses
CC-BY-NC 4.0 defines NonCommercial as "not primarily intended for or directed towards commercial advantage or monetary compensation." A paid app or an ad-supported site is clearly on the wrong side of that line, and production use inside a for-profit company is a gray area at best.
Common non-commercial weights:
- FLUX.1 [dev] and its Fill, Depth, Canny, Redux, Kontext and Krea variants use the FLUX.1 [dev] Non-Commercial License, which grants "non-commercial and non-production use." Outputs may be used for any purpose, including commercial, except training, fine-tuning or distilling a competing model. Running [dev] as the engine of a paid service needs a license from Black Forest Labs. FLUX.1 [schnell] and the autoencoder are Apache-2.0.
- Depth Anything V2 Small is Apache-2.0, while Base, Large and Giant are CC-BY-NC-4.0.
- CodeFormer uses the S-Lab License 1.0, which permits use "for non-commercial purpose" only.
- Fish Speech uses the Fish Audio Research License: research and non-commercial use free, commercial use requires a separate license.
- XTTS-v2 uses the Coqui Public Model License, which "allows only non-commercial use of a machine learning model and its outputs." Note the last two words: the audio you generate is covered too.
Our guides to the open source image generation stack and open-weight text-to-speech models list commercially clean alternatives in each category.
The trap: permissive code shipped with restricted weights
This is the most common trap. The GitHub badge says MIT or Apache, the code really is permissive, and the checkpoint the README tells you to download is not.
| Project | Code license | Weights license | Where the restriction comes from |
|---|---|---|---|
| F5-TTS | MIT | CC-BY-NC | Trained on Emilia, an in-the-wild dataset |
| AudioCraft (MusicGen) | MIT | CC-BY-NC 4.0, in a separate LICENSE_weights file | Meta's release terms |
| InstantID | Apache-2.0 | Own checkpoints research only; also needs InsightFace face models, non-commercial research only | The README terms plus an auxiliary model pulled in at setup |
| InsightFace | MIT | Pretrained models non-commercial research only, including auto-downloads | Training data terms |
| Coqui TTS with XTTS-v2 | MPL-2.0 | Coqui Public Model License, non-commercial | Model license |
| Marker 2.x | Apache-2.0 | Modified OpenRAIL-M, free under USD 5M funding or revenue | Vendor's commercial model |
InstantID is the sneakiest because the restriction sits in two places. Its README says its own released checkpoints are for research purposes only, and adds that "both manual-downloading and auto-downloading face models from insightface are for non-commercial research purposes only." InsightFace's README applies the same policy to models its Python package fetches automatically, and since November 2025 it routes commercial licensing of its open face recognition packs, such as buffalo_l, to a licensing contact. Many face swap tools load those same models.
The fix is boring: for every model file your code loads, find the license of that specific file. Sometimes both sides are clean. Kokoro, an 82 million parameter TTS model, ships Apache-licensed weights. Whisper's code and weights are MIT on GitHub. gpt-oss is Apache-2.0, as are Qwen3 and Gemma 4, and DeepSeek V4 is MIT.
Source-available app licenses that look open
Several popular self-hosted AI apps use modified licenses that are free to self-host but reserve the commercial patterns a startup might want.
- Open WebUI moved to the Open WebUI License in April 2025: BSD-3 terms plus a clause that bars removing or altering "Open WebUI" branding unless your deployment has 50 or fewer end users in any rolling 30-day period, you have written permission, or you hold an enterprise license.
- Dify uses a modified Apache-2.0. Commercial use as a backend or internal platform is allowed, but operating a multi-tenant environment, where one tenant is one workspace, needs written authorization, and you may not remove the logo or copyright information from the Dify frontend.
- n8n uses the Sustainable Use License, which limits use to "your own internal business purposes or for non-commercial or personal use." Files with
.eein the name require an enterprise license. - LobeChat uses the LobeHub Community License: commercial use without modifying the source is fine, but developing and distributing a derivative work requires a commercial license.
White-labeling, reselling seats, or running a multi-tenant SaaS on one of these is the exact case the license reserves. Self-hosting for your own team is almost always fine. For clean product terms, AnythingLLM is MIT and Jan's LICENSE file is Apache-2.0. Our self-hosted stack guide for small teams compares them in daily use.
Open source AI license commercial use: the decision table
Use this as a first pass, then read the actual license for anything below the first row.
| License family | Examples | Sell or embed commercially | Hosted SaaS | Ship inside an app you distribute | What to do |
|---|---|---|---|---|---|
| MIT, Apache-2.0, BSD | vLLM, Ollama, Kokoro weights | Yes | Yes | Yes | Keep LICENSE and NOTICE files |
| GPL-3.0 | ComfyUI | Yes | Yes, no source duty | Combined work is GPL | Run it as a separate service |
| AGPL-3.0 | Ultralytics YOLO, PyMuPDF | Yes, with source duties | Modified versions must offer source | Combined work is AGPL | Buy a commercial license or swap |
| OpenRAIL-M, RAIL++-M | SDXL, SD 1.5 | Yes | Yes, restrictions in your terms | Yes, pass restrictions on | Check your use is not banned |
| Modified OpenRAIL with caps | Marker, Surya, Chandra weights | Under USD 2M to 5M | Same cap | Same cap | Budget for a license |
| Community, user or revenue caps | Llama 4, Qwen 2.5 72B, SD 3.5, MinerU | Under threshold | Yes | Yes, with attribution | Track MAU and revenue |
| Community, territory limits | HunyuanVideo, HunyuanImage 3.0 | Outside EU, UK, South Korea | Geo-restrict | Same | Geo-block or switch models |
| Non-commercial | F5-TTS weights, FLUX.1 [dev], XTTS-v2 | No | No | No | Swap the model or buy a license |
| Source-available apps | Open WebUI, Dify, n8n | Limited | Branding, tenancy or resale limits | Limited | Self-host internally |
Three commands cover most of the audit: the repository license as GitHub sees it, the weights license from the Hugging Face model card, and a build check that fails on any GPL-family Python dependency.
gh api repos/ultralytics/ultralytics/license --jq .license.spdx_id
curl -s https://huggingface.co/api/models/SWivid/F5-TTS | jq -r .cardData.license
pip install pip-licenses
pip-licenses --partial-match --fail-on="General Public License;GPL"
The first returns AGPL-3.0, the second cc-by-nc-4.0. The partial match also catches LGPL, so use --allow-only with an explicit list if you want to permit it. None of these see weights downloaded at runtime, which is why the model-by-model check still matters.
Common questions
Can I use MIT-licensed AI code in a commercial product? Yes, as long as you keep the copyright and license notice. Check the weights separately: an MIT repository can download non-commercial checkpoints, as F5-TTS and AudioCraft do.
Does the AGPL apply if I only call the software over an API? Section 13 applies to modified versions that users interact with over a network. Calling an unmodified AGPL service from a separate process is generally treated as lower risk than importing an AGPL library into your own code; for a closed hosted product, the commercial license or a swap is the safe route.
Can I sell images made with FLUX.1 [dev]? The license lets you use outputs for any purpose, including commercial, except training a competing model. Running FLUX.1 [dev] itself as a commercial or production service falls outside the non-commercial grant, while FLUX.1 [schnell] is Apache-2.0.
Is Llama open source for commercial use? Llama allows commercial use under the Llama Community License up to 700 million monthly active users, with "Built with Llama" attribution and naming rules. The Open Source Initiative has said the license does not meet the Open Source Definition, and EU-based companies do not receive rights to multimodal versions, which includes Llama 4.
Do model licenses cover the outputs? Sometimes. OpenRAIL++-M and FLUX.1 [dev] disclaim ownership of outputs, Tencent and Stability restrict using outputs to improve other models, and the Coqui Public Model License limits outputs to non-commercial use.
Related Tools
Dify
Open-source platform for building LLM applications with visual workflow editor.
F5-TTS
Diffusion transformer text-to-speech model using flow matching for fluent, faithful speech.
FLUX.1
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
HunyuanVideo
Open-source video generation model by Tencent with text and image conditioning.
InstantID
Zero-shot identity-preserving image generation from a single face photo.
Llama 4
Latest Llama model family by Meta with Mixture-of-Experts architecture.
Marker
Converts PDFs and office documents into Markdown or JSON, handling tables, equations, and OCR locally.
MinerU
Local parser that turns PDFs, scans and Office files into Markdown or JSON ready for LLM pipelines.
Open WebUI
Self-hosted browser chat app for local Ollama models or OpenAI-compatible APIs, with RAG, plugins, and user management.
Qwen
Alibaba's hosted chat interface to the Qwen model family with document, image, and coding support.
Stable Diffusion 3.5
Latest Stable Diffusion model with improved text rendering and composition.
Ultralytics YOLO
State-of-the-art real-time object detection supporting YOLOv5 through v11.
More Articles
How to Evaluate Open Source AI Tools: The Checklist We Use
How BuilderAI.tools vets and maintains its 787 listings, plus a reusable checklist for maintenance, licenses, hardware floors, bus factor, and security.
The Best Self-Hosted AI Stack for Small Teams in 2026
An opinionated reference architecture for self-hosted team AI in 2026: vLLM, Open WebUI, LiteLLM, Qdrant, and Langfuse on one GPU box, with honest alternatives at every layer.
Best Local LLM Setups by GPU Budget: 8GB, 16GB, 24GB, Multi-GPU
Exact model, quant, and engine picks for every VRAM tier in 2026, from an 8GB laptop GPU to a dual-3090 tower, with honest notes on what each setup cannot do.