open-source-licensescommercial-usemodel-weightsagplopenrailai-compliance

Open Source AI License Commercial Use: What You Can Ship in 2026

BuilderAI Editorial Team•

This guide covers open source AI license commercial use for people shipping products: which licenses on AI code and model weights let you sell, host, or embed, and which quietly do not. It is written for developers and indie makers who pull tools from GitHub and checkpoints from Hugging Face and need to know, before launch, whether the stack is actually clear to ship.

Every license cited below was checked against the project's LICENSE file or model card in early October 2026, and licenses change between versions, so recheck the release you pin. This is not legal advice.

Code and weights carry separate licenses

An AI project usually has at least two licenses. The repository license covers the source code. The model weights, often hosted separately on Hugging Face, carry their own terms, and the training data can push those terms somewhere the code author never intended.

GitHub's license badge reports the repository license and nothing else. When GitHub cannot match a file to a standard license, its API reports the license as "Other" with NOASSERTION as the SPDX ID. Treat that as an instruction to read the file: MinerU, Dify and Open WebUI all show up that way. On Hugging Face, a model card with license: other means the same thing. Llama 4, Qwen 2.5 72B, Tencent's Hunyuan video and 3D releases and Stability's newer models are all tagged other.

So the audit question is not "is this open source?" but "what are the terms on each artifact I ship or run": the code, the weights, and any auxiliary models downloaded at runtime.

Permissive code licenses: MIT, Apache-2.0, BSD

Permissive licenses are the easy case. MIT, Apache-2.0 and BSD let you use, modify, embed and sell the software in closed products, hosted or distributed, as long as you keep the copyright and license notices.

  • vLLM is Apache-2.0, as are Qdrant and Unsloth.
  • Ollama and llama.cpp are MIT, as are LangChain and Docling.
  • PyTorch uses a BSD-style 3-clause license.

The differences show up at the edges:

LicensePatent grantNotice dutiesExtra condition
MITNone explicitKeep copyright and license textNone
Apache-2.0Explicit grant from contributors, ends if you sue over patents in the workKeep the license and pass along any NOTICE fileMark files you changed
BSD-3-ClauseNone explicitKeep copyright and license textNo using contributor names to endorse your product

Apache-2.0's main advantage is the explicit patent grant in Section 3. If a dependency ships a NOTICE file, Section 4(d) requires you to carry its attribution text in your distribution, a step often skipped when vendoring.

The catch with permissive runtimes is what they load. Ollama's MIT license covers the Ollama binary, not the models you pull through it. Each model keeps its own license, and Ollama prints it for any model you have pulled:

ollama show llama3.1 --license

A permissive engine serving a non-commercial checkpoint is a non-commercial deployment.

Copyleft: GPL, AGPL, and why AGPL matters for hosted apps

Copyleft licenses allow commercial use but require derivative works to carry the same license once a trigger is met. The trigger is what separates GPL from AGPL.

GPL-3.0 triggers on distribution, which the license calls conveying. ComfyUI is GPL-3.0. Running it on your own server behind an API and selling the images does not oblige you to publish anything, because you are not distributing ComfyUI. Shipping a desktop app that bundles a modified ComfyUI is distribution, and the GPL then covers the combined work.

AGPL-3.0 closes that network gap. Section 13 says that if you modify the program, "your modified version must prominently offer all users interacting with it remotely through a computer network" access to the corresponding source. For a hosted product, users of your SaaS can ask for the source of the modified AGPL component, and if your own code is combined with it into one program, your code too.

AGPL turns up in places developers do not expect:

  • Ultralytics YOLO is AGPL-3.0. The alternative is the Ultralytics Enterprise License, which the README positions as the route for business products, internal tools and production deployments, covering the company's AI models as well as its code.
  • PyMuPDF, the fast PDF library inside many RAG pipelines, is AGPL-3.0, with commercial licenses sold by Artifex for proprietary applications.
  • text-generation-webui, KoboldCpp, Khoj and the AUTOMATIC1111 Stable Diffusion web UI are all AGPL-3.0.

A common practical line: running an unmodified AGPL server as a separate process and talking to it over HTTP is lower risk. Importing an AGPL library into your own application, say from ultralytics import YOLO in your backend, makes that backend part of the combined work. For a closed hosted product, your options are to open-source it under AGPL, buy the commercial license, or swap the dependency. For detection, RF-DETR's core package and its Apache-designated models are Apache-2.0, though its smallest and largest detection sizes ship as Plus models under a separate PML 1.0 license. For PDFs, Docling is MIT; our PDF parsing comparison covers the quality tradeoffs.

Licenses also move. MinerU was AGPL-3.0 through the 3.0.x releases. From 3.1.0, released in April 2026, it ships under the "MinerU Open Source License": Apache-2.0, plus a separate commercial license above 100 million monthly active users or USD 20 million in monthly revenue, plus a duty to state in your interface or public docs that an online service uses MinerU. Pinning an old release and upgrading are different legal positions.

OpenRAIL and use-based restrictions

Responsible AI Licenses (RAIL) look permissive and grant commercial rights, then add a list of banned uses that must flow down to everyone downstream. Stable Diffusion XL ships under the CreativeML Open RAIL++-M License. Stable Diffusion 1.5 uses the earlier CreativeML OpenRAIL-M, and StarCoder2 uses BigCode OpenRAIL-M.

What the SDXL license actually requires:

  • You may host the model as a service and sell outputs. The licensor "claims no rights in the Output You generate."
  • Attachment A bans specific uses, including providing medical advice or interpreting medical results, fully automated decision making that adversely impacts someone's legal rights, and use in law enforcement, immigration or asylum processes such as predicting that an individual will commit fraud.
  • The use-based restrictions "MUST be included as an enforceable provision" in any agreement governing use or distribution, and the license names software-as-a-service explicitly. In practice, your terms of service have to carry them.
  • Fine-tunes may be released under different terms, but must keep at least the same use restrictions.

For a consumer image app this is paperwork. For a health, HR, lending or legal product it can be a hard stop, because the banned use is the product.

The OpenRAIL name has also been borrowed for licenses that are far from permissive. Marker 2.0, released in July 2026, moved its code from GPL to Apache-2.0, but its model weights use a "modified AI Pubs Open Rail-M" license that is free only for research, personal use, and startups under USD 5 million in funding or revenue. Surya OCR uses the same split. Datalab's Chandra sets its cap at USD 2 million and adds that the weights cannot be used competitively with Datalab's API. Read the modifications, not the acronym.

Meta's SAM License for SAM 3 is a use-based variant with no revenue cap: commercial use is allowed, but ITAR-regulated, military, nuclear, espionage and weapons uses are excluded. SAM 2 was plain Apache-2.0.

Community licenses: user caps, revenue caps and regional exclusions

The big-lab community license is a custom agreement: free commercial use up to a threshold, plus attribution, naming and sometimes territorial rules. These carry the most clauses that bite as you grow.

Model familyLicenseFree-use ceilingConditions that catch people
Llama 4, Llama 3.1Llama Community License700M monthly active usersDisplay "Built with Llama"; distributed fine-tunes must start their name with "Llama"; Llama 4 multimodal rights not granted to EU-based individuals or companies
Qwen 2.5 72B, Qwen2.5-VL 72BQwen License100M monthly active usersShow "Built with Qwen" or "Improved using Qwen" in the docs of any model you train with it or its outputs and make available
HunyuanVideo, Hunyuan3D 2, HunyuanImage 3.0Tencent Hunyuan community licenses100M monthly active users (1M for Hunyuan3D 2)No rights in the EU, UK or South Korea; outputs may not improve other AI models
Stable Diffusion 3.5, Stable Audio OpenStability AI Community LicenseUSD 1M annual revenueDisplay "Powered by Stability AI"; no using outputs to build other foundation models
Kimi K2Modified MITNoneShow "Kimi K2" in the UI above 100M monthly active users or USD 20M monthly revenue
LTX-2.xLTX-2.x Community LicenseEntities under USD 10M annual revenuePaid license above the threshold

Regional exclusions. The acceptable use policies attached to the Llama 3.2 and Llama 4 licenses both state that the Section 1(a) rights for multimodal models are not granted to an individual domiciled in, or a company with its principal place of business in, the European Union. Llama 4 Scout and Maverick are natively multimodal, so an EU-headquartered startup does not get those rights. The restriction does not apply to end users of a product that incorporates the model. Tencent goes further: the Hunyuan license "DOES NOT APPLY IN THE EUROPEAN UNION, UNITED KINGDOM AND SOUTH KOREA", and Section 5(c) forbids using or displaying the works or their output outside the licensed territory. A global SaaS built on HunyuanVideo is hard to square with that text without geo-blocking.

Output restrictions. Tencent bars using outputs to improve any other AI model, and Stability bars using them to create or improve any foundational generative model. Generating synthetic training data for your own model is exactly what these clauses forbid.

Thresholds. Llama's 700 million user test uses your user count in the month before the version's release date. Stability's USD 1 million test is annual revenue, and your rights "shall terminate as of such date" when you cross it, so arrange an enterprise license before you get there.

The 2026 trend runs toward plain permissive weights, which makes per-checkpoint checking more important, not less. Qwen3 models are Apache-2.0 while the 72B Qwen 2.5 models stay on the Qwen License and the 3B Qwen 2.5 sits under a research-only license. Gemma 4 is Apache-2.0, while Gemma 3 still uses Google's Gemma terms with gated access. The original DeepSeek-V3 release shipped weights under a custom DeepSeek License; V3-0324, R1 and the V4 family are MIT. Same family name, different rules.

Non-commercial and research-only licenses

CC-BY-NC 4.0 defines NonCommercial as "not primarily intended for or directed towards commercial advantage or monetary compensation." A paid app or an ad-supported site is clearly on the wrong side of that line, and production use inside a for-profit company is a gray area at best.

Common non-commercial weights:

  • FLUX.1 [dev] and its Fill, Depth, Canny, Redux, Kontext and Krea variants use the FLUX.1 [dev] Non-Commercial License, which grants "non-commercial and non-production use." Outputs may be used for any purpose, including commercial, except training, fine-tuning or distilling a competing model. Running [dev] as the engine of a paid service needs a license from Black Forest Labs. FLUX.1 [schnell] and the autoencoder are Apache-2.0.
  • Depth Anything V2 Small is Apache-2.0, while Base, Large and Giant are CC-BY-NC-4.0.
  • CodeFormer uses the S-Lab License 1.0, which permits use "for non-commercial purpose" only.
  • Fish Speech uses the Fish Audio Research License: research and non-commercial use free, commercial use requires a separate license.
  • XTTS-v2 uses the Coqui Public Model License, which "allows only non-commercial use of a machine learning model and its outputs." Note the last two words: the audio you generate is covered too.

Our guides to the open source image generation stack and open-weight text-to-speech models list commercially clean alternatives in each category.

The trap: permissive code shipped with restricted weights

This is the most common trap. The GitHub badge says MIT or Apache, the code really is permissive, and the checkpoint the README tells you to download is not.

ProjectCode licenseWeights licenseWhere the restriction comes from
F5-TTSMITCC-BY-NCTrained on Emilia, an in-the-wild dataset
AudioCraft (MusicGen)MITCC-BY-NC 4.0, in a separate LICENSE_weights fileMeta's release terms
InstantIDApache-2.0Own checkpoints research only; also needs InsightFace face models, non-commercial research onlyThe README terms plus an auxiliary model pulled in at setup
InsightFaceMITPretrained models non-commercial research only, including auto-downloadsTraining data terms
Coqui TTS with XTTS-v2MPL-2.0Coqui Public Model License, non-commercialModel license
Marker 2.xApache-2.0Modified OpenRAIL-M, free under USD 5M funding or revenueVendor's commercial model

InstantID is the sneakiest because the restriction sits in two places. Its README says its own released checkpoints are for research purposes only, and adds that "both manual-downloading and auto-downloading face models from insightface are for non-commercial research purposes only." InsightFace's README applies the same policy to models its Python package fetches automatically, and since November 2025 it routes commercial licensing of its open face recognition packs, such as buffalo_l, to a licensing contact. Many face swap tools load those same models.

The fix is boring: for every model file your code loads, find the license of that specific file. Sometimes both sides are clean. Kokoro, an 82 million parameter TTS model, ships Apache-licensed weights. Whisper's code and weights are MIT on GitHub. gpt-oss is Apache-2.0, as are Qwen3 and Gemma 4, and DeepSeek V4 is MIT.

Source-available app licenses that look open

Several popular self-hosted AI apps use modified licenses that are free to self-host but reserve the commercial patterns a startup might want.

  • Open WebUI moved to the Open WebUI License in April 2025: BSD-3 terms plus a clause that bars removing or altering "Open WebUI" branding unless your deployment has 50 or fewer end users in any rolling 30-day period, you have written permission, or you hold an enterprise license.
  • Dify uses a modified Apache-2.0. Commercial use as a backend or internal platform is allowed, but operating a multi-tenant environment, where one tenant is one workspace, needs written authorization, and you may not remove the logo or copyright information from the Dify frontend.
  • n8n uses the Sustainable Use License, which limits use to "your own internal business purposes or for non-commercial or personal use." Files with .ee in the name require an enterprise license.
  • LobeChat uses the LobeHub Community License: commercial use without modifying the source is fine, but developing and distributing a derivative work requires a commercial license.

White-labeling, reselling seats, or running a multi-tenant SaaS on one of these is the exact case the license reserves. Self-hosting for your own team is almost always fine. For clean product terms, AnythingLLM is MIT and Jan's LICENSE file is Apache-2.0. Our self-hosted stack guide for small teams compares them in daily use.

Open source AI license commercial use: the decision table

Use this as a first pass, then read the actual license for anything below the first row.

License familyExamplesSell or embed commerciallyHosted SaaSShip inside an app you distributeWhat to do
MIT, Apache-2.0, BSDvLLM, Ollama, Kokoro weightsYesYesYesKeep LICENSE and NOTICE files
GPL-3.0ComfyUIYesYes, no source dutyCombined work is GPLRun it as a separate service
AGPL-3.0Ultralytics YOLO, PyMuPDFYes, with source dutiesModified versions must offer sourceCombined work is AGPLBuy a commercial license or swap
OpenRAIL-M, RAIL++-MSDXL, SD 1.5YesYes, restrictions in your termsYes, pass restrictions onCheck your use is not banned
Modified OpenRAIL with capsMarker, Surya, Chandra weightsUnder USD 2M to 5MSame capSame capBudget for a license
Community, user or revenue capsLlama 4, Qwen 2.5 72B, SD 3.5, MinerUUnder thresholdYesYes, with attributionTrack MAU and revenue
Community, territory limitsHunyuanVideo, HunyuanImage 3.0Outside EU, UK, South KoreaGeo-restrictSameGeo-block or switch models
Non-commercialF5-TTS weights, FLUX.1 [dev], XTTS-v2NoNoNoSwap the model or buy a license
Source-available appsOpen WebUI, Dify, n8nLimitedBranding, tenancy or resale limitsLimitedSelf-host internally

Three commands cover most of the audit: the repository license as GitHub sees it, the weights license from the Hugging Face model card, and a build check that fails on any GPL-family Python dependency.

gh api repos/ultralytics/ultralytics/license --jq .license.spdx_id
curl -s https://huggingface.co/api/models/SWivid/F5-TTS | jq -r .cardData.license
pip install pip-licenses
pip-licenses --partial-match --fail-on="General Public License;GPL"

The first returns AGPL-3.0, the second cc-by-nc-4.0. The partial match also catches LGPL, so use --allow-only with an explicit list if you want to permit it. None of these see weights downloaded at runtime, which is why the model-by-model check still matters.

Common questions

Can I use MIT-licensed AI code in a commercial product? Yes, as long as you keep the copyright and license notice. Check the weights separately: an MIT repository can download non-commercial checkpoints, as F5-TTS and AudioCraft do.

Does the AGPL apply if I only call the software over an API? Section 13 applies to modified versions that users interact with over a network. Calling an unmodified AGPL service from a separate process is generally treated as lower risk than importing an AGPL library into your own code; for a closed hosted product, the commercial license or a swap is the safe route.

Can I sell images made with FLUX.1 [dev]? The license lets you use outputs for any purpose, including commercial, except training a competing model. Running FLUX.1 [dev] itself as a commercial or production service falls outside the non-commercial grant, while FLUX.1 [schnell] is Apache-2.0.

Is Llama open source for commercial use? Llama allows commercial use under the Llama Community License up to 700 million monthly active users, with "Built with Llama" attribution and naming rules. The Open Source Initiative has said the license does not meet the Open Source Definition, and EU-based companies do not receive rights to multimodal versions, which includes Llama 4.

Do model licenses cover the outputs? Sometimes. OpenRAIL++-M and FLUX.1 [dev] disclaim ownership of outputs, Tencent and Stability restrict using outputs to improve other models, and the Coqui Public Model License limits outputs to non-commercial use.

Related Tools

More Articles