open-source-aiai-licensingtool-evaluationself-hostingai-securityeditorial

How to Evaluate Open Source AI Tools: The Checklist We Use

BuilderAI Editorial Team•

This guide explains how BuilderAI.tools decides what gets listed and how we keep listings accurate, then turns that process into a checklist you can reuse. It is for developers and technical founders who want to know how to evaluate open source AI tools before a dependency, a model license, or an open port becomes their problem.

What it takes to get listed

A tool gets a page in the directory only after review, and the review starts from primary sources. For every listing we work from the project repository, the LICENSE file, the model card when weights are involved, and the official documentation. Vendor landing pages, launch threads, and aggregator blurbs are not sources. At best they are leads that point to one.

Submissions go through the same review before anything is published. Two kinds are rejected: thin submissions, where there is not enough substance behind the project to say anything useful, and uncategorized ones, where the tool does not fit a category a reader would actually browse. A directory that accepts everything becomes a pile of links, and a pile of links does not help anyone choose.

The directory currently lists 787 tools. Each page is a researched assessment built from those primary sources, not a claim that we have benchmarked every project on our own hardware. That distinction matters, and it is one you should apply to any review site, including this one.

How listings stay accurate

Three habits do most of the work.

Licenses come from the license file, not the badge. We check the actual license text, and we check code and weights separately, because AI projects split them constantly. A repository can be MIT while the checkpoint it downloads is non-commercial. The license section below shows how common that is.

Descriptions are written fresh. Every description is written from scratch and then mechanically checked for phrasing overlap with the upstream README. This is not about style. A paraphrased README inherits the README's claims, and READMEs are written to persuade.

Every listing ends in a decision. Each page carries a "should you use it" verdict and a set of alternatives with their tradeoffs, because knowing that a tool exists is not the same as knowing whether to pick it.

Accuracy also means removal. In October 2026 we retired 85 listings that fell into three buckets: archived repositories, projects inactive for about two years, and paid services that had been presented as open tools. A few historically important inactive projects stayed, with honest notes about what replaced them, because people still search for them and should land on a page that says where the field moved.

How to evaluate open source AI tools: five questions

The rest of this article is the reusable part. Before adopting any open source AI tool, answer these five questions in order. Each one can end the evaluation early, so the cheap checks come first.

  1. Is anyone still maintaining it? Look at the default branch, the last tagged release, and any status notice in the README.
  2. What license governs the part you will ship? Code, weights, and sometimes training data each carry their own terms.
  3. Can your hardware run it at the precision the project measured? Model card figures come with assumptions attached.
  4. How many people would have to leave before it stalls? That is the bus factor.
  5. What does it expose when you self-host it? Bind addresses, authentication, plugins, and model file formats.

For commercial and hosted products, our broader guide on how to evaluate AI developer tools is the better starting point.

Maintenance signals that actually predict abandonment

Stars are the weakest signal on a repository page. Researchers from Carnegie Mellon, North Carolina State University, and Socket built a detector called StarScout and flagged roughly six million suspected fake GitHub stars between 2019 and 2024, in a paper accepted to ICSE 2026. Most of those stars promoted short-lived phishing and malware repositories, and AI and LLM projects were one of the main categories promoted with the rest. Stars measure attention, not upkeep.

Better signals, roughly in order of usefulness:

  • The README status line. Projects increasingly say plainly when they are winding down. AutoGen states that it is in maintenance mode, will not receive new features, and points new users to Microsoft Agent Framework. Fooocus says it is in limited long-term support with bug fixes only and has no current plans to adopt newer model architectures. DeepSpeech was archived in June 2025 under a one-line "discontinued" notice.
  • Commits on the default branch, not "last pushed". GitHub's push timestamp counts activity on any branch. AUTOMATIC1111 shows why that matters: its master branch has not moved since the v1.10.1 tag in July 2024, while the dev branch still takes occasional merges, most recently in March 2026. A glance at "updated recently" hides the first fact.
  • Where development actually moved. When the company behind a project disappears, work often continues in a fork under a new package name. Coqui shut down in early 2024, the original repository's last commit dates from February 2024, and the maintained continuation of Coqui TTS is the Idiap fork, published on PyPI as coqui-tts. Run pip install TTS today and you still get version 0.22.0 from December 2023.
  • OpenSSF Scorecard. Its Maintained check awards full marks only for at least one commit per week over the previous 90 days and gives archived projects the lowest score. Other checks cover code review, signed releases, and a published security policy. Pre-computed results exist for only some repositories, so for most AI projects you will run it yourself.
# Archive status, last push, and detected license
gh api repos/coqui-ai/TTS --jq '{archived, pushed_at, license: .license.spdx_id}'

# Last commit on the default branch (more honest than pushed_at)
gh api 'repos/coqui-ai/TTS/commits?per_page=1' --jq '.[0].commit.committer.date'

# Scorecard checks that matter most for an adoption decision
export GITHUB_AUTH_TOKEN=<your token>
scorecard --repo=github.com/ollama/ollama --checks=Maintained,Code-Review,Signed-Releases,Security-Policy

Our retirement pass used roughly two years of inactivity as its cutoff, which suits a directory that also serves people researching older work. For a dependency inside your own product, set the bar higher: a year without a release on the default branch is a reason to inspect the fork network before you commit.

License traps: read the file, not the badge

GitHub's sidebar license is detected from a single file, and AI projects routinely spread code, weights, and data across several. For many important projects GitHub simply reports "Other". Sometimes it is actively misleading: the AutoGen repository shows CC-BY-4.0 because its LICENSE file covers documentation, while the code is MIT under a separate LICENSE-CODE file. The real terms live in the license file next to whatever you will actually run.

ProjectWhat the repo page suggestsWhat governs the part you shipPractical effect
F5-TTSMITPretrained models are CC-BY-NC because of the Emilia training dataCode is free to reuse; the released checkpoints are non-commercial
FLUX.1 [dev]Apache-2.0 on the inference repoFLUX.1 [dev] Non-Commercial License on the weightsFLUX.1 [schnell] is Apache 2.0 and the safer commercial choice
XTTS via Coqui TTSMPL-2.0Coqui Public Model License on the XTTS weights, non-commercialThe company that sold commercial licenses no longer exists
HunyuanVideo 1.5OtherTencent Hunyuan Community LicenseDoes not apply in the EU, UK, or South Korea; separate license above 100M monthly active users
Devstral 2 (123B)Modified MITNo rights at all if your company's monthly revenue exceeds USD 20 millionDevstral Small 2 is plain Apache 2.0
Llama 4Llama 4 Community License700M monthly active user threshold, "Built with Llama" display, "Llama" prefix on derivative model namesMultimodal models carry an EU restriction for developers, not end users
Ultralytics YOLOAGPL-3.0AGPL-3.0, or a paid Enterprise licenseClosed commercial products need the Enterprise license or full AGPL compliance
Open WebUIOtherBSD-3 plus a branding clause since v0.6.6 (May 2025)Removing branding needs 50 or fewer users in 30 days, written permission, or an enterprise license
PiperMIT, archived October 2025Maintained line moved to piper1-gpl under GPL-3.0Upgrading to the maintained code changes your obligations

Three patterns cover most of the traps:

  • Split licensing. Code under a permissive license, weights under a non-commercial or custom one. Always open the model card and the LICENSE file in the weights repository, not just the code repository.
  • Thresholds and territories. Revenue caps (Devstral 2), user caps (Llama 4, Tencent Hunyuan), and geographic exclusions (Hunyuan, Llama 4 multimodal) never show up in a one-word license label. Kimi K2 uses a modified MIT license that requires a visible "Kimi K2" credit in products above 100 million monthly active users or 20 million US dollars in monthly revenue.
  • Source-available platforms. n8n ships under its Sustainable Use License, which limits use to internal business, personal, or non-commercial purposes. Dify's license is Apache 2.0 with added conditions: no multi-tenant operation without written authorization, and no removing the logo or copyright information from the console. Both are fine to self-host for your own team and a problem if you plan to resell them.

Licenses also move in the other direction. Gemma 4 dropped Google's custom Gemma terms for Apache 2.0 in April 2026. Recheck the license at every major version, because it can loosen as easily as it tightens.

To pull a license file straight from Hugging Face without opening a browser:

curl -sL https://huggingface.co/mistralai/Devstral-2-123B-Instruct-2512/raw/main/LICENSE | head -20

Hardware floors: the number comes with assumptions

Every hardware figure on a model card assumes a precision, often a resolution or context length, and sometimes CPU offloading. Read the assumption before the number.

ModelFloor stated by the projectAssumption attached
Whisper largeAbout 10 GB VRAMOriginal implementation; the turbo model needs about 6 GB
Wan 2.1 T2V-1.3B8.19 GB VRAM480P output; about 4 minutes per 5-second clip on an RTX 4090 without quantization
HunyuanVideo 1.5 (8.3B)14 GB GPU memoryModel offloading enabled
gpt-oss-20b16 GB of memoryMoE weights quantized to MXFP4
gpt-oss-120bOne 80 GB GPUSame MXFP4 quantization
Llama 4 Scout (109B total, 17B active)One H100On-the-fly int4 quantization of the BF16 weights

Two lessons fall out of the table. First, "fits" and "usable" are different floors: the small Wan 2.1 model fits on most consumer cards but still needs minutes per clip on a 4090. Second, mixture-of-experts models are sized by total parameters and timed by active ones. Llama 4 Scout computes with 17B parameters per token but has to hold all 109B in memory.

When a project gives no figure, estimate the weights yourself: parameters times bytes per parameter, then add headroom for context, activations, and any text encoders. FLUX.1 [dev] has 12B parameters, so its transformer alone is about 24 GB in BF16 before the rest of the pipeline loads, which is why consumer setups lean on quantized variants.

python -c "p=12e9; print({k: p*b/1e9 for k,b in {'bf16':2,'int8':1,'int4':0.5}.items()})"
# {'bf16': 24.0, 'int8': 12.0, 'int4': 6.0}

For matching models to a specific card, see our guide to local LLM setups by GPU budget.

Bus factor: count people, not commits

CHAOSS, the Linux Foundation project that defines open source health metrics, renamed the bus factor to the Contributor Absence Factor: the smallest number of contributors responsible for half of all contributions. A factor of one means a single person carries the project.

text-generation-webui, which now lives at oobabooga/textgen on GitHub, is a clear case. Of roughly 5,700 commits credited to its listed contributors, about 4,700 come from the founder's account, around 82 percent. That is not a reason to avoid it; it has been in active development since December 2022. It is a reason to know your exit path, such as which other front ends load the same model formats.

gh api 'repos/oobabooga/textgen/contributors?per_page=5' --jq '.[] | [.login, .contributions] | @tsv'

Company-backed projects fail differently. The risk is less one person leaving than a priority change: Coqui shutting down, or Microsoft moving AutoGen to maintenance mode in favor of a successor. For those, the useful questions are whether the license lets the community fork (Coqui's MPL-2.0 code is what made the Idiap fork possible), whether outside maintainers have merge rights, and whether releases are cut by more than one person.

Security of self-hosted AI stacks

Self-hosting moves the attack surface onto your own machines. Five checks catch most of the real incidents:

  • Bind address and authentication. Ollama binds to 127.0.0.1:11434 by default and has no built-in authentication; its FAQ shows a reverse proxy for remote access. Setting OLLAMA_HOST=0.0.0.0 without one is how SentinelLABS and Censys came to report about 175,000 publicly reachable Ollama hosts across 130 countries in January 2026, more than 48 percent of them advertising tool calling. ComfyUI also defaults to 127.0.0.1, but its --listen flag with no argument binds every interface.
  • Unauthenticated code paths in dev tools. Langflow before 1.3.0 executed user-supplied Python through an unauthenticated validation endpoint (CVE-2025-3248, CVSS 9.8), and CISA added it to the Known Exploited Vulnerabilities catalog in May 2025. MCP Inspector before 0.14.1 let a malicious web page run commands on a developer's machine (CVE-2025-49596, CVSS 9.4). Tools built for localhost debugging are rarely hardened for anything else.
  • Plugins are code. In June 2024 a ComfyUI custom node called ComfyUI_LLMVISION shipped code that stole browser credentials and sent them to an attacker-controlled Discord server. Treat any plugin ecosystem as arbitrary code execution and install only what you have checked.
  • Model files can be code too. Pickle-based checkpoints can execute arbitrary functions when loaded. PyTorch 2.6 changed torch.load to default to weights_only=True. Prefer safetensors files, and do not flip that flag back for files you did not produce.
  • Supply chain. In December 2024, four Ultralytics releases on PyPI (8.3.41, 8.3.42, 8.3.45, and 8.3.46) shipped an XMRig cryptominer after attackers abused the project's GitHub Actions build cache and a leftover PyPI token. The malicious code never appeared in the GitHub source. Pin versions and verify hashes.
# What is listening beyond localhost? (Ollama, Gradio-style UIs, ComfyUI)
ss -ltnp | grep -E ':(11434|7860|8188)'

# Install only exact, hash-verified dependencies
pip install --require-hashes -r requirements.txt

For agents and MCP servers, isolate rather than trust. ToolHive runs each MCP server in its own container with secrets management and network isolation, and garak probes a model endpoint for prompt injection, data leakage, and jailbreaks before you put it in front of users. Our longer piece on securing self-hosted AI stacks goes further on the rest of the setup.

Common questions

Is an AI model with an MIT-licensed repository safe for commercial use? Not automatically. The repository license usually covers the code, while the weights often carry their own license in the model card or weights repository. F5-TTS, for example, is MIT code with CC-BY-NC checkpoints.

How can I tell if an open source AI project is abandoned? Check the README for a status notice, the date of the last commit on the default branch, the last tagged release, and whether development has moved to a fork. A project with no default-branch commits for a year and unanswered security issues is a migration candidate.

What is the difference between open weights and open source AI? The Open Source Initiative published its Open Source AI Definition 1.0 on October 28, 2024. It requires the freedom to use, study, modify, and share the system for any purpose, along with code, parameters, and detailed information about the training data. Models released with usage, revenue, or territory limits are open weights, not open source in that sense.

Is it safe to expose Ollama to the internet? Not on its own. Ollama has no built-in authentication, so keep it on 127.0.0.1 and put an authenticating reverse proxy or a VPN in front of it if you need remote access.

How much VRAM do I need to run a model locally? Start with parameters times bytes per parameter: 2 bytes for BF16, 1 for 8-bit, and about 0.5 for 4-bit. Add headroom for context and activations, and check which precision the model card's own figure assumed.

Related Tools

More Articles