Featured Tool

llama.cpp

Port of Meta's LLaMA model in C/C++ for efficient CPU inference

Open SourceSelf HostedOffline Capable
0.0 (0)

About

llama.cpp by Georgi Gerganov is a C and C++ inference engine for LLaMA-family and many other transformer language models, designed to run with minimal setup on a wide range of hardware including CPU-only laptops. It supports the GGUF quantized model format, multiple backends (CUDA, Metal, Vulkan, ROCm, BLAS), a server with an OpenAI-compatible API, and bindings for many languages. MIT licensed; the substrate for much of the local LLM ecosystem.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MIT
Added
Jan 29, 2026

Related Tools

Open-source ChatGPT alternative that runs 100% offline on your computer.

Open SourceSelf HostedOffline
Beginner

Fast LLM inference on consumer GPUs using neuron-aware sparse computation.

Open SourceSelf HostedOfflineGPU 4GB+
Advanced
Featured

High-throughput LLM serving engine with PagedAttention

Open SourceSelf HostedOfflineGPU 16GB+
Intermediate

Easy-to-use local AI inference with built-in web UI and API.

Open SourceSelf HostedOffline
Beginner

High-performance LLM inference engine forked from vLLM with extra features.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate

Minimalist machine learning framework for Rust focused on performance and serverless inference.

Open SourceSelf HostedOffline
Intermediate
Browse all LLM Inference & Serving tools