Local & Uncensored · verified 2026-07-27
llama.cpp
The inference engine everything else is built on.
Georgi Gerganov's C/C++ implementation of LLM inference. Quantised GGUF models run on laptops, phones and servers, and most local tools embed it under the hood.
OPEN LLAMA.CPP- Rating
- 4.7 / 5
- Pricing
- Open source
- Best for
- Self-hosting, Offline, Cost control
- Platforms
- Linux, macOS, Windows, CLI, API
- Runs locally
- Yes
- Content policy
- Standard filters
Matching filters
What stands out
- GGUF quantisation
- CPU, CUDA, Metal and Vulkan back ends
- MIT licence