All payments in the preview are in test mode. Read more
← Index

Local & Uncensored · verified 2026-07-27

llama.cpp

The inference engine everything else is built on.

Georgi Gerganov's C/C++ implementation of LLM inference. Quantised GGUF models run on laptops, phones and servers, and most local tools embed it under the hood.

OPEN LLAMA.CPP
Rating
4.7 / 5
Pricing
Open source
Best for
Self-hosting, Offline, Cost control
Platforms
Linux, macOS, Windows, CLI, API
Runs locally
Yes
Content policy
Standard filters

Matching filters

What stands out

  • GGUF quantisation
  • CPU, CUDA, Metal and Vulkan back ends
  • MIT licence

Also in Local & Uncensored