llama.cpp
Category: Local
| Detail | Value |
|---|---|
| Models | C++ inference engine |
| Context | Varies |
| Pricing | Free |
Overview
Highly optimized C++ implementation for running LLMs on consumer hardware. Supports quantized models (GGUF format) for reduced memory usage. The foundation that most local AI tools build on.
