llama.cpp
Lightweight, pure C/C++ LLM inference with minimal setup and top performance on any hardware.
High-performance LLM inference engine in C/C++ with minimal dependencies, supporting quantized models (1.5–8 bit) and diverse hardware (Apple Silicon, CUDA, Vulkan, etc.).