llama.cpp
LLM inference in C/C++ with minimal setup and state-of-the-art performance on a wide range of hardware, locally and in the cloud.
Visit ggml-org/llama.cppOverview
llama.cpp enables LLM and VLM inference with minimal setup and state-of-the-art performance on a wide range of hardware, locally and in the cloud. It is built on top of the ggml library.
Key Features
- Plain C/C++ implementation without any dependencies
- Optimized for Apple silicon via ARM NEON, Accelerate, and Metal frameworks
- Integer quantization from 1.5-bit to 8-bit for faster inference and reduced memory use
- Custom CUDA kernels for NVIDIA GPUs, with AMD via HIP and Moore Threads via MUSA
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference for models larger than VRAM
Use Cases
- Run LLM and VLM models locally from Hugging Face
- Launch an OpenAI-compatible API server
- Use the built-in web UI for interactive sessions
- Accelerate large models with hybrid CPU+GPU inference
Getting Started
- Visit https://llama\.app and follow the instructions
- Run with Docker (see Docker documentation)
- Download pre-built binaries from the releases page
- Build from source by cloning the repository (see build guide)
- Example: llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
Deployment & Requirements
- Supports a wide range of hardware backends including CPU, NVIDIA GPU (CUDA), AMD GPU (HIP), Apple Silicon (Metal), Intel GPU (SYCL), Vulkan, WebGPU, and more
- No external dependencies for the core implementation
- Can be deployed via Docker or pre-built binaries
Before You Adopt
- License: MIT. Review its terms before using, modifying, or distributing the project.
- The OpenVINO backend is marked as in progress, indicating incomplete support.