Skip to main content
AdOpenFree logoPromote your productReach more potential users and drive product growth and revenue.Advertise
Favicon of llama.cpp

llama.cpp

Free Listing

LLM inference in C/C++ with minimal setup and state-of-the-art performance on a wide range of hardware, locally and in the cloud.

Visit ggml-org/llama.cpp

Overview

llama.cpp enables LLM and VLM inference with minimal setup and state-of-the-art performance on a wide range of hardware, locally and in the cloud. It is built on top of the ggml library.

Key Features

  • Plain C/C++ implementation without any dependencies
  • Optimized for Apple silicon via ARM NEON, Accelerate, and Metal frameworks
  • Integer quantization from 1.5-bit to 8-bit for faster inference and reduced memory use
  • Custom CUDA kernels for NVIDIA GPUs, with AMD via HIP and Moore Threads via MUSA
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference for models larger than VRAM

Use Cases

  • Run LLM and VLM models locally from Hugging Face
  • Launch an OpenAI-compatible API server
  • Use the built-in web UI for interactive sessions
  • Accelerate large models with hybrid CPU+GPU inference

Getting Started

  • Visit https://llama\.app and follow the instructions
  • Run with Docker (see Docker documentation)
  • Download pre-built binaries from the releases page
  • Build from source by cloning the repository (see build guide)
  • Example: llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

Deployment & Requirements

  • Supports a wide range of hardware backends including CPU, NVIDIA GPU (CUDA), AMD GPU (HIP), Apple Silicon (Metal), Intel GPU (SYCL), Vulkan, WebGPU, and more
  • No external dependencies for the core implementation
  • Can be deployed via Docker or pre-built binaries

Before You Adopt

  • License: MIT. Review its terms before using, modifying, or distributing the project.
  • The OpenVINO backend is marked as in progress, indicating incomplete support.

Comments

Sign in to leave a comment.

More like llama.cpp

Favicon of DeepSeek-V3

DeepSeek-V3

Free ListingStars: 104.5K

Mixture-of-Experts language model with 671B parameters

Language Models

DeepSeek-V3 is a Mixture-of-Experts language model with 671B total parameters and 37B activated per token. It supports a 128K context and is available in Base and Chat variants with open weights for local deployment and API use.

Favicon of Qwen3-VL

Qwen3-VL

Free ListingStars: 20K

Multimodal large language model series by Qwen team

Language Models

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud, offering dense and MoE architectures, advanced visual reasoning, long context, and expanded OCR.

Favicon of Qwen3

Qwen3

Free ListingStars: 27.7K

Large language model series by the Qwen team, Alibaba Cloud

Language Models

Qwen3 is the large language model series from Alibaba Cloud's Qwen team. It offers dense and Mixture-of-Experts models with thinking and non-thinking modes, multilingual, long-context, reasoning, coding, and agent capabilities.