Technology · LLMs → Serving · Open Source · wiki:stub
This pack is still a wiki stub. Opening text below if present; full article bands appear after deepen (`Wiki-status: deep`).
llama.cpp runs LLMs efficiently on consumer and edge hardware via quantized C/C++ inference. Underpins many local model deployments. Enables private/on-prem workers without GPU clouds — trade throughput for data locality. LLM inference in C/C++ ggml / ops / maintainer PRs%20sort%3Aupdated-desc) / dev stats / lib llama API / llama-server REST API - Visit https://llama.app and follow the instructions - Run with Docker - see our Docker documentation - Download pre-built binaries from the releases page - Build from source by cloning this repository.
9 tags · 10 out · 11 in · 3 artifacts · 1 gaps · 14 corpus docs
map_edge · 5
tech_features · 1
tech_quote · 7
tech_readme · 1