Technology subcategory

Serving

Inference servers and engines that run LLMs in production.

Serving packs (vLLM, TGI, TensorRT-LLM, Ollama, …) own tokens/second, batching, and hardware utilization.

Apps and gateways call them; agents usually should not replace them. See LLM Gateway, Fine-tuning.

See also

Technologies

Features

Stacks

GitHub in this term

Primary repositories linked from member packs and devices.