Technology subcategory
Serving
Inference servers and engines that run LLMs in production.
Serving packs (vLLM, TGI, TensorRT-LLM, Ollama, …) own tokens/second, batching, and hardware utilization.
Apps and gateways call them; agents usually should not replace them. See LLM Gateway, Fine-tuning.
See also
Technologies
Features
Stacks
GitHub in this term
Primary repositories linked from member packs and devices.