Technology · LLMs → Serving · Open Source · wiki:stub
This pack is still a wiki stub. Opening text below if present; full article bands appear after deepen (`Wiki-status: deep`).
Text Generation Inference (TGI) is Hugging Face’s production server for serving HF models with batching, streaming, and tracing. Standard self-host option in HF-centric stacks. Serve open models behind HF/OpenAI-compatible APIs for agents. Large Language Model Text Generation Inference [!CAUTION] text-generation-inference is now in maintenance mode. Going forward, we will accept pull requests for minor bug fixes, documentation improvements and lightweight maintenance tasks. TGI has initiated the movement for optimized inference engines to rely on a transformers model architectures. This approach is now adopted by downstream inference engines, which we contribute to and recommend using going forward: vllm, SGLang, as well as local engines with inter-compatibility such as llama.cpp or MLX. A Rust, Python and gRPC server for text generation inference. Used in production.
9 tags · 10 out · 11 in · 3 artifacts · 1 gaps · 16 corpus docs
map_edge · 5
tech_features · 1
tech_quote · 9
tech_readme · 1