Technology · LLMs → Serving · Open Source · wiki:stub
This pack is still a wiki stub. Opening text below if present; full article bands appear after deepen (`Wiki-status: deep`).
TensorRT-LLM (NVIDIA) optimizes LLM inference on NVIDIA GPUs with highly tuned kernels and serving integrations. Performance path for NVIDIA GPU estates. Choose when maximizing tokens/sec on datacenter GPUs matters more than portability. TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a.
9 tags · 10 out · 11 in · 3 artifacts · 1 gaps · 13 corpus docs
map_edge · 5
tech_features · 1
tech_quote · 6
tech_readme · 1