Technology · LLMs → Serving · Open Source · wiki:stub

TensorRT-LLM

This pack is still a wiki stub. Opening text below if present; full article bands appear after deepen (`Wiki-status: deep`).

TensorRT-LLM (NVIDIA) optimizes LLM inference on NVIDIA GPUs with highly tuned kernels and serving integrations. Performance path for NVIDIA GPU estates. Choose when maximizing tokens/sec on datacenter GPUs matters more than portability. TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a.

Research inventory

9 tags · 10 out · 11 in · 3 artifacts · 1 gaps · 13 corpus docs

Catalog tags

landscape.layer
LLMs
landscape.subcategory
Serving
license_tag
Open Source
maps.dm
absent
maps.features
5
one_liner
LLMs
review.depth
science
slug
tensorrt-llm
title
TensorRT-LLM

Artifacts

  • dm_map · absent
  • features_map · present · technologies/tensorrt-llm/features.md
  • readme · present · technologies/tensorrt-llm/README.md

Out · alternative_to

Out · maps_to

In · alternative_to

In · in_stack

In · maps_to

Corpus tags

category
LLMs → Serving
dedication
open-source
feature
bring-your-own-model
cost-latency-ops
deployment-data-residency
model-flexibility-routing
self-host
kind
map_edge
tech_features
tech_quote
tech_readme
needs_deepen
false
quality
ok
slug
tensorrt-llm
source_id
trtllm-docs
trtllm-github
vllm-contrast
technology
tensorrt-llm

Corpus documents (13)

map_edge · 5

  • tensorrt-llm → bring-your-own-model
  • tensorrt-llm → cost-latency-ops
  • tensorrt-llm → deployment-data-residency
  • tensorrt-llm → model-flexibility-routing
  • tensorrt-llm → self-host

tech_features · 1

  • TensorRT-LLM · features

tech_quote · 6

  • TensorRT-LLM · trtllm-docs
  • TensorRT-LLM · trtllm-docs
  • TensorRT-LLM · trtllm-github
  • TensorRT-LLM · trtllm-github
  • TensorRT-LLM · vllm-contrast
  • TensorRT-LLM · vllm-contrast

tech_readme · 1

  • TensorRT-LLM