Technology · LLMs → Serving · Open Source · wiki:stub

vLLM

This pack is still a wiki stub. Opening text below if present; full article bands appear after deepen (`Wiki-status: deep`).

vLLM is a leading open-source high-throughput LLM serving engine (PagedAttention, continuous batching) with broad model support. Often the default for self-hosted agent backends. Pair with an LLM gateway so agents stay provider-agnostic. Teams adopting vLLM should verify auth, routing, and tool/provider boundaries against their real traffic shape, including failure modes when upstream models or tools stall.

Research inventory

9 tags · 11 out · 12 in · 3 artifacts · 1 gaps · 19 corpus docs

Catalog tags

landscape.layer
LLMs
landscape.subcategory
Serving
license_tag
Open Source
maps.dm
absent
maps.features
6
one_liner
LLMs
review.depth
science
slug
vllm
title
vLLM

Artifacts

  • dm_map · absent
  • features_map · present · technologies/vllm/features.md
  • readme · present · technologies/vllm/README.md

Out · alternative_to

Out · maps_to

In · alternative_to

In · in_stack

In · maps_to

Corpus tags

category
LLMs → Serving
dedication
open-source
feature
bring-your-own-model
cost-latency-ops
deployment-data-residency
model-flexibility-routing
open-source
self-host
kind
map_edge
tech_features
tech_quote
tech_readme
needs_deepen
false
quality
ok
slug
vllm
source_id
acm-3600006
pagedattention-sosp23
vllm-github
technology
vllm

Corpus documents (19)

map_edge · 6

  • vllm → bring-your-own-model
  • vllm → cost-latency-ops
  • vllm → deployment-data-residency
  • vllm → model-flexibility-routing
  • vllm → open-source
  • vllm → self-host

tech_features · 1

  • vLLM · features

tech_quote · 11

  • vLLM · acm-3600006
  • vLLM · acm-3600006
  • vLLM · acm-3600006
  • vLLM · pagedattention-sosp23
  • vLLM · pagedattention-sosp23
  • vLLM · pagedattention-sosp23
  • vLLM · pagedattention-sosp23
  • vLLM · vllm-github
  • … +3 more

tech_readme · 1

  • vLLM