If you're choosing between vLLM and SGLang, start with vLLM unless your traffic looks like an agent: lots of requests that share a long prefix, such as a system prompt, tool definitions or a growing conversation. That's where SGLang's design pays off. For everything else the two are close enough that published benchmarks disagree about which one wins, and the honest advice is to test your own workload.

Both are open source under Apache 2.0, both expose an OpenAI-compatible API, and both ship a new release roughly every two weeks. As of September 28, 2026 the current versions are vLLM 0.30.0 and SGLang 0.5.20. Below is what each one does, how they handle prompt caching differently, what the benchmarks really show (and why I wouldn't trust any single one), the commands to run each, and where I'd start.

The Quick Verdict

vLLM

The safe default. It documents 200+ model architectures and the longest hardware list, and it has the larger contributor base.

SGLang

Worth it for agents, multi-turn chat and RAG with long shared prefixes, where prefix reuse decides your latency.

Mostly unique prompts

Either one. In one benchmark the two landed within 4% of each other on throughput. Pick on ecosystem and hardware.

Not sure?

Run both on your real prompts. Each is a single install away, and an afternoon of testing beats any blog post, including this one.

Apache 2.0 Licence for both vLLM and SGLang
~2 weeks Gap between releases, for both projects
3 tests Published benchmarks below, and no single winner across them

How to Choose

Five questions decide most of it, and you can answer them from your own logs before you run a single benchmark.

  1. How much of each prompt is shared between requests? A fixed system prompt, tool definitions or a RAG document that repeats across calls favours SGLang's radix tree. Mostly unique prompts mean caching barely matters.
  2. Is it one long conversation or many short calls? Multi-turn chats and agent loops keep growing a prefix that could be reused. One-shot completions don't.
  3. What hardware do you have? Both cover NVIDIA and AMD GPUs. vLLM's list is longer, with plugins for Intel Gaudi and Apple Silicon. SGLang lists Intel Xeon CPUs, Google TPUs and Ascend NPUs.
  4. Is your model supported today? vLLM documents 200+ architectures. SGLang lists the major families and says it works with most Hugging Face models. Check your exact model on both, especially a brand-new release. If you are still picking a model, our list of the best open-source LLMs to self-host is a good start.
  5. How much churn can you take? Both release about every two weeks, and both remove features. If you can't afford to retest after an upgrade, pin a version and upgrade on your own schedule.

1. vLLM: The Broad Default

  • Inference server
  • Apache 2.0
  • Self-hostable

vLLM started at UC Berkeley's Sky Computing Lab and is now maintained by a community of academic institutions and companies, with more than 2,000 contributors. It's built around PagedAttention, which manages the KV cache in fixed-size blocks, plus continuous batching and chunked prefill. Our Ollama vs vLLM comparison covers where it fits against a single-user tool.

Why choose it

  • Documents support for 200+ model architectures on Hugging Face, including mixture-of-experts, multimodal and embedding models.
  • Runs on NVIDIA, AMD and Intel GPUs and on CPUs, with plugins for Google TPUs, Intel Gaudi, Huawei Ascend and Apple Silicon.
  • Serves the OpenAI API, and also the Anthropic Messages API and gRPC.
  • Structured output through xgrammar or guidance, multi-LoRA, and tensor, pipeline, data, expert and context parallelism.

Watch out for

  • Prefix caching works on full blocks only. A shared prefix that ends mid-block gets partly recomputed.
  • Releases move fast and break things. Version 0.30.0 removed GPTQ activation ordering and made scale-out endpoints opt-in behind a flag.

Licence

Apache 2.0, free to use commercially. You pay for the GPUs, not per token. vLLM docs

2. SGLang: Built Around Prefix Reuse

  • Inference server
  • Apache 2.0
  • Self-hostable

SGLang is a serving framework for large language and multimodal models, developed under the LMSYS non-profit. Its centrepiece is RadixAttention, which keeps cached prefixes in a radix tree, alongside a zero-overhead CPU scheduler and prefill-decode disaggregation. That design suits agent workloads like the ones in our WebRTC for AI voice agents guide.

Why choose it

  • RadixAttention tracks shared prefixes automatically, including branching conversations.
  • Runs on NVIDIA (including GB200, B300, H100 and A100), AMD (MI355 and MI300), Intel Xeon CPUs, Google TPUs and Ascend NPUs.
  • Supports Llama, Qwen, DeepSeek, Kimi, GLM, GPT, Gemma and Mistral, plus embedding and diffusion models.
  • Structured outputs, speculative decoding, multi-LoRA batching and tensor, pipeline, expert and data parallelism.

Watch out for

  • The install docs say it requires CUDA 13. CUDA 12 wheels ended after version 0.5.19, so check your driver first.
  • Also fast-moving. Version 0.5.20 removed the first version of prefill context parallelism.
  • The adoption numbers on its README (trillions of tokens a day, 400,000+ GPUs) are the project's own claims.

Licence

Apache 2.0, free to use commercially. You pay for the GPUs. SGLang docs

How They Cache Prompts

This is the real technical difference, and the reason the two are usually pitched at different workloads.

vLLM hashes each block of KV cache using the tokens in the block and the tokens before it. Its prefix caching design docs describe it as a hash-based system with least-recently-used eviction, and they're explicit about a limit: only full blocks are cached. SGLang's RadixAttention stores cached sequences in a radix tree instead, so a new request finds the longest matching prefix wherever it branches, without waiting for a block boundary.

In practice that matters when many requests start the same way. An agent that sends the same 2,000-token system prompt and tool definitions on every call, or a chat that appends a turn each time, reuses a lot of work. If most of your prompts are unique, there's little to reuse and the difference mostly disappears.

What the Benchmarks Say

I looked at three published head-to-head tests and SGLang's own release notes. They don't agree, and the reasons why are more useful than any single number.

Source Setup Result Caveat
Spheron Llama 3.3 70B FP8, H100, vLLM 0.18.0 vs SGLang 0.5.9 Unique prompts: 1,850 vs 1,920 tok/s. With an 80% shared prefix, median time to first token 310 ms vs 195 ms GPU cloud vendor. Main runs didn't have vLLM prefix caching on
Runpod DeepSeek-R1-Distill-Llama-70B, 2x H100, multi-turn at 7k context SGLang 35.0 vs vLLM 32.8 tok/s. On a single-turn prompt vLLM was 1.1x faster GPU cloud vendor. Doesn't say whether vLLM prefix caching was on
Jarvislabs Qwen models, H100, May 2026, ShareGPT chat workload Qwen2.5-7B: vLLM 23,523 vs SGLang 16,787 tok/s. Qwen3-30B-A3B: 8,607 vs 7,247 GPU cloud vendor. Framework versions not recorded
SGLang v0.5.20 notes SGLang against itself, workload not stated Unified radix tree: cache hit rate 43.8% to 60.8%, mean time to first token 1.57 s to 1.07 s Self-reported. Compares SGLang versions, not the two engines

Figures as published by each source. I haven't re-run any of them.

Splitno single winner

SGLang came out ahead where prompts shared a prefix or the conversation ran over several turns. vLLM came out ahead on the Qwen throughput tests and on a single-turn prompt. The Jarvislabs authors say it themselves: these are results from the configurations tested, not a universal ranking.

Two more caveats. Every test above ran on versions that are now well behind the current 0.30.0 and 0.5.20, and every source sells GPU time. Treat the numbers as a hint about which workloads to test first, not as a verdict.

vLLM vs SGLang Side by Side

Facts from each project's own README and docs, checked on September 28, 2026.

Feature vLLM SGLang
LicenceApache 2.0Apache 2.0
Current version0.30.00.5.20
Origin and stewardUC Berkeley Sky Computing Lab, community-maintainedLMSYS non-profit
Prompt cache designHash-based blocks, full blocks only, LRU evictionRadixAttention, radix tree of cached prefixes
QuantizationFP8, INT8, INT4, GPTQ, AWQ, GGUF and moreFP4, FP8, INT4, AWQ, GPTQ
Speculative decodingYes (n-gram, suffix, EAGLE, DFlash)Yes
Prefill and decode disaggregationYesYes
Structured outputsYes (xgrammar or guidance)Yes
Multi-LoRAYesYes
APIsOpenAI, Anthropic Messages, gRPCOpenAI-compatible
HardwareNVIDIA, AMD, Intel GPUs, CPUs, plus TPU, Gaudi, Ascend, Apple pluginsNVIDIA, AMD, Intel Xeon, TPU, Ascend
Model coverage200+ architectures documentedMajor families listed, "compatible with most Hugging Face models"
Default port800030000
Install requirementPython 3.10 to 3.13Python 3.10+, CUDA 13

Most of the feature rows are ticks on both sides, and that's the point. The differences that matter are the cache design, hardware and how each behaves on your traffic.

Running Each One

Both install with pip or uv and start an OpenAI-compatible server with one command. This is vLLM, from its quickstart:

uv pip install vllm --torch-backend=auto
vllm serve Qwen/Qwen2.5-1.5B-Instruct
# serves on http://localhost:8000

And this is SGLang, from its quickstart. It also publishes an official lmsysorg/sglang Docker image if you'd rather not manage CUDA wheels yourself:

uv pip install --prerelease=allow sglang
sglang serve --model-path qwen/qwen2.5-0.5b-instruct --host 0.0.0.0 --port 30000
# serves on http://localhost:30000

Because both speak the OpenAI API, a client only needs a different base URL. That makes a side-by-side test cheap:

from openai import OpenAI

# vLLM listens on 8000, SGLang on 30000
client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    messages=[{"role": "user", "content": "List 3 countries and their capitals."}],
    max_tokens=64,
)
print(response.choices[0].message.content)

When you compare them, replay your real prompts at your real concurrency, and make sure prefix caching is actually on for vLLM. One of the benchmarks above ran vLLM without it in its main tests and another doesn't say, so it's a real variable.

Track four numbers: median and p95 time to first token, tokens per second at your peak concurrency, the prefix cache hit rate if the engine exposes it, and GPU memory headroom. Tokens per second alone hides the thing that agent users actually feel, which is how long they wait before anything appears.

Keeping Up With Releases

Both projects release about every two weeks, and both break things between releases. Version 0.30.0 of vLLM removed GPTQ activation ordering and put its scale-out endpoints behind a flag. SGLang 0.5.20 retired the first version of prefill context parallelism, and its CUDA 12 builds ended after 0.5.19.

None of that is a criticism. It's what a fast-moving inference stack looks like. It does mean you should pin an exact version in production, read the release notes (vLLM, SGLang) before every upgrade, and keep your previous image around so you can roll back.

Self-Hosting With Meetrix

  • vLLM on AWS and Google Cloud
  • NVIDIA drivers and CUDA included
  • OpenAI-compatible API

To be upfront: we package vLLM, not SGLang. Our vLLM image launches on an AWS or Google Cloud GPU with the NVIDIA drivers, CUDA and an OpenAI-compatible API already set up, so you pay for the GPU instead of per token. The vLLM developer guide for AWS and the GCP version walk through it.

If SGLang turns out to fit your workload better, nothing here stops you running it. It's Apache 2.0, it has its own official Docker image, and the same GPU instance works for either. If you're still choosing a model to serve, our list of the best self-hosted LLMs is the place to start, and the self-hosted voice AI agent stack shows where an inference engine fits in a full agent build.

Where I'd Start

vLLM

If you serve a mix of workloads, run on more than one kind of hardware, or just want the option with the largest ecosystem: start here.

SGLang

If you're building agents, multi-turn chat or RAG with long shared prompts: test SGLang first, on prefix-heavy traffic.

Both, briefly

If latency is your product: run both for a day on real traffic and compare median and p95 time to first token, not just tokens per second.

Neither

If you're the only user on a laptop: Ollama is simpler. See the Ollama vs vLLM comparison for when to move.

Already running one of them and happy with the latency? Stay. Switching inference engines means retesting your models, your quantization and your client settings, and the gap between these two rarely justifies that unless prefix reuse is your bottleneck. Revisit when your traffic changes shape or your GPU bill does.

Frequently Asked Questions

What is the difference between vLLM and SGLang?

Both are Apache 2.0 inference servers with OpenAI-compatible APIs. The main difference is prompt caching: vLLM hashes fixed-size blocks, while SGLang's RadixAttention uses a radix tree that finds the longest shared prefix. That suits agents and multi-turn chat.

Is SGLang faster than vLLM?

It depends on the workload. Published benchmarks show SGLang ahead on prefix-heavy and multi-turn traffic, and vLLM ahead on some Qwen throughput and single-turn tests. Test your own prompts on current versions before deciding.

Which is better for production, vLLM or SGLang?

Both run in production. vLLM documents broader model and hardware support, while SGLang reports large-scale deployments. Pin an exact version, read the release notes before upgrading, and choose based on load tests, not reputation.

Is SGLang better for AI agents and multi-turn chat?

Often, when requests share long prefixes such as system prompts and tool definitions, because RadixAttention reuses them automatically. Published tests show gains there, but vLLM has prefix caching too, so measure both.

Do vLLM and SGLang have an OpenAI-compatible API?

Yes. vLLM serves on port 8000 and SGLang on port 30000 by default. A client only needs a different base URL, so you can run both side by side for testing.

Can I run SGLang on AMD GPUs?

Yes. SGLang's README lists AMD MI355 and MI300 support, alongside NVIDIA GPUs, Intel Xeon CPUs, Google TPUs and Ascend NPUs. vLLM also supports AMD GPUs.

Are vLLM and SGLang free for commercial use?

Yes. Both are Apache 2.0 licensed with no per-token or per-seat fee. Your cost is the GPUs, plus the engineering time to run and upgrade them.

Can I switch from vLLM to SGLang without changing client code?

Mostly. Both expose OpenAI-compatible endpoints, so clients mainly need a new base URL and port. Re-test model names, sampling parameters and quantization, since behaviour can differ between engines.

Run vLLM on Your Own GPU

A pre-configured vLLM inference server for AWS and Google Cloud, with NVIDIA drivers, CUDA and an OpenAI-compatible API ready to go.

Launch vLLM on Your GPU