> Source: https://meetrix.io/blogs/vllm-eks-gpu-deployment/
> Markdown copy of that page. Cite the URL above, not this file.

LLMs & APIs

# Deploying vLLM on Amazon EKS with GPUs

[By Shalomi Umeshika](https://meetrix.io/blogs/authors/shalomi-umeshika/) • September 29, 2026 • 10 min read

Getting vLLM running on Amazon EKS is mostly about getting the GPU node right before you write a single line of YAML for the pod itself. Pick the wrong AMI and your pods sit in `Pending` forever because nothing ever tells Kubernetes the node has a GPU (our [GPU scheduling on Kubernetes](https://meetrix.io/blogs/gpu-scheduling-kubernetes/) article covers why that mechanism works the way it does). Get that part right and the rest is an ordinary Deployment.

This walks through both GPU node paths on EKS, the AMI choice that trips people up, a working `vllm-openai` Deployment, and what to reach for once one node isn't enough. If you just need one GPU instance and don't need Kubernetes at all, our [vLLM developer guide](https://meetrix.io/blogs/vllm-developer-guide/) gets you there faster with CloudFormation instead. Running on Google Cloud instead? Our [GKE version of this guide](https://meetrix.io/blogs/open-weight-llms-gke/) covers the same ground, with one node-level step this article spends a whole section on that GKE mostly handles for you.

The Quick Verdict

Testing a model

Bottlerocket on a single L4. Skips the device plugin step entirely, and EKS Auto Mode skips the node group too.

Production, exact config needed

Self-managed node group. Pick your AMI, instance type and driver version instead of Auto Mode's defaults.

Model too big for one GPU

LeaderWorkerSet across an 8-GPU node type, only once a single node genuinely can't fit the model.

Steady, high-utilization traffic

Self-managed, tuned to run hot, with KEDA scale-to-zero for whatever idle windows you do have.

$0.10/hr EKS cluster fee before any GPU compute, for as long as your version is on standard support

60% Off Auto Mode's GPU management fee for P-series and Trainium, effective July 2026

1 of 2 AMI families that skip the separate NVIDIA device plugin step entirely

Prerequisites

-   An EKS cluster on a currently supported version, 1.33 (released May 30, 2025) or later, plus `kubectl`, `eksctl` and `helm` installed locally.
-   GPU quota for the instance family you're launching. New AWS accounts often start with none; see our guide to [increasing your AWS vCPU quota](https://meetrix.io/blogs/increase-aws-vcpu-quota/).
-   IAM permissions to create node groups and, if you're using S3 or EFS for model weights, an IAM role for the service account (IRSA) to access them.

## Two Ways to Get GPU Nodes

EKS gives you two real options for provisioning the GPU capacity itself.

**EKS Auto Mode** is the low-effort path. AWS runs Karpenter for you off-cluster, provisions GPU nodes on demand from a NodePool you define, and includes built-in support for NVIDIA GPU plugins, so you don't install a device plugin yourself. It also adds Node Monitoring Agent and Node Auto Repair, which detect a failed GPU and cordon or replace the node automatically, and AWS cut its GPU management fee from July 1, 2026: 35% off G-series, 60% off P-series and Trainium, applied automatically in every region Auto Mode supports. You give up some control over the exact AMI and driver version in exchange.

**Self-managed node groups**, with either a managed node group or your own Karpenter install, give you full control of the AMI, the instance types and the taints, at the cost of installing the device plugin and handling node health yourself. This is the path the rest of this article walks through, since it's the one where the AMI choice below actually matters.

## The AMI Choice That Trips People Up

Both EKS-optimized accelerated AMI families, AL2023 NVIDIA and Bottlerocket NVIDIA, ship the NVIDIA driver, CUDA user-mode driver and container toolkit already installed. Where they differ is the one piece that actually gets your pods scheduled:

| Component | AL2023 NVIDIA AMI | Bottlerocket NVIDIA AMI |
| --- | --- | --- |
| NVIDIA driver, CUDA, container toolkit | Included | Included |
| NVIDIA Kubernetes device plugin | Not included, install separately | Included |
| NVIDIA GPU Operator compatible | Yes, disable its driver/toolkit install | Yes, disable its driver/toolkit/device-plugin install |

Miss this and the symptom is confusing: the node joins the cluster, `nvidia-smi` works fine over SSH, but `kubectl describe node` shows no `nvidia.com/gpu` in `Allocatable`, so your pod's GPU resource request can never be satisfied and it stays `Pending` with no useful error. If you're on AL2023, install the device plugin as a DaemonSet before you deploy vLLM. If you're on Bottlerocket, you can skip that step.

1.63.0min Bottlerocket for DRA

Two more things worth knowing before you pick a side. If you're on multi-node EFA instances, AL2023 has an edge Bottlerocket doesn't advertise as loudly: it does topology-aligned allocation of GPUs and EFA interfaces automatically. Bottlerocket and custom AMIs don't, so you'd need the separate EFA DRA driver to get the same alignment, and NVIDIA's own DRA driver isn't supported on Bottlerocket at all.

And if you do want to move Bottlerocket onto the newer DRA driver instead of its bundled device plugin, you need Bottlerocket 1.63.0 or later, which is where `settings.kubelet-device-plugins.nvidia.enabled = false` first became a valid setting.

## Create the GPU Node Group

Using Bottlerocket here to skip the device plugin step, on a single L4 (24 GB), enough to serve an 8B-class model for testing:

```plaintext
eksctl create nodegroup \
  --cluster your-cluster \
  --name gpu-workers \
  --node-type g6.xlarge \
  --nodes 1 --nodes-min 0 --nodes-max 3 \
  --node-ami-family Bottlerocket \
  --node-labels "nvidia.com/gpu=true" \
  --node-taints "nvidia.com/gpu=true:NoSchedule"
```

The taint matters. Without it, ordinary CPU-only pods can land on your expensive GPU node while it waits for a device plugin or scheduler pressure to notice. The `tolerations` block in the Deployment below is what lets vLLM's pod past that taint.

If you'd rather stay on AL2023, drop `--node-ami-family Bottlerocket` and install the device plugin once the node group is up:

```plaintext
kubectl apply -f https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/main/deployments/static/nvidia-device-plugin.yml

# confirm the node advertises GPU capacity
kubectl get nodes -o json | jq '.items[].status.capacity."nvidia.com/gpu"'
```

## Deploy vLLM

A minimal Deployment and Service for a single-GPU model. This follows [vLLM's own Kubernetes deployment guidance](https://docs.vllm.ai/en/latest/deployment/k8s.html): request GPU capacity through `nvidia.com/gpu`, cache Hugging Face weights on a PVC so a pod restart doesn't re-download them, and mount `/dev/shm` as memory-backed for tensor-parallel setups.

```plaintext
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-server
spec:
  replicas: 1
  selector:
    matchLabels: { app: vllm-server }
  template:
    metadata:
      labels: { app: vllm-server }
    spec:
      tolerations:
        - key: "nvidia.com/gpu"
          operator: "Exists"
          effect: "NoSchedule"
      containers:
        - name: vllm
          image: vllm/vllm-openai:latest
          args:
            - "--model=meta-llama/Llama-3.1-8B-Instruct"
            - "--max-model-len=8192"
          ports:
            - containerPort: 8000
          resources:
            limits: { nvidia.com/gpu: "1" }
            requests: { nvidia.com/gpu: "1" }
          volumeMounts:
            - { name: hf-cache, mountPath: /root/.cache/huggingface }
            - { name: shm, mountPath: /dev/shm }
          readinessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 120
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 120
            periodSeconds: 10
      volumes:
        - name: hf-cache
          persistentVolumeClaim: { claimName: hf-cache-pvc }
        - name: shm
          emptyDir: { medium: Memory, sizeLimit: 4Gi }
---
apiVersion: v1
kind: Service
metadata:
  name: vllm-server
spec:
  selector: { app: vllm-server }
  ports:
    - { port: 8000, targetPort: 8000 }
```

The readiness and liveness probes both point at vLLM's `/health` endpoint with a 120-second initial delay. That delay is not padding: `/health` returns 200 the moment the server process binds to port 8000, not once the model is actually loaded, so a delay under about two minutes can add the pod to the Service before it can serve a request, or kill it mid-download and restart the whole thing in a loop. Give a larger model even more room. If you want a probe that reflects the model actually being ready rather than just the process being alive, point it at `/v1/models` instead.

Create the PVC before applying this (a 50 to 100 GB volume is enough for most single models), then confirm the pod is scheduled onto the GPU node and passes its probes:

```plaintext
kubectl get pods -o wide
kubectl logs -f deployment/vllm-server
```

## Expose and Test the API

The Service above is `ClusterIP`, which is fine for testing from inside the cluster or through a port-forward. For real traffic, put an Ingress with an AWS Load Balancer Controller in front of it, or a plain internal `LoadBalancer` Service if the callers are all inside your VPC.

```plaintext
kubectl port-forward svc/vllm-server 8000:8000

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}]
  }'
```

vLLM's server speaks the OpenAI chat completions format, so anything already built against the OpenAI SDK works here by changing the `base_url`. Our guide to [self-hosting an OpenAI-compatible API](https://meetrix.io/blogs/openai-compatible-api-self-hosted/) covers that swap in more detail if you're pointing an existing app at it.

If hand-writing this YAML for every model starts to feel like a chore, vLLM's own [production-stack Helm chart](https://github.com/vllm-project/production-stack) wraps it: a model-download init container, PVC-backed weight storage, and KEDA-based autoscaling that can scale down to zero replicas, which a plain Deployment can't do on its own. Worth switching to once you're running more than one model.

## Autoscaling GPU Capacity

Two or three layers do different jobs here, and it's worth keeping them separate in your head. Karpenter (or EKS Auto Mode's built-in Karpenter) watches for unschedulable pods, picks the lowest-cost instance type that fits, and adds GPU nodes to match, retrying other instance types and zones if one hits an availability error. It removes nodes again once they're idle. A Horizontal Pod Autoscaler, on top, scales the number of vLLM replicas based on load, using vLLM's own Prometheus metrics (queue depth and GPU cache usage are the useful ones) rather than plain CPU, which barely moves on a GPU-bound workload.

One gap worth knowing about: HPA can scale replicas up and down, but it can't scale a deployment to zero. If you want vLLM pods to disappear entirely during a predictable idle window (overnight, over a weekend) and that matters for your GPU bill, that needs KEDA instead, driving off the same GPU metrics through a DCGM exporter and Prometheus.

Most teams don't start with all of this. Get a fixed number of replicas on a fixed node group working and tested first, then add node-level autoscaling once you know your real traffic pattern, then HPA once that traffic is variable enough to justify it, then KEDA if you actually have idle windows worth scaling to zero for. Autoscaling a GPU fleet that's already misconfigured just makes the misconfiguration bigger, faster.

## Serving Models Too Big for One GPU

An 8B model fits on one L4. A 70B-class model in FP8 needs an 80 GB H100, and anything larger than that needs more than one GPU, sometimes more than one node. vLLM's documented way to do this on Kubernetes is [LeaderWorkerSet (LWS)](https://lws.sigs.k8s.io/), a Kubernetes API built for exactly this shape of workload: one group of pods that has to be scheduled and scaled together. One pod runs as the leader with `vllm serve` and the OpenAI-compatible endpoint, the rest run `vllm serve --headless` as workers, and tensor or pipeline parallelism spans the whole group. vLLM's own reference example uses a group of two pods, each requesting all 8 GPUs on its node, which in practice means two 8-GPU nodes for that one replica; scale the group size and per-pod GPU count to whatever your model actually needs.

This is a step up in operational complexity, not something to reach for by default. If a single 8-GPU node (a `p5.48xlarge`, for instance) fits your model, that's simpler than spreading it across nodes with LWS. Save multi-node for models that genuinely don't fit on the biggest single node you're willing to run.

## What This Actually Costs

The node group above bills like any other EC2 GPU instance, plus the EKS cluster fee. We've already worked through GPU hourly prices, cost per million tokens and the utilization math in detail in our [GPU costs for LLM inference](https://meetrix.io/blogs/gpu-costs-llm-inference-aws-vs-gcp/) article, so I won't repeat it here. The one EKS-specific addition: an idle GPU node group costs the same whether Kubernetes is scheduling pods onto it or not, so the autoscaling in the section above isn't just an ops nicety, it's the thing that keeps a quiet weekend from billing like a busy Tuesday.

6xon an outdated cluster

The EKS cluster fee itself is $0.10 an hour, about $73 a month, for as long as your cluster's Kubernetes version is on standard support (14 months from release). Ride that out to extended support and the same cluster jumps to $0.60 an hour, six times as much, which is one more reason not to let a cluster drift too far behind.

## Troubleshooting

-   **Pod stuck in Pending, no events.** Almost always the AL2023-without-device-plugin gap above, or the pod's missing the toleration for the GPU node's taint. Check `kubectl describe node` for `nvidia.com/gpu` under `Allocatable` first.
-   **CrashLoopBackOff on a fresh deploy.** The pod is still downloading the model when the liveness probe's initial delay runs out. Raise `initialDelaySeconds`, especially for anything above 8B parameters.
-   **OOMKilled or a CUDA out-of-memory error.** The PVC for the Hugging Face cache is too small, or `/dev/shm` is undersized for tensor-parallel. Both show up as the pod dying partway through startup, not immediately.
-   **Works on one node, fails after Karpenter scales out.** The new node came up on a different AMI variant or instance family than you tested. Pin the AMI family and instance types explicitly rather than leaving Karpenter to pick freely across a wide GPU category.
-   **GPU pod can't claim its EFA device on a multi-node setup.** NVIDIA's device plugin versions 0.19.0 through 0.19.2 default to mounting all `/dev/infiniband/uverbs*` devices into any container that requests a GPU, which collides with the EFA device plugin trying to manage the same devices. Disable MOFED explicitly on managed or self-managed nodes, or upgrade past 0.19.2, where this default was reverted. EKS Auto Mode was never affected.

For a from-scratch reference setup rather than a production build, AWS publishes a [GenAI-on-EKS starter kit](https://github.com/aws-samples/sample-genai-on-eks-starter-kit) that wires up vLLM behind a LiteLLM gateway, with Langfuse for observability and Open WebUI or Qdrant available as optional pieces, all through Terraform. It's explicitly labelled for demonstration and learning, not production, but it's a fast way to see the whole stack running before you build your own.

## Where I'd Start

Testing a model

Bottlerocket, a single L4. Skip the device plugin step and get straight to the Deployment.

Production, one node fits

Self-managed node group, pinned AMI family and instance type, with the 120-second probe delay above.

Model needs multiple nodes

LeaderWorkerSet across 8-GPU nodes, and watch for the EFA/MOFED conflict if you're on device plugin 0.19.0 through 0.19.2.

Don't want to hand-tune any of this

Start with vLLM's production-stack Helm chart. It wraps the Deployment, storage and KEDA-based autoscaling into one install.

Running on Google Cloud instead? The device plugin and scheduler mechanism is identical, but GKE handles more of the node-level setup for you automatically. Our [GKE version of this guide](https://meetrix.io/blogs/open-weight-llms-gke/) covers the specifics, and [GPU scheduling on Kubernetes](https://meetrix.io/blogs/gpu-scheduling-kubernetes/) covers the vendor-neutral mechanism underneath both.

## Technical Support

Reach out to Meetrix Support ([support@meetrix.io](mailto:support@meetrix.io)) for help with vLLM deployment issues, on EKS or otherwise.

## Frequently Asked Questions

Does Amazon EKS support GPU nodes for vLLM?

Yes. EKS runs GPU worker nodes on standard NVIDIA-family EC2 instances (g5, g6, g6e, p4d, p5 and others), using either the EKS-optimized accelerated AL2023 or Bottlerocket AMI. Both are officially supported node operating systems for accelerated workloads.

Do I need to install the NVIDIA device plugin on EKS?

Depends on the path. The AL2023 NVIDIA AMI ships the driver and CUDA toolkit but not the device plugin, so you install it yourself. The Bottlerocket NVIDIA AMI includes it already. EKS Auto Mode handles GPU plugins for you on either.

Which GPU instance should I use for vLLM on EKS?

A single L4 or A10G (g6.xlarge, g5.xlarge) is enough to test an 8B-class model. Production 70B-class serving typically needs an H100 (p5.4xlarge) or a multi-GPU shape. Our GPU costs article has the current hourly prices for each.

Can I autoscale vLLM pods based on GPU load?

Yes, in layers. Karpenter or EKS Auto Mode adds and removes GPU nodes as pods are scheduled. A Horizontal Pod Autoscaler scales replicas using vLLM's Prometheus metrics, but can't go to zero. KEDA can, if you have idle windows worth the extra setup.

How do I serve a model too large for one GPU on EKS?

Use LeaderWorkerSet (LWS), the Kubernetes API vLLM's own docs recommend for multi-node serving. One pod runs as the leader with vllm serve, the rest run vllm serve --headless as workers, and tensor or pipeline parallelism spans the group.

Is EKS Auto Mode a good fit for running vLLM?

For most teams, yes. Auto Mode provisions and repairs GPU nodes without you installing a device plugin or running Karpenter yourself, and AWS cut its GPU management fee from July 2026: 35% off G-series, 60% off P-series and Trainium. The trade-off is less control over the exact AMI and driver version.

Do I need vLLM's production-stack project, or is a plain Deployment enough?

A plain Deployment with a PersistentVolumeClaim, the right GPU resource request, and health probes is enough for one model on one node type. Reach for production-stack or a Helm chart once you're running several models or need built-in routing and autoscaling policies.

## Don't Need Kubernetes? Deploy vLLM on One GPU Instance

If EKS is more than your workload needs, our CloudFormation-based vLLM developer guide gets an OpenAI-compatible endpoint running on a single AWS GPU instance in minutes.

[See the Single-Instance Guide](https://meetrix.io/blogs/vllm-developer-guide/)

Meetrix Store

vLLM

Serve open models on your own GPU

[Deploy it](https://meetrix.io/store/vllm/)

Meetrix Store New

Deploy what this guide covers, pre-configured.

-    [vLLM Serve open models on your own GPU](https://meetrix.io/store/vllm/)
-    [OpenWebUI A private ChatGPT-style assistant](https://meetrix.io/store/openwebui/)
-    [Llama 4 Scout Mixture-of-Experts, 17B active parameters](https://meetrix.io/store/llama-4-scout/)
-    [Jitsi Meet Self-hosted video calls for 50 to 500 users](https://meetrix.io/store/jitsi-meet/)
-    [RustDesk Remote desktop AMI, a TeamViewer alternative](https://meetrix.io/store/rustdesk/)
-    [Coturn TURN/STUN for WebRTC, no per-minute relay fees](https://meetrix.io/store/coturn/)

[Browse all products](https://meetrix.io/store/)
