Kubernetes doesn't know what a GPU is. It never has. What it schedules is an integer count of an opaque resource that something else, a device plugin, told it about. Every GPU platform you've used, from a managed EKS node group to a bare-metal cluster, is built on that one fact plus a handful of add-ons that make it usable in practice.

This article covers the mechanism itself: how a GPU becomes visible to the scheduler at all, the two competing APIs for requesting one, how to share a single card between pods, and why the scheduler that ships with Kubernetes still can't run a distributed training job on its own. If you want the applied version on a specific cloud, our guide to deploying vLLM on Amazon EKS with GPUs walks through the same concepts in a real cluster.

The Quick Verdict

Same model, many replicas

Time-slicing. No isolation, but you don't need to protect replicas of the same trusted model from each other.

Untrusted, multi-tenant workloads

MIG. Hardware-level isolation, at the cost of fixed, static profiles instead of flexible slices.

Multi-GPU job, one team

Kueue alone, gang admission with no second scheduler binary to run.

Multi-GPU job, strict pod placement

Kueue plus Volcano underneath, for fabric-aware, pod-level gang scheduling.

1.26 Kubernetes version the device plugin framework has been stable since
1.35 Kubernetes version where Dynamic Resource Allocation reached general availability
7 Maximum hardware-isolated instances MIG can carve from one supported GPU

The Short Version

  • A device plugin, not Kubernetes itself, tells the kubelet a node has GPUs.
  • The scheduler treats a GPU as an integer resource: one whole unit per request, no fractions, no overcommit.
  • Dynamic Resource Allocation (DRA) is the more expressive alternative, now GA. It supports device filtering and cross-container sharing, but preemption isn't supported yet.
  • Sharing one physical GPU needs MIG, time-slicing, or MPS on top. Each trades isolation for flexibility differently. Kata Containers adds a hardware-virtualization option.
  • Plain kube-scheduler doesn't gang-schedule. Kueue, Volcano, or KAI Scheduler handle that today. Native gang scheduling is now Beta but disabled by default.
  • Karpenter is the most common EKS GPU autoscaler, but as of v1.14 it still can't provision a new node for a pod whose GPU request is DRA-only.

How the Scheduler Sees a GPU

A GPU becomes visible to Kubernetes through the device plugin framework, stable since Kubernetes 1.26. A device plugin runs as a DaemonSet on each GPU node, registers with the kubelet over a local gRPC socket, and reports what it finds under a vendor-namespaced extended resource name, nvidia.com/gpu being the one you'll see everywhere. The kubelet folds that into the node's Allocatable capacity, and from that point the scheduler treats it exactly like CPU or memory: a number to bin-pack pods against.

The catch is how narrow that number is. Extended resources are integer-only, can't be overcommitted, and can't be split across containers in the same pod. A request for nvidia.com/gpu: "1" reserves one whole card. There's no way to ask for a quarter of a GPU through this path. That limitation is exactly what MIG, time-slicing, MPS, and DRA each exist to work around, in different ways.

apiVersion: v1
kind: Pod
metadata:
  name: gpu-pod
spec:
  tolerations:
    - key: "nvidia.com/gpu"
      operator: "Exists"
      effect: "NoSchedule"
  containers:
    - name: app
      image: your-image
      resources:
        limits:
          nvidia.com/gpu: "1"   # integer only, no overcommit, no fractions

The toleration matters as much as the resource request. GPU nodes are almost always tainted on purpose, so an ordinary CPU-only pod doesn't accidentally land on hardware that costs ten times as much per hour. Without a matching toleration, your own GPU pod would be blocked by the same taint.

Extended Resources vs Dynamic Resource Allocation

Extended resources have one real weakness: a bare integer can't express "give me an H100, not an A10G" or "these two containers can share one GPU." Dynamic Resource Allocation is Kubernetes' answer to that.

DRA is now generally available, with the feature gate locked on, so you can't disable it even if you try. It first appeared as alpha in Kubernetes 1.30.

Instead of a plain number, DRA asks you to write what you actually need. A cluster admin defines a DeviceClass that filters available hardware using CEL expressions: by GPU model, memory size, whatever attributes the driver exposes. Workloads reference that class through a ResourceClaim or a ResourceClaimTemplate, and a DRA-aware driver publishes what's available on each node through a ResourceSlice.

apiVersion: resource.k8s.io/v1
kind: DeviceClass
metadata:
  name: high-performance-gpu
spec:
  selectors:
    - cel:
        expression: "device.attributes['gpu.nvidia.com'].productName == 'H100'"
---
apiVersion: v1
kind: Pod
metadata:
  name: gpu-pod-dra
spec:
  containers:
    - name: app
      image: your-image
      resources:
        claims:
          - name: gpu
  resourceClaims:
    - name: gpu
      resourceClaimTemplateName: h100-claim-template

DRA also supports sharing a claimed device across multiple containers or even multiple pods, something extended resources flatly can't do. NVIDIA moved its GPU DRA driver into Kubernetes SIGs, dropping the Beta label, which is a reasonable signal that this is where GPU scheduling is heading rather than a side experiment.

The honest caveat: the scheduler still doesn't support preemption for DRA claims. A high-priority pod that needs a DRA-claimed device just waits if a lower-priority pod is holding it, even if it's otherwise entitled to bump it. This is a documented limitation, not a temporary bug.

On EKS specifically, Karpenter added DRA support in v1.14.0 (July 2026), and it correctly schedules DRA pods onto GPU nodes that already exist. What it still can't do is provision a brand new node for a pod whose GPU request is expressed purely as a ResourceClaim, with no matching node already sitting there. Since provisioning capacity on demand is one of Karpenter's main reasons for existing on GPU workloads, that gap is not a footnote, it can decide your whole architecture. A managed node group or self-managed nodes running the DRA driver directly still cover that case.

For most teams today, extended resources remain the simpler default. DRA is worth adopting once you specifically need its device filtering or cross-container sharing.

Sharing One GPU Between Pods

All three of the common options come from NVIDIA's stack, and they sit at genuinely different points on the isolation-versus-flexibility line, not just different configuration flags on the same thing.

ApproachIsolationGranularityGood for
Time-slicingNone. Shared memory and fault domain.Software-scheduled turns on one physical GPUMany replicas of the same trusted model, cost over isolation
MIGHardware. Each slice acts as its own GPU.Fixed, static profiles, up to 7 instances on supported GPUsMulti-tenant clusters where a noisy or crashing neighbor can't be tolerated
MPSSoft, spatial. One fault can still affect other clients.Concurrent CUDA kernels from multiple processesTrusted internal workloads, training experiments, higher throughput than time-slicing

A rough decision rule: the same model behind many replicas leans toward time-slicing, because you don't need to protect replicas from each other. Genuinely mixed, untrusted workloads on shared infrastructure lean toward MIG, because a bad actor or a crash in one slice can't reach another. MPS sits in between, for workloads you trust but want more real concurrency from than time-slicing gives you.

One detail worth knowing about MPS: the device plugin advertises MPS capacity under nvidia.com/gpu resources, and only with full GPUs. There's no fractional MPS resource you can request through the standard plugin path.

A fourth option: hardware-level isolation. If MIG's hardware partitioning isn't enough (say, you're in a regulated environment with hard multi-tenancy requirements), Kata Containers with GPU support runs each pod in a dedicated microVM, with the GPU exposed via VFIO passthrough. Isolation is enforced at the hardware virtualization boundary rather than the driver layer. This is heavier than the other three options and worth it only when the isolation guarantee itself is the requirement.

DRA can express any of these as a device attribute once your driver supports it, which is part of why it's the direction NVIDIA and the Kubernetes project are both investing in.

Queueing and Gang Scheduling for Multi-GPU Jobs

Plain kube-scheduler places pods one at a time, independently. That's fine for a stateless web service. It's a real problem for a distributed training job or a multi-node inference deployment that needs, say, 8 GPUs across 2 nodes to do anything useful: the scheduler can happily place half the pods, leave the rest pending, and let the job sit there holding GPUs it can't use while doing no useful work.

Kueue runs as CRDs on top of the existing scheduler, no second scheduler binary needed, and handles quota and admission: a job either gets all the resources it needs admitted together, or it waits in a queue. Kueue's native support for JobSet and LeaderWorkerSet means it can gang-admit the kind of multi-pod groups our vLLM on EKS guide covers for multi-node serving.

1.37K8s native gang scheduling, beta

Volcano goes further, replacing pod-level scheduling decisions entirely with gang scheduling and fairness policies built for HPC-style workloads, at the cost of running as its own scheduler component. Volcano 1.15 strengthened this further with gang-aware preemption and resource reclaim: preemption decisions get evaluated at the gang level, and surplus replicas get evicted ahead of random pods, so you don't end up in the "released a bunch of pods but nobody can run" trap. Plenty of production clusters run Kueue for organizational quota and admission, with Volcano underneath for the pod-placement details Kueue doesn't touch.

Native gang scheduling is arriving. Kubernetes 1.37 graduated the Workload and PodGroup APIs to beta behind the GenericWorkload feature gate, alongside workload-aware preemption and a CompositePodGroup API for hierarchical topology constraints. All of it stays disabled by default for now, but the direction is clear: track it rather than assuming Kueue or Volcano will always be the only answer.

KAI Scheduler has been accepted as a CNCF Sandbox project, moving from an NVIDIA-governed tool toward community-developed standard. It handles DRA for GPUs, topology-aware placement, and hierarchical PodGroups for gang scheduling, using dominant resource fairness to share GPUs across teams. It's worth watching if Kueue and Volcano together feel like more than your cluster needs.

Autoscaling the Nodes Underneath

Everything above assumes the GPU nodes already exist. Getting them to exist, and disappear again when idle, is a separate job.

On EKS specifically, Karpenter is the autoscaler most teams reach for. Cluster Autoscaler struggles with instance heterogeneity, Spot diversification and consolidation at GenAI scale, which is exactly the territory GPU fleets live in. Karpenter provisions GPU nodes on demand from a NodePool you define, then removes them when idle. It works with the Kubernetes scheduler rather than replacing it: the scheduler places pods on nodes, and Karpenter provisions capacity for the pods the scheduler can't fit. GKE and AKS solve the same problem with their own tools, not Karpenter; our guide to open-weight LLMs on GKE covers GKE's version of this.

A few operational notes that matter more than they look. NodePool limits are your cost guardrail: always set spec.limits (e.g., nvidia.com/gpu) plus a per-namespace ResourceQuota, so a runaway Deployment or a misconfigured HPA can't scale GPUs without bound. Inference pools and training pools want different consolidation policies too: inference does well with consolidationPolicy: WhenEmptyOrUnderutilized and a short consolidateAfter, cheap to lose a replica, expensive to keep an idle node, while training usually wants a longer window or a WhenEmpty policy so Karpenter doesn't evict mid-run and destroy hours of work. And if you're mixing GPU and Neuron hardware, provision two NodePools from day one rather than one; future hardware migration becomes a cost experiment rather than a re-architecture.

None of this changes anything above: the device plugin and scheduler layer works identically no matter which autoscaler put the node there.

What This Costs

Scheduling decisions and cost decisions are more connected than they look. An idle GPU node costs the same whether kube-scheduler put a pod on it or not, so getting the scheduling layer right (tight autoscaling, sensible sharing where isolation allows it, gang admission so jobs don't sit half-started) is directly what keeps a GPU fleet from quietly overspending. We've worked through the actual GPU pricing and utilization math in our GPU costs for LLM inference article, and the enterprise buy-versus-build version of that question in self-hosted LLMs vs API costs.

Where I'd Start

Sharing a single GPU

Time-slicing by default. Only reach for MIG once you're running genuinely untrusted, multi-tenant workloads.

A multi-GPU job today

Start with Kueue alone. Add Volcano underneath only once you need pod-level gang scheduling or fabric awareness.

Running on EKS

Karpenter, with explicit NodePool limits from day one. Know the DRA-provisioning gap before you design around it.

Running on GKE or AKS

Use their own autoscaler tooling. Our GKE guide covers that side in detail.

Native Kubernetes gang scheduling is real but still beta and disabled by default as of 1.37. Worth tracking for a platform you're designing now, not worth migrating onto yet if Kueue or Volcano are already doing the job.

Frequently Asked Questions

How does Kubernetes know a node has a GPU?

A device plugin runs as a DaemonSet on the node, registers with the kubelet over a local gRPC socket, and reports GPU capacity under a vendor-namespaced extended resource name like nvidia.com/gpu. The kubelet adds that to the node's Allocatable capacity, and the scheduler treats it like any other resource count.

Can a pod request half a GPU in Kubernetes?

Not through extended resources. They're integer-only and can't be split. To share a GPU you need MIG, time-slicing, or MPS on top of the device plugin. DRA also supports device sharing across containers or pods, and Kata Containers can provide hardware-virtualization isolation for the same physical card.

What is Dynamic Resource Allocation (DRA) in Kubernetes?

DRA is a newer API that replaces bare integer resource requests with declarative claims. You define a DeviceClass with CEL-based filters, reference it through a ResourceClaim, and the scheduler allocates a matching device. It's now GA, with the feature gate locked on.

What's the difference between MIG and time-slicing for GPU sharing?

MIG partitions a GPU at the hardware level, so each slice acts as its own GPU with its own memory and fault domain. Time-slicing shares a GPU at the software level, with multiple pods taking turns and no isolation. MIG is for untrusted multi-tenant workloads; time-slicing is for many replicas of the same trusted model.

Does the default Kubernetes scheduler support gang scheduling for multi-GPU jobs?

Not natively today, though native gang scheduling is now Beta and disabled by default. Plain kube-scheduler places pods one at a time. Kueue, Volcano, and KAI Scheduler are the production-ready options right now.

Should I use Kueue or Volcano for GPU workloads?

They occupy different layers. Kueue handles quota and admission on top of the existing scheduler, with no second scheduler binary needed. Volcano replaces pod-level scheduling entirely with gang scheduling and fairness policies. Many clusters run both: Kueue for organizational quota, Volcano for pod-placement details.

Is GPU scheduling different on EKS, GKE, and AKS?

The core mechanism, device plugin, kubelet, extended resources, is identical everywhere. What differs is the autoscaler. Karpenter is the AWS option; GKE and AKS ship their own equivalents built on the same idea. DRA support and managed-service maturity still vary by provider.

Does Karpenter support DRA?

Karpenter added DRA support in v1.14.0, so it can schedule DRA pods onto GPU nodes that already exist. It still can't provision a brand new node for a pod whose GPU request is a bare ResourceClaim with nothing matching already running, so that case needs a managed node group or self-managed nodes with the DRA driver instead.

Run a GPU Workload on Infrastructure You Control

A pre-configured vLLM inference server for AWS and Google Cloud, with an OpenAI-compatible API, ready to schedule behind whatever setup this article walked through.

Launch vLLM on Your GPU