Skip to main content
Version: 1.3.3 (Latest)

Prepare GPU nodes

A platform administrator prepares GPU nodes before Kamiwaza serves models on them. Kamiwaza installs nothing on a node: no driver, no container runtime configuration, and no GPU operator. This page covers each node-side requirement for NVIDIA GPUs, with a check and its expected output.

A default install cannot use GPUs even on prepared nodes. It installs without node reads and with tenant-mode inference off, and the install prints WARNING: core cannot see GPUs on this install. Prepare the nodes here, then turn on GPU serving with Tenant-mode inference, which uses a profile that the cluster owner signs.

Before you start​

You needNotes
Root access on each GPU nodeThe commands are for Ubuntu 22.04 and 24.04 nodes. On another distribution, use its packages; the checks are the same.
Cluster-admin accessTo install the GPU Operator and create the nvidia RuntimeClass.
Packages and charts from inside your networkThe node packages come from your mirror of the Ubuntu archive and of NVIDIA's container toolkit repository, and the GPU Operator chart and its images come from your registry.
kubectl, helm, and jqOn the machine you run the cluster checks from.
Two public test imagesThe checks run docker.io/library/ubuntu:24.04 and docker.io/nvidia/cuda:13.0.3-base-ubuntu24.04. With internet access, use them as named. In an air gap, copy them into your registry first; the checks write them as <YOUR_REGISTRY>/library/ubuntu:24.04 and <YOUR_REGISTRY>/nvidia/cuda:13.0.3-base-ubuntu24.04. If your registry needs credentials, the check pods run in the default namespace: create a pull Secret there and add imagePullSecrets: [{name: <SECRET_NAME>}] to each check pod's spec.

These steps were verified on a four-GPU NVIDIA A100 80GB host running k0s 1.36, the R580 server driver, and GPU Operator v26.7.1. Where a step differs on k0s, the step says so.

Summary​

RequirementCheck
Supported GPU and drivernvidia-smi reports the driver and CUDA version the engines need
Driver installedThe driver loads, with Secure Boot on if the node uses it
Persistence modenvidia-smi answers in well under a second
Container runtimeThe node advertises an nvidia runtime handler
nvidia RuntimeClassA pod with runtimeClassName: nvidia sees the GPUs
GPU OperatorNodes advertise nvidia.com/gpu, and nvidia.com/mig-* for MIG
MIG layout (optional)Nodes advertise the MIG resources you configured
No taints on GPU nodesNo GPU node carries a NoSchedule or NoExecute taint
Pinned GPU stackDriver, toolkit, and kernel packages are held
VRAM plugin (optional)Nodes advertise kamiwaza.ai/vram-gb-gpu-<N> for each shared GPU

Supported GPUs and driver floors​

Kamiwaza serves models on NVIDIA GPUs. Each engine image is built against one CUDA major, and the host driver must support that major. nvidia-smi prints the highest CUDA version the installed driver supports in its header, as CUDA Version: 13.0.

Engine imageEnginesMinimum host driverLowest GPU compute capability
CUDA 13llama.cpp, vLLM580.65.06 (R580)7.5 for llama.cpp; 8.0 for vLLM on amd64 and 12.1 (GB10) on arm64, and vLLM requires bf16
CUDA 12llama.cpp525.60.13; R570 or later for Blackwell GPUs7.0
CUDA 12, built with CUDA 12.8DiffusionR570 or later recommended; not enforced8.0, and the engine requires bf16

A GPU older than Volta (compute capability 7.0) has no engine image. Diffusion ships for amd64 only. vLLM ships for amd64, and for arm64 on GB10 only. llama.cpp ships for amd64 and arm64. See Engines for what each engine serves.

Check on each GPU node:

nvidia-smi --query-gpu=index,name,driver_version,compute_cap --format=csv,noheader
nvidia-smi | grep -o 'CUDA Version: [0-9.]*'
0, NVIDIA A100 80GB PCIe, 580.178.04, 8.0
1, NVIDIA A100 80GB PCIe, 580.178.04, 8.0
CUDA Version: 13.0

Expect every GPU on a node to report the same name and compute capability. A profile binding cannot single out one kind of GPU on a node that has two (see Tenant-mode inference), so keep such a node out of tenant-mode GPU serving.

Keep a node with two kinds of GPU out​

A binding selects nodes by architecture only, so on a cluster with other GPU nodes a tenant-mode GPU pod could land on this one. After you install the GPU Operator in step 5, turn off the operator's components on the node:

kubectl label nodes <NODE_NAME> nvidia.com/gpu.deploy.operands=false --overwrite

The operator then removes its device plugin, validator, GPU feature discovery, DCGM exporter and mig-manager from the node, so the scheduler can no longer give any pod one of its GPUs. That holds for every pod, not only tenant-mode ones; pods already running keep the GPUs they have. The host's driver and container toolkit stay, so the step 4 check still passes there. In the step 5 check, the node shows 0 for its NVIDIA resources, or none at all. To undo it, remove the label: kubectl label nodes <NODE_NAME> nvidia.com/gpu.deploy.operands-.

A cluster whose only GPU node has two kinds of GPU, such as a single-node cluster, needs none of this: leave tenant-mode GPU serving off there, and publish no GPU binding for it.

Step 1: Install the driver​

Upgrade the node and reboot before you install the driver. The kernel meta-package can be ahead of the kernel the node booted, and the driver's kernel modules track the meta-package:

export DEBIAN_FRONTEND=noninteractive
sudo -E apt-get update
sudo -E apt-get -y full-upgrade
sudo systemctl reboot

List the driver branches the archive offers for the node's GPUs:

sudo apt-get install -y ubuntu-drivers-common
sudo ubuntu-drivers list --gpgpu
nvidia-driver-580-server, (kernel modules provided by linux-modules-nvidia-580-server-generic)

Install the driver together with the prebuilt kernel module package that the list names. The flavor at the end of its name, -generic above, matches your kernel. Naming the module package stops apt from building the module with DKMS:

sudo DEBIAN_FRONTEND=noninteractive apt-get install -y \
linux-modules-nvidia-580-server-generic nvidia-driver-580-server
sudo systemctl reboot

Check:

nvidia-smi -L
GPU 0: NVIDIA A100 80GB PCIe (UUID: GPU-8d898926-...)

Expect one line per GPU. Right after a driver load, nvidia-smi can print couldn't communicate with the NVIDIA driver for a few seconds while each GPU initializes.

Secure Boot​

With UEFI Secure Boot on, the kernel loads only signed modules. Check the state:

mokutil --sb-state
SecureBoot enabled

The prebuilt module packages from step 1 are signed by Canonical and load without further steps. A module that DKMS builds is signed with a Machine Owner Key (MOK) on the node, and the firmware must trust that key. Enroll it once per node, from a console that reaches the pre-boot screen (physical, or a BMC or serial console); SSH cannot do the enrollment:

  1. Create the key, if the node has none, and stage its enrollment. mokutil asks for a one-time password twice:

    sudo update-secureboot-policy --new-key
    sudo mokutil --import /var/lib/shim-signed/mok/MOK.der
  2. Reboot. On the blue MOK manager screen, which waits about 10 seconds for a key press, select Enroll MOK, Continue, Yes, enter the one-time password, and select Reboot.

  3. Rebuild the module so DKMS signs it with the enrolled key. Name the DKMS package of your driver branch, nvidia-dkms-580-server for the R580 server driver:

    sudo dpkg-reconfigure nvidia-dkms-580-server
    sudo systemctl reboot

Check:

mokutil --test-key /var/lib/shim-signed/mok/MOK.der
modinfo -F signer nvidia
nvidia-smi -L

Expect is already enrolled, a signer name, and one line per GPU. If the enrollment screen timed out, the request stays pending; reboot again. If the password is lost, clear the request with sudo mokutil --revoke-import and repeat step 1.

Step 2: Turn on persistence mode​

Without persistence mode, every nvidia-smi call initializes each GPU again. On a four-GPU host that took about 13 seconds per call, long enough for the GPU Operator's validator to exceed the kubelet's 2-minute container start timeout (RunContainerError: context deadline exceeded), and the operator never became ready. Ubuntu's nvidia-persistenced service starts with --no-persistence-mode; override it:

sudo mkdir -p /etc/systemd/system/nvidia-persistenced.service.d
sudo tee /etc/systemd/system/nvidia-persistenced.service.d/10-persistence-mode.conf >/dev/null <<'EOF'
[Service]
ExecStart=
ExecStart=/usr/bin/nvidia-persistenced --user nvidia-persistenced --persistence-mode --verbose
EOF
sudo systemctl daemon-reload
sudo systemctl restart nvidia-persistenced

Check:

nvidia-smi --query-gpu=index,persistence_mode --format=csv,noheader
/usr/bin/time -f "%es" nvidia-smi -L >/dev/null

Expect Enabled for every GPU, and a time well under a second (0.04 seconds on the verification host).

Step 3: Configure the container runtime​

Install the NVIDIA Container Toolkit from your mirror of NVIDIA's container toolkit repository:

sudo DEBIAN_FRONTEND=noninteractive apt-get install -y nvidia-container-toolkit
nvidia-ctk --version
NVIDIA Container Toolkit CLI version 1.20.1

Then add an nvidia runtime to containerd. Keep runc as the default runtime: only pods that name the nvidia RuntimeClass use the NVIDIA runtime.

On a node whose containerd configuration you manage, let the toolkit write it, then restart containerd:

sudo nvidia-ctk runtime configure --runtime=containerd
sudo systemctl restart containerd

On k0s, k0s owns /etc/k0s/containerd.toml, so write the runtime as a drop-in by hand instead. k0s 1.36 runs containerd 2, whose drop-ins use config version 3:

sudo tee /etc/k0s/containerd.d/nvidia.toml >/dev/null <<'EOF'
version = 3

[plugins."io.containerd.cri.v1.runtime".containerd.runtimes.nvidia]
privileged_without_host_devices = false
runtime_type = "io.containerd.runc.v2"

[plugins."io.containerd.cri.v1.runtime".containerd.runtimes.nvidia.options]
BinaryName = "/usr/bin/nvidia-container-runtime"
EOF
sudo systemctl restart k0sworker # k0scontroller on a controller that runs workloads

Check from the cluster, for each GPU node. The node also lists an unnamed default handler, which the check leaves out:

kubectl get node <NODE_NAME> -o json | jq -r '[.status.runtimeHandlers[]?.name | select(. != "")] | join(" ")'
nvidia runc

If your nodes do not report runtime handlers, this prints nothing; the pod check below is then the proof.

Step 4: Create the nvidia RuntimeClass​

cat <<'EOF' | kubectl apply -f -
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: nvidia
handler: nvidia
EOF

Check that a pod on the nvidia RuntimeClass sees the GPUs. The runtime adds nvidia-smi to the container, so a plain Ubuntu image from your registry is enough. The pod is pinned to one node; set GPU_NODE to its name, and run the check once for each GPU node. The pod has no tolerations, so a NoSchedule or NoExecute taint on the node keeps it Pending. If your GPU nodes carry either, do step 7 first, or rerun this check after it.

GPU_NODE="<NODE_NAME>"
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: nvidia-runtime-check
namespace: default
spec:
runtimeClassName: nvidia
restartPolicy: Never
nodeSelector:
kubernetes.io/hostname: ${GPU_NODE}
containers:
- name: smi
image: <YOUR_REGISTRY>/library/ubuntu:24.04
command: ["nvidia-smi", "-L"]
env:
- {name: NVIDIA_VISIBLE_DEVICES, value: all}
- {name: NVIDIA_DRIVER_CAPABILITIES, value: utility}
EOF
kubectl -n default wait --for=jsonpath='{.status.phase}'=Succeeded pod/nvidia-runtime-check --timeout=5m
kubectl -n default logs nvidia-runtime-check
kubectl -n default delete pod nvidia-runtime-check

Expect one GPU n: ... line per GPU on the node.

Use only the nvidia RuntimeClass. The GPU Operator in step 5 also creates nvidia-cdi and nvidia-legacy RuntimeClasses; on a node where the operator does not manage the toolkit, those have no containerd handler, and pods that name them fail to start. The operator takes ownership of the nvidia RuntimeClass, and uninstalling the operator deletes it.

Step 5: Install the GPU Operator​

The NVIDIA GPU Operator's device plugin advertises GPUs to Kubernetes as nvidia.com/gpu, and MIG partitions as nvidia.com/mig-<profile>. With the driver and toolkit on the host, the operator must not install its own.

Save these values as gpu-operator-values.yaml. The mig and migManager blocks are needed only for MIG; keep them anyway, since they are harmless without it:

driver:
enabled: false # the host driver from step 1
toolkit:
enabled: false # the host toolkit and runtime from step 3
mig:
strategy: mixed # whole GPUs as nvidia.com/gpu, partitions as nvidia.com/mig-<profile>
# On k0s with its default kubelet root, uncomment the next two lines. Leave them
# commented on a k0s started with --kubelet-root-dir=/var/lib/kubelet, as
# Kamiwaza's setup tooling starts it and as the VRAM plugin needs.
# hostPaths:
# kubeletRootDir: /var/lib/k0s/kubelet
migManager:
# Persistence mode (step 2) keeps the GPUs open, so mig-manager must stop
# nvidia-persistenced before a MIG change and start it again afterwards.
gpuClientsConfig:
name: gpu-clients
extraObjects:
- apiVersion: v1
kind: ConfigMap
metadata:
name: gpu-clients
data:
clients.yaml: |
version: v1
systemd-services:
- nvidia-persistenced.service
- nvidia-dcgm.service
- dcgm.service
- dcgm-exporter.service

Install the operator chart from your registry, with its images mirrored there. With the values above, GPU Operator v26.7.1 runs these images; for another version, read the tags from helm show values:

ImageValues key to point at your copy
nvcr.io/nvidia/gpu-operator:v26.7.1operator.repository and validator.repository
nvcr.io/nvidia/k8s-device-plugin:v0.20.1devicePlugin.repository and gfd.repository
nvcr.io/nvidia/cloud-native/k8s-mig-manager:v0.15.1migManager.repository
nvcr.io/nvidia/k8s/dcgm-exporter:4.6.1-4.8.4-distrolessdcgmExporter.repository
registry.k8s.io/nfd/node-feature-discovery:v0.19.0node-feature-discovery.image.repository

Copy each image to your registry under the same path, e.g. <YOUR_REGISTRY>/nvidia/gpu-operator:v26.7.1. Then point the chart at your copies: save these values as gpu-operator-images.yaml. Each repository value is the image's path without its last segment, except node-feature-discovery.image.repository, which is the full image path. The chart supplies the tags in the table:

operator:
repository: <YOUR_REGISTRY>/nvidia
validator:
repository: <YOUR_REGISTRY>/nvidia
devicePlugin:
repository: <YOUR_REGISTRY>/nvidia
gfd:
repository: <YOUR_REGISTRY>/nvidia
migManager:
repository: <YOUR_REGISTRY>/nvidia/cloud-native
dcgmExporter:
repository: <YOUR_REGISTRY>/nvidia/k8s
node-feature-discovery:
image:
repository: <YOUR_REGISTRY>/nfd/node-feature-discovery

If your registry needs credentials, create a pull Secret in the gpu-operator namespace and list it under each component's imagePullSecrets.

Install with both values files. Helm merges them, so the migManager keys in each file apply together:

helm upgrade --install gpu-operator <GPU_OPERATOR_CHART> --version <GPU_OPERATOR_VERSION> \
-n gpu-operator --create-namespace -f gpu-operator-values.yaml -f gpu-operator-images.yaml \
--wait --timeout 10m

Check:

kubectl get clusterpolicy -o jsonpath='{.items[0].status.state}'; echo
kubectl get node <NODE_NAME> -o jsonpath='{.status.allocatable}' | tr ',' '\n' | grep nvidia
ready
"nvidia.com/gpu":"4"
"nvidia.com/mig-3g.40gb":"0"

With mig.strategy: mixed, a node whose GPUs support MIG can also list MIG resources with a count of 0 until you partition a GPU in step 6. A node that has never had MIG instances may not list them at all.

Check that a pod gets exactly the GPUs it requests. Set COUNT to the number of GPUs to request. Use a CUDA base image from your registry whose CUDA version is at or below the CUDA Version from the driver check; the container refuses to start on an older driver. Like the step 4 check, the pod cannot run on a node with a NoSchedule or NoExecute taint. If your GPU nodes carry either, do step 7 first, or rerun this check after it.

COUNT=2
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: gpu-allocation-check
namespace: default
spec:
runtimeClassName: nvidia
restartPolicy: Never
containers:
- name: cuda
image: <YOUR_REGISTRY>/nvidia/cuda:13.0.3-base-ubuntu24.04
command: ["bash", "-c", "nvidia-smi -L; nvidia-smi topo -p2p r"]
resources:
limits:
nvidia.com/gpu: $COUNT
EOF
kubectl -n default wait --for=jsonpath='{.status.phase}'=Succeeded pod/gpu-allocation-check --timeout=5m
kubectl -n default logs gpu-allocation-check
kubectl -n default delete pod gpu-allocation-check

Expect COUNT GPU lines, even though CUDA images ask for every GPU.

Step 6: Partition GPUs with MIG (optional)​

MIG divides a supported GPU into hardware-isolated instances. Describe the layouts the node can take in the mig-manager configuration, then select one with a node label. Add a configuration to gpu-operator-values.yaml under migManager, e.g., one that leaves GPUs 0 and 1 whole and partitions GPUs 2 and 3 into two 3g.40gb instances each. On a host bridged in pairs, keeping one bridged pair whole is what lets the whole GPUs serve tensor parallelism; see MIG and tensor parallelism.

migManager:
config:
create: true
name: custom-mig-config
default: all-disabled
data:
config.yaml: |-
version: v1
mig-configs:
all-disabled:
- devices: all
mig-enabled: false
gpu23-2x3g40gb:
- devices: [0, 1]
mig-enabled: false
- devices: [2, 3]
mig-enabled: true
mig-devices:
"3g.40gb": 2
gpuClientsConfig:
name: gpu-clients

Rerun the step 5 install with the updated values. mig-manager reads the new configuration only after the ConfigMap reaches its pod, and it acts only when the node label changes, so a label set too early fails with selected mig-config not present and is not retried. Confirm the ConfigMap holds the layout, then restart mig-manager and wait for it to be ready:

kubectl -n gpu-operator get configmap custom-mig-config \
-o jsonpath='{.data.config\.yaml}' | grep -c 'gpu23-2x3g40gb:'
1

kubectl wait fails with no matching resources found until the DaemonSet has created the new pod, so the restart first waits for the pod to exist:

kubectl -n gpu-operator delete pod -l app=nvidia-mig-manager
for i in $(seq 30); do # up to a minute for the new pod to appear
kubectl -n gpu-operator get pod -l app=nvidia-mig-manager -o name | grep -q . && break
sleep 2
done
kubectl -n gpu-operator wait --for=condition=Ready pod -l app=nvidia-mig-manager --timeout=5m

Select the layout and wait for mig-manager. It evicts the GPU workloads on the node, changes the MIG mode, and creates the instances; on the verification host this took about 4 minutes:

kubectl label node <NODE_NAME> nvidia.com/mig.config=gpu23-2x3g40gb --overwrite
sleep 30 # mig-manager marks the new layout pending before it applies it
while :; do
state=$(kubectl get node <NODE_NAME> -o jsonpath='{.metadata.labels.nvidia\.com/mig\.config\.state}')
echo "mig.config.state=$state"
case "$state" in success|failed) break ;; esac
sleep 15
done

Expect mig.config.state=success. On failed, read why, then restart mig-manager, which applies the labeled layout again when it starts:

kubectl -n gpu-operator logs -l app=nvidia-mig-manager --tail=30
kubectl -n gpu-operator delete pod -l app=nvidia-mig-manager
for i in $(seq 30); do # up to a minute for the new pod to appear
kubectl -n gpu-operator get pod -l app=nvidia-mig-manager -o name | grep -q . && break
sleep 2
done
kubectl -n gpu-operator wait --for=condition=Ready pod -l app=nvidia-mig-manager --timeout=5m

Then rerun the block that selects the layout, including its sleep 30. The label reads failed until the new pod updates it, so the loop alone stops at once.

Check:

nvidia-smi --query-gpu=index,mig.mode.current --format=csv,noheader
kubectl get node <NODE_NAME> -o jsonpath='{.status.allocatable}' | tr ',' '\n' | grep nvidia
0, Disabled
1, Disabled
2, Enabled
3, Enabled
"nvidia.com/gpu":"2"
"nvidia.com/mig-3g.40gb":"4"

MIG instances do not survive a reboot, and on a virtual machine stopped and deallocated the GPUs' MIG mode is reset too. mig-manager applies the labeled layout again when it starts after boot, and pods that request a MIG resource stay Pending until it finishes. After each boot, check that nvidia.com/mig.config.state is success; if it is failed, see Troubleshooting.

MIG and tensor parallelism​

A deployment that spans several GPUs needs a peer-to-peer path between every pair of GPUs the device plugin can give it. Check the paths between GPUs:

nvidia-smi topo -p2p r
GPU0 GPU1 GPU2 GPU3
GPU0 X OK NS NS
GPU1 OK X NS NS
GPU2 NS NS X OK
GPU3 NS NS OK X

Only OK is a peer path. On this host the GPUs are bridged in pairs, 0 with 1 and 2 with 3, and have no path between the pairs. A multi-GPU request can land on any whole GPUs, so this node can offer tensor parallelism only if its whole GPUs are exactly one bridged pair. Partition the GPUs outside that pair with MIG, here GPUs 2 and 3, as the step 6 layout does. The owner's detector claims the p2p feature only when every pair of whole GPUs reports OK; see One replica across several devices.

Step 7: Remove taints from GPU nodes​

Tenant-mode engine pods carry no tolerations, so they never schedule on a node with a NoSchedule or NoExecute taint. GPU nodes used for tenant-mode inference must carry neither.

Check every node that advertises an NVIDIA resource:

kubectl get nodes -o json | jq -r '.items[]
| select(.status.allocatable | keys | any(startswith("nvidia.com/")))
| [.metadata.name, ([.spec.taints[]? | select(.effect == "NoSchedule" or .effect == "NoExecute")
| .key + ":" + .effect] | join(","))]
| @tsv'
gpu-node-1

Expect each GPU node's name with nothing after it. For a node listed with a taint, remove it, repeating the key and effect the check printed, e.g., nvidia.com/gpu:NoSchedule:

kubectl taint nodes <NODE_NAME> <TAINT_KEY>:<EFFECT>-

A taint that your cluster needs, such as one that reserves GPU nodes for other workloads, cannot stay on the nodes Kamiwaza serves from.

Step 8: Pin the GPU stack​

An unattended upgrade of the NVIDIA libraries while the old kernel module is still loaded breaks the driver until a reboot, and nvidia-smi then fails with Failed to initialize NVML: Driver/library version mismatch. Hold the driver, toolkit, and kernel packages, and upgrade them together in a maintenance window:

PKGS=$(dpkg-query -W -f='${db:Status-Abbrev} ${Package}\n' | awk '$1=="ii"{print $2}' \
| grep -E '^(nvidia-|libnvidia-|linux-modules-nvidia-|linux-objects-nvidia-|linux-signatures-nvidia-|linux-image-|linux-generic)')
sudo apt-mark hold $PKGS

Check:

apt-mark showhold | grep -E '^(nvidia-driver|nvidia-container-toolkit|linux-modules-nvidia)'

Expect the driver, the container toolkit, and the kernel module package. Upgrade with sudo apt-mark unhold $PKGS, upgrade, reboot, and hold again.

Share GPUs with the VRAM plugin (optional)​

The VRAM plugin lets several deployments share one GPU by its memory. Install it only if the cluster owner's signed profile has a binding with the vram-plugin allocator; see Tenant-mode inference. It accounts for memory and does not isolate it.

The plugin is a DaemonSet, kamiwaza-vram-plugin, with a ServiceAccount and a ClusterRole. It advertises each GPU it shares as kamiwaza.ai/vram-gb-gpu-<N>, where <N> is the GPU's index and the count is its memory in decimal GB. A cluster administrator installs it once per cluster; the Kamiwaza release does not. Its chart is published with the release from Kamiwaza 1.3.2; 1.3.0 and 1.3.1 publish no chart for it. The chart is for Kubernetes clusters: it ships no OpenShift SecurityContextConstraints.

Install it from the release:

helm upgrade --install vram-plugin oci://ghcr.io/kamiwaza-ai/releases/kamiwaza/charts/vram-plugin \
--version 1.3.3 -n kamiwaza

The chart pulls its image from the release, ghcr.io/kamiwaza-ai/releases/kamiwaza/images/vram-plugin, without credentials. In an air gap, install it from your registry instead; see Install Without Internet Access. In the kamiwaza namespace, Helm may print Pod Security restricted warnings for the plugin; they are expected.

By default the plugin runs on every node that GPU feature discovery from step 5 labels nvidia.com/gpu.present=true, tolerates the nvidia.com/gpu:NoSchedule taint, sizes each GPU from the nvidia.com/gpu.memory label, and, unless you set gpuIndices, shares every GPU, MIG-enabled ones included. It sees neither MIG slices nor whole-device allocations, so keep two kinds of GPU out of it:

  • A MIG-enabled GPU. The plugin advertises the card's full memory, and a deployment it places there finds no CUDA device and fails to start.
  • A GPU that something allocates whole. A GPU the plugin shares is still offered as nvidia.com/gpu, and a whole-device binding covers every node of its architecture, so it can take that GPU while accounted deployments use it. Do not publish a whole-device binding for an architecture whose nodes the plugin shares GPUs on, and run nothing else there that requests nvidia.com/gpu.

List the GPUs to share, by their indices as nvidia-smi prints them, in gpuIndices. For example, on a node whose GPU 0 is MIG-partitioned and GPU 1 is whole, --set 'gpuIndices={1}' shares GPU 1, and GPU 0's slices stay available to MIG bindings. The step 6 example node partitions GPUs 2 and 3 and keeps GPUs 0 and 1 for whole-device bindings, so the plugin has no GPU to share there. Every node the plugin runs on must have each listed index; on a node without one, the plugin exits.

A node you kept out still carries nvidia.com/gpu.present, so the default placement would run the plugin there too. Keep it off such a node: label the node kamiwaza.ai/gpu-vendor=none, which the default placement excludes, or set affinity yourself.

The plugin registers with the kubelet only through /var/lib/kubelet/device-plugins, and the chart has no value to move it. On a node whose kubelet root is elsewhere, such as k0s's default /var/lib/k0s/kubelet, its pod does not become ready: it waits on the missing directory, or restarts when it cannot reach the kubelet. On k0s, install the plugin only on nodes whose k0s runs with --kubelet-root-dir=/var/lib/kubelet, as Kamiwaza's own setup tooling installs it.

Install it as release vram-plugin in the kamiwaza namespace, as above: the same release name and namespace Kamiwaza's setup tooling uses, so the two do not contend for it. Its ServiceAccount may get, watch, and patch every Node, and anyone who can create pods in the kamiwaza namespace can run a pod as that ServiceAccount. Limit who can create pods there.

Check:

kubectl -n kamiwaza rollout status daemonset/kamiwaza-vram-plugin --timeout=5m
kubectl get node <NODE_NAME> -o jsonpath='{.status.allocatable}' | tr ',' '\n' | grep vram-gb

Expect successfully rolled out, then one kamiwaza.ai/vram-gb-gpu-<N> line for each GPU the plugin shares on the node.

Next step​

Give the cluster owner the node facts the signed profile attests. They run the detector from the Kamiwaza core image on each GPU node; see Tenant-mode inference.

Troubleshooting​

GPU Operator pods fail with RunContainerError: context deadline exceeded, and the operator never reports ready. Persistence mode is off; see step 2. Time nvidia-smi -L on the node: several seconds means persistence mode is off.

GPU pods stay Pending with Insufficient nvidia.com/gpu. Compare the request with the node's allocatable resources, from step 5. No NVIDIA resources at all means the device plugin is not running on the node, as expected on a node you kept out; otherwise check kubectl -n gpu-operator get pods -o wide. On k0s, a device plugin that runs but registers nothing means hostPaths.kubeletRootDir does not match the kubelet root: /var/lib/k0s/kubelet on a default k0s, and the GPU Operator's default, /var/lib/kubelet, on a k0s started with --kubelet-root-dir=/var/lib/kubelet.

MIG resources read 0 after a reboot, and nvidia.com/mig.config.state is failed. mig-manager could not apply the layout when the node started. Read its log:

kubectl get node <NODE_NAME> -o jsonpath='{.metadata.labels.nvidia\.com/mig\.config\.state}'; echo
kubectl get node <NODE_NAME> -o jsonpath='{.status.allocatable}' | tr ',' '\n' | grep mig
kubectl -n gpu-operator logs -l app=nvidia-mig-manager --tail=50 | grep -i 'fail\|error'

failed to stop service nvidia-persistenced.service: timeout stopping service means stopping the persistence daemon, which mig-manager does before a MIG change, took longer than mig-manager's 30-second limit. On the verification host it took 32 seconds while the node was still starting. Once the node is up, restart mig-manager, which applies the layout again:

kubectl -n gpu-operator delete pod -l app=nvidia-mig-manager
for i in $(seq 30); do # up to a minute for the new pod to appear
kubectl -n gpu-operator get pod -l app=nvidia-mig-manager -o name | grep -q . && break
sleep 2
done
kubectl -n gpu-operator wait --for=condition=Ready pod -l app=nvidia-mig-manager --timeout=5m

Then wait for success with the step 6 block that selects the layout, including its sleep 30: the label reads failed until the new pod updates it. selected mig-config not present means the label names a layout mig-manager has not loaded; see step 6.

GPU pods stay Pending with had untolerated taint. A GPU node is tainted; see step 7.

GPU pods fail with no runtime for "nvidia" is configured. The node's containerd has no nvidia runtime; check runtimeHandlers as in step 3. On k0s, check the drop-in's version = 3 line and restart k0s.

nvidia-smi prints Driver/library version mismatch. The libraries were upgraded while the old kernel module is loaded. Reboot the node, then pin the stack as in step 8.

An engine exits at start with CUDA driver version is insufficient for CUDA runtime version. The host driver is older than the engine image's CUDA major. Upgrade the driver to the floor in Supported GPUs and driver floors, or have the owner attest the driver's CUDA major, so the platform picks an image the driver can run; see Attest the NVIDIA driver's CUDA major.

nvidia-smi finds no GPU with Secure Boot on. The kernel refused an unsigned module. Check dmesg | grep -i 'module verification failed', then use the prebuilt, signed modules from step 1 or enroll a MOK as in Secure Boot.