Skip to main content
Version: 1.3.3 (Latest)

Tenant-mode inference

Support scope

This page documents tenant-mode inference for Kamiwaza 1.3.3. Support covers the configurations marked Supported in Compatibility scope, qualified on k0s Kubernetes 1.36.

Tenant-mode inference is how a namespaced Kamiwaza install serves local models on GPUs and CPUs. Kamiwaza installs into one namespace and cannot read the cluster's Nodes, so it cannot discover GPUs itself. The cluster owner describes the hardware instead, in a profile that the owner signs, and Kamiwaza places each deployment on hardware the signed profile vouches for.

A default install has tenant-mode inference off and cannot use GPUs: a GPU deployment is refused with HTTP 422. To serve models on GPUs:

  1. Prepare the GPU nodes; see GPU Nodes.
  2. Create the owner key, detect each GPU host's facts, and sign a profile, with the owner tool in the Kamiwaza core image. See Turn on tenant-mode inference.
  3. Point the install at the signed profile, in four Helm values.

Compatibility scope​

Live qualification for this release targets k0s Kubernetes 1.36 only. Chart, API, and client contract checks cover Kubernetes 1.28 through 1.36, but passing a contract check is not a live support claim for that minor version: only 1.36 has been exercised end to end. The allocation classes are:

ClassEnginesKubernetes resourceIsolation statementRelease status
CPUllama.cpp, whisper.cppcpu and memory requestsKubernetes scheduler allocationSupported on k0s 1.36
Whole NVIDIA GPUllama.cpp, vLLM, diffusion (CUDA)nvidia.com/gpuWhole-device allocationSupported on k0s 1.36, one device per replica only. A deployment that spans more than one device is implemented but not qualified live
Whole AMD GPUNone in this release; no ROCm engine image shipsamd.com/gpuWhole-device allocationNot qualified in this release. Do not deploy against a support claim
NVIDIA MIGllama.cpp, vLLM, diffusion (CUDA)Qualified nvidia.com/mig-*Hardware partition, exact shape onlySupported on k0s 1.36, recorded slice shape only
VRAM-plugin sharingllama.cpp, vLLM, diffusion (CUDA)Owner-advertised kamiwaza.ai/vram-gb-*Accounted sharing, not hard isolationSupported on k0s 1.36, recorded plugin and profile only. Optional: the cluster owner installs the VRAM plugin

Engines​

Tenant mode serves only the engines in the platform engine catalog that ships inside the release's core image:

EngineServesPlacementImages
llama.cppText generation; embeddings when the model configuration sets embedding and states its poolingGPU or CPUNVIDIA CUDA 13 and CUDA 12, amd64 and arm64; CPU amd64 and arm64
vLLMText generationGPU onlyNVIDIA CUDA 13, amd64 and arm64 (GB10 only)
DiffusionImage generationGPU only, one device per replicaNVIDIA CUDA 12 (built with 12.8), amd64 only
whisper.cppTranscriptionCPU onlyCPU amd64 and arm64

The engine catalog has no arm64 image for diffusion in this release. A binding on an arm64 node therefore cannot serve diffusion, and the platform refuses the deployment with no diffusion image is published for nvidia/arm64.

vLLM's arm64 image serves one card, the GB10 (DGX Spark, compute capability 12.1). The platform refuses vLLM on any other arm64 NVIDIA card, such as GH200, GB200, Jetson Thor, or a discrete Blackwell card in an arm64 server, with a reason such as vllm requires compute capability 12.1; this binding declares 10.0.

A GPU binding serves an engine only when the engine publishes an image for the binding's vendor and architecture, and the hardware its accelerator block attests meets the engine's requirements. Each NVIDIA image is built against one CUDA major and carries kernels down to a lowest compute capability, so the card and the host driver both limit which image can run:

EngineNVIDIA imageArchitecturesLowest compute capabilityHost driver
llama.cppCUDA 13amd64, arm647.5 (Turing)580.65.06 (R580) or later
llama.cppCUDA 12amd64, arm647.0 (Volta)525.60.13 or later; R570 or later for Blackwell GPUs
vLLMCUDA 13amd647.5, and the engine requires bf16 (8.0 or later)580.65.06 (R580) or later
vLLMCUDA 13arm6412.1 (GB10 only)580.65.06 (R580) or later
DiffusionCUDA 12 (built with 12.8)amd648.0, and the engine requires bf16R570 or later recommended, not enforced

llama.cpp therefore requires compute capability 7.0 or later: a Volta card runs the CUDA 12 image, and no image serves Pascal or older. vLLM requires a CUDA 13 driver. For how the platform picks between the llama.cpp images, see Attest the NVIDIA driver's CUDA major.

The platform compares only the driver's CUDA major, so it does not refuse the diffusion image on an older CUDA 12 driver. The image is built with PyTorch for CUDA 12.8. Unlike the llama.cpp CUDA 12 image, which runs on older CUDA 12 drivers under NVIDIA's minor-version compatibility, the diffusion image is not qualified below R570, so run diffusion on R570 or later.

A diffusion deployment serves one diffusers pipeline on one device through /v1/images/generations. The model must be a diffusers repository with model_index.json at its top; a single-file checkpoint is refused. A pipeline has no context length, so the platform sizes the device from the pipeline's weights. Tenant mode refuses a diffusion deployment that configures LoRA or Lightning adapters, a stream profile, trust_remote_code, DFloat11 settings, or more than one device (tensor_parallel_size or gpu_count above 1). When a configuration key causes the refusal, the reason names that key. A DFloat11 model that the platform recognizes by its name is also refused; that reason says DFloat11 models need more than one device and names no key. A stream profile that the platform detects from the model's name is kept, not refused.

When the release sets core.inferenceResources.enabled: true, every local deployment takes the tenant path, including a request without an inferenceResources block, which the platform derives. There is no fallback to the standard deployment path. An engine that the engine catalog does not list is refused. MLX is not in the catalog, so MLX is not supported in tenant mode in this release. A deployment without a block that selects such an engine receives HTTP 422, and the detail ends with the engine and the reason:

cannot derive inference resources for this deployment: engine 'mlx' is not qualified in this release

A request that carries its own block is accepted and then fails at dispatch with the same reason.

External endpoints are not refused. An external endpoint proxies a remote service and allocates no local resources, so on a tenant cluster the platform does not size it and creates no workload in the tenant namespace. On a tenant cluster, two external-endpoint requests receive HTTP 422:

  • A request that supplies an inferenceResources block:

    external endpoints allocate no local resources; omit inferenceResources
  • A request that sets a nonzero gpu_allocation or vram_allocation. A cluster without tenant resources still honors these overrides:

    external endpoints allocate no local resources; omit gpu_allocation and vram_allocation

One replica across several devices​

A deployment can span up to eight whole devices in a single Pod through tensor parallelism. This is implemented but not qualified live in this release; see Compatibility scope. The device count comes from tensor_parallel_size in the model configuration, or from a signed recipe when one pins it. An absent value is one device — never the node's device count, so a first deployment does not take a whole multi-GPU node.

Tensor parallelism needs a peer-to-peer path between the devices, so it is admitted only on a whole-device binding whose accelerator block attests the p2p feature. The owner tool's detector claims p2p only when every pair of whole-mode GPUs on the host reports a peer path, because the device plugin can give a multi-GPU request any whole GPUs. MIG-mode GPUs are left out: a whole-device binding is never given one. On a host whose GPUs are bridged in pairs, partition the GPUs outside one bridged pair with MIG, so its whole GPUs are exactly that pair; see MIG and tensor parallelism.

A MIG partition does not have a peer path, and an accounted VRAM share is a slice of a device the workload does not exclusively hold, so both refuse a count above one rather than quietly serving a single device to an engine told to shard across several. A count that is not a whole number, is below one, or exceeds eight is refused for the same reason: the request and the Pod must agree.

Pipeline and data parallelism are not part of this release.

Ownership boundary​

The cluster owner supplies drivers, device plugins, optional VRAM sharing, RuntimeClasses, quotas and policy, and the signed profile. Every GPU binding in a profile attests its hardware in an accelerator block, which the owner tool detects on the host; a CPU binding needs only its architecture. Recipe catalogs are optional, GPU and CPU alike: without one, the platform derives each deployment, and a published recipe pins a specific model and configuration. The tenant supplies only namespaced Helm and serving inputs.

Kubernetes namespace-admin permissions do not grant the Kamiwaza admin role. Model deployment through the API, SDK, or UI remains Kamiwaza-admin-only.

A deployment request needs no inferenceResources block: the platform derives one. For a GPU deployment it sizes memory from the model's own metadata and the context the model configuration sets (max_model_len), or else the context the model declares, and places the request on the smallest binding that holds it. A deployment goes to CPU instead when it sets force_cpu, when its engine ships only a CPU build (whisper.cpp), or when the owner declared no GPU hardware; it is sized from the platform's CPU floor and the model's weights, and runs the platform's CPU build for the binding's architecture. A request that cannot be served as asked -- no context declared anywhere, or an embedding configuration that does not state its pooling -- is refused with HTTP 422 and the reason; one the platform cannot evaluate, such as an unreadable profile bundle, with HTTP 503.

The tenant chart creates no cluster-scoped GPU resource and does not require nodes/get, nodes/list, or nodes/watch. A profile or allocator failure is reported as a failure; the platform does not silently fall back to CPU or a weaker isolation class.

Turn on tenant-mode inference​

Before you start​

You needNotes
Prepared GPU nodesEvery requirement on GPU Nodes passes. Skip for CPU-only serving.
The release's core image, by digestThe owner tool ships in it. Sign with the core image of the release you install, so the release that verifies the profile is the one that signed it.
Docker or PodmanOn the machine where you sign, not on the GPU nodes: step 3 runs the detector as a pod. The owner tool runs in the core image; it contacts no cluster and no network.
kubectl and jqOn the machine where you publish the profile.
Do not install Docker on a Kubernetes node

Docker sets the host's iptables FORWARD policy to DROP, which breaks pod networking on the node. Run the owner tool on a separate machine, and the detector as a pod.

Read the core image from a running install:

kubectl -n kamiwaza get deploy core-scheduler \
-o jsonpath='{.spec.template.spec.containers[?(@.name=="core")].image}'; echo
<YOUR_REGISTRY>/kamiwaza-ai/releases/kamiwaza/images/core@sha256:<DIGEST>

Before the install, take the same reference from the release's image list; see Install Without Internet Access.

If your registry requires credentials, log the container engine in to it first: docker login <REGISTRY_HOST>, or podman login <REGISTRY_HOST>. Podman also reads an existing credentials file named with --authfile <FILE> or the REGISTRY_AUTH_FILE environment variable, in the same format as the install's pull Secret.

Set the image, and define a shell function that runs the owner tool with the current directory mounted. With Docker:

export CORE_IMAGE="<YOUR_REGISTRY>/kamiwaza-ai/releases/kamiwaza/images/core@sha256:<DIGEST>"
owner_tool() {
docker run --rm -i --user "$(id -u):$(id -g)" -v "$PWD:/work" -w /work \
"$CORE_IMAGE" python -m kamiwaza.serving.inference_resources.owner_tool "$@"
}

With rootless Podman, which needs --userns=keep-id to write files you own, and :Z on the mount to use the directory on a host with SELinux, such as RHEL or Fedora:

export CORE_IMAGE="<YOUR_REGISTRY>/kamiwaza-ai/releases/kamiwaza/images/core@sha256:<DIGEST>"
owner_tool() {
podman run --rm -i --userns=keep-id --user "$(id -u):$(id -g)" -v "$PWD:/work:Z" -w /work \
"$CORE_IMAGE" python -m kamiwaza.serving.inference_resources.owner_tool "$@"
}

:Z relabels the current directory for the container, and Podman ignores it on a host without SELinux. Podman refuses to relabel your home directory, so run the function from a directory of its own.

The steps below use the function; the GPU detector in step 3 runs on the GPU node instead, because it must see the GPUs.

Step 1: Create the owner signing key​

owner_tool keygen --private-key-out owner-ed25519-private.pem
{
"keyId": "sha256:6e1f...",
"publicKey": "8Ljv..."
}

The tool writes an unencrypted Ed25519 private key readable only by you, and refuses to replace an existing file. It prints the public key, which goes into the Helm values as ownerPublicKey, and the keyId your profile names. Keep the private key out of source control, out of the tenant namespace, and out of any evidence you record. To print the public half again:

owner_tool pubkey --private-key owner-ed25519-private.pem

Step 2: Choose a cluster binding ID​

Pick a clusterBindingId for this cluster: 1 to 128 characters, starting with a letter or digit, then letters, digits, ., _, :, or -. The profile and the Helm values must carry the same value; Kamiwaza refuses a profile signed for another binding.

Step 3: Detect each GPU host's facts​

Run the detector on every GPU node, with the node's GPUs visible to it. Run it as a one-shot pod on the nvidia RuntimeClass: the node then needs no container engine of its own. The pod runs in the Kamiwaza namespace, so it pulls the core image with the install's pull Secret; if your install has no pull Secret, delete the two imagePullSecrets lines. Set NODE_NAME to the GPU node:

NODE_NAME="<NODE_NAME>"
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: owner-tool-detect
namespace: kamiwaza
annotations:
sidecar.istio.io/inject: "false"
spec:
runtimeClassName: nvidia
restartPolicy: Never
nodeSelector:
kubernetes.io/hostname: $NODE_NAME
imagePullSecrets:
- name: kamiwaza-pull
containers:
- name: detect
image: $CORE_IMAGE
command: ["python", "-m", "kamiwaza.serving.inference_resources.owner_tool", "detect"]
env:
- {name: NVIDIA_VISIBLE_DEVICES, value: all}
- {name: NVIDIA_DRIVER_CAPABILITIES, value: "compute,utility"}
EOF
kubectl -n kamiwaza wait --for=jsonpath='{.status.phase}'=Succeeded pod/owner-tool-detect --timeout=5m
kubectl -n kamiwaza logs owner-tool-detect
kubectl -n kamiwaza delete pod owner-tool-detect

The detector prints the accelerator block for the node's GPUs:

{
"accelerator": {
"computeCapability": "8.0",
"cudaDriverMajor": 13,
"features": [
"bf16",
"mig",
"nvlink",
"p2p"
],
"vendor": "nvidia"
}
}

A GPU node that already has a container engine can run the detector directly instead. With Podman, as root, through the NVIDIA runtime; add --authfile <FILE> if your registry requires credentials:

sudo podman run --rm --runtime /usr/bin/nvidia-container-runtime \
-e NVIDIA_VISIBLE_DEVICES=all "$CORE_IMAGE" \
python -m kamiwaza.serving.inference_resources.owner_tool detect

podman run --device nvidia.com/gpu=all does not work with Podman 4.9, the version Ubuntu 24.04 ships: it cannot read the device description that current NVIDIA container toolkits write, and fails with unresolvable CDI devices nvidia.com/gpu=all. On a node that already runs Docker with the NVIDIA runtime, docker run --rm --gpus all "$CORE_IMAGE" python -m kamiwaza.serving.inference_resources.owner_tool detect also works. Do not install Docker on a Kubernetes node to run it.

The detector reads the facts from nvidia-smi instead of expecting you to know them:

  • computeCapability is the GPUs' compute capability. The detector refuses a node with two kinds of GPU: a binding cannot single out one kind on such a node, so keep it out of tenant-mode GPU serving; see GPU Nodes.
  • cudaDriverMajor is the major of the CUDA Version in nvidia-smi's header. The detector exits with an error rather than guess when it cannot read it.
  • bf16 and fp8 follow from the compute capability, and nvlink from the links the GPUs report.
  • p2p appears only when every pair of whole-mode GPUs has a peer path; see One replica across several devices.
  • mig appears when the GPUs can be partitioned, whether or not a GPU is in MIG mode now. It is informational: no engine requires it, and a MIG binding does not need it.

Copy the block into the binding for those nodes. A binding covers every node of its architecture that advertises its resource (see nodeSelector in step 4), so when those nodes differ, attest only what holds on all of them: the lowest compute capability and cudaDriverMajor, and only the features every node has. Drop p2p and nvlink from a MIG or VRAM-share binding: a slice has neither.

Step 4: Write the profile envelope​

The envelope lists the profiles the cluster offers. Each profile has one guarantee and one or more bindings; each binding names the Kubernetes resource a deployment requests and the nodes it runs on. This envelope offers CPU serving on amd64 nodes and whole NVIDIA GPUs. Save it as owner-profile-envelope.json:

{
"schemaVersion": 1,
"clusterBindingId": "<CLUSTER_BINDING_ID>",
"namespace": "kamiwaza",
"version": 1,
"keyId": "<KEY_ID>",
"issuedAt": "<ISSUED_AT>",
"expiresAt": "<EXPIRES_AT>",
"evidenceId": "gpu-nodes-checked-2026-10-01",
"profiles": [
{
"name": "cpu-standard",
"capability": "cpu",
"guarantees": [],
"compatibleVariantIds": ["cpu-default"],
"bindings": [{
"id": "cpu-amd64",
"allocator": "cpu",
"architecture": "amd64",
"runtime": {"kind": "cpu", "variantId": "cpu-default", "image": "<LLAMACPP_CPU_IMAGE>"},
"capacity": {"kind": "scheduler"},
"nodeSelector": {"kubernetes.io/arch": "amd64"}
}]
},
{
"name": "nvidia-whole",
"capability": "gpu",
"guarantees": ["whole-device"],
"compatibleVariantIds": ["llamacpp-nvidia-amd64", "vllm-nvidia-amd64", "diffusion-nvidia-amd64"],
"bindings": [{
"id": "nvidia-whole-amd64",
"allocator": "extended-resource",
"architecture": "amd64",
"accelerator": {
"vendor": "nvidia",
"computeCapability": "8.0",
"cudaDriverMajor": 13,
"features": ["bf16", "mig", "nvlink", "p2p"]
},
"capacity": {"kind": "scheduler", "usableMemory": "79Gi"},
"resource": {"name": "nvidia.com/gpu", "quantity": 1},
"runtime": {"kind": "cuda", "variantId": "llamacpp-nvidia-amd64", "image": "<LLAMACPP_CUDA_IMAGE>"},
"runtimeClassName": "nvidia",
"nodeSelector": {"kubernetes.io/arch": "amd64"}
}]
}
]
}

Fill in the placeholders:

FieldValue
clusterBindingIdThe ID from step 2.
namespaceThe namespace Kamiwaza is installed in.
versionAn integer you raise with every new revision. Kamiwaza refuses a revision below core.inferenceResources.minimumBundleVersion (default 1).
keyIdThe keyId from step 1.
issuedAt, expiresAtUTC timestamps with a Z suffix, such as 2026-10-01T00:00:00Z. Kamiwaza refuses the profile outside that window, so publish a new revision before expiresAt.
evidenceIdA label for the checks behind this revision, for your records.
acceleratorThe block from step 3.
nodeSelectorkubernetes.io/arch only; this release accepts no other key. A binding therefore covers every node of that architecture that advertises its resource, so its accelerator facts and usableMemory must hold on all of them.
usableMemoryThe GPU memory a deployment may plan on, less what the driver and CUDA context hold back. Write a whole number of Gi, Mi or Ki; the profile schema refuses a fraction such as 121.7Gi. The right value depends on the binding; see the list below this table.
runtime.imageA digest-pinned engine image. Derived deployments take their image from the engine catalog; this value is used only for a deployment that a recipe pins, which runs it as written. In a mirrored or air-gapped install, name your mirror's copy. Print the catalog's images with the command below.

Set usableMemory for each kind of binding as follows:

  • Whole device or MIG partition on a discrete card: the device's or the slice's memory less the driver's and CUDA context's share, 79Gi on an 80 GB card. Understating it is the safe direction.
  • VRAM plugin: state it accurately. vLLM reserves a fraction of the whole card: 90%, the engine catalog default, times the request divided by usableMemory. Understating usableMemory therefore raises what vLLM reserves, and the reservation passes the grant once usableMemory is about a tenth below the card's memory.
  • Whole-device GB10: its GPU memory is the host's, so set usableMemory to the device total rounded down, 121Gi on a DGX Spark, whose total is about 121.69 GiB. vLLM's budget is the request divided by usableMemory, rounded up to a hundredth, applied to that total. At 121Gi vLLM already reserves slightly more than the request: 121.69/121 of it, plus up to 0.01 of the total from the rounding. A smaller value reserves further above the request. The host's own processes use the same memory, so a request must also fit in what the host leaves free. A VRAM-plugin binding on a GB10 follows the VRAM-plugin rule instead.

A GPU profile's compatibleVariantIds must list <engine>-nvidia-<architecture> for each engine it may serve, and each binding's runtime.variantId must be in its profile's list. To print the engine images the core image carries:

docker run --rm "$CORE_IMAGE" python -c \
"from importlib.resources import files; print(files('kamiwaza.serving.inference_resources').joinpath('engine-catalog.json').read_text())" \
| jq -r '.engines[] | .engine as $e | .images[] | [$e, .vendor, .architecture, (.cuda // "-" | tostring), .image] | @tsv'
llamacpp nvidia amd64 13 ghcr.io/kamiwaza-internal/containers/images/llamacpp-cuda13@sha256:a482...
llamacpp cpu amd64 - ghcr.io/kamiwaza-internal/containers/images/llamacpp-cpu@sha256:ccbd...

For MIG, add a profile whose guarantee is mig and whose binding requests the partition the node advertises. This profile offers the 3g.40gb slices from GPU Nodes; add it to the profiles list:

{
"name": "nvidia-mig-3g40gb",
"capability": "gpu",
"guarantees": ["mig"],
"compatibleVariantIds": ["llamacpp-nvidia-amd64", "vllm-nvidia-amd64", "diffusion-nvidia-amd64"],
"bindings": [{
"id": "nvidia-mig-3g40gb-amd64",
"allocator": "extended-resource",
"architecture": "amd64",
"accelerator": {
"vendor": "nvidia",
"computeCapability": "8.0",
"cudaDriverMajor": 13,
"features": ["bf16", "mig"]
},
"capacity": {"kind": "scheduler", "usableMemory": "38Gi"},
"resource": {"name": "nvidia.com/mig-3g.40gb", "quantity": 1},
"runtime": {"kind": "cuda", "variantId": "llamacpp-nvidia-amd64", "image": "<LLAMACPP_CUDA_IMAGE>"},
"runtimeClassName": "nvidia",
"nodeSelector": {"kubernetes.io/arch": "amd64"}
}]
}

The accelerator block is the node's, from step 3, without p2p and nvlink: a slice has neither. Set usableMemory to the slice's memory less headroom for the CUDA context. On the verification host that GPU Nodes describes, a 3g.40gb slice of an A100 80GB reports 40192 MiB in nvidia-smi, so the example uses 38Gi. Binding id values must be unique across the whole envelope.

An accounted-share profile, for the VRAM plugin, uses the vram-plugin allocator and {"kind": "accounted", ...} capacity, and reserves a bookkeeping bucket; it does not provide hard VRAM or compute isolation.

Step 5: Sign and publish the profile​

owner_tool sign --envelope owner-profile-envelope.json \
--private-key owner-ed25519-private.pem > owner-profile-configmap.json

The tool runs Kamiwaza's own verifier over what it signs, so it refuses a profile this release would reject, and names the check that failed:

error: signed bundle would be refused: invalid inference profile authorization: bundle is expired or its window is empty

It emits an immutable ConfigMap named kamiwaza-inference-profiles. Give each revision its own name, then apply it:

export PROFILE_CONFIGMAP=kamiwaza-inference-profiles-r1
jq --arg name "$PROFILE_CONFIGMAP" '.metadata.name = $name' owner-profile-configmap.json \
| kubectl apply --server-side --field-manager=owner-inference-publisher -f -

Check:

kubectl -n kamiwaza get configmap "$PROFILE_CONFIGMAP" -o jsonpath='{.immutable} {.data.bundle\.sha256}'; echo
true 9c1d...

Step 6: Set the Helm values​

Add these values to the install's values file and upgrade the release with it:

core:
inferenceResources:
enabled: true
bundleConfigMap: kamiwaza-inference-profiles-r1 # the name from step 5
ownerPublicKey: <OWNER_PUBLIC_KEY> # publicKey from step 1
clusterBindingId: <CLUSTER_BINDING_ID> # from step 2
Helm valueNotes
bundleConfigMapKamiwaza reads the profile only from the ConfigMap with this name. Left at the default, kamiwaza-inference-profiles, it cannot read a renamed revision, and if an older ConfigMap with the default name still exists, Kamiwaza uses that older profile, with no error, for as long as it verifies.
ownerPublicKeyThe canonical base64 of the raw 32-byte public key, as keygen and pubkey print it.
clusterBindingIdMust equal the profile's clusterBindingId.

On an install whose global.imagePullSecrets names the pull Secret, these four values are enough; an install that pulls every image without credentials needs no pull Secret. The staging values default as follows, and an explicit value always wins over its default:

ValueDefault with tenant-mode inference on
gpuRegistry, cpuRegistryThe in-cluster model registry, which the model downloader already pushes to. It serves plain HTTP, so gpuRegistryInsecure and cpuRegistryInsecure default to true with it.
gpuStagingImage, cpuStagingImageThe chart's own model-fetch image, relocated like every chart image.
gpuRegistryCredentialsSecret, cpuRegistryCredentialsSecretThe first entry of global.imagePullSecrets, the install's pull Secret. Not defaulted when global.openshift.serviceAccountImagePullSecretsOnly is true.

Engine pods carry only this staging credential, as their image pull Secret: the engine and model-fetch images pull with it. Set gpuRegistryCredentialsSecret and cpuRegistryCredentialsSecret yourself only when your mirror needs different credentials from the install's pull Secret.

Deployments created before you turn tenant-mode inference on carry no tenant resources, so they cannot start again once it is on. After the upgrade, stop each one and deploy it again (see Deploy through the API); the new deployment takes its resources from the signed profile and its image from the engine catalog. While core.context.embedding.autoProvision is on (the default, except on global.platform: osgm), the platform's embedding model needs only the stop: the platform deploys it again on the next embedding request. With it off, deploy the embedding model again yourself.

While core.context.embedding.autoProvision is on, turning tenant-mode inference on for an install with no earlier embedding deployment also deploys the platform's bundled embedding model, all-MiniLM-L6-v2, on CPU, so expect one CPU engine pod shortly after the upgrade. On an install that already served the embedding model before you turned tenant-mode inference on, the platform finds that earlier deployment, still DEPLOYED, and deploys nothing until you stop it and an embedding request arrives. With it off, the platform does not deploy the embedding model on its own.

Step 7: Verify​

Deploy a model that fits a GPU, from the Models page, then check the engine pods. The columns are the pod, its node, its RuntimeClass, the NVIDIA resources it was granted, and whether the grant was verified:

kubectl -n kamiwaza get pods -l kamiwaza.ai/deployment-id -o json | jq -r '.items[]
| [.metadata.name, .spec.nodeName, (.spec.runtimeClassName // "-"),
([.spec.containers[].resources.limits // {} | to_entries[]
| select(.key | startswith("nvidia.com/")) | "\(.key)=\(.value)"]
| join(",") | if . == "" then "-" else . end),
([.status.conditions[]? | select(.type == "kamiwaza.ai/AllocationVerified") | .status]
| first // "-")]
| @tsv'
tmi-3f2a9c... gpu-node-1 nvidia nvidia.com/gpu=1 True
tmi-8b1e04... gpu-node-1 nvidia nvidia.com/mig-3g.40gb=1 True
tmi-51c7aa... gpu-node-1 - - True

Expect the nvidia RuntimeClass and the GPU or MIG resource the deployment asked for on a GPU pod, and no RuntimeClass or NVIDIA resource on a CPU pod. Every pod should read True in the last column. Readiness alone is not proof of the requested grant: the kamiwaza.ai/AllocationVerified condition is.

A CPU pod requests the engine's CPU floor from the engine catalog, 100m for llama.cpp and 500m for whisper.cpp, and memory of at least 1.25 times the model's weights, rounded up to a whole GiB and never below the floor of 1Gi for llama.cpp or 2Gi for whisper.cpp. It sets requests only, no limits.

Recipe catalogs (optional)​

A recipe catalog pins a specific model and configuration to exact settings. Without one, every deployment is derived. Render each catalog with the owner tool, which validates it with the parser Kamiwaza runs when it mounts it and emits an immutable ConfigMap whose name carries the catalog's digest:

owner_tool render-catalog --catalog gpu-recipes.json \
--namespace kamiwaza --capability gpu > gpu-catalog.json
kubectl apply --dry-run=server -f gpu-catalog.json
kubectl apply -f gpu-catalog.json
export GPU_CATALOG_CONFIGMAP="$(jq -r '.metadata.name' gpu-catalog.json)"

Use --capability cpu for a CPU catalog. Point the release at the exact name with core.inferenceResources.gpuRecipeCatalogConfigMap or cpuRecipeCatalogConfigMap. Publish a new revision for every change and never mutate a published ConfigMap. Keep the previous revision until every consumer has rolled out; roll back by selecting the previous name.

A pinned deployment runs the image its recipe variant names, and the profile binding's runtime.image must be the same reference. Kamiwaza does not relocate it: the bundle's image map, global.imageRegistryPrefix, and global.airgap.registry.url apply only to the engine catalog's images. In a mirrored or air-gapped install, name your mirror's copy, by digest, in both the recipe variant and the binding.

Engine images​

The engine catalog in the core image pins every engine image by digest. Kamiwaza relocates those references once, when it loads the catalog, using the same values the chart relocates its own images with:

  1. core.inferenceResources.engineImageMap, when set. The tenant OCI bundle writes it for you; see Bundle installs.
  2. Otherwise global.imageRegistryPrefix: <prefix>/images/<name>@<digest>, where <name> is the last segment of the catalog's repository.
  3. Otherwise global.airgap.registry.url: <registry>/<catalog repository>@<digest>.
  4. Otherwise the reference as the catalog pins it.

An engine image must therefore be in your registry by digest, at the location the rule produces. Most catalog pins are untagged manifests, so a copy by tag does not bring them. Print the pins with the catalog command in step 4, and see Install Without Internet Access for the copy.

Engine pods pull with the staging registry Secret, core.inferenceResources.gpuRegistryCredentialsSecret and cpuRegistryCredentialsSecret, which default to the install's pull Secret; see step 6. There is no separate engine pull-secret setting. Set those two values only when your mirror needs different credentials.

A deployment that a recipe pins is the exception to the relocation above: it runs the image the owner signed, exactly as written. See Recipe catalogs.

After you change any of these values, redeploy the affected inference deployments: stop each one and deploy it again (see Deploy through the API). A running deployment keeps the engine reference it was planned with, and so does the platform's embedding model. Stop the embedding model's deployment too. While core.context.embedding.autoProvision is on, leave the redeploy to the platform: it deploys the model again on the next embedding request, such as POST /api/embedding/generate, and not before. That request returns HTTP 503 with the code embedding_deploying and a retry_after_seconds value, and embedding requests return it until the new deployment is ready, about 75 seconds on an air-gapped A100 install; retry them. With it off, the platform does not redeploy the model, and embedding requests return HTTP 503 with the code embedding_unavailable until you deploy it again yourself. On an upgrade, keep the previous release's engine images in your registry until every deployment has been redeployed; removing them breaks rescheduling.

Bundle installs​

A tenant OCI bundle carries the engine images for one architecture, and writes values/99-engine-images.yaml with the map that tells Kamiwaza where it imported each one. Keep that file last among the install's values files. A cluster with amd64 and arm64 nodes needs a bundle for each architecture.

A bundle can leave engines out. An engine the bundle does not carry is refused at admission, before the deployment record exists:

cannot derive inference resources for this deployment: engine 'vllm' is not qualified in this release

Attest the NVIDIA driver's CUDA major​

For each NVIDIA binding, the platform selects the newest image of the engine that both the card and the host driver can run. The card's limit is the computeCapability the binding attests. The driver's limit is cudaDriverMajor, an optional integer from 11 to 99 in the same accelerator block: the major of the CUDA Version that nvidia-smi prints in its header on the host (e.g., 12 on an R570 driver, which reports CUDA 12.8, and 13 on an R580 driver, which reports CUDA 13.0). An image is selected only when its CUDA major is at or below the attested value.

The detector in step 3 attests cudaDriverMajor by default. Without it, the platform does not judge the driver: a Turing or newer card on an R570 host is given the CUDA 13 llama.cpp image, which exits at startup with CUDA driver version is insufficient for CUDA runtime version. Re-detect and publish a new signed profile revision after any driver change, including a driver rollback.

A binding without cudaDriverMajor keeps verifying and selects the newest image its card can run:

  • llama.cpp resolves its CUDA 13 image on compute capability 7.5 or later, and its CUDA 12 image on Volta (7.0).
  • vLLM resolves its CUDA 13 image on amd64 on compute capability 8.0 or later, because the engine also requires bf16, and on arm64 only on a GB10 (12.1).
  • Diffusion resolves its one image, CUDA 12, on compute capability 8.0 or later, on amd64 only.
  • A Pascal (6.x) or older card is refused.

vLLM ships only a CUDA 13 image. On a binding that attests cudaDriverMajor: 12 the platform refuses it:

host driver supports CUDA 12 (attested cudaDriverMajor); engine vllm ships nvidia/amd64 images for CUDA 13+

On a binding that does not attest the fact, the platform cannot tell, and the vLLM image fails at startup on a CUDA 12 driver. When the card is older than every image, the reason names the compute capability instead:

llamacpp requires compute capability 7.0; this binding declares 6.1

A deployment without an inferenceResources block is refused with HTTP 422 before the deployment record exists, with the reason after cannot derive inference resources for this deployment:. This happens only when the driver or the card rules out every image on every GPU binding in the profile. A request that carries its own block is refused at binding selection: the deployment fails with NoFit and the reason in its error message. To resolve a driver refusal, upgrade the host driver, re-detect, and publish a new signed profile revision. No owner action resolves a card refusal, because the release ships no image for that card.

Before you roll Kamiwaza back to a release earlier than 1.3.0, publish a new signed profile revision without cudaDriverMajor and point the release at it. Earlier releases refuse the unknown key and reject the whole profile, so every tenant deployment, CPU deployments included, fails authorization.

Requests that name their own resources​

A deployment from the Models page needs no inferenceResources block: the platform derives one. To choose the profile, device count or memory yourself, deploy through the API, POST /api/serving/deploy_model, with the block in the request body. The block names the logical profile and capability rather than a Node or a provider-specific inventory:

{
"inferenceResources": {
"schemaVersion": 1,
"accelerator": {
"capability": "gpu",
"count": 1,
"memory": {"minimum": "1Gi"},
"isolation": "any-qualified",
"profile": "nvidia-whole"
},
"runtime": {"selection": "automatic"}
}
}

For cpu-standard, use the CPU request shape instead:

{
"inferenceResources": {
"schemaVersion": 1,
"cpu": {
"architecture": "amd64",
"requests": {"cpu": "2", "memory": "4Gi"}
},
"runtime": {"selection": "automatic"}
}
}

Do not add Node names, Node selectors owned by the tenant, or cluster-scoped resources to the tenant request.

Deploy through the API​

Get a token as an administrator. Set KAMIWAZA_URL to your Kamiwaza domain; if the edge certificate is issued by a private authority, add --cacert <CA_FILE> to each curl command. The admin password is the one from your install page's verify step:

KAMIWAZA_URL="https://<KAMIWAZA_DOMAIN>"
printf 'admin password: '; read -rs ADMIN_PASSWORD; echo # typed, not echoed
# curl reads the password from stdin, so it never appears on a command line.
TOKEN=$(printf '%s' "$ADMIN_PASSWORD" | curl -sS -X POST "$KAMIWAZA_URL/api/auth/token" \
-H "Content-Type: application/x-www-form-urlencoded" \
--data-urlencode "username=admin" --data-urlencode "password@-" \
| jq -r '.access_token')
unset ADMIN_PASSWORD

Find the model, its configuration and its weights file. A model is in the list once it has been downloaded; see Downloading Models:

curl -sS -H "Authorization: Bearer $TOKEN" "$KAMIWAZA_URL/api/models/" \
| jq -r '.[] | [.id, .name] | @tsv'
MODEL_ID="<MODEL_ID>"
curl -sS -H "Authorization: Bearer $TOKEN" "$KAMIWAZA_URL/api/model_configs/?model_id=$MODEL_ID" \
| jq -r '.[] | [.id, .name, .default] | @tsv'
curl -sS -H "Authorization: Bearer $TOKEN" "$KAMIWAZA_URL/api/models/$MODEL_ID" \
| jq -r '.m_files[] | [.id, .name] | @tsv'

Write the request, with the configuration marked true in the second list and the weights file to serve, then deploy:

cat > deploy.json <<'JSON'
{
"m_id": "<MODEL_ID>",
"m_config_id": "<CONFIG_ID>",
"m_file_id": "<FILE_ID>",
"inferenceResources": {
"schemaVersion": 1,
"accelerator": {
"capability": "gpu",
"count": 1,
"memory": {"minimum": "8Gi"},
"isolation": "any-qualified",
"profile": "nvidia-whole"
},
"runtime": {"selection": "automatic"}
}
}
JSON
curl -sS -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
--data @deploy.json "$KAMIWAZA_URL/api/serving/deploy_model"
"631e7142-3417-487c-be65-d7715ba6cbf2"

The response is the deployment's ID, returned before the launch finishes. Follow its status until DEPLOYED; see Model Deployment for the statuses:

curl -sS -H "Authorization: Bearer $TOKEN" \
"$KAMIWAZA_URL/api/serving/deployment/<DEPLOYMENT_ID>/status"; echo

To stop a deployment, e.g. before deploying it again after a registry change:

curl -sS -X DELETE -H "Authorization: Bearer $TOKEN" \
"$KAMIWAZA_URL/api/serving/deployment/<DEPLOYMENT_ID>"; echo

The response is true, and the deployment's status becomes STOPPED.

Upgrade and recover​

To change the profile, publish a new revision under a new name with a higher version, point bundleConfigMap at it, upgrade the release, and repeat the checks in step 7. To roll back, point bundleConfigMap at the previous revision. If a create, readiness, authorization, or deadline step fails, the platform marks the allocation failed and removes only resources owned by that attempt; keep the prior known-good workload until the replacement passes its checks.

Upgrading Kamiwaza, including its core image, does not require a new signed profile. Sign a new revision only when your profiles change, or when the release notes say the profile schema changed.

To rotate the owner key, create a new key, sign a new revision with it, and set ownerPublicKey and bundleConfigMap together in one upgrade.

Troubleshooting​

A GPU deployment returns HTTP 422 with Core cannot see a GPU on this install. Tenant-mode inference is off. Turn it on as described in Turn on tenant-mode inference.

A deployment returns HTTP 503 with owner profile bundle: invalid inference profile authorization: and a reason. Kamiwaza found the profile but refused it. The reason names the check:

ReasonFix
bundle did not arrive from the configured ConfigMapbundleConfigMap does not name the published revision, or it is in another namespace.
bundle targets another cluster binding, namespace, or a version below the configured floorCompare the profile's clusterBindingId, namespace, and version with the Helm values.
bundle is signed by a key this cluster does not trustownerPublicKey is not the key that signed the profile. Print it with owner_tool pubkey.
bundle is expired or its window is emptyPublish a new revision with a later expiresAt.

A deployment returns HTTP 422 with engine '<name>' is not qualified in this release. The engine is not in the engine catalog on this install: MLX is never in it, and a bundle install may have left the engine out.

A deployment returns HTTP 422 naming the driver's CUDA major or the card's compute capability. See Attest the NVIDIA driver's CUDA major.

A request with its own block fails with NoFit. No binding in the profile holds the request: compare its memory, count, and isolation with the profile's bindings.

A deployment stays REQUESTED with no error, and its engine pod is Pending. The pod cannot be scheduled, and the deployment reports no error until the per-attempt scheduling deadline, schedulingTimeoutSeconds, passes. A tensor-parallel deployment that needs two whole GPUs while another deployment holds one of them is a common cause. Find the pod and read its events:

kubectl -n kamiwaza get pods -l kamiwaza.ai/deployment-id --field-selector=status.phase=Pending
kubectl -n kamiwaza describe pod <POD_NAME> | sed -n '/Events:/,$p'

Insufficient nvidia.com/gpu means no node has enough free GPUs of the requested kind; stop a deployment that holds one, or add GPUs. had untolerated taint means a GPU node is tainted; see GPU Nodes.

An engine pod fails with ImagePullBackOff. Read the image it pulls:

kubectl -n kamiwaza get pod <POD_NAME> -o jsonpath='{.spec.containers[*].image}'; echo

An image that names a registry you do not mirror means the relocation values are not set; see Engine images. An image in your registry means that digest was not copied there, or the staging registry Secret does not hold its credentials.

A deployment returns HTTP 503 with engine image map carries no image from this release's catalog. The bundle's map was built for another release. Build the bundle for the core image you install.

Support status​

Support covers the configurations marked Supported in Compatibility scope, qualified on k0s Kubernetes 1.36, within the limits each row states. Anything the table does not mark Supported, such as AMD GPUs or a deployment that spans more than one device, is not supported in this release.