Tenant-mode inference
This page documents tenant-mode inference for Kamiwaza 1.3.3. Support covers the configurations marked Supported in Compatibility scope, qualified on k0s Kubernetes 1.36.
Tenant-mode inference is how a namespaced Kamiwaza install serves local models on GPUs and CPUs. Kamiwaza installs into one namespace and cannot read the cluster's Nodes, so it cannot discover GPUs itself. The cluster owner describes the hardware instead, in a profile that the owner signs, and Kamiwaza places each deployment on hardware the signed profile vouches for.
A default install has tenant-mode inference off and cannot use GPUs: a GPU deployment is refused with HTTP 422. To serve models on GPUs:
- Prepare the GPU nodes; see GPU Nodes.
- Create the owner key, detect each GPU host's facts, and sign a profile, with the owner tool in the Kamiwaza core image. See Turn on tenant-mode inference.
- Point the install at the signed profile, in four Helm values.
Compatibility scope
Live qualification for this release targets k0s Kubernetes 1.36 only. Chart, API, and client contract checks cover Kubernetes 1.28 through 1.36, but passing a contract check is not a live support claim for that minor version: only 1.36 has been exercised end to end. The allocation classes are:
| Class | Engines | Kubernetes resource | Isolation statement | Release status |
|---|---|---|---|---|
| CPU | llama.cpp, whisper.cpp | cpu and memory requests | Kubernetes scheduler allocation | Supported on k0s 1.36 |
| Whole NVIDIA GPU | llama.cpp, vLLM, diffusion (CUDA) | nvidia.com/gpu | Whole-device allocation | Supported on k0s 1.36, one device per replica only. A deployment that spans more than one device is implemented but not qualified live |
| Whole AMD GPU | None in this release; no ROCm engine image ships | amd.com/gpu | Whole-device allocation | Not qualified in this release. Do not deploy against a support claim |
| NVIDIA MIG | llama.cpp, vLLM, diffusion (CUDA) | Qualified nvidia.com/mig-* | Hardware partition, exact shape only | Supported on k0s 1.36, recorded slice shape only |
| VRAM-plugin sharing | llama.cpp, vLLM, diffusion (CUDA) | Owner-advertised kamiwaza.ai/vram-gb-* | Accounted sharing, not hard isolation | Supported on k0s 1.36, recorded plugin and profile only. Optional: the cluster owner installs the VRAM plugin |
Engines
Tenant mode serves only the engines in the platform engine catalog that ships inside the release's core image:
| Engine | Serves | Placement | Images |
|---|---|---|---|
| llama.cpp | Text generation; embeddings when the model configuration sets embedding and states its pooling | GPU or CPU | NVIDIA CUDA 13 and CUDA 12, amd64 and arm64; CPU amd64 and arm64 |
| vLLM | Text generation | GPU only | NVIDIA CUDA 13, amd64 and arm64 (GB10 only) |
| Diffusion | Image generation | GPU only, one device per replica | NVIDIA CUDA 12 (built with 12.8), amd64 only |
| whisper.cpp | Transcription | CPU only | CPU amd64 and arm64 |
The engine catalog has no arm64 image for diffusion in this release. A binding
on an arm64 node therefore cannot serve diffusion, and the platform refuses the
deployment with no diffusion image is published for nvidia/arm64.
vLLM's arm64 image serves one card, the GB10 (DGX Spark, compute capability
12.1). The platform refuses vLLM on any other arm64 NVIDIA card, such as GH200,
GB200, Jetson Thor, or a discrete Blackwell card in an arm64 server, with a
reason such as vllm requires compute capability 12.1; this binding declares 10.0.
A GPU binding serves an engine only when the engine publishes an image for the
binding's vendor and architecture, and the hardware its accelerator block
attests meets the engine's requirements. Each NVIDIA image is built against one
CUDA major and carries kernels down to a lowest compute capability, so the card
and the host driver both limit which image can run:
| Engine | NVIDIA image | Architectures | Lowest compute capability | Host driver |
|---|---|---|---|---|
| llama.cpp | CUDA 13 | amd64, arm64 | 7.5 (Turing) | 580.65.06 (R580) or later |
| llama.cpp | CUDA 12 | amd64, arm64 | 7.0 (Volta) | 525.60.13 or later; R570 or later for Blackwell GPUs |
| vLLM | CUDA 13 | amd64 | 7.5, and the engine requires bf16 (8.0 or later) | 580.65.06 (R580) or later |
| vLLM | CUDA 13 | arm64 | 12.1 (GB10 only) | 580.65.06 (R580) or later |
| Diffusion | CUDA 12 (built with 12.8) | amd64 | 8.0, and the engine requires bf16 | R570 or later recommended, not enforced |
llama.cpp therefore requires compute capability 7.0 or later: a Volta card runs the CUDA 12 image, and no image serves Pascal or older. vLLM requires a CUDA 13 driver. For how the platform picks between the llama.cpp images, see Attest the NVIDIA driver's CUDA major.
The platform compares only the driver's CUDA major, so it does not refuse the diffusion image on an older CUDA 12 driver. The image is built with PyTorch for CUDA 12.8. Unlike the llama.cpp CUDA 12 image, which runs on older CUDA 12 drivers under NVIDIA's minor-version compatibility, the diffusion image is not qualified below R570, so run diffusion on R570 or later.
A diffusion deployment serves one diffusers pipeline on one device through
/v1/images/generations. The model must be a diffusers repository with
model_index.json at its top; a single-file checkpoint is refused. A pipeline
has no context length, so the platform sizes the device from the pipeline's
weights. Tenant mode refuses a diffusion deployment that configures LoRA or
Lightning adapters, a stream profile, trust_remote_code, DFloat11 settings,
or more than one device (tensor_parallel_size or gpu_count above 1). When
a configuration key causes the refusal, the reason names that key. A DFloat11
model that the platform recognizes by its name is also refused; that reason
says DFloat11 models need more than one device and names no key. A stream
profile that the platform detects from the model's name is kept, not refused.
When the release sets core.inferenceResources.enabled: true, every local
deployment takes the tenant path, including a request without an
inferenceResources block, which the platform derives. There is no fallback to
the standard deployment path. An engine that the engine catalog does not list is
refused. MLX is not in the catalog, so MLX is not supported in tenant mode in
this release. A deployment without a block that selects such an engine receives
HTTP 422, and the detail ends with the engine and the reason:
cannot derive inference resources for this deployment: engine 'mlx' is not qualified in this release
A request that carries its own block is accepted and then fails at dispatch with the same reason.
External endpoints are not refused. An external endpoint proxies a remote service and allocates no local resources, so on a tenant cluster the platform does not size it and creates no workload in the tenant namespace. On a tenant cluster, two external-endpoint requests receive HTTP 422:
-
A request that supplies an
inferenceResourcesblock:external endpoints allocate no local resources; omit inferenceResources -
A request that sets a nonzero
gpu_allocationorvram_allocation. A cluster without tenant resources still honors these overrides:external endpoints allocate no local resources; omit gpu_allocation and vram_allocation
One replica across several devices
A deployment can span up to eight whole devices in a single Pod through tensor
parallelism. This is implemented but not qualified live in this release; see
Compatibility scope. The device count comes from
tensor_parallel_size in the model configuration, or from a signed recipe when
one pins it. An absent value is one device — never the node's device count, so
a first deployment does not take a whole multi-GPU node.
Tensor parallelism needs a peer-to-peer path between the devices, so it is
admitted only on a whole-device binding whose accelerator block attests the
p2p feature. The owner tool's detector claims p2p only when every pair of
whole-mode GPUs on the host reports a peer path, because the device plugin can
give a multi-GPU request any whole GPUs. MIG-mode GPUs are left out: a
whole-device binding is never given one. On a host whose GPUs are bridged in
pairs, partition the GPUs outside one bridged pair with MIG, so its whole GPUs
are exactly that pair; see
MIG and tensor parallelism.
A MIG partition does not have a peer path, and an accounted VRAM share is a slice of a device the workload does not exclusively hold, so both refuse a count above one rather than quietly serving a single device to an engine told to shard across several. A count that is not a whole number, is below one, or exceeds eight is refused for the same reason: the request and the Pod must agree.
Pipeline and data parallelism are not part of this release.
Ownership boundary
The cluster owner supplies drivers, device plugins, optional VRAM sharing,
RuntimeClasses, quotas and policy, and the signed profile. Every GPU binding in
a profile attests its hardware in an accelerator block, which the owner tool
detects on the host; a CPU binding needs only its architecture. Recipe catalogs
are optional, GPU and CPU alike: without one, the platform derives each
deployment, and a published recipe pins a specific model and configuration. The
tenant supplies only namespaced Helm and serving inputs.
Kubernetes namespace-admin permissions do not grant the Kamiwaza admin role.
Model deployment through the API, SDK, or UI remains Kamiwaza-admin-only.
A deployment request needs no inferenceResources block: the platform derives
one. For a GPU deployment it sizes memory from the model's own metadata and the
context the model configuration sets (max_model_len), or else the context the
model declares, and places the request on the smallest binding that holds it.
A deployment goes to CPU instead when it sets force_cpu, when its engine
ships only a CPU build (whisper.cpp), or when the owner declared no GPU
hardware; it is sized from the platform's CPU floor and the model's weights,
and runs the platform's CPU build for the binding's architecture. A request
that cannot be served as asked -- no context declared anywhere, or an
embedding configuration that does not state its pooling -- is refused with
HTTP 422 and the reason; one the platform cannot evaluate, such as an
unreadable profile bundle, with HTTP 503.
The tenant chart creates no cluster-scoped GPU resource and does not require
nodes/get, nodes/list, or nodes/watch. A profile or allocator failure is
reported as a failure; the platform does not silently fall back to CPU or a
weaker isolation class.
Turn on tenant-mode inference
Before you start
| You need | Notes |
|---|---|
| Prepared GPU nodes | Every requirement on GPU Nodes passes. Skip for CPU-only serving. |
| The release's core image, by digest | The owner tool ships in it. Sign with the core image of the release you install, so the release that verifies the profile is the one that signed it. |
| Docker or Podman | On the machine where you sign, not on the GPU nodes: step 3 runs the detector as a pod. The owner tool runs in the core image; it contacts no cluster and no network. |
kubectl and jq | On the machine where you publish the profile. |
Docker sets the host's iptables FORWARD policy to DROP, which breaks pod
networking on the node. Run the owner tool on a separate machine, and the
detector as a pod.
Read the core image from a running install:
kubectl -n kamiwaza get deploy core-scheduler \
-o jsonpath='{.spec.template.spec.containers[?(@.name=="core")].image}'; echo
<YOUR_REGISTRY>/kamiwaza-ai/releases/kamiwaza/images/core@sha256:<DIGEST>
Before the install, take the same reference from the release's image list; see Install Without Internet Access.
If your registry requires credentials, log the container engine in to it first:
docker login <REGISTRY_HOST>, or podman login <REGISTRY_HOST>. Podman also
reads an existing credentials file named with --authfile <FILE> or the
REGISTRY_AUTH_FILE environment variable, in the same format as the install's
pull Secret.
Set the image, and define a shell function that runs the owner tool with the current directory mounted. With Docker:
export CORE_IMAGE="<YOUR_REGISTRY>/kamiwaza-ai/releases/kamiwaza/images/core@sha256:<DIGEST>"
owner_tool() {
docker run --rm -i --user "$(id -u):$(id -g)" -v "$PWD:/work" -w /work \
"$CORE_IMAGE" python -m kamiwaza.serving.inference_resources.owner_tool "$@"
}
With rootless Podman, which needs --userns=keep-id to write files you own, and
:Z on the mount to use the directory on a host with SELinux, such as RHEL or
Fedora:
export CORE_IMAGE="<YOUR_REGISTRY>/kamiwaza-ai/releases/kamiwaza/images/core@sha256:<DIGEST>"
owner_tool() {
podman run --rm -i --userns=keep-id --user "$(id -u):$(id -g)" -v "$PWD:/work:Z" -w /work \
"$CORE_IMAGE" python -m kamiwaza.serving.inference_resources.owner_tool "$@"
}
:Z relabels the current directory for the container, and Podman ignores it
on a host without SELinux. Podman refuses to relabel your home directory, so
run the function from a directory of its own.
The steps below use the function; the GPU detector in step 3 runs on the GPU node instead, because it must see the GPUs.
Step 1: Create the owner signing key
owner_tool keygen --private-key-out owner-ed25519-private.pem
{
"keyId": "sha256:6e1f...",
"publicKey": "8Ljv..."
}
The tool writes an unencrypted Ed25519 private key readable only by you, and
refuses to replace an existing file. It prints the public key, which goes into
the Helm values as ownerPublicKey, and the keyId your profile names. Keep the
private key out of source control, out of the tenant namespace, and out of any
evidence you record. To print the public half again:
owner_tool pubkey --private-key owner-ed25519-private.pem
Step 2: Choose a cluster binding ID
Pick a clusterBindingId for this cluster: 1 to 128 characters, starting with a
letter or digit, then letters, digits, ., _, :, or -. The profile and
the Helm values must carry the same value; Kamiwaza refuses a profile signed for
another binding.
Step 3: Detect each GPU host's facts
Run the detector on every GPU node, with the node's GPUs visible to it. Run it as
a one-shot pod on the nvidia RuntimeClass: the node then needs no container
engine of its own. The pod runs in the Kamiwaza namespace, so it pulls the core
image with the install's pull Secret; if your install has no pull Secret,
delete the two imagePullSecrets lines. Set NODE_NAME to the GPU node:
NODE_NAME="<NODE_NAME>"
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: owner-tool-detect
namespace: kamiwaza
annotations:
sidecar.istio.io/inject: "false"
spec:
runtimeClassName: nvidia
restartPolicy: Never
nodeSelector:
kubernetes.io/hostname: $NODE_NAME
imagePullSecrets:
- name: kamiwaza-pull
containers:
- name: detect
image: $CORE_IMAGE
command: ["python", "-m", "kamiwaza.serving.inference_resources.owner_tool", "detect"]
env:
- {name: NVIDIA_VISIBLE_DEVICES, value: all}
- {name: NVIDIA_DRIVER_CAPABILITIES, value: "compute,utility"}
EOF
kubectl -n kamiwaza wait --for=jsonpath='{.status.phase}'=Succeeded pod/owner-tool-detect --timeout=5m
kubectl -n kamiwaza logs owner-tool-detect
kubectl -n kamiwaza delete pod owner-tool-detect
The detector prints the accelerator block for the node's GPUs:
{
"accelerator": {
"computeCapability": "8.0",
"cudaDriverMajor": 13,
"features": [
"bf16",
"mig",
"nvlink",
"p2p"
],
"vendor": "nvidia"
}
}
A GPU node that already has a container engine can run the detector directly
instead. With Podman, as root, through the NVIDIA runtime; add
--authfile <FILE> if your registry requires credentials:
sudo podman run --rm --runtime /usr/bin/nvidia-container-runtime \
-e NVIDIA_VISIBLE_DEVICES=all "$CORE_IMAGE" \
python -m kamiwaza.serving.inference_resources.owner_tool detect
podman run --device nvidia.com/gpu=all does not work with Podman 4.9, the
version Ubuntu 24.04 ships: it cannot read the device description that current
NVIDIA container toolkits write, and fails with
unresolvable CDI devices nvidia.com/gpu=all. On a node that already runs
Docker with the NVIDIA runtime, docker run --rm --gpus all "$CORE_IMAGE" python -m kamiwaza.serving.inference_resources.owner_tool detect also works. Do not
install Docker on a Kubernetes node to run it.
The detector reads the facts from nvidia-smi instead of expecting you to know
them:
computeCapabilityis the GPUs' compute capability. The detector refuses a node with two kinds of GPU: a binding cannot single out one kind on such a node, so keep it out of tenant-mode GPU serving; see GPU Nodes.cudaDriverMajoris the major of theCUDA Versioninnvidia-smi's header. The detector exits with an error rather than guess when it cannot read it.bf16andfp8follow from the compute capability, andnvlinkfrom the links the GPUs report.p2pappears only when every pair of whole-mode GPUs has a peer path; see One replica across several devices.migappears when the GPUs can be partitioned, whether or not a GPU is in MIG mode now. It is informational: no engine requires it, and a MIG binding does not need it.
Copy the block into the binding for those nodes. A binding covers every node of
its architecture that advertises its resource (see nodeSelector in step 4),
so when those nodes differ, attest only what holds on all of them: the lowest
compute capability and cudaDriverMajor, and only the features every node has.
Drop p2p and nvlink from a MIG or VRAM-share binding: a slice has neither.
Step 4: Write the profile envelope
The envelope lists the profiles the cluster offers. Each profile has one
guarantee and one or more bindings; each binding names the Kubernetes resource
a deployment requests and the nodes it runs on. This envelope offers CPU serving
on amd64 nodes and whole NVIDIA GPUs. Save it as owner-profile-envelope.json:
{
"schemaVersion": 1,
"clusterBindingId": "<CLUSTER_BINDING_ID>",
"namespace": "kamiwaza",
"version": 1,
"keyId": "<KEY_ID>",
"issuedAt": "<ISSUED_AT>",
"expiresAt": "<EXPIRES_AT>",
"evidenceId": "gpu-nodes-checked-2026-10-01",
"profiles": [
{
"name": "cpu-standard",
"capability": "cpu",
"guarantees": [],
"compatibleVariantIds": ["cpu-default"],
"bindings": [{
"id": "cpu-amd64",
"allocator": "cpu",
"architecture": "amd64",
"runtime": {"kind": "cpu", "variantId": "cpu-default", "image": "<LLAMACPP_CPU_IMAGE>"},
"capacity": {"kind": "scheduler"},
"nodeSelector": {"kubernetes.io/arch": "amd64"}
}]
},
{
"name": "nvidia-whole",
"capability": "gpu",
"guarantees": ["whole-device"],
"compatibleVariantIds": ["llamacpp-nvidia-amd64", "vllm-nvidia-amd64", "diffusion-nvidia-amd64"],
"bindings": [{
"id": "nvidia-whole-amd64",
"allocator": "extended-resource",
"architecture": "amd64",
"accelerator": {
"vendor": "nvidia",
"computeCapability": "8.0",
"cudaDriverMajor": 13,
"features": ["bf16", "mig", "nvlink", "p2p"]
},
"capacity": {"kind": "scheduler", "usableMemory": "79Gi"},
"resource": {"name": "nvidia.com/gpu", "quantity": 1},
"runtime": {"kind": "cuda", "variantId": "llamacpp-nvidia-amd64", "image": "<LLAMACPP_CUDA_IMAGE>"},
"runtimeClassName": "nvidia",
"nodeSelector": {"kubernetes.io/arch": "amd64"}
}]
}
]
}
Fill in the placeholders:
| Field | Value |
|---|---|
clusterBindingId | The ID from step 2. |
namespace | The namespace Kamiwaza is installed in. |
version | An integer you raise with every new revision. Kamiwaza refuses a revision below core.inferenceResources.minimumBundleVersion (default 1). |
keyId | The keyId from step 1. |
issuedAt, expiresAt | UTC timestamps with a Z suffix, such as 2026-10-01T00:00:00Z. Kamiwaza refuses the profile outside that window, so publish a new revision before expiresAt. |
evidenceId | A label for the checks behind this revision, for your records. |
accelerator | The block from step 3. |
nodeSelector | kubernetes.io/arch only; this release accepts no other key. A binding therefore covers every node of that architecture that advertises its resource, so its accelerator facts and usableMemory must hold on all of them. |
usableMemory | The GPU memory a deployment may plan on, less what the driver and CUDA context hold back. Write a whole number of Gi, Mi or Ki; the profile schema refuses a fraction such as 121.7Gi. The right value depends on the binding; see the list below this table. |
runtime.image | A digest-pinned engine image. Derived deployments take their image from the engine catalog; this value is used only for a deployment that a recipe pins, which runs it as written. In a mirrored or air-gapped install, name your mirror's copy. Print the catalog's images with the command below. |
Set usableMemory for each kind of binding as follows:
- Whole device or MIG partition on a discrete card: the device's or the
slice's memory less the driver's and CUDA context's share,
79Gion an 80 GB card. Understating it is the safe direction. - VRAM plugin: state it accurately. vLLM reserves a fraction of the whole
card: 90%, the engine catalog default, times the request divided
by
usableMemory. UnderstatingusableMemorytherefore raises what vLLM reserves, and the reservation passes the grant onceusableMemoryis about a tenth below the card's memory. - Whole-device GB10: its GPU memory is the host's, so set
usableMemoryto the device total rounded down,121Gion a DGX Spark, whose total is about 121.69 GiB. vLLM's budget is the request divided byusableMemory, rounded up to a hundredth, applied to that total. At121GivLLM already reserves slightly more than the request: 121.69/121 of it, plus up to 0.01 of the total from the rounding. A smaller value reserves further above the request. The host's own processes use the same memory, so a request must also fit in what the host leaves free. A VRAM-plugin binding on a GB10 follows the VRAM-plugin rule instead.
A GPU profile's compatibleVariantIds must list <engine>-nvidia-<architecture>
for each engine it may serve, and each binding's runtime.variantId must be in
its profile's list. To print the engine images the core image carries:
docker run --rm "$CORE_IMAGE" python -c \
"from importlib.resources import files; print(files('kamiwaza.serving.inference_resources').joinpath('engine-catalog.json').read_text())" \
| jq -r '.engines[] | .engine as $e | .images[] | [$e, .vendor, .architecture, (.cuda // "-" | tostring), .image] | @tsv'
llamacpp nvidia amd64 13 ghcr.io/kamiwaza-internal/containers/images/llamacpp-cuda13@sha256:a482...
llamacpp cpu amd64 - ghcr.io/kamiwaza-internal/containers/images/llamacpp-cpu@sha256:ccbd...
For MIG, add a profile whose guarantee is mig and whose binding requests the
partition the node advertises. This profile offers the 3g.40gb slices from
GPU Nodes;
add it to the profiles list:
{
"name": "nvidia-mig-3g40gb",
"capability": "gpu",
"guarantees": ["mig"],
"compatibleVariantIds": ["llamacpp-nvidia-amd64", "vllm-nvidia-amd64", "diffusion-nvidia-amd64"],
"bindings": [{
"id": "nvidia-mig-3g40gb-amd64",
"allocator": "extended-resource",
"architecture": "amd64",
"accelerator": {
"vendor": "nvidia",
"computeCapability": "8.0",
"cudaDriverMajor": 13,
"features": ["bf16", "mig"]
},
"capacity": {"kind": "scheduler", "usableMemory": "38Gi"},
"resource": {"name": "nvidia.com/mig-3g.40gb", "quantity": 1},
"runtime": {"kind": "cuda", "variantId": "llamacpp-nvidia-amd64", "image": "<LLAMACPP_CUDA_IMAGE>"},
"runtimeClassName": "nvidia",
"nodeSelector": {"kubernetes.io/arch": "amd64"}
}]
}
The accelerator block is the node's, from step 3, without p2p and nvlink:
a slice has neither. Set usableMemory to the slice's memory less headroom for
the CUDA context. On the verification host that
GPU Nodes describes, a
3g.40gb slice of an A100 80GB reports 40192 MiB in nvidia-smi, so the
example uses 38Gi. Binding id values must be unique
across the whole envelope.
An accounted-share profile, for the
VRAM plugin,
uses the vram-plugin allocator and {"kind": "accounted", ...} capacity,
and reserves a bookkeeping bucket; it does not provide hard VRAM or compute
isolation.
Step 5: Sign and publish the profile
owner_tool sign --envelope owner-profile-envelope.json \
--private-key owner-ed25519-private.pem > owner-profile-configmap.json
The tool runs Kamiwaza's own verifier over what it signs, so it refuses a profile this release would reject, and names the check that failed:
error: signed bundle would be refused: invalid inference profile authorization: bundle is expired or its window is empty
It emits an immutable ConfigMap named kamiwaza-inference-profiles. Give each
revision its own name, then apply it:
export PROFILE_CONFIGMAP=kamiwaza-inference-profiles-r1
jq --arg name "$PROFILE_CONFIGMAP" '.metadata.name = $name' owner-profile-configmap.json \
| kubectl apply --server-side --field-manager=owner-inference-publisher -f -
Check:
kubectl -n kamiwaza get configmap "$PROFILE_CONFIGMAP" -o jsonpath='{.immutable} {.data.bundle\.sha256}'; echo
true 9c1d...
Step 6: Set the Helm values
Add these values to the install's values file and upgrade the release with it:
core:
inferenceResources:
enabled: true
bundleConfigMap: kamiwaza-inference-profiles-r1 # the name from step 5
ownerPublicKey: <OWNER_PUBLIC_KEY> # publicKey from step 1
clusterBindingId: <CLUSTER_BINDING_ID> # from step 2
| Helm value | Notes |
|---|---|
bundleConfigMap | Kamiwaza reads the profile only from the ConfigMap with this name. Left at the default, kamiwaza-inference-profiles, it cannot read a renamed revision, and if an older ConfigMap with the default name still exists, Kamiwaza uses that older profile, with no error, for as long as it verifies. |
ownerPublicKey | The canonical base64 of the raw 32-byte public key, as keygen and pubkey print it. |
clusterBindingId | Must equal the profile's clusterBindingId. |
On an install whose global.imagePullSecrets names the pull Secret, these
four values are enough; an install that pulls every image without credentials
needs no pull Secret. The staging values default as follows, and an explicit
value always wins over its default:
| Value | Default with tenant-mode inference on |
|---|---|
gpuRegistry, cpuRegistry | The in-cluster model registry, which the model downloader already pushes to. It serves plain HTTP, so gpuRegistryInsecure and cpuRegistryInsecure default to true with it. |
gpuStagingImage, cpuStagingImage | The chart's own model-fetch image, relocated like every chart image. |
gpuRegistryCredentialsSecret, cpuRegistryCredentialsSecret | The first entry of global.imagePullSecrets, the install's pull Secret. Not defaulted when global.openshift.serviceAccountImagePullSecretsOnly is true. |
Engine pods carry only this staging credential, as their image pull Secret:
the engine and model-fetch images pull with it. Set
gpuRegistryCredentialsSecret and cpuRegistryCredentialsSecret yourself only
when your mirror needs different credentials from the install's pull Secret.
Deployments created before you turn tenant-mode inference on carry no tenant
resources, so they cannot start again once it is on. After the upgrade, stop
each one and deploy it again (see Deploy through the API);
the new deployment takes its resources from the signed profile and its image
from the engine catalog. While core.context.embedding.autoProvision is on
(the default, except on global.platform: osgm), the platform's embedding
model needs only the stop: the platform deploys it again on the next embedding
request. With it off, deploy the embedding model again yourself.
While core.context.embedding.autoProvision is on, turning tenant-mode
inference on for an install with no earlier embedding deployment also deploys
the platform's bundled embedding model, all-MiniLM-L6-v2, on CPU, so expect one
CPU engine pod shortly after the upgrade. On an install that already served the
embedding model before you turned tenant-mode inference on, the platform finds
that earlier deployment, still DEPLOYED, and deploys nothing until you stop it
and an embedding request arrives. With it off, the platform does not deploy the
embedding model on its own.
Step 7: Verify
Deploy a model that fits a GPU, from the Models page, then check the engine pods. The columns are the pod, its node, its RuntimeClass, the NVIDIA resources it was granted, and whether the grant was verified:
kubectl -n kamiwaza get pods -l kamiwaza.ai/deployment-id -o json | jq -r '.items[]
| [.metadata.name, .spec.nodeName, (.spec.runtimeClassName // "-"),
([.spec.containers[].resources.limits // {} | to_entries[]
| select(.key | startswith("nvidia.com/")) | "\(.key)=\(.value)"]
| join(",") | if . == "" then "-" else . end),
([.status.conditions[]? | select(.type == "kamiwaza.ai/AllocationVerified") | .status]
| first // "-")]
| @tsv'
tmi-3f2a9c... gpu-node-1 nvidia nvidia.com/gpu=1 True
tmi-8b1e04... gpu-node-1 nvidia nvidia.com/mig-3g.40gb=1 True
tmi-51c7aa... gpu-node-1 - - True
Expect the nvidia RuntimeClass and the GPU or MIG resource the deployment
asked for on a GPU pod, and no RuntimeClass or NVIDIA resource on a CPU pod.
Every pod should read True in the last column. Readiness alone is not proof of
the requested grant: the kamiwaza.ai/AllocationVerified condition is.
A CPU pod requests the engine's CPU floor from the engine catalog, 100m for
llama.cpp and 500m for whisper.cpp, and memory of at least 1.25 times the
model's weights, rounded up to a whole GiB and never below the floor of 1Gi
for llama.cpp or 2Gi for whisper.cpp. It sets requests only, no limits.
Recipe catalogs (optional)
A recipe catalog pins a specific model and configuration to exact settings. Without one, every deployment is derived. Render each catalog with the owner tool, which validates it with the parser Kamiwaza runs when it mounts it and emits an immutable ConfigMap whose name carries the catalog's digest:
owner_tool render-catalog --catalog gpu-recipes.json \
--namespace kamiwaza --capability gpu > gpu-catalog.json
kubectl apply --dry-run=server -f gpu-catalog.json
kubectl apply -f gpu-catalog.json
export GPU_CATALOG_CONFIGMAP="$(jq -r '.metadata.name' gpu-catalog.json)"
Use --capability cpu for a CPU catalog. Point the release at the exact name
with core.inferenceResources.gpuRecipeCatalogConfigMap or
cpuRecipeCatalogConfigMap. Publish a new revision for every change and never
mutate a published ConfigMap. Keep the previous revision until every consumer
has rolled out; roll back by selecting the previous name.
A pinned deployment runs the image its recipe variant names, and the profile
binding's runtime.image must be the same reference. Kamiwaza does not relocate
it: the bundle's image map, global.imageRegistryPrefix, and
global.airgap.registry.url apply only to the engine catalog's images. In a
mirrored or air-gapped install, name your mirror's copy, by digest, in both the
recipe variant and the binding.
Engine images
The engine catalog in the core image pins every engine image by digest. Kamiwaza relocates those references once, when it loads the catalog, using the same values the chart relocates its own images with:
core.inferenceResources.engineImageMap, when set. The tenant OCI bundle writes it for you; see Bundle installs.- Otherwise
global.imageRegistryPrefix:<prefix>/images/<name>@<digest>, where<name>is the last segment of the catalog's repository. - Otherwise
global.airgap.registry.url:<registry>/<catalog repository>@<digest>. - Otherwise the reference as the catalog pins it.
An engine image must therefore be in your registry by digest, at the location the rule produces. Most catalog pins are untagged manifests, so a copy by tag does not bring them. Print the pins with the catalog command in step 4, and see Install Without Internet Access for the copy.
Engine pods pull with the staging registry Secret,
core.inferenceResources.gpuRegistryCredentialsSecret and
cpuRegistryCredentialsSecret, which default to the install's pull Secret; see
step 6. There is no separate engine pull-secret
setting. Set those two values only when your mirror needs different
credentials.
A deployment that a recipe pins is the exception to the relocation above: it runs the image the owner signed, exactly as written. See Recipe catalogs.
After you change any of these values, redeploy the affected inference
deployments: stop each one and deploy it again (see
Deploy through the API). A running deployment keeps
the engine reference it was planned with, and so does the platform's embedding
model. Stop the embedding model's deployment too. While
core.context.embedding.autoProvision is on, leave the redeploy to the
platform: it deploys the model again on the next embedding request, such as
POST /api/embedding/generate, and not before. That request returns
HTTP 503 with the code embedding_deploying and a retry_after_seconds value,
and embedding requests return it until the new deployment is ready, about 75
seconds on an air-gapped A100 install; retry them. With it off, the platform
does not redeploy the model, and embedding requests return HTTP 503 with the
code embedding_unavailable until you deploy it again yourself. On an upgrade,
keep the previous release's engine images in your registry until every
deployment has been redeployed; removing them breaks rescheduling.
Bundle installs
A tenant OCI bundle carries the engine images for one architecture, and writes
values/99-engine-images.yaml with the map that tells Kamiwaza where it
imported each one. Keep that file last among the install's values files. A
cluster with amd64 and arm64 nodes needs a bundle for each architecture.
A bundle can leave engines out. An engine the bundle does not carry is refused at admission, before the deployment record exists:
cannot derive inference resources for this deployment: engine 'vllm' is not qualified in this release
Attest the NVIDIA driver's CUDA major
For each NVIDIA binding, the platform selects the newest image of the engine
that both the card and the host driver can run. The card's limit is the
computeCapability the binding attests. The driver's limit is
cudaDriverMajor, an optional integer from 11 to 99 in the same accelerator
block: the major of the CUDA Version that nvidia-smi prints in its header on
the host (e.g., 12 on an R570 driver, which reports CUDA 12.8, and 13 on an
R580 driver, which reports CUDA 13.0). An image is selected only when its CUDA
major is at or below the attested value.
The detector in step 3 attests
cudaDriverMajor by default. Without it, the platform does not judge the
driver: a Turing or newer card on an R570 host is given the CUDA 13 llama.cpp
image, which exits at startup with
CUDA driver version is insufficient for CUDA runtime version. Re-detect and
publish a new signed profile revision after any driver change, including a
driver rollback.
A binding without cudaDriverMajor keeps verifying and selects the newest
image its card can run:
- llama.cpp resolves its CUDA 13 image on compute capability 7.5 or later, and its CUDA 12 image on Volta (7.0).
- vLLM resolves its CUDA 13 image on amd64 on compute capability 8.0 or later,
because the engine also requires
bf16, and on arm64 only on a GB10 (12.1). - Diffusion resolves its one image, CUDA 12, on compute capability 8.0 or later, on amd64 only.
- A Pascal (6.x) or older card is refused.
vLLM ships only a CUDA 13 image. On a binding that attests cudaDriverMajor: 12
the platform refuses it:
host driver supports CUDA 12 (attested cudaDriverMajor); engine vllm ships nvidia/amd64 images for CUDA 13+
On a binding that does not attest the fact, the platform cannot tell, and the vLLM image fails at startup on a CUDA 12 driver. When the card is older than every image, the reason names the compute capability instead:
llamacpp requires compute capability 7.0; this binding declares 6.1
A deployment without an inferenceResources block is refused with HTTP 422
before the deployment record exists, with the reason after
cannot derive inference resources for this deployment:. This happens only
when the driver or the card rules out every image on every GPU binding in the
profile. A request that carries its own block is refused at binding selection:
the deployment fails with NoFit and the reason in its error message. To
resolve a driver refusal, upgrade the host driver, re-detect, and publish a new
signed profile revision. No owner action resolves a card refusal, because the
release ships no image for that card.
Before you roll Kamiwaza back to a release earlier than 1.3.0, publish a new
signed profile revision without cudaDriverMajor and point the release at it.
Earlier releases refuse the unknown key and reject the whole profile, so every
tenant deployment, CPU deployments included, fails authorization.
Requests that name their own resources
A deployment from the Models page needs no inferenceResources block: the
platform derives one. To choose the profile, device count or memory yourself,
deploy through the API, POST /api/serving/deploy_model, with the block in the
request body. The block names the logical profile and capability rather than a
Node or a provider-specific inventory:
{
"inferenceResources": {
"schemaVersion": 1,
"accelerator": {
"capability": "gpu",
"count": 1,
"memory": {"minimum": "1Gi"},
"isolation": "any-qualified",
"profile": "nvidia-whole"
},
"runtime": {"selection": "automatic"}
}
}
For cpu-standard, use the CPU request shape instead:
{
"inferenceResources": {
"schemaVersion": 1,
"cpu": {
"architecture": "amd64",
"requests": {"cpu": "2", "memory": "4Gi"}
},
"runtime": {"selection": "automatic"}
}
}
Do not add Node names, Node selectors owned by the tenant, or cluster-scoped resources to the tenant request.
Deploy through the API
Get a token as an administrator. Set KAMIWAZA_URL to your Kamiwaza domain; if
the edge certificate is issued by a private authority, add --cacert <CA_FILE>
to each curl command. The admin password is the one from your install page's
verify step:
KAMIWAZA_URL="https://<KAMIWAZA_DOMAIN>"
printf 'admin password: '; read -rs ADMIN_PASSWORD; echo # typed, not echoed
# curl reads the password from stdin, so it never appears on a command line.
TOKEN=$(printf '%s' "$ADMIN_PASSWORD" | curl -sS -X POST "$KAMIWAZA_URL/api/auth/token" \
-H "Content-Type: application/x-www-form-urlencoded" \
--data-urlencode "username=admin" --data-urlencode "password@-" \
| jq -r '.access_token')
unset ADMIN_PASSWORD
Find the model, its configuration and its weights file. A model is in the list once it has been downloaded; see Downloading Models:
curl -sS -H "Authorization: Bearer $TOKEN" "$KAMIWAZA_URL/api/models/" \
| jq -r '.[] | [.id, .name] | @tsv'
MODEL_ID="<MODEL_ID>"
curl -sS -H "Authorization: Bearer $TOKEN" "$KAMIWAZA_URL/api/model_configs/?model_id=$MODEL_ID" \
| jq -r '.[] | [.id, .name, .default] | @tsv'
curl -sS -H "Authorization: Bearer $TOKEN" "$KAMIWAZA_URL/api/models/$MODEL_ID" \
| jq -r '.m_files[] | [.id, .name] | @tsv'
Write the request, with the configuration marked true in the second list and
the weights file to serve, then deploy:
cat > deploy.json <<'JSON'
{
"m_id": "<MODEL_ID>",
"m_config_id": "<CONFIG_ID>",
"m_file_id": "<FILE_ID>",
"inferenceResources": {
"schemaVersion": 1,
"accelerator": {
"capability": "gpu",
"count": 1,
"memory": {"minimum": "8Gi"},
"isolation": "any-qualified",
"profile": "nvidia-whole"
},
"runtime": {"selection": "automatic"}
}
}
JSON
curl -sS -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
--data @deploy.json "$KAMIWAZA_URL/api/serving/deploy_model"
"631e7142-3417-487c-be65-d7715ba6cbf2"
The response is the deployment's ID, returned before the launch finishes.
Follow its status until DEPLOYED; see
Model Deployment for the
statuses:
curl -sS -H "Authorization: Bearer $TOKEN" \
"$KAMIWAZA_URL/api/serving/deployment/<DEPLOYMENT_ID>/status"; echo
To stop a deployment, e.g. before deploying it again after a registry change:
curl -sS -X DELETE -H "Authorization: Bearer $TOKEN" \
"$KAMIWAZA_URL/api/serving/deployment/<DEPLOYMENT_ID>"; echo
The response is true, and the deployment's status becomes STOPPED.
Upgrade and recover
To change the profile, publish a new revision under a new name with a higher
version, point bundleConfigMap at it, upgrade the release, and repeat the
checks in step 7. To roll back, point bundleConfigMap at the
previous revision. If a create, readiness, authorization, or deadline step
fails, the platform marks the allocation failed and removes only resources
owned by that attempt; keep the prior known-good workload until the replacement
passes its checks.
Upgrading Kamiwaza, including its core image, does not require a new signed profile. Sign a new revision only when your profiles change, or when the release notes say the profile schema changed.
To rotate the owner key, create a new key, sign a new revision with it, and set
ownerPublicKey and bundleConfigMap together in one upgrade.
Troubleshooting
A GPU deployment returns HTTP 422 with Core cannot see a GPU on this install. Tenant-mode inference is off. Turn it on as described in
Turn on tenant-mode inference.
A deployment returns HTTP 503 with owner profile bundle: invalid inference profile authorization: and a reason. Kamiwaza found the profile but refused
it. The reason names the check:
| Reason | Fix |
|---|---|
bundle did not arrive from the configured ConfigMap | bundleConfigMap does not name the published revision, or it is in another namespace. |
bundle targets another cluster binding, namespace, or a version below the configured floor | Compare the profile's clusterBindingId, namespace, and version with the Helm values. |
bundle is signed by a key this cluster does not trust | ownerPublicKey is not the key that signed the profile. Print it with owner_tool pubkey. |
bundle is expired or its window is empty | Publish a new revision with a later expiresAt. |
A deployment returns HTTP 422 with engine '<name>' is not qualified in this release. The engine is not in the engine catalog on this install: MLX is
never in it, and a bundle install may have left the engine out.
A deployment returns HTTP 422 naming the driver's CUDA major or the card's compute capability. See Attest the NVIDIA driver's CUDA major.
A request with its own block fails with NoFit. No binding in the profile
holds the request: compare its memory, count, and isolation with the profile's
bindings.
A deployment stays REQUESTED with no error, and its engine pod is
Pending. The pod cannot be scheduled, and the deployment reports no error
until the per-attempt scheduling deadline, schedulingTimeoutSeconds, passes.
A tensor-parallel deployment that needs two whole GPUs while another deployment
holds one of them is a common cause. Find the pod and read its events:
kubectl -n kamiwaza get pods -l kamiwaza.ai/deployment-id --field-selector=status.phase=Pending
kubectl -n kamiwaza describe pod <POD_NAME> | sed -n '/Events:/,$p'
Insufficient nvidia.com/gpu means no node has enough free GPUs of the
requested kind; stop a deployment that holds one, or add GPUs. had untolerated taint means a GPU node is tainted; see
GPU Nodes.
An engine pod fails with ImagePullBackOff. Read the image it pulls:
kubectl -n kamiwaza get pod <POD_NAME> -o jsonpath='{.spec.containers[*].image}'; echo
An image that names a registry you do not mirror means the relocation values are not set; see Engine images. An image in your registry means that digest was not copied there, or the staging registry Secret does not hold its credentials.
A deployment returns HTTP 503 with engine image map carries no image from this release's catalog. The bundle's map was built for another release. Build
the bundle for the core image you install.
Support status
Support covers the configurations marked Supported in Compatibility scope, qualified on k0s Kubernetes 1.36, within the limits each row states. Anything the table does not mark Supported, such as AMD GPUs or a deployment that spans more than one device, is not supported in this release.