Skip to main content
Version: 1.3.3 (Latest)

Model serving 1.3.0 release notes

Kamiwaza 1.3.0 changes four model-serving behaviors on every cluster, including clusters that do not enable tenant-mode inference. No setting turns them on or off. Each one corrects an estimate, a status, or a node label that earlier releases got wrong. The node label needs one step after the upgrade: re-label existing NVIDIA nodes.

ChangeApplies toVisible in
vLLM estimates without a GPU inventoryvLLM, and estimate requests that name no engineThe VRAM estimate only
KV cache sizing for GGUF modelsGGUF models whose header declares an attention head widthThe VRAM estimate and the GPU memory a deployment reserves
A stop clears the deployment's instancesDeployments stopped after the upgradeGET /api/serving/deployment/{deployment_id}
NVIDIA nodes get the correct CUDA labelNVIDIA nodes with an R570 or R575 driver (570.x to 579.x), and pre-Turing GPUs with R580 or laterThe kamiwaza.ai/gpu-accel-version node label and the llama.cpp image a deployment on the node runs

Behavior changes​

vLLM estimates without a GPU inventory​

Where Kamiwaza finds no GPU in the cluster's node inventory, it now estimates a vLLM deployment's memory as if the deployment will run on an NVIDIA CUDA GPU. Earlier releases estimated it as a CPU deployment, which left out the CUDA graph pool and the GPU runtime's memory. The CPU estimate also assumed flash attention and so counted no attention workspace. Without a GPU inventory, Kamiwaza cannot tell whether the GPU supports flash attention, so the CUDA estimate includes the workspace. vLLM cannot run without a GPU or other accelerator, so the CPU figure was too low. For an 8B model with 16 GiB of weights, the estimate rises from 19.07 GiB to 22.70 GiB at an 8,192-token context, and from 18.48 GiB to 22.12 GiB at a 4,096-token context.

The estimator treats a request that names no engine as vLLM. On such a cluster, POST /api/serving/estimate_model_vram (the SDK's client.serving.estimate_model_vram()) without engine_name therefore also returns the CUDA figure, whichever engine the model would deploy with. The VRAM button in a model's Model Configurations list sends no engine, so it shows the CUDA figure too, and it offers no way to name one.

This applies to CPU-only clusters, to clusters whose node inventory cannot be read, and to Apple silicon hosts, which run Kamiwaza in a managed Lima virtual machine. A cluster whose inventory reports NVIDIA, AMD, Intel, or Gaudi GPUs is estimated as before. Only the estimate changes: on a cluster with no GPU inventory a deployment reserves no GPU memory, so placement is unaffected.

To compare an estimate with the deployment it describes, pass the deployment's engine_name in an API or SDK estimate request.

KV cache sizing for GGUF models​

Kamiwaza now reads the attention head width, <architecture>.attention.key_length, from a GGUF file's header and sizes the KV cache from it. Earlier releases derived the width by dividing the embedding length by the attention head count, and the KV cache grows in proportion to the width.

The estimate changes wherever the declared width differs from the derived one. It rises where the declared width is larger: Qwen3-4B heads are 128 wide against a derived 80, so earlier releases sized its KV cache at 80/128 of what the model uses. It falls where the declared width is smaller: Mistral-Nemo-Instruct-2407 heads are 128 wide against a derived 160. Where the two agree, as for Qwen3-8B, nothing changes.

Kamiwaza reads the header only when it has no other source for the model's architecture. GGUF models whose architecture Kamiwaza reads from a config.json, on disk or from Hugging Face, are sized as before, as are GGUF files that do not declare the width. A head_dim set in the model configuration still takes precedence over the header.

A higher estimate raises the GPU memory the deployment reserves. On a GPU shared by several models, a deployment that fit before can now fail with a no-fit error. Before 1.3.0 the same deployment reserved less memory than its KV cache needs and could run out of memory when the model loaded. If a deployment no longer fits, reduce its context length, place it on a GPU with more free memory, or stop another deployment on that GPU.

A stop clears the deployment's instances​

When a stop succeeds and the deployment reaches STOPPED, Kamiwaza now deletes the deployment's instance records. Earlier releases deleted them only when a forced stop had to clean up after a stop that failed. A stop that succeeded left them in place, so a stopped deployment still listed the instance it had while running.

GET /api/serving/deployment/{deployment_id} (the SDK's client.serving.get_deployment()) now returns an empty instances list for a stopped deployment. The response no longer carries the instance's host_name, listen_port, or node_id, so read those fields while the deployment is running.

Deployments stopped before the upgrade keep their instance records until they are deleted.

NVIDIA nodes get the correct CUDA label​

The kamiwaza.ai/gpu-accel-version label now reads cuda13 only with an R580 or later driver, and cuda12 with an older one. A GPU older than Turing (compute capability below 7.5: Maxwell, Pascal, and Volta) has no kernels in a CUDA 13 image, so labeling also caps it at cuda12 on any driver. The cap gives a Volta GPU (compute capability 7.0) a working llama.cpp image, because the CUDA 12 image carries kernels from compute capability 7.0 up. Maxwell and Pascal GPUs have no llama.cpp CUDA image in 1.3.0, whatever the label says.

Earlier releases labeled an NVIDIA node cuda13 from R570 up, with no GPU cap. A node with an R570 or R575 driver (570.x to 579.x) therefore received a CUDA 13 llama.cpp image that fails with "CUDA driver version is insufficient", and a pre-Turing GPU with R580 received one with no kernels for the GPU. The label changes on those nodes only; a node with an older driver, such as R535 or R550, was already cuda12. A node keeps its old label until it is labeled again, so correct it after the upgrade, as described in Re-label existing NVIDIA nodes.

On a node that the NVIDIA GPU Operator labels and that carries no kamiwaza.ai/gpu-vendor label, Kamiwaza derives the CUDA generation from the operator's labels each time it deploys, with the same compute-capability cap. Such a node needs no re-label.

The label selects the llama.cpp image for a deployment on the standard deployment path, where Kamiwaza can read the cluster's Nodes. An install that does not grant Kamiwaza read access to Nodes cannot see the label, so the label does not select the image there. A tenant-mode deployment does not read node labels: it selects an image from the facts in the owner's signed profile, as described in tenant-mode inference.

A Kamiwaza chart install with default values does not grant read access to Nodes. To check whether your install has it, run the following as a cluster administrator, with <NAMESPACE> set to the namespace Kamiwaza is installed in:

for sa in core-api core-ray core-scheduler; do
kubectl auth can-i list nodes --as="system:serviceaccount:<NAMESPACE>:${sa}"
done

If every line prints yes, Kamiwaza can read Nodes, and the label selects the llama.cpp image on the standard deployment path. If they print no, Kamiwaza cannot read Nodes, and the label does not select the image.

Re-label existing NVIDIA nodes​

Tenant-mode installs. No re-label is needed: tenant-mode inference does not read node labels. The driver's CUDA major that selects the image comes from the owner's signed profile. After a driver change, re-detect the node's facts and publish a new signed profile revision; see Attest the NVIDIA driver's CUDA major.

Installs that grant Kamiwaza read access to Nodes. Set the label on each NVIDIA node that carries kamiwaza.ai/gpu-vendor. On the node, read the driver version and each GPU's compute capability:

nvidia-smi --query-gpu=driver_version,compute_cap --format=csv,noheader
570.172.08, 8.0

The label is cuda13 when the driver is 580 or later and every GPU's compute capability is 7.5 or later, and cuda12 otherwise. The node above, with an R570 driver, is cuda12. Set it, with <CUDA_LABEL> as cuda12 or cuda13:

kubectl label node <NODE_NAME> kamiwaza.ai/gpu-accel-version=<CUDA_LABEL> --overwrite

Confirm the new value on every NVIDIA node:

kubectl get nodes -L kamiwaza.ai/gpu-accel-version

A manual label does not help a Maxwell or Pascal GPU (compute capability below 7.0): no llama.cpp CUDA image in this release has kernels for it.

A running deployment keeps the image it started with. Redeploy a llama.cpp model on a re-labeled node to move it to the image the new label selects.

Keep core.gpu.nodeFeatureRule.enabled at its default, false: when it is true on a cluster with Node Feature Discovery installed, the chart's NodeFeatureRule writes sm_XX values such as sm_80 to kamiwaza.ai/gpu-accel-version on the GPU models it lists, replacing cuda12 and cuda13.