Skip to main content
Version: 1.3.3 (Latest)

Platform Prerequisites for Kubernetes with Istio

Kamiwaza installs into a single namespace on a Kubernetes cluster that you already run. Before the install, a platform administrator (someone with cluster-admin rights) prepares the cluster so it provides everything on this page. The Kamiwaza chart does not create or repair any of it: it creates no cluster-scoped objects, installs no operators, and does not configure your service mesh. If a prerequisite is missing, the install fails or the platform comes up unable to serve traffic.

Work through each section and run its check. Each section also says what to record; the install needs those facts.

Hardware sizing and supported operating systems are covered separately, in System Requirements.

Summary​

PrerequisiteWhat the cluster must provide
NamespaceA pre-created namespace, labeled for Istio sidecar injection
Installer identityA namespace-scoped identity to run Helm as. Never cluster-admin
Istio and the ingress gatewayAn Istio mesh in sidecar mode, with an ingress gateway serving the Kamiwaza domain on HTTPS
Edge certificateA TLS certificate for the domain on that gateway
DNSThe domain resolves to the gateway, from clients and from inside the cluster
Block storageAn RWO StorageClass whose volumes keep their data when pods are replaced
Object storageA decision: the built-in object store, or your own S3-compatible service
Kubernetes API addressesThe API server's Service and endpoint addresses, recorded for network policy
Registry accessNodes can pull Kamiwaza images
Node kernel limitsfs.inotify limits at or above the floor on every node
Outbound network accessHTTPS to the image registry and to any model sources you use

GPU inference and multi-node clusters add requirements; see GPU nodes and Multi-node clusters.

Kamiwaza installs into a namespace named kamiwaza; use that name. The examples use kamiwaza.example.com as the domain; substitute your own. Run the checks with the administrator's kubeconfig unless a check says otherwise.

Namespace​

Create the kamiwaza namespace. The installer does not create it, and the install never uses --create-namespace.

Label it for Istio sidecar injection, and to warn and audit against the Pod Security restricted profile. Do not enforce restricted: some Kamiwaza components, and Istio's init container, do not meet it yet.

kubectl create namespace kamiwaza
kubectl label namespace kamiwaza \
istio-injection=enabled \
pod-security.kubernetes.io/warn=restricted \
pod-security.kubernetes.io/audit=restricted

If your mesh uses revision-based injection (istio.io/rev=<revision>), use that label instead of istio-injection=enabled.

Check:

kubectl get namespace kamiwaza --show-labels
# Expect istio-injection=enabled (or istio.io/rev=...), and the pod-security labels

Installer identity​

Helm runs as a namespace-scoped identity, never as cluster-admin. It needs:

  • the built-in admin ClusterRole, bound in the namespace only with a RoleBinding;
  • create and manage rights on the Istio resources Kamiwaza writes into its namespace (VirtualService, DestinationRule, EnvoyFilter, AuthorizationPolicy, PeerAuthentication);
  • rights to manage endpoints and patch pods/status. The chart grants these to its own service accounts, and Kubernetes only lets an identity grant permissions it holds itself.

This manifest creates a service account with those rights. You can bind the same Roles to a user or group from your identity provider instead.

apiVersion: v1
kind: ServiceAccount
metadata:
name: kamiwaza-installer
namespace: kamiwaza
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kamiwaza-installer-admin
namespace: kamiwaza
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: admin
subjects:
- kind: ServiceAccount
name: kamiwaza-installer
namespace: kamiwaza
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: kamiwaza-installer-istio
namespace: kamiwaza
rules:
- apiGroups: ["networking.istio.io"]
resources: ["destinationrules", "envoyfilters", "virtualservices"]
verbs: ["create", "delete", "get", "list", "patch", "update", "watch"]
- apiGroups: ["security.istio.io"]
resources: ["authorizationpolicies", "peerauthentications"]
verbs: ["create", "delete", "get", "list", "patch", "update", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kamiwaza-installer-istio
namespace: kamiwaza
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: kamiwaza-installer-istio
subjects:
- kind: ServiceAccount
name: kamiwaza-installer
namespace: kamiwaza
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: kamiwaza-installer-delegation
namespace: kamiwaza
rules:
- apiGroups: [""]
resources: ["endpoints"]
verbs: ["create", "delete", "patch", "update"]
- apiGroups: [""]
resources: ["pods/status"]
verbs: ["patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kamiwaza-installer-delegation
namespace: kamiwaza
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: kamiwaza-installer-delegation
subjects:
- kind: ServiceAccount
name: kamiwaza-installer
namespace: kamiwaza

Save it as kamiwaza-installer.yaml and apply it as the administrator:

kubectl apply -f kamiwaza-installer.yaml

Give the installer a kubeconfig for this identity, following your cluster's normal practice for issuing credentials. For example, with a service account token, run as the administrator:

(
set -e
SERVER=$(kubectl config view --minify -o jsonpath='{.clusters[0].cluster.server}')
kubectl config view --minify --raw \
-o jsonpath='{.clusters[0].cluster.certificate-authority-data}' | base64 -d > cluster-ca.crt
TOKEN=$(kubectl -n kamiwaza create token kamiwaza-installer --duration=24h)

export KUBECONFIG=kamiwaza-installer.kubeconfig
kubectl config set-cluster kamiwaza --server="$SERVER" \
--certificate-authority=cluster-ca.crt --embed-certs
kubectl config set-credentials kamiwaza-installer --token="$TOKEN"
kubectl config set-context kamiwaza --cluster=kamiwaza --user=kamiwaza-installer --namespace=kamiwaza
kubectl config use-context kamiwaza
)

The parentheses keep these commands from changing your own shell's KUBECONFIG. If your kubeconfig names a certificate-authority file instead of embedding certificate-authority-data, copy that file to cluster-ca.crt instead. The token expires, and the API server may cap the duration you ask for. Issue a new one before each later upgrade or uninstall.

Check that the identity can create namespaced resources and cannot create cluster-scoped ones:

SA=system:serviceaccount:kamiwaza:kamiwaza-installer
kubectl auth can-i create deployments -n kamiwaza --as=$SA # yes
kubectl auth can-i create clusterroles --as=$SA # no
kubectl auth can-i create customresourcedefinitions --as=$SA # no
kubectl auth can-i list nodes --as=$SA # no

The installer's Istio rights are checked in the next section, once Istio's resource types exist.

Istio and the ingress gateway​

Kamiwaza routes all traffic through an Istio mesh that the platform owns. Install and operate Istio the way your organization does; Kamiwaza does not install or configure it. The mesh must provide:

  • istiod, in sidecar mode. Ambient mode is not supported.
  • An ingress gateway exposed to your users on port 443, for example through a LoadBalancer Service or your existing load balancer.
  • An Istio Gateway resource (networking.istio.io) that selects that gateway, serves the Kamiwaza domain on HTTPS port 443 with the edge certificate, and serves or redirects port 80. Its hosts must include the Kamiwaza domain: Istio programs a route only where the Gateway's hosts match, so a mismatch leaves the platform unreachable while the install reports success.

Kamiwaza attaches its routes to this Gateway from its own namespace. It writes nothing into the gateway's namespace.

A cluster without Istio. If the cluster has no mesh yet, for example a fresh k0s node with no load balancer, this example sets up one that meets the requirements above. It installs Istio 1.30.3 with Istio's Helm charts, and exposes the gateway on ports 80 and 443 of a node address through the Service's externalIPs instead of a LoadBalancer. Replace <node-ip> with the node address your users reach:

helm repo add istio https://istio-release.storage.googleapis.com/charts
helm repo update
helm upgrade --install istio-base istio/base --version 1.30.3 \
--namespace istio-system --create-namespace --wait
helm upgrade --install istiod istio/istiod --version 1.30.3 \
--namespace istio-system --wait
helm upgrade --install istio-ingressgateway istio/gateway --version 1.30.3 \
--namespace istio-system --wait \
--set service.type=ClusterIP --set 'service.externalIPs={<node-ip>}'

Kubernetes deprecates externalIPs as of 1.36, so expect the API server to warn when the gateway Service is created. The kube-proxy that k0s runs by default still routes it. If the cluster has a load balancer, leave out both --set options: the chart then creates a LoadBalancer Service.

Then create the Gateway. Save this as kamiwaza-gateway.yaml, with your domain in both hosts lists, and apply it with kubectl apply -f kamiwaza-gateway.yaml:

apiVersion: networking.istio.io/v1
kind: Gateway
metadata:
name: kamiwaza-gateway
namespace: istio-system
spec:
selector:
istio: ingressgateway
servers:
- port:
number: 80
name: http
protocol: HTTP
hosts:
- kamiwaza.example.com
tls:
httpsRedirect: true
- port:
number: 443
name: https
protocol: HTTPS
hosts:
- kamiwaza.example.com
tls:
mode: SIMPLE
credentialName: kamiwaza-edge-tls

kamiwaza-edge-tls is the edge certificate Secret, which you create in istio-system. The release name istio-ingressgateway and this Gateway match the gateway defaults in the Record table below: name, namespace, selector, service account, and Service. In DNS, point your domain at <node-ip>.

Check:

kubectl -n istio-system get deploy istiod
kubectl get crd virtualservices.networking.istio.io authorizationpolicies.security.istio.io
kubectl auth can-i create virtualservices.networking.istio.io -n kamiwaza \
--as=system:serviceaccount:kamiwaza:kamiwaza-installer # yes
kubectl get gateways.networking.istio.io -A
kubectl -n <gateway-namespace> get gateways.networking.istio.io <gateway-name> -o yaml
# Expect a server on port 443, protocol HTTPS, whose hosts include your domain,
# and a tls.credentialName naming your edge certificate Secret

Record the gateway's coordinates. The install uses the defaults in the right column unless you record something different.

FactHow to find itDefault
Gateway name and namespacekubectl get gateways.networking.istio.io -Akamiwaza-gateway in istio-system
Mesh namespace (where istiod runs)kubectl get deploy -A | grep istiodistio-system
Gateway pod selectorspec.selector of the Gateway resource. Record every label as-is.istio: ingressgateway
Gateway workload's service accountkubectl -n <ns> get pod -l <selector> -o jsonpath='{.items[0].spec.serviceAccountName}'. A gateway installed with istioctl uses istio-ingressgateway-service-account.istio-ingressgateway
Gateway Service's in-cluster hostname<service-name>.<namespace>.svc.cluster.localistio-ingressgateway.istio-system.svc.cluster.local
Sidecar styleSee belowNative sidecars

Sidecar style. Kamiwaza expects Istio to inject istio-proxy as a Kubernetes native sidecar (an init container with restartPolicy: Always). If your mesh injects it as a regular container, record that: the install sets a value for it. To tell which, start a short-lived pod in the kamiwaza namespace and look for istio-proxy:

kubectl -n kamiwaza run kamiwaza-sidecar-check --image=busybox:1.37 --restart=Never -- sleep 300
kubectl -n kamiwaza wait --for=condition=Ready pod/kamiwaza-sidecar-check --timeout=120s
kubectl -n kamiwaza get pod kamiwaza-sidecar-check \
-o jsonpath='{range .spec.initContainers[*]}init {.name} {.restartPolicy}{"\n"}{end}{range .spec.containers[*]}container {.name}{"\n"}{end}'
# "init istio-proxy Always": native sidecars.
# "container istio-proxy": classic injection.
# No istio-proxy line at all: the pod was not injected; check the namespace label.
kubectl -n kamiwaza delete pod kamiwaza-sidecar-check

A Pod Security restricted warning when this pod starts is expected; the namespace only warns and audits.

Edge certificate​

The gateway serves the Kamiwaza domain with a TLS certificate that you provide and rotate. Create it as a kubernetes.io/tls Secret in the gateway's namespace and name it in the Gateway's tls.credentialName. The certificate must cover the Kamiwaza domain in its subject alternative names.

If the certificate is not issued by a publicly trusted CA (a private CA, or self-signed), Kamiwaza workloads need its CA to verify the gateway. Publish the CA certificate, or for a self-signed certificate the certificate itself, in the Kamiwaza namespace:

kubectl -n kamiwaza create configmap kamiwaza-edge-trust \
--from-file=ca-certificates.crt=./edge-ca.crt

This ConfigMap holds public certificates only. Never put a private key in it. Kamiwaza reads it by name when it exists; no install value refers to it.

Check:

kubectl -n <gateway-namespace> get secret <edge-secret> -o jsonpath='{.data.tls\.crt}' \
| base64 -d | openssl x509 -noout -subject -issuer -ext subjectAltName -enddate
# Expect your domain under subjectAltName and an end date in the future

Record: whether the certificate is self-signed (subject and issuer are the same). The install sets a value for a self-signed certificate.

DNS​

The Kamiwaza domain must resolve to the ingress gateway in two places:

  • For users: your DNS points the domain at the gateway's external address.
  • Inside the cluster: Kamiwaza services and apps call the platform at its public domain. Pods must resolve the domain to an address that reaches the gateway. If your cluster DNS forwards to a resolver that already answers for the domain, and pods can reach that address, nothing more is needed. Otherwise, add a cluster DNS entry that maps the domain to the gateway Service's in-cluster hostname.

With CoreDNS, add a rewrite line to the Corefile, above the kubernetes plugin:

rewrite name exact kamiwaza.example.com istio-ingressgateway.istio-system.svc.cluster.local

Some distributions, k0s among them, manage the CoreDNS configuration and can put it back on restart or upgrade. On those, make the change through the distribution's own configuration. On k0s, add a patch under spec.network.coreDNS.patches in the k0s configuration that replaces the coredns ConfigMap's Corefile with your edited copy (see k0s component patches), then restart k0s so it reads the configuration again. Make the same change on every controller node.

First find the node's configuration file and the datastore k0s runs on. A restart with a configuration that names a different datastore brings the cluster up empty:

systemctl cat k0scontroller | grep ExecStart
# --config=<file> names the configuration file. Without --config, k0s reads
# /etc/k0s/k0s.yaml if it exists, and otherwise runs on built-in defaults
sudo grep -E '^ +type: (etcd|kine)$' /run/k0s/k0s.yaml
# The datastore the running k0s uses: etcd or kine
  • If the configuration file exists, add the patch to it and leave spec.storage as it is.
  • If it does not, create /etc/k0s/k0s.yaml with the patch and a spec.storage.type equal to the running datastore; k0s fills in defaults for everything else. Do not generate the file with k0s config create. Its output sets type: etcd, while a node started with --single runs on kine, and restarting that node with type: etcd brings it up on a new, empty etcd datastore.

Start from the current Corefile (kubectl -n kube-system get configmap coredns -o jsonpath='{.data.Corefile}'), add your line, and put the whole file in the patch. This is a complete file for a node that has none; to an existing file, add only spec.network.coreDNS:

# /etc/k0s/k0s.yaml
apiVersion: k0s.k0sproject.io/v1beta1
kind: ClusterConfig
spec:
storage:
type: kine # the datastore the check above printed
network:
coreDNS:
patches:
- target:
kind: ConfigMap
name: coredns
namespace: kube-system
patch:
type: StrategicMergePatch
content: |
data:
Corefile: |
.:53 {
# your complete edited Corefile
}

type: kine on its own selects k0s's default SQLite database under /var/lib/k0s, which is what --single uses. If ExecStart also sets --data-dir, copy kine.dataSource from /run/k0s/k0s.yaml into the file too.

Validate the file, take a backup, and restart k0s. k0s backup copies the configuration file, so it fails on a node that has none until you create it. The archive holds the datastore and the cluster's private keys: keep it off the node, somewhere safe. If the restart goes wrong, restore it as described in k0s backup and restore.

sudo "$(command -v k0s)" config validate --config /etc/k0s/k0s.yaml # or the file --config names
sudo "$(command -v k0s)" backup --save-path "$HOME"
sudo systemctl restart k0scontroller

After any cluster DNS change, allow a minute or two for CoreDNS to reload before running the check.

Check from your workstation, and from a pod:

curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/
# Any HTTP status proves DNS and TLS reach the gateway. Before Kamiwaza is
# installed, expect 404.

kubectl -n default run kamiwaza-dns-check --restart=Never \
--image=curlimages/curl --annotations=sidecar.istio.io/inject=false \
--overrides='{"spec":{"securityContext":{"runAsNonRoot":true,"runAsUser":100,"seccompProfile":{"type":"RuntimeDefault"}},"containers":[{"name":"kamiwaza-dns-check","image":"curlimages/curl","args":["-sk","-o","/dev/null","-w","%{http_code}\\n","https://kamiwaza.example.com/"],"securityContext":{"allowPrivilegeEscalation":false,"capabilities":{"drop":["ALL"]}}}]}}'
kubectl -n default wait --for=jsonpath='{.status.phase}'=Succeeded \
pod/kamiwaza-dns-check --timeout=120s
kubectl -n default logs kamiwaza-dns-check
# Expect the same status from inside the cluster
kubectl -n default delete pod kamiwaza-dns-check

Record: the domain.

Block storage​

Kamiwaza's databases and caches use ReadWriteOnce persistent volumes. The cluster must have a StorageClass that provisions them. Record its name; the install names it explicitly.

Size the storage provider for Kamiwaza's volume requests plus your provider's replication and free-space reserves. A StorageClass existing does not prove its volumes can be provisioned; the write test below does. Node-local storage, such as local-path, works on a single node but does not replicate data.

Any reclaimPolicy works, including Delete, the local-path default. The policy decides only what happens to a volume once its claim is deleted, as at uninstall.

Kamiwaza requests these volumes by default, about 624 GiB in total. Apps you deploy later add their own.

VolumeDefault requestHoldsChange it with
kamiwaza-registry-data500 GiBModel and inference-engine cacheregistry.persistence.size
seaweedfs-data100 GiBBuilt-in object storeseaweedfs.persistence.size
core-postgres-data10 GiBPlatform database—
keycloak-postgres-data10 GiBIdentity database—
core-spicedb-postgres-data1 GiBAuthorization database—
core-etcd-data-core-etcd-0 to -21 GiB eachConfiguration store—

The model registry cache holds inference-engine images and the model files your users deploy. If you serve only external models (for example, through a cloud provider), or your storage is small, set registry.persistence.size lower, for example 100Gi, in your install values before the first install. An existing volume keeps its size on upgrade.

Check the classes, then run a write test in the Kamiwaza namespace: write a marker, replace the pod, and read the marker back.

kubectl get storageclass
# Expect one class marked (default), or note the class you will name at install

cat > kamiwaza-storage-check.yaml <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: kamiwaza-storage-check
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
# storageClassName: <class> # uncomment if you will not use the default
---
apiVersion: v1
kind: Pod
metadata:
name: kamiwaza-storage-check
annotations:
sidecar.istio.io/inject: "false"
spec:
securityContext:
runAsNonRoot: true
runAsUser: 1000
fsGroup: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: check
image: busybox:1.37
command: ["sh", "-c", "sleep 3600"]
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
volumeMounts:
- name: data
mountPath: /data
volumes:
- name: data
persistentVolumeClaim:
claimName: kamiwaza-storage-check
EOF

kubectl -n kamiwaza apply -f kamiwaza-storage-check.yaml
kubectl -n kamiwaza wait --for=condition=Ready pod/kamiwaza-storage-check --timeout=5m
kubectl -n kamiwaza exec kamiwaza-storage-check -- sh -c 'echo ok > /data/marker'

# Replace the pod; the claim and its volume stay
kubectl -n kamiwaza delete pod kamiwaza-storage-check
kubectl -n kamiwaza apply -f kamiwaza-storage-check.yaml
kubectl -n kamiwaza wait --for=condition=Ready pod/kamiwaza-storage-check --timeout=5m
kubectl -n kamiwaza exec kamiwaza-storage-check -- cat /data/marker # expect: ok

kubectl -n kamiwaza delete -f kamiwaza-storage-check.yaml

A class with volumeBindingMode: WaitForFirstConsumer binds the claim only once the pod is scheduled, so a Pending claim before the pod starts is normal.

Record: the StorageClass name.

Object storage​

Workrooms, the Skills Library, and other features store objects. Choose one:

  • The built-in object store (default). Kamiwaza runs SeaweedFS in its namespace, backed by a ReadWriteOnce volume from your block storage. Nothing to prepare.
  • Your own S3-compatible service. See S3 workroom storage.

Kubernetes API addresses​

Some Kamiwaza apps call the Kubernetes API from inside the namespace, and Kamiwaza allows that traffic with NetworkPolicy. Many CNIs, including the ones k0s and OpenShift use, evaluate NetworkPolicy after a Service address is translated to the API server's real address. So the policy needs both the kubernetes Service address and every API server endpoint address. With only the first, apps install cleanly and then fail at runtime.

Record both sets of addresses, as /32 for IPv4 or /128 for IPv6:

kubectl -n default get service kubernetes -o jsonpath='{.spec.clusterIPs[*]}{"\n"}'
kubectl -n default get endpointslices -l kubernetes.io/service-name=kubernetes \
-o jsonpath='{.items[*].endpoints[*].addresses[*]}{"\n"}'

Copy the addresses exactly. A wrong address still installs cleanly, and the apps then fail at runtime. If the API server's addresses change (for example, a control-plane node is replaced), update them in the install values and upgrade.

If your CNI enforces NetworkPolicy before Service address translation, the endpoint addresses are not needed; the Service address alone works. Record it anyway.

Registry access​

Kamiwaza's charts and images are published to a public registry and pull without credentials. Nodes must be able to pull from it; see Outbound network access.

If you mirror images into your own registry and it requires credentials, create a kubernetes.io/dockerconfigjson pull Secret in the Kamiwaza namespace. Kamiwaza uses only namespaced pull Secrets, never node-wide registry credentials. Build the Secret from a file, so the password never appears in a command line or your shell history:

read -rs REGISTRY_PASSWORD # typed, not echoed
(umask 077; d=$(mktemp -d)
printf '{"auths":{"%s":{"auth":"%s"}}}' '<registry-host>' \
"$(printf '%s:%s' '<username>' "$REGISTRY_PASSWORD" | base64 | tr -d '\n')" > "$d/config.json"
kubectl -n kamiwaza create secret generic kamiwaza-pull \
--type=kubernetes.io/dockerconfigjson --from-file=.dockerconfigjson="$d/config.json"
rm -rf "$d")
unset REGISTRY_PASSWORD

<registry-host> is the registry's host name, with its port if it has one, as it appears in the image references.

Record: the pull Secret's name, kamiwaza-pull above. The install lists it in global.imagePullSecrets. That list does not reach the engine pods of tenant-mode inference: they carry only the staging credential, which defaults to the first entry in global.imagePullSecrets.

Node kernel limits​

Every Istio sidecar watches files with inotify. Distribution defaults are too low for the number of pods Kamiwaza runs, and sidecars then crash with too many open files. Set these floors on every node, persistently:

cat <<'EOF' | sudo tee /etc/sysctl.d/99-kamiwaza-inotify.conf
fs.inotify.max_user_instances = 8192
fs.inotify.max_user_watches = 1048576
fs.inotify.max_queued_events = 16384
EOF
sudo sysctl --system

Check on each node:

sysctl fs.inotify.max_user_instances fs.inotify.max_user_watches fs.inotify.max_queued_events
# Expect values at or above 8192, 1048576, and 16384

Managed platforms that do not give you node access usually expose these limits through a node configuration resource; use your platform's mechanism.

Outbound network access​

Nodes and pods need outbound HTTPS (port 443) to:

  • ghcr.io and pkg-containers.githubusercontent.com, for Kamiwaza charts and images;
  • the model sources you use, for example huggingface.co and its content hosts, or your cloud model provider's endpoint.

The checks on this page use the public busybox and curlimages/curl images from Docker Hub. If your cluster cannot reach Docker Hub, mirror them to your registry.

For installs with no outbound access, see Install Without Internet Access.

GPU nodes​

GPU inference is optional. To serve models on NVIDIA GPUs, prepare each GPU node as described in GPU Nodes: a driver that meets the engine images' floor, persistence mode, the NVIDIA container runtime and the nvidia RuntimeClass, the NVIDIA GPU Operator, and no NoSchedule or NoExecute taint on the node. Kamiwaza installs none of these.

Prepared nodes are not enough on their own: a default install cannot use GPUs. After the install, the cluster owner turns GPU serving on with Tenant-mode inference.

Check:

kubectl get nodes -o custom-columns='NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu'
# Expect a count on each GPU node
kubectl get runtimeclass nvidia
# Expect a RuntimeClass named nvidia, with handler nvidia

Record: which nodes have GPUs, and the GPU model on each. The cluster owner needs them to publish a signed profile.

Multi-node clusters​

On a cluster with more than one node, also provide:

  • Storage that survives a node loss. Use a StorageClass whose volumes are replicated or network-attached, such as a CSI driver for your SAN or cloud block storage, or a replicated provider like Longhorn. Node-local storage ties each database to one node: if that node is lost, so is the data.
  • Every GPU node prepared. Prepare each GPU node as in GPU nodes, including nodes added later.
  • A CNI that enforces NetworkPolicy. Kamiwaza isolates apps and jobs from each other with NetworkPolicy. A CNI that accepts policies but does not enforce them leaves that isolation off without any error.
  • The inotify floor on every node, including nodes added later.
  • The API endpoint addresses of every control-plane node in the recorded Kubernetes API addresses.

Next step​

When every check passes, hand the recorded facts and the installer kubeconfig to whoever runs the install, and continue to Install on Kubernetes with Istio.