Skip to main content
Version: 1.3.3 (Latest)

Platform Prerequisites for OpenShift with GreyMatter

Kamiwaza installs into a single namespace on an OpenShift cluster that you already run, with GreyMatter as its service mesh. Before the install, a platform administrator (someone with cluster-admin rights, and credentials for your GreyMatter platform) prepares the cluster so it provides everything on this page. The Kamiwaza chart does not create or repair any of it: it creates no cluster-scoped objects, installs no operators, and does not configure OpenShift or GreyMatter. If a prerequisite is missing, the install fails or the platform comes up unable to serve traffic.

Work through each section and run its check. Each section also says what to record; the install needs those facts.

Hardware sizing and supported operating systems are covered separately, in System Requirements.

Summary​

PrerequisiteWhat the cluster must provide
NamespaceA pre-created project
Installer identityA namespace-scoped identity to run Helm as. Never cluster-admin
Security context constraintsOpenShift's default restricted-v2 for Kamiwaza's service accounts, and no custom SCC
GreyMatterThe namespace enrolled in GreyMatter, and a Git repository for Kamiwaza's GreyMatter configuration
Edge certificateA TLS certificate for the domain, issued by your GreyMatter mesh CA and valid for server and client authentication
DNSThe domain resolves to the OpenShift router, from clients and from inside the cluster
Block storageAn RWO StorageClass that provisions durable volumes
Object storageA decision: the built-in object store, or your own S3-compatible service
Kubernetes API addressesThe API server's Service and endpoint addresses, recorded for network policy
Registry accessNodes can pull Kamiwaza images and the GreyMatter proxy image
Node kernel limitsfs.inotify limits at or above the floor on every node
Outbound network accessHTTPS to the image registry, your Git server, and any model sources you use

GPU inference and multi-node clusters add requirements; see GPU nodes and Multi-node clusters.

Kamiwaza installs into a namespace named kamiwaza; use that name. The examples use kamiwaza.example.com as the domain; substitute your own. Run the checks with the administrator's kubeconfig unless a check says otherwise. oc works everywhere this page uses kubectl.

Namespace​

Create the kamiwaza namespace. The installer does not create it, and the install never uses --create-namespace.

Do not label it for Istio injection: GreyMatter injects its own sidecars once the namespace is enrolled (see GreyMatter). OpenShift applies pod security through security context constraints, covered below.

oc new-project kamiwaza
oc label namespace kamiwaza app.kubernetes.io/part-of=kamiwaza

oc new-project also switches your kubeconfig's current namespace to kamiwaza. Switch back with oc project <previous namespace> if you need to.

Check:

oc get namespace kamiwaza --show-labels
# Expect app.kubernetes.io/part-of=kamiwaza, and no istio-injection label

Installer identity​

Helm runs as a namespace-scoped identity, never as cluster-admin. It needs:

  • the built-in admin ClusterRole, bound in the namespace only with a RoleBinding;
  • rights to manage endpoints, patch pods/status, and read GreyMatter's sidecarresources. The chart grants these to its own service accounts, and Kubernetes only lets an identity grant permissions it holds itself. Without the sidecarresources read, the install stops with attempting to grant RBAC permissions not currently held.

This manifest creates a service account with those rights. You can bind the same Roles to a user or group from your identity provider instead.

apiVersion: v1
kind: ServiceAccount
metadata:
name: kamiwaza-installer
namespace: kamiwaza
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kamiwaza-installer-admin
namespace: kamiwaza
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: admin
subjects:
- kind: ServiceAccount
name: kamiwaza-installer
namespace: kamiwaza
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: kamiwaza-installer-delegation
namespace: kamiwaza
rules:
- apiGroups: [""]
resources: ["endpoints"]
verbs: ["create", "delete", "patch", "update"]
- apiGroups: [""]
resources: ["pods/status"]
verbs: ["patch"]
- apiGroups: ["tenant.greymatter.io"]
resources: ["sidecarresources"]
verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kamiwaza-installer-delegation
namespace: kamiwaza
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: kamiwaza-installer-delegation
subjects:
- kind: ServiceAccount
name: kamiwaza-installer
namespace: kamiwaza

Save it as kamiwaza-installer.yaml and apply it as the administrator:

kubectl apply -f kamiwaza-installer.yaml

Give the installer a kubeconfig for this identity, following your cluster's normal practice for issuing credentials. For example, with a service account token, run as the administrator:

(
set -e
SERVER=$(kubectl config view --minify -o jsonpath='{.clusters[0].cluster.server}')
kubectl config view --minify --raw \
-o jsonpath='{.clusters[0].cluster.certificate-authority-data}' | base64 -d > cluster-ca.crt
TOKEN=$(kubectl -n kamiwaza create token kamiwaza-installer --duration=24h)

export KUBECONFIG=kamiwaza-installer.kubeconfig
kubectl config set-cluster kamiwaza --server="$SERVER" \
--certificate-authority=cluster-ca.crt --embed-certs
kubectl config set-credentials kamiwaza-installer --token="$TOKEN"
kubectl config set-context kamiwaza --cluster=kamiwaza --user=kamiwaza-installer --namespace=kamiwaza
kubectl config use-context kamiwaza
)

The parentheses keep these commands from changing your own shell's KUBECONFIG. If your kubeconfig names a certificate-authority file instead of embedding certificate-authority-data, copy that file to cluster-ca.crt instead. The token expires, and the API server may cap the duration you ask for. Issue a new one before each later upgrade or uninstall.

Check that the identity can create namespaced resources and cannot create cluster-scoped ones:

SA=system:serviceaccount:kamiwaza:kamiwaza-installer
kubectl auth can-i create deployments -n kamiwaza --as=$SA # yes
kubectl auth can-i patch pods/status -n kamiwaza --as=$SA # yes
kubectl auth can-i get sidecarresources.tenant.greymatter.io -n kamiwaza --as=$SA # yes
kubectl auth can-i create clusterroles --as=$SA # no
kubectl auth can-i create customresourcedefinitions --as=$SA # no
kubectl auth can-i list nodes --as=$SA # no
kubectl auth can-i create securitycontextconstraints --as=$SA # no

When your kubeconfig sets a namespace, the cluster-scoped checks also print Warning: resource '...' is not namespace scoped before their answer. The yes or no is the result. The security context constraint check below passes -n itself, so it always prints that warning.

Security context constraints​

Kamiwaza runs under OpenShift's default restricted-v2 security context constraint (SCC), which assigns each pod a user ID from the namespace's range. The chart creates no SCCs and needs none.

Don't create or bind a custom SCC for Kamiwaza, and don't grant anyuid, nonroot, or a similar SCC to its service accounts. Kamiwaza's pods rely on restricted-v2 assigning their user ID; under an SCC that doesn't, they fail to start with errors such as container has runAsNonRoot and image has non-numeric user or image will run as root.

Check that no permissive SCC is granted to the whole namespace or to all service accounts. It works before the install creates Kamiwaza's service accounts:

for scc in anyuid nonroot nonroot-v2 privileged; do
printf '%s: ' "$scc"
oc auth can-i use "securitycontextconstraints/$scc" -n kamiwaza \
--as=system:serviceaccount:kamiwaza:core-scheduler \
--as-group=system:serviceaccounts --as-group=system:serviceaccounts:kamiwaza
done
# Expect "no" for each

A grant to a single service account shows up in the install's verification, which checks that every Kamiwaza pod runs under restricted-v2.

GreyMatter​

Kamiwaza routes all traffic through a GreyMatter mesh that the platform owns. GreyMatter Core 2.4 must already be installed and healthy; Kamiwaza does not install or configure it. The steps below change GreyMatter configuration, so they need your GreyMatter credentials.

GreyMatter configures Kamiwaza's routes from configuration files (GSL) in a Git repository. Kamiwaza supplies those files: you publish them at install time, and Kamiwaza adds routes to them at runtime as apps and models are deployed. The browser-facing entry point is a GreyMatter edge in the Kamiwaza namespace, published through an OpenShift Route with TLS passthrough.

Enroll the namespace​

In your GreyMatter Core configuration, add the namespace to tenant_namespaces in config.cue. If your configuration also sets ftr_namespaces, add it there too. Leave ftr_namespaces unset if it is: when it is absent, forced traffic redirection already covers every tenant namespace, and setting it limits redirection to the namespaces it lists.

config: {
tenant_namespaces: [
"kamiwaza",
]
}

The namespace must exist before GreyMatter receives this change. Apply it through your normal GreyMatter change process.

Check that GreyMatter has acted on the namespace. It starts its tenant controllers there and copies its image pull Secret in:

oc -n kamiwaza get deployments
# Expect the GreyMatter tenant controllers, for example greymatter-tcm
oc -n kamiwaza get secret greymatter-image-pull -o jsonpath='{.type}{"\n"}'
# Expect kubernetes.io/dockerconfigjson

The tenant controllers stay in ContainerCreating until the edge certificate Secret exists: they mount it. Check that they are ready after that section.

Record:

  • the GreyMatter image pull Secret's name, if it is not greymatter-image-pull;

  • the GreyMatter proxy image your platform runs. Kamiwaza runs the same proxy beside its Ray workloads:

    oc get deployments -A \
    -o jsonpath='{range .items[*].spec.template.spec.containers[*]}{.image}{"\n"}{end}' \
    | grep -i greymatter-proxy | sort -u

Tenant configuration repository​

Provide a Git repository for Kamiwaza's GreyMatter configuration, with a branch that GreyMatter Sync watches. Kamiwaza pushes route changes to that branch while it runs, so the branch must accept direct pushes, not only pull requests. Kamiwaza's files live under _tenants/kamiwaza in the repository.

Create the GreyMatter project there with the greymatter CLI, and push it. Use the CLI that matches your GreyMatter platform: the .greymatter file at the root of your GreyMatter Core configuration records the platform and CLI versions. GreyMatter distributes the CLI to its customers; if the exact CLI build isn't offered, use the newest patch release of the same minor version.

The CLI writes the project into the current directory, so create the project directory first:

git clone https://git.example.com/org/kamiwaza-gsl.git && cd kamiwaza-gsl
git switch -c tenant/kamiwaza
mkdir -p _tenants/kamiwaza && cd _tenants/kamiwaza
greymatter create project --security pki --openshift kamiwaza
cd ../..
git add _tenants/kamiwaza
git commit -m "Create the Kamiwaza GreyMatter project"
git push -u origin tenant/kamiwaza

The CLI writes .greymatter and cue.mod/, including the GSL packages that match your GreyMatter version, and a starter greymatter/ directory. Commit them unchanged. Kamiwaza's own files are added to greymatter/ at install time, replacing the starter files.

Then create the Secret that connects GreyMatter Sync and Kamiwaza to the repository. It holds the repository's location and a credential that can push to the branch. Read the credential without leaving it in your shell history:

printf 'Git username: '; read -r GIT_USER
printf 'Git token: '; read -rs GIT_TOKEN; echo
(umask 077; d=$(mktemp -d)
printf '%s' "$GIT_TOKEN" > "$d/token"
oc -n kamiwaza create secret generic greymatter-tenant-repo \
--type=greymatter.io/repo \
--from-literal=auth_type=https \
--from-literal=url=https://git.example.com/org/kamiwaza-gsl.git \
--from-literal=branch=tenant/kamiwaza \
--from-literal=relative_path=_tenants/kamiwaza \
--from-literal=http_username="$GIT_USER" \
--from-file=http_password="$d/token" \
--from-literal=tls_insecure_verify=false
rm -rf "$d")
unset GIT_USER GIT_TOKEN

If your Git server uses a private CA, add --from-file=tls_remote_ca=<ca-bundle>. Keep tls_insecure_verify=false.

Check:

oc -n kamiwaza get secret greymatter-tenant-repo -o json \
| jq -r '.type, (.data | keys | join(" "))'
# Expect greymatter.io/repo, then the keys
# auth_type branch http_password http_username relative_path tls_insecure_verify url
# (and tls_remote_ca, if you added it)

Record: the repository URL and branch. Whoever publishes Kamiwaza's files at install time needs push access to them.

Network policy​

If the kamiwaza namespace denies traffic by default, allow:

  • DNS to OpenShift DNS;
  • the Kubernetes API, at the addresses you record below;
  • HTTPS to your Git server, from the GreyMatter tenant controllers and from Kamiwaza's core-scheduler and core-raycluster pods;
  • GreyMatter control-plane traffic to and from the namespace;
  • the OpenShift router to the GreyMatter edge;
  • HTTPS to the OpenShift router, for calls to the Kamiwaza domain from inside the cluster (see DNS);
  • HTTPS to the hosts in Outbound network access;
  • traffic within the namespace.

Scope each egress rule to the addresses it needs, not to every address.

Edge certificate​

The Kamiwaza domain is served with a TLS certificate that you provide and rotate. The certificate must cover the Kamiwaza domain in its subject alternative names.

The GreyMatter edge serves it; the OpenShift router passes the connection through without decrypting it. The same certificate is also the mesh identity in the Kamiwaza namespace: GreyMatter's tenant controllers and Kamiwaza's Ray sidecar present it when they connect to other mesh services, including GreyMatter Core. So it must:

  • be issued by the CA that your GreyMatter Core trusts for mesh (PKI) traffic. That CA is the ca.crt of Core's PKI certificate Secret, by default greymatter-edge-ingress in the greymatter namespace. A certificate from any other CA, or a self-signed one, leaves the tenant controllers unable to connect to GreyMatter Core, and GreyMatter never applies Kamiwaza's configuration;
  • be valid for both server and client authentication (extended key usages serverAuth and clientAuth).

Publicly trusted CAs don't issue certificates with client authentication, and browsers don't trust your mesh CA, so users see a certificate warning unless their machines trust that CA.

Create it in the Kamiwaza namespace, named greymatter-edge-ingress, with the server certificate and any intermediates in tls.crt and your mesh CA in ca.crt:

oc -n kamiwaza create secret generic greymatter-edge-ingress \
--from-file=ca.crt=./edge-ca.crt \
--from-file=tls.crt=./edge.crt \
--from-file=tls.key=./edge.key

Kamiwaza workloads need the mesh CA to verify the edge, since it is not publicly trusted. Publish the CA certificate in the Kamiwaza namespace:

kubectl -n kamiwaza create configmap kamiwaza-edge-trust \
--from-file=ca-certificates.crt=./edge-ca.crt

This ConfigMap holds public certificates only. Never put a private key in it. Kamiwaza reads it by name when it exists; no install value refers to it.

Check:

oc -n kamiwaza get secret greymatter-edge-ingress -o jsonpath='{.data.tls\.crt}' \
| base64 -d | openssl x509 -noout -subject -issuer -ext subjectAltName -enddate -purpose \
| grep -E 'subject|issuer|DNS:|notAfter|^SSL (client|server) :'
# Expect your domain under subjectAltName, an end date in the future,
# "SSL client : Yes" and "SSL server : Yes"
oc -n kamiwaza get secret greymatter-edge-ingress -o json | jq -r '.data | keys | join(" ")'
# Expect ca.crt tls.crt tls.key

# The certificate chains to the CA GreyMatter Core trusts
# (use your Core PKI Secret if it isn't the default):
oc -n greymatter get secret greymatter-edge-ingress -o jsonpath='{.data.ca\.crt}' \
| base64 -d > greymatter-mesh-ca.crt
oc -n kamiwaza get secret greymatter-edge-ingress -o jsonpath='{.data.tls\.crt}' \
| base64 -d > kamiwaza-edge.crt
openssl verify -CAfile greymatter-mesh-ca.crt -untrusted kamiwaza-edge.crt kamiwaza-edge.crt
# Expect kamiwaza-edge.crt: OK

Once this Secret exists, the GreyMatter tenant controllers from Enroll the namespace start. Check that they are ready:

oc -n kamiwaza get deployments
# Expect greymatter-tcm and the other tenant controllers ready

If they were already running with a different certificate, restart them so they load this one: oc -n kamiwaza rollout restart deployment <name> for each.

DNS​

The Kamiwaza domain must resolve to the OpenShift router's external address, the same way as other passthrough Routes on the cluster:

  • For users: your DNS points the domain at the router's load balancer. Find its address with:

    oc -n openshift-ingress get service router-default \
    -o jsonpath='{.status.loadBalancer.ingress[0].hostname}{.status.loadBalancer.ingress[0].ip}{"\n"}'

    A host name (as on AWS) needs a CNAME record; an IP address needs an A record. The router passes TLS through to the GreyMatter edge, so don't put a proxy or CDN that terminates TLS in front of the domain: publish the record as plain DNS.

  • Inside the cluster: Kamiwaza services and apps call the platform at its public domain, so pods must resolve it and reach the router the same way. OpenShift's cluster DNS forwards to your upstream resolvers, so this usually needs nothing more.

The router routes on the TLS server name, so clients must connect using the Kamiwaza domain. The Route itself is created when Kamiwaza's GreyMatter configuration is published during the install; until then the router answers for the domain with its own 503 page.

Check from your workstation, and from a pod:

curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/
# Any HTTP status proves DNS reaches the router. Before Kamiwaza is
# installed, expect 503.

(
export KUBECONFIG=kamiwaza-installer.kubeconfig
oc -n kamiwaza run kamiwaza-dns-check --restart=Never --image=curlimages/curl \
-- -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/
oc -n kamiwaza wait --for=jsonpath='{.status.phase}'=Succeeded \
pod/kamiwaza-dns-check --timeout=60s
oc -n kamiwaza logs kamiwaza-dns-check
# Expect the same status from inside the cluster
oc -n kamiwaza delete pod kamiwaza-dns-check
)

The pod commands run with the installer kubeconfig for the same reason as the block storage write test: as cluster-admin, OpenShift can admit the pod under a security context constraint that assigns it no user ID.

Record: the domain.

Block storage​

Kamiwaza's databases and caches use ReadWriteOnce persistent volumes. The cluster must have a StorageClass that provisions them. Record its name; the install names it explicitly.

Size the storage provider for Kamiwaza's volume requests plus your provider's replication and free-space reserves. A StorageClass existing does not prove its volumes can be provisioned; the write test below does. Node-local storage, such as local-path, works on a single node but does not replicate data.

Kamiwaza requests these volumes by default, about 624 GiB in total. Apps you deploy later add their own.

VolumeDefault requestHoldsChange it with
kamiwaza-registry-data500 GiBModel and inference-engine cacheregistry.persistence.size
seaweedfs-data100 GiBBuilt-in object storeseaweedfs.persistence.size
core-postgres-data10 GiBPlatform database—
keycloak-postgres-data10 GiBIdentity database—
core-spicedb-postgres-data1 GiBAuthorization database—
core-etcd-data-core-etcd-0 to -21 GiB eachConfiguration store—

The model registry cache holds inference-engine images and the model files your users deploy. If you serve only external models (for example, through a cloud provider), or your storage is small, set registry.persistence.size lower, for example 100Gi, in your install values before the first install. An existing volume keeps its size on upgrade.

Check the classes, then run a write test in the Kamiwaza namespace: write a marker, replace the pod, and read the marker back. Run the test with the installer kubeconfig from Installer identity, as below. As cluster-admin, OpenShift can admit the test pod under a security context constraint that assigns it no user ID, and the pod then fails to start.

kubectl get storageclass
# Expect one class marked (default), or note the class you will name at install

(
export KUBECONFIG=kamiwaza-installer.kubeconfig
cat > kamiwaza-storage-check.yaml <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: kamiwaza-storage-check
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
# storageClassName: <class> # uncomment if you will not use the default
---
apiVersion: v1
kind: Pod
metadata:
name: kamiwaza-storage-check
spec:
securityContext:
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
containers:
- name: check
image: busybox:1.37
command: ["sh", "-c", "sleep 3600"]
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
volumeMounts:
- name: data
mountPath: /data
volumes:
- name: data
persistentVolumeClaim:
claimName: kamiwaza-storage-check
EOF

kubectl -n kamiwaza apply -f kamiwaza-storage-check.yaml
kubectl -n kamiwaza wait --for=condition=Ready pod/kamiwaza-storage-check --timeout=5m
kubectl -n kamiwaza exec kamiwaza-storage-check -- sh -c 'echo ok > /data/marker'

# Replace the pod; the claim and its volume stay
kubectl -n kamiwaza delete pod kamiwaza-storage-check
kubectl -n kamiwaza apply -f kamiwaza-storage-check.yaml
kubectl -n kamiwaza wait --for=condition=Ready pod/kamiwaza-storage-check --timeout=5m
kubectl -n kamiwaza exec kamiwaza-storage-check -- cat /data/marker # expect: ok

kubectl -n kamiwaza delete -f kamiwaza-storage-check.yaml
)

A class with volumeBindingMode: WaitForFirstConsumer binds the claim only once the pod is scheduled, so a Pending claim before the pod starts is normal.

Record: the StorageClass name.

Object storage​

Workrooms, the Skills Library, and other features store objects. Choose one:

  • The built-in object store (default). Kamiwaza runs SeaweedFS in its namespace, backed by a ReadWriteOnce volume from your block storage. Nothing to prepare.
  • Your own S3-compatible service. See S3 workroom storage.

Kubernetes API addresses​

Some Kamiwaza apps call the Kubernetes API from inside the namespace, and Kamiwaza allows that traffic with NetworkPolicy. Many CNIs, including the ones k0s and OpenShift use, evaluate NetworkPolicy after a Service address is translated to the API server's real address. So the policy needs both the kubernetes Service address and every API server endpoint address. With only the first, apps install cleanly and then fail at runtime.

While OpenShift rolls out the API server or restarts a control-plane node, for example on a new cluster or after a configuration change, the endpoint list leaves out the server being restarted. Record the addresses only when no rollout is in progress:

oc get clusteroperator kube-apiserver
# Expect PROGRESSING False. If the cluster has no kube-apiserver cluster
# operator, as can happen when the provider hosts the control plane, skip this.
oc get machineconfigpool master
# Expect UPDATED True and UPDATING False. Skip this too if the provider hosts
# the control plane.

Record both sets of addresses, as /32 for IPv4 or /128 for IPv6:

kubectl -n default get service kubernetes -o jsonpath='{.spec.clusterIPs[*]}{"\n"}'
kubectl -n default get endpointslices -l kubernetes.io/service-name=kubernetes \
-o jsonpath='{.items[*].endpoints[*].addresses[*]}{"\n"}'
# Expect one endpoint address for each API server: three on a cluster with
# three control-plane nodes

Copy the addresses exactly. A wrong address still installs cleanly, and the apps then fail at runtime. If the API server's addresses change (for example, a control-plane node is replaced), update them in the install values and upgrade.

If your CNI enforces NetworkPolicy before Service address translation, the endpoint addresses are not needed; the Service address alone works. Record it anyway.

Registry access​

Kamiwaza's charts and images are published to a public registry and pull without credentials. Nodes must be able to pull from it; see Outbound network access.

If you mirror images into your own registry and it requires credentials, create a kubernetes.io/dockerconfigjson pull Secret in the Kamiwaza namespace. Kamiwaza uses only namespaced pull Secrets, never node-wide registry credentials. Build the Secret from a file, so the password never appears in a command line or your shell history:

read -rs REGISTRY_PASSWORD # typed, not echoed
(umask 077; d=$(mktemp -d)
printf '{"auths":{"%s":{"auth":"%s"}}}' '<registry-host>' \
"$(printf '%s:%s' '<username>' "$REGISTRY_PASSWORD" | base64 | tr -d '\n')" > "$d/config.json"
kubectl -n kamiwaza create secret generic kamiwaza-pull \
--type=kubernetes.io/dockerconfigjson --from-file=.dockerconfigjson="$d/config.json"
rm -rf "$d")
unset REGISTRY_PASSWORD

<registry-host> is the registry's host name, with its port if it has one, as it appears in the image references.

Record: the pull Secret's name, kamiwaza-pull above. The install lists it in global.imagePullSecrets.

Nodes must also pull the GreyMatter proxy image. GreyMatter provides its pull Secret; see Enroll the namespace.

Node kernel limits​

Every GreyMatter sidecar watches files with inotify. Node defaults can be too low for the number of pods Kamiwaza runs, and sidecars then crash with too many open files. Set these floors on every worker node.

On OpenShift, set them with a MachineConfig for each machine pool that runs Kamiwaza. Whether a custom MachineConfig is supported depends on your provider's support policy. If your platform doesn't let you create one (for example, ROSA with hosted control planes), set the same limits through your provider's node tuning mechanism, or ask your provider's support to set them.

Red Hat Enterprise Linux CoreOS already sets fs.inotify.max_user_watches to 65536 in /etc/sysctl.d/inotify.conf. Files in /etc/sysctl.d apply in name order and the last one wins, so Kamiwaza's file must sort after that one; the zz- prefix below does that.

This MachineConfig covers the worker pool, which on ROSA also includes the infra nodes. Save it as kamiwaza-inotify-machineconfig.yaml:

apiVersion: machineconfiguration.openshift.io/v1
kind: MachineConfig
metadata:
name: 99-worker-kamiwaza-inotify
labels:
machineconfiguration.openshift.io/role: worker
spec:
config:
ignition:
version: 3.2.0
storage:
files:
- path: /etc/sysctl.d/zz-kamiwaza-inotify.conf
mode: 0644
overwrite: true
contents:
source: data:,fs.inotify.max_user_instances%20%3D%208192%0Afs.inotify.max_user_watches%20%3D%201048576%0Afs.inotify.max_queued_events%20%3D%2016384%0A
oc apply -f kamiwaza-inotify-machineconfig.yaml
oc get machineconfigpool worker -w # wait until UPDATED is True

Applying a MachineConfig restarts the pool's nodes one at a time, a few minutes each. Schedule it with your platform's change process. If Kamiwaza is already installed, expect a brief interruption while the node running its edge restarts.

Check on each worker node:

oc debug node/<node> -- chroot /host \
sysctl fs.inotify.max_user_instances fs.inotify.max_user_watches fs.inotify.max_queued_events
# Expect values at or above 8192, 1048576, and 16384

Outbound network access​

Nodes and pods need outbound HTTPS (port 443) to:

  • ghcr.io and pkg-containers.githubusercontent.com, for Kamiwaza charts and images;
  • the model sources you use, for example huggingface.co and its content hosts, or your cloud model provider's endpoint;
  • your Git server, for the tenant configuration repository;
  • the registry that serves the GreyMatter proxy image.

The checks on this page use the public busybox and curlimages/curl images from Docker Hub. If your cluster cannot reach Docker Hub, mirror them to your registry.

For installs with no outbound access, see Install Without Internet Access.

GPU nodes​

GPU inference is optional. To serve models on NVIDIA GPUs, prepare each GPU node as described in GPU Nodes: a driver that meets the engine images' floor, persistence mode, the NVIDIA container runtime and the nvidia RuntimeClass, the NVIDIA GPU Operator, and no NoSchedule or NoExecute taint on the node. Kamiwaza installs none of these. The node commands on that page are for Ubuntu nodes. On OpenShift, install the NVIDIA GPU Operator with your platform's usual method, then run the checks on that page that apply to your nodes.

Prepared nodes are not enough on their own: a default install cannot use GPUs. After the install, the cluster owner turns GPU serving on with Tenant-mode inference.

Check:

kubectl get nodes -o custom-columns='NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu'
# Expect a count on each GPU node
kubectl get runtimeclass nvidia
# Expect a RuntimeClass named nvidia, with handler nvidia

Record: which nodes have GPUs, and the GPU model on each. The cluster owner needs them to publish a signed profile.

Multi-node clusters​

On a cluster with more than one node, also provide:

  • Storage that survives a node loss. Use a StorageClass whose volumes are replicated or network-attached, such as a CSI driver for your SAN or cloud block storage, or a replicated provider like Longhorn. Node-local storage ties each database to one node: if that node is lost, so is the data.
  • Every GPU node prepared. Prepare each GPU node as in GPU nodes, including nodes added later.
  • A CNI that enforces NetworkPolicy. Kamiwaza isolates apps and jobs from each other with NetworkPolicy. A CNI that accepts policies but does not enforce them leaves that isolation off without any error.
  • The inotify floor on every node, including nodes added later.
  • The API endpoint addresses of every control-plane node in the recorded Kubernetes API addresses.

Next step​

When every check passes, hand the recorded facts and the installer kubeconfig to whoever runs the install, and continue to Install on OpenShift with GreyMatter.