Platform Prerequisites for OpenShift with GreyMatter
Kamiwaza installs into a single namespace on an OpenShift cluster that you already run, with GreyMatter as its service mesh. Before the install, a platform administrator (someone with cluster-admin rights, and credentials for your GreyMatter platform) prepares the cluster so it provides everything on this page. The Kamiwaza chart does not create or repair any of it: it creates no cluster-scoped objects, installs no operators, and does not configure OpenShift or GreyMatter. If a prerequisite is missing, the install fails or the platform comes up unable to serve traffic.
Work through each section and run its check. Each section also says what to record; the install needs those facts.
Hardware sizing and supported operating systems are covered separately, in System Requirements.
Summary
| Prerequisite | What the cluster must provide |
|---|---|
| Namespace | A pre-created project |
| Installer identity | A namespace-scoped identity to run Helm as. Never cluster-admin |
| Security context constraints | OpenShift's default restricted-v2 for Kamiwaza's service accounts, and no custom SCC |
| GreyMatter | The namespace enrolled in GreyMatter, and a Git repository for Kamiwaza's GreyMatter configuration |
| Edge certificate | A TLS certificate for the domain, issued by your GreyMatter mesh CA and valid for server and client authentication |
| DNS | The domain resolves to the OpenShift router, from clients and from inside the cluster |
| Block storage | An RWO StorageClass that provisions durable volumes |
| Object storage | A decision: the built-in object store, or your own S3-compatible service |
| Kubernetes API addresses | The API server's Service and endpoint addresses, recorded for network policy |
| Registry access | Nodes can pull Kamiwaza images and the GreyMatter proxy image |
| Node kernel limits | fs.inotify limits at or above the floor on every node |
| Outbound network access | HTTPS to the image registry, your Git server, and any model sources you use |
GPU inference and multi-node clusters add requirements; see GPU nodes and Multi-node clusters.
Kamiwaza installs into a namespace named kamiwaza; use that name. The examples
use kamiwaza.example.com as the domain; substitute your own. Run the checks with
the administrator's kubeconfig unless a check says otherwise. oc works everywhere
this page uses kubectl.
Namespace
Create the kamiwaza namespace. The installer does not create it, and the
install never uses --create-namespace.
Do not label it for Istio injection: GreyMatter injects its own sidecars once the namespace is enrolled (see GreyMatter). OpenShift applies pod security through security context constraints, covered below.
oc new-project kamiwaza
oc label namespace kamiwaza app.kubernetes.io/part-of=kamiwaza
oc new-project also switches your kubeconfig's current namespace to
kamiwaza. Switch back with oc project <previous namespace> if you need to.
Check:
oc get namespace kamiwaza --show-labels
# Expect app.kubernetes.io/part-of=kamiwaza, and no istio-injection label
Installer identity
Helm runs as a namespace-scoped identity, never as cluster-admin. It needs:
- the built-in
adminClusterRole, bound in the namespace only with a RoleBinding; - rights to manage
endpoints, patchpods/status, and read GreyMatter'ssidecarresources. The chart grants these to its own service accounts, and Kubernetes only lets an identity grant permissions it holds itself. Without thesidecarresourcesread, the install stops withattempting to grant RBAC permissions not currently held.
This manifest creates a service account with those rights. You can bind the same Roles to a user or group from your identity provider instead.
apiVersion: v1
kind: ServiceAccount
metadata:
name: kamiwaza-installer
namespace: kamiwaza
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kamiwaza-installer-admin
namespace: kamiwaza
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: admin
subjects:
- kind: ServiceAccount
name: kamiwaza-installer
namespace: kamiwaza
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: kamiwaza-installer-delegation
namespace: kamiwaza
rules:
- apiGroups: [""]
resources: ["endpoints"]
verbs: ["create", "delete", "patch", "update"]
- apiGroups: [""]
resources: ["pods/status"]
verbs: ["patch"]
- apiGroups: ["tenant.greymatter.io"]
resources: ["sidecarresources"]
verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kamiwaza-installer-delegation
namespace: kamiwaza
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: kamiwaza-installer-delegation
subjects:
- kind: ServiceAccount
name: kamiwaza-installer
namespace: kamiwaza
Save it as kamiwaza-installer.yaml and apply it as the administrator:
kubectl apply -f kamiwaza-installer.yaml
Give the installer a kubeconfig for this identity, following your cluster's normal practice for issuing credentials. For example, with a service account token, run as the administrator:
(
set -e
SERVER=$(kubectl config view --minify -o jsonpath='{.clusters[0].cluster.server}')
kubectl config view --minify --raw \
-o jsonpath='{.clusters[0].cluster.certificate-authority-data}' | base64 -d > cluster-ca.crt
TOKEN=$(kubectl -n kamiwaza create token kamiwaza-installer --duration=24h)
export KUBECONFIG=kamiwaza-installer.kubeconfig
kubectl config set-cluster kamiwaza --server="$SERVER" \
--certificate-authority=cluster-ca.crt --embed-certs
kubectl config set-credentials kamiwaza-installer --token="$TOKEN"
kubectl config set-context kamiwaza --cluster=kamiwaza --user=kamiwaza-installer --namespace=kamiwaza
kubectl config use-context kamiwaza
)
The parentheses keep these commands from changing your own shell's KUBECONFIG.
If your kubeconfig names a certificate-authority file instead of embedding
certificate-authority-data, copy that file to cluster-ca.crt instead. The
token expires, and the API server may cap the duration you ask for. Issue a new
one before each later upgrade or uninstall.
Check that the identity can create namespaced resources and cannot create cluster-scoped ones:
SA=system:serviceaccount:kamiwaza:kamiwaza-installer
kubectl auth can-i create deployments -n kamiwaza --as=$SA # yes
kubectl auth can-i patch pods/status -n kamiwaza --as=$SA # yes
kubectl auth can-i get sidecarresources.tenant.greymatter.io -n kamiwaza --as=$SA # yes
kubectl auth can-i create clusterroles --as=$SA # no
kubectl auth can-i create customresourcedefinitions --as=$SA # no
kubectl auth can-i list nodes --as=$SA # no
kubectl auth can-i create securitycontextconstraints --as=$SA # no
When your kubeconfig sets a namespace, the cluster-scoped checks also print
Warning: resource '...' is not namespace scoped before their answer. The
yes or no is the result. The security context constraint check below
passes -n itself, so it always prints that warning.
Security context constraints
Kamiwaza runs under OpenShift's default restricted-v2 security context
constraint (SCC), which assigns each pod a user ID from the namespace's range.
The chart creates no SCCs and needs none.
Don't create or bind a custom SCC for Kamiwaza, and don't grant anyuid,
nonroot, or a similar SCC to its service accounts. Kamiwaza's pods rely on
restricted-v2 assigning their user ID; under an SCC that doesn't, they fail
to start with errors such as container has runAsNonRoot and image has non-numeric user or image will run as root.
Check that no permissive SCC is granted to the whole namespace or to all service accounts. It works before the install creates Kamiwaza's service accounts:
for scc in anyuid nonroot nonroot-v2 privileged; do
printf '%s: ' "$scc"
oc auth can-i use "securitycontextconstraints/$scc" -n kamiwaza \
--as=system:serviceaccount:kamiwaza:core-scheduler \
--as-group=system:serviceaccounts --as-group=system:serviceaccounts:kamiwaza
done
# Expect "no" for each
A grant to a single service account shows up in the install's
verification, which checks that every Kamiwaza pod
runs under restricted-v2.
GreyMatter
Kamiwaza routes all traffic through a GreyMatter mesh that the platform owns. GreyMatter Core 2.4 must already be installed and healthy; Kamiwaza does not install or configure it. The steps below change GreyMatter configuration, so they need your GreyMatter credentials.
GreyMatter configures Kamiwaza's routes from configuration files (GSL) in a Git repository. Kamiwaza supplies those files: you publish them at install time, and Kamiwaza adds routes to them at runtime as apps and models are deployed. The browser-facing entry point is a GreyMatter edge in the Kamiwaza namespace, published through an OpenShift Route with TLS passthrough.
Enroll the namespace
In your GreyMatter Core configuration, add the namespace to
tenant_namespaces in config.cue. If your configuration also sets
ftr_namespaces, add it there too. Leave ftr_namespaces unset if it is: when
it is absent, forced traffic redirection already covers every tenant namespace,
and setting it limits redirection to the namespaces it lists.
config: {
tenant_namespaces: [
"kamiwaza",
]
}
The namespace must exist before GreyMatter receives this change. Apply it through your normal GreyMatter change process.
Check that GreyMatter has acted on the namespace. It starts its tenant controllers there and copies its image pull Secret in:
oc -n kamiwaza get deployments
# Expect the GreyMatter tenant controllers, for example greymatter-tcm
oc -n kamiwaza get secret greymatter-image-pull -o jsonpath='{.type}{"\n"}'
# Expect kubernetes.io/dockerconfigjson
The tenant controllers stay in ContainerCreating until the
edge certificate Secret exists: they mount it. Check that
they are ready after that section.
Record:
-
the GreyMatter image pull Secret's name, if it is not
greymatter-image-pull; -
the GreyMatter proxy image your platform runs. Kamiwaza runs the same proxy beside its Ray workloads:
oc get deployments -A \-o jsonpath='{range .items[*].spec.template.spec.containers[*]}{.image}{"\n"}{end}' \| grep -i greymatter-proxy | sort -u
Tenant configuration repository
Provide a Git repository for Kamiwaza's GreyMatter configuration, with a branch
that GreyMatter Sync watches. Kamiwaza pushes route changes to that branch while
it runs, so the branch must accept direct pushes, not only pull requests.
Kamiwaza's files live under _tenants/kamiwaza in the repository.
Create the GreyMatter project there with the greymatter CLI, and push it. Use
the CLI that matches your GreyMatter platform: the .greymatter file at the root
of your GreyMatter Core configuration records the platform and CLI versions.
GreyMatter distributes the CLI to its customers; if the exact CLI build isn't
offered, use the newest patch release of the same minor version.
The CLI writes the project into the current directory, so create the project directory first:
git clone https://git.example.com/org/kamiwaza-gsl.git && cd kamiwaza-gsl
git switch -c tenant/kamiwaza
mkdir -p _tenants/kamiwaza && cd _tenants/kamiwaza
greymatter create project --security pki --openshift kamiwaza
cd ../..
git add _tenants/kamiwaza
git commit -m "Create the Kamiwaza GreyMatter project"
git push -u origin tenant/kamiwaza
The CLI writes .greymatter and cue.mod/, including the GSL packages that
match your GreyMatter version, and a starter greymatter/ directory. Commit
them unchanged. Kamiwaza's own files are added to greymatter/ at install
time, replacing the starter files.
Then create the Secret that connects GreyMatter Sync and Kamiwaza to the repository. It holds the repository's location and a credential that can push to the branch. Read the credential without leaving it in your shell history:
printf 'Git username: '; read -r GIT_USER
printf 'Git token: '; read -rs GIT_TOKEN; echo
(umask 077; d=$(mktemp -d)
printf '%s' "$GIT_TOKEN" > "$d/token"
oc -n kamiwaza create secret generic greymatter-tenant-repo \
--type=greymatter.io/repo \
--from-literal=auth_type=https \
--from-literal=url=https://git.example.com/org/kamiwaza-gsl.git \
--from-literal=branch=tenant/kamiwaza \
--from-literal=relative_path=_tenants/kamiwaza \
--from-literal=http_username="$GIT_USER" \
--from-file=http_password="$d/token" \
--from-literal=tls_insecure_verify=false
rm -rf "$d")
unset GIT_USER GIT_TOKEN
If your Git server uses a private CA, add
--from-file=tls_remote_ca=<ca-bundle>. Keep tls_insecure_verify=false.
Check:
oc -n kamiwaza get secret greymatter-tenant-repo -o json \
| jq -r '.type, (.data | keys | join(" "))'
# Expect greymatter.io/repo, then the keys
# auth_type branch http_password http_username relative_path tls_insecure_verify url
# (and tls_remote_ca, if you added it)
Record: the repository URL and branch. Whoever publishes Kamiwaza's files at install time needs push access to them.
Network policy
If the kamiwaza namespace denies traffic by default, allow:
- DNS to OpenShift DNS;
- the Kubernetes API, at the addresses you record below;
- HTTPS to your Git server, from the GreyMatter tenant controllers and from
Kamiwaza's
core-schedulerandcore-rayclusterpods; - GreyMatter control-plane traffic to and from the namespace;
- the OpenShift router to the GreyMatter edge;
- HTTPS to the OpenShift router, for calls to the Kamiwaza domain from inside the cluster (see DNS);
- HTTPS to the hosts in Outbound network access;
- traffic within the namespace.
Scope each egress rule to the addresses it needs, not to every address.
Edge certificate
The Kamiwaza domain is served with a TLS certificate that you provide and rotate. The certificate must cover the Kamiwaza domain in its subject alternative names.
The GreyMatter edge serves it; the OpenShift router passes the connection through without decrypting it. The same certificate is also the mesh identity in the Kamiwaza namespace: GreyMatter's tenant controllers and Kamiwaza's Ray sidecar present it when they connect to other mesh services, including GreyMatter Core. So it must:
- be issued by the CA that your GreyMatter Core trusts for mesh (PKI) traffic.
That CA is the
ca.crtof Core's PKI certificate Secret, by defaultgreymatter-edge-ingressin thegreymatternamespace. A certificate from any other CA, or a self-signed one, leaves the tenant controllers unable to connect to GreyMatter Core, and GreyMatter never applies Kamiwaza's configuration; - be valid for both server and client authentication (extended key usages
serverAuthandclientAuth).
Publicly trusted CAs don't issue certificates with client authentication, and browsers don't trust your mesh CA, so users see a certificate warning unless their machines trust that CA.
Create it in the Kamiwaza namespace, named greymatter-edge-ingress, with the
server certificate and any intermediates in tls.crt and your mesh CA in
ca.crt:
oc -n kamiwaza create secret generic greymatter-edge-ingress \
--from-file=ca.crt=./edge-ca.crt \
--from-file=tls.crt=./edge.crt \
--from-file=tls.key=./edge.key
Kamiwaza workloads need the mesh CA to verify the edge, since it is not publicly trusted. Publish the CA certificate in the Kamiwaza namespace:
kubectl -n kamiwaza create configmap kamiwaza-edge-trust \
--from-file=ca-certificates.crt=./edge-ca.crt
This ConfigMap holds public certificates only. Never put a private key in it. Kamiwaza reads it by name when it exists; no install value refers to it.
Check:
oc -n kamiwaza get secret greymatter-edge-ingress -o jsonpath='{.data.tls\.crt}' \
| base64 -d | openssl x509 -noout -subject -issuer -ext subjectAltName -enddate -purpose \
| grep -E 'subject|issuer|DNS:|notAfter|^SSL (client|server) :'
# Expect your domain under subjectAltName, an end date in the future,
# "SSL client : Yes" and "SSL server : Yes"
oc -n kamiwaza get secret greymatter-edge-ingress -o json | jq -r '.data | keys | join(" ")'
# Expect ca.crt tls.crt tls.key
# The certificate chains to the CA GreyMatter Core trusts
# (use your Core PKI Secret if it isn't the default):
oc -n greymatter get secret greymatter-edge-ingress -o jsonpath='{.data.ca\.crt}' \
| base64 -d > greymatter-mesh-ca.crt
oc -n kamiwaza get secret greymatter-edge-ingress -o jsonpath='{.data.tls\.crt}' \
| base64 -d > kamiwaza-edge.crt
openssl verify -CAfile greymatter-mesh-ca.crt -untrusted kamiwaza-edge.crt kamiwaza-edge.crt
# Expect kamiwaza-edge.crt: OK
Once this Secret exists, the GreyMatter tenant controllers from Enroll the namespace start. Check that they are ready:
oc -n kamiwaza get deployments
# Expect greymatter-tcm and the other tenant controllers ready
If they were already running with a different certificate, restart them so
they load this one: oc -n kamiwaza rollout restart deployment <name> for each.
DNS
The Kamiwaza domain must resolve to the OpenShift router's external address, the same way as other passthrough Routes on the cluster:
-
For users: your DNS points the domain at the router's load balancer. Find its address with:
oc -n openshift-ingress get service router-default \-o jsonpath='{.status.loadBalancer.ingress[0].hostname}{.status.loadBalancer.ingress[0].ip}{"\n"}'A host name (as on AWS) needs a
CNAMErecord; an IP address needs anArecord. The router passes TLS through to the GreyMatter edge, so don't put a proxy or CDN that terminates TLS in front of the domain: publish the record as plain DNS. -
Inside the cluster: Kamiwaza services and apps call the platform at its public domain, so pods must resolve it and reach the router the same way. OpenShift's cluster DNS forwards to your upstream resolvers, so this usually needs nothing more.
The router routes on the TLS server name, so clients must connect using the
Kamiwaza domain. The Route itself is created when Kamiwaza's GreyMatter
configuration is published during the install; until then the router answers
for the domain with its own 503 page.
Check from your workstation, and from a pod:
curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/
# Any HTTP status proves DNS reaches the router. Before Kamiwaza is
# installed, expect 503.
(
export KUBECONFIG=kamiwaza-installer.kubeconfig
oc -n kamiwaza run kamiwaza-dns-check --restart=Never --image=curlimages/curl \
-- -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/
oc -n kamiwaza wait --for=jsonpath='{.status.phase}'=Succeeded \
pod/kamiwaza-dns-check --timeout=60s
oc -n kamiwaza logs kamiwaza-dns-check
# Expect the same status from inside the cluster
oc -n kamiwaza delete pod kamiwaza-dns-check
)
The pod commands run with the installer kubeconfig for the same reason as the block storage write test: as cluster-admin, OpenShift can admit the pod under a security context constraint that assigns it no user ID.
Record: the domain.
Block storage
Kamiwaza's databases and caches use ReadWriteOnce persistent volumes. The
cluster must have a StorageClass that provisions them. Record its name; the install
names it explicitly.
Size the storage provider for Kamiwaza's volume requests plus your provider's
replication and free-space reserves. A StorageClass existing does not prove its
volumes can be provisioned; the write test below does. Node-local storage, such
as local-path, works on a single node but does not replicate data.
Kamiwaza requests these volumes by default, about 624 GiB in total. Apps you deploy later add their own.
| Volume | Default request | Holds | Change it with |
|---|---|---|---|
kamiwaza-registry-data | 500 GiB | Model and inference-engine cache | registry.persistence.size |
seaweedfs-data | 100 GiB | Built-in object store | seaweedfs.persistence.size |
core-postgres-data | 10 GiB | Platform database | — |
keycloak-postgres-data | 10 GiB | Identity database | — |
core-spicedb-postgres-data | 1 GiB | Authorization database | — |
core-etcd-data-core-etcd-0 to -2 | 1 GiB each | Configuration store | — |
The model registry cache holds inference-engine images and the model files your
users deploy. If you serve only external models (for example, through a cloud
provider), or your storage is small, set registry.persistence.size lower, for
example 100Gi, in your install values before the first install. An existing
volume keeps its size on upgrade.
Check the classes, then run a write test in the Kamiwaza namespace: write a marker, replace the pod, and read the marker back. Run the test with the installer kubeconfig from Installer identity, as below. As cluster-admin, OpenShift can admit the test pod under a security context constraint that assigns it no user ID, and the pod then fails to start.
kubectl get storageclass
# Expect one class marked (default), or note the class you will name at install
(
export KUBECONFIG=kamiwaza-installer.kubeconfig
cat > kamiwaza-storage-check.yaml <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: kamiwaza-storage-check
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
# storageClassName: <class> # uncomment if you will not use the default
---
apiVersion: v1
kind: Pod
metadata:
name: kamiwaza-storage-check
spec:
securityContext:
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
containers:
- name: check
image: busybox:1.37
command: ["sh", "-c", "sleep 3600"]
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
volumeMounts:
- name: data
mountPath: /data
volumes:
- name: data
persistentVolumeClaim:
claimName: kamiwaza-storage-check
EOF
kubectl -n kamiwaza apply -f kamiwaza-storage-check.yaml
kubectl -n kamiwaza wait --for=condition=Ready pod/kamiwaza-storage-check --timeout=5m
kubectl -n kamiwaza exec kamiwaza-storage-check -- sh -c 'echo ok > /data/marker'
# Replace the pod; the claim and its volume stay
kubectl -n kamiwaza delete pod kamiwaza-storage-check
kubectl -n kamiwaza apply -f kamiwaza-storage-check.yaml
kubectl -n kamiwaza wait --for=condition=Ready pod/kamiwaza-storage-check --timeout=5m
kubectl -n kamiwaza exec kamiwaza-storage-check -- cat /data/marker # expect: ok
kubectl -n kamiwaza delete -f kamiwaza-storage-check.yaml
)
A class with volumeBindingMode: WaitForFirstConsumer binds the claim only once
the pod is scheduled, so a Pending claim before the pod starts is normal.
Record: the StorageClass name.
Object storage
Workrooms, the Skills Library, and other features store objects. Choose one:
- The built-in object store (default). Kamiwaza runs SeaweedFS in its
namespace, backed by a
ReadWriteOncevolume from your block storage. Nothing to prepare. - Your own S3-compatible service. See S3 workroom storage.
Kubernetes API addresses
Some Kamiwaza apps call the Kubernetes API from inside the namespace, and
Kamiwaza allows that traffic with NetworkPolicy. Many CNIs, including the ones
k0s and OpenShift use, evaluate NetworkPolicy after a Service address is
translated to the API server's real address. So the policy needs both the
kubernetes Service address and every API server endpoint address. With only
the first, apps install cleanly and then fail at runtime.
While OpenShift rolls out the API server or restarts a control-plane node, for example on a new cluster or after a configuration change, the endpoint list leaves out the server being restarted. Record the addresses only when no rollout is in progress:
oc get clusteroperator kube-apiserver
# Expect PROGRESSING False. If the cluster has no kube-apiserver cluster
# operator, as can happen when the provider hosts the control plane, skip this.
oc get machineconfigpool master
# Expect UPDATED True and UPDATING False. Skip this too if the provider hosts
# the control plane.
Record both sets of addresses, as /32 for IPv4 or /128 for IPv6:
kubectl -n default get service kubernetes -o jsonpath='{.spec.clusterIPs[*]}{"\n"}'
kubectl -n default get endpointslices -l kubernetes.io/service-name=kubernetes \
-o jsonpath='{.items[*].endpoints[*].addresses[*]}{"\n"}'
# Expect one endpoint address for each API server: three on a cluster with
# three control-plane nodes
Copy the addresses exactly. A wrong address still installs cleanly, and the apps then fail at runtime. If the API server's addresses change (for example, a control-plane node is replaced), update them in the install values and upgrade.
If your CNI enforces NetworkPolicy before Service address translation, the endpoint addresses are not needed; the Service address alone works. Record it anyway.
Registry access
Kamiwaza's charts and images are published to a public registry and pull without credentials. Nodes must be able to pull from it; see Outbound network access.
If you mirror images into your own registry and it requires credentials, create
a kubernetes.io/dockerconfigjson pull Secret in the Kamiwaza namespace.
Kamiwaza uses only namespaced pull Secrets, never node-wide registry
credentials. Build the Secret from a file, so the password never appears in a
command line or your shell history:
read -rs REGISTRY_PASSWORD # typed, not echoed
(umask 077; d=$(mktemp -d)
printf '{"auths":{"%s":{"auth":"%s"}}}' '<registry-host>' \
"$(printf '%s:%s' '<username>' "$REGISTRY_PASSWORD" | base64 | tr -d '\n')" > "$d/config.json"
kubectl -n kamiwaza create secret generic kamiwaza-pull \
--type=kubernetes.io/dockerconfigjson --from-file=.dockerconfigjson="$d/config.json"
rm -rf "$d")
unset REGISTRY_PASSWORD
<registry-host> is the registry's host name, with its port if it has one, as it
appears in the image references.
Record: the pull Secret's name, kamiwaza-pull above. The install lists it in
global.imagePullSecrets.
Nodes must also pull the GreyMatter proxy image. GreyMatter provides its pull Secret; see Enroll the namespace.
Node kernel limits
Every GreyMatter sidecar watches files with inotify. Node defaults can be too
low for the number of pods Kamiwaza runs, and sidecars then crash with
too many open files. Set these floors on every worker node.
On OpenShift, set them with a MachineConfig for each machine pool that runs
Kamiwaza. Whether a custom MachineConfig is supported depends on your provider's
support policy. If your platform doesn't let you create one (for example, ROSA
with hosted control planes), set the same limits through your provider's node
tuning mechanism, or ask your provider's support to set them.
Red Hat Enterprise Linux CoreOS already sets fs.inotify.max_user_watches to
65536 in /etc/sysctl.d/inotify.conf. Files in /etc/sysctl.d apply in name
order and the last one wins, so Kamiwaza's file must sort after that one; the
zz- prefix below does that.
This MachineConfig covers the worker pool, which on ROSA also includes the
infra nodes. Save it as kamiwaza-inotify-machineconfig.yaml:
apiVersion: machineconfiguration.openshift.io/v1
kind: MachineConfig
metadata:
name: 99-worker-kamiwaza-inotify
labels:
machineconfiguration.openshift.io/role: worker
spec:
config:
ignition:
version: 3.2.0
storage:
files:
- path: /etc/sysctl.d/zz-kamiwaza-inotify.conf
mode: 0644
overwrite: true
contents:
source: data:,fs.inotify.max_user_instances%20%3D%208192%0Afs.inotify.max_user_watches%20%3D%201048576%0Afs.inotify.max_queued_events%20%3D%2016384%0A
oc apply -f kamiwaza-inotify-machineconfig.yaml
oc get machineconfigpool worker -w # wait until UPDATED is True
Applying a MachineConfig restarts the pool's nodes one at a time, a few
minutes each. Schedule it with your platform's change process. If Kamiwaza is
already installed, expect a brief interruption while the node running its edge
restarts.
Check on each worker node:
oc debug node/<node> -- chroot /host \
sysctl fs.inotify.max_user_instances fs.inotify.max_user_watches fs.inotify.max_queued_events
# Expect values at or above 8192, 1048576, and 16384
Outbound network access
Nodes and pods need outbound HTTPS (port 443) to:
ghcr.ioandpkg-containers.githubusercontent.com, for Kamiwaza charts and images;- the model sources you use, for example
huggingface.coand its content hosts, or your cloud model provider's endpoint; - your Git server, for the tenant configuration repository;
- the registry that serves the GreyMatter proxy image.
The checks on this page use the public busybox and curlimages/curl images from
Docker Hub. If your cluster cannot reach Docker Hub, mirror them to your registry.
For installs with no outbound access, see Install Without Internet Access.
GPU nodes
GPU inference is optional. To serve models on NVIDIA GPUs, prepare each GPU node
as described in GPU Nodes: a driver that meets the engine
images' floor, persistence mode, the NVIDIA container runtime and the nvidia
RuntimeClass, the NVIDIA GPU Operator, and no NoSchedule or NoExecute taint
on the node. Kamiwaza installs none of these.
The node commands on that page are for Ubuntu nodes. On OpenShift, install the
NVIDIA GPU Operator with your platform's usual method, then run the checks on
that page that apply to your nodes.
Prepared nodes are not enough on their own: a default install cannot use GPUs. After the install, the cluster owner turns GPU serving on with Tenant-mode inference.
Check:
kubectl get nodes -o custom-columns='NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu'
# Expect a count on each GPU node
kubectl get runtimeclass nvidia
# Expect a RuntimeClass named nvidia, with handler nvidia
Record: which nodes have GPUs, and the GPU model on each. The cluster owner needs them to publish a signed profile.
Multi-node clusters
On a cluster with more than one node, also provide:
- Storage that survives a node loss. Use a StorageClass whose volumes are replicated or network-attached, such as a CSI driver for your SAN or cloud block storage, or a replicated provider like Longhorn. Node-local storage ties each database to one node: if that node is lost, so is the data.
- Every GPU node prepared. Prepare each GPU node as in GPU nodes, including nodes added later.
- A CNI that enforces NetworkPolicy. Kamiwaza isolates apps and jobs from each other with NetworkPolicy. A CNI that accepts policies but does not enforce them leaves that isolation off without any error.
- The inotify floor on every node, including nodes added later.
- The API endpoint addresses of every control-plane node in the recorded Kubernetes API addresses.
Next step
When every check passes, hand the recorded facts and the installer kubeconfig to whoever runs the install, and continue to Install on OpenShift with GreyMatter.