Install on OpenShift with GreyMatter
This page installs Kamiwaza into one namespace of an OpenShift cluster whose
service mesh is GreyMatter. You publish Kamiwaza's GreyMatter configuration,
then run a single helm upgrade --install as a namespace-scoped identity. The
chart creates only namespaced objects.
Before you start
| You need | Notes |
|---|---|
| A prepared cluster | Every check in Platform Prerequisites passes, and you have the facts it told you to record. |
| The installer kubeconfig | The namespace-scoped identity from Installer identity. Never install with a cluster-admin kubeconfig. |
oc or kubectl, and Helm 3.14 or later | Helm pulls the chart from an OCI registry. With Helm 4, add --server-side=false to every helm upgrade --install on this page; see step 4. |
jq, curl, and git | Used by the configuration and verification steps. |
| Push access to the tenant configuration repository | The repository and branch from Tenant configuration repository. If the platform administrator keeps that access, they run step 3. |
| A license file | license.lic, issued by Kamiwaza. See License File. |
No internet access? Do Install Without Internet Access first. It copies the release into your registry and gives you values to merge into your values file in step 2. In steps 3 and 4, use the chart from your registry in place of the release location.
Point kubectl and helm at the installer identity for every command on this
page:
export KUBECONFIG=<path to the installer kubeconfig>
Kamiwaza installs into the kamiwaza namespace. The examples use
kamiwaza.example.com as the domain; substitute your own.
Step 1: Add the license
Create a Secret from the license file. The key must be license.lic:
kubectl -n kamiwaza create secret generic kamiwaza-license \
--from-file=license.lic=./license.lic --dry-run=client -o yaml \
| kubectl apply -f -
This also works on a reinstall: uninstalling keeps this Secret, and the command updates it in place.
The chart refuses to render until a license is configured and the EULA is accepted. You do both in the values file in the next step.
Step 2: Write your values file
Your values file selects the OpenShift with GreyMatter platform and carries this cluster's facts and your acceptance of the EULA. Everything else is the chart's defaults for that platform. Keep the file somewhere safe and outside any shared repository; you will need it again for upgrades.
Required values
| Value | Set it to |
|---|---|
global.platform | osgm. This switches the chart to GreyMatter routing and OpenShift's security model. |
global.eula.accepted | true, after reading the EULA. The render fails without it. |
core.license.existingSecret | kamiwaza-license, the Secret from step 1. The render fails without it. |
global.domain | The Kamiwaza domain. |
global.namespaces.platform, .extensions, .sandboxes, .observability, .ca, .registry, .system | kamiwaza, for each one. Everything runs in that one namespace. |
global.storage.stateful.standard, .postgres, .etcd, .registry, .extensions | The StorageClass you recorded, for each one. |
core.extensionRuntime.kubernetesApi.egressCidrs | Every Kubernetes API Service address and every endpoint address you recorded, each as /32 (IPv4) or /128 (IPv6). Copy them exactly: a wrong address installs cleanly and then fails at runtime. The render fails without this value. |
global.greymatter.proxy.image | The GreyMatter proxy image you recorded, split into registry, repository, and tag. It must match the proxy version your GreyMatter platform runs. |
Values that depend on your cluster
| Value | Set it when |
|---|---|
global.greymatter.proxy.imagePullSecret | GreyMatter's image pull Secret in the namespace is not named greymatter-image-pull. |
global.imagePullSecrets | Kamiwaza images come from a registry that needs credentials. List the pull Secret you created, for example [{name: kamiwaza-pull}]. |
core.inferenceResources.enabled, .bundleConfigMap, .ownerPublicKey, .clusterBindingId | You serve models on GPUs. Set them after the cluster owner publishes a signed profile, as in Tenant-mode inference. Without them, a GPU deployment is refused. Prepare the nodes first; see GPU nodes. |
Optional values
| Value | Notes |
|---|---|
registry.persistence.size, seaweedfs.persistence.size | Volume sizes for the model registry cache (default 500Gi) and the built-in object store (default 100Gi). Lower them when storage is small; see Block storage. Set them before the first install; they can't be reduced later. |
core.userDefaults.adminPassword | The password for the admin user. Leave it unset and Kamiwaza generates one, which you read back in step 5. Setting it here stores it in the Helm release. It applies at first install only; changing it later does not rotate the password. |
Example
global:
platform: osgm
domain: kamiwaza.example.com
eula:
accepted: true
namespaces:
platform: kamiwaza
extensions: kamiwaza
sandboxes: kamiwaza
observability: kamiwaza
ca: kamiwaza
registry: kamiwaza
system: kamiwaza
storage:
stateful:
standard: gp3-csi
postgres: gp3-csi
etcd: gp3-csi
registry: gp3-csi
extensions: gp3-csi
greymatter:
proxy:
image: # the proxy image you recorded
registry: oci.download.greymatter.io
repository: greymatter-proxy
tag: 1.17.2
core:
license:
existingSecret: kamiwaza-license
extensionRuntime:
kubernetesApi:
egressCidrs:
- 172.30.0.1/32 # kubectl -n default get service kubernetes
- 10.0.1.10/32 # the kubernetes EndpointSlice addresses
- 10.0.1.11/32
- 10.0.1.12/32
Save it as kamiwaza-values.yaml.
Step 3: Publish the GreyMatter configuration
GreyMatter needs Kamiwaza's configuration before the install. It defines the edge that serves the Kamiwaza domain, the routes to Kamiwaza's services, and the sidecars GreyMatter adds to them. The chart generates this configuration from your values file, so it always matches the release you install. You render it, commit it to the tenant configuration repository, and GreyMatter Sync applies it.
-
Clone the tenant configuration repository on its branch, and go to the project the platform administrator created:
git clone -b tenant/kamiwaza https://git.example.com/org/kamiwaza-gsl.git kamiwaza-gslcd kamiwaza-gsl/_tenants/kamiwazaIf you already have a clone, run
git switch tenant/kamiwazaandgit pullin it instead. -
Render the configuration and write its files into the project. This does not change the cluster.
helm template kamiwaza \oci://ghcr.io/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza \--version 1.3.3 \--namespace kamiwaza \--values <absolute path to kamiwaza-values.yaml> \--show-only charts/network/templates/greymatter/gsl-bundle.yaml \| kubectl create --dry-run=client --validate=false -f - -o json > gsl-bundle.jsonjq -r '.data["file-list.txt"]' gsl-bundle.json | grep '^greymatter/' \| while read -r f; docase "$f" in greymatter/generated/*) [ -e "$f" ] && continue ;; esacmkdir -p "$(dirname "$f")"jq -r --arg k "${f//\//__}" '.data[$k]' gsl-bundle.json > "$f"donerm gsl-bundle.jsonThis writes the
greymatter/directory. It leaves the.greymatterfile andcue.mod/directory that the GreyMatter CLI created unchanged. It creates the files undergreymatter/generated/only if they do not exist: Kamiwaza records its app and model routes there while it runs, so an upgrade must keep them. -
Check that the configuration names your domain, then commit and push:
grep -rh 'api_endpoint' greymatter | sort -u# Expect https://kamiwaza.example.com, and paths under itgit add greymattergit diff --cached --quiet || git commit -m "Kamiwaza 1.3.3 GreyMatter configuration"git pull --rebasegit pushOn a reinstall or an upgrade that changes nothing here, there is nothing to commit, and the push is a no-op.
Check that GreyMatter applied it and the edge serves the domain:
kubectl -n kamiwaza get route \
-o custom-columns='NAME:.metadata.name,HOST:.spec.host,TLS:.spec.tls.termination'
# Expect a Route for kamiwaza.example.com with TLS "passthrough"
kubectl -n kamiwaza get pods -l app.kubernetes.io/name=edge
# Expect the edge pod Running and ready
curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/
# Any HTTP status from the edge proves the Route and edge certificate work.
# Before Kamiwaza is installed, expect 404 or 503.
GreyMatter can take a few minutes to sync. A connection that fails, or ends during the TLS handshake, is not a sync delay: see Troubleshooting.
Step 4: Install
Record the start time first. The verification step uses it to ignore events from before this install.
INSTALL_START=$(date -u +%Y-%m-%dT%H:%M:%SZ)
helm upgrade --install kamiwaza \
oci://ghcr.io/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza \
--version 1.3.3 \
--namespace kamiwaza \
--values <absolute path to kamiwaza-values.yaml> \
--wait --wait-for-jobs --timeout 30m
- Do not add
--create-namespace. The namespace belongs to the platform administrator. - With Helm 4, add
--server-side=false, on the first install and every upgrade. OpenShift adds its own pull Secret to each of Kamiwaza's service accounts, and Helm 4's default server-side apply then refuses every later upgrade withconflict with "openshift.io/image-registry-pull-secrets_service-account-controller". Helm 3 is not affected. - Helm prints
WARNING: core cannot see GPUs on this install.until tenant-mode inference is on. It is expected; see Tenant-mode inference. - A typical install takes 5 to 15 minutes, most of it pulling images.
Install and upgrade with Helm, not rendered manifests
Always install and upgrade with helm upgrade --install against the cluster.
Rendering the chart with helm template or --dry-run and applying the output,
for example through Argo CD or Flux, is not supported. On first install the chart
generates the platform's encryption keys, signing keys, and database passwords,
and on each upgrade it reads them back from the cluster to keep them. A rendered
chart cannot read the cluster, so every render generates new ones. Applying them
makes existing encrypted data and access tokens unreadable.
Step 3 renders only the GreyMatter configuration files, which contain no generated secrets, and applies nothing to the cluster.
To upgrade to a later release in the same line, repeat step 3 with the new
version, then rerun the command in this step with the new --version and the
same values file.
Step 5: Verify
Check what the platform does, not only its status fields.
Pods are running and ready:
kubectl -n kamiwaza get pods --no-headers \
| awk '$3 != "Completed" { split($2, r, "/"); if ($3 != "Running" || r[1] != r[2]) print }'
# Expect no output
An Error pod whose Job later completed is a retried attempt, not a failure.
Confirm with kubectl -n kamiwaza get jobs: every Job should show all completions.
GreyMatter sidecars are in place. The chart adds one to the core-raycluster
pods; GreyMatter adds one to frontend and keycloak. A pod without it cannot
receive traffic from the edge.
kubectl -n kamiwaza get pods \
-o custom-columns='POD:.metadata.name,INIT:.spec.initContainers[*].name' \
| grep -E '^(core-raycluster|frontend|keycloak)-' | grep -v '^keycloak-postgres'
# Expect greymatter-sidecar listed for each pod
Every pod runs under OpenShift's default security context constraint:
kubectl -n kamiwaza get pods \
-o jsonpath='{range .items[*]}{.metadata.name} {.metadata.annotations.openshift\.io/scc}{"\n"}{end}' \
| grep -v -E ' restricted-v2$' | grep -v -E '^greymatter-'
# Expect no output
GreyMatter's own controllers in the namespace (greymatter-*) use GreyMatter's
SCCs and are left out.
No image pulls failed during this install. Older events in the namespace are
ignored. Run this in the same shell as step 4, or set INSTALL_START to the time you
started the install, in the same format:
kubectl -n kamiwaza get events -o json | jq --arg start "$INSTALL_START" '
[.items[]
| select((.lastTimestamp // .eventTime // .metadata.creationTimestamp) >= $start)
| select((.message // "") | test("ErrImagePull|ImagePullBackOff|Failed to pull image|Back-off pulling image"))
] | length'
# Expect 0
The platform answers at its domain:
curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/ # expect 200
curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/api/ # expect 401
Log in. If you did not set an admin password, read the generated one:
kubectl -n kamiwaza get secret kamiwaza-user-admin \
-o jsonpath='{.data.password}' | base64 -d; echo
Open https://kamiwaza.example.com/ and log in as admin. With a certificate
your browser does not trust, accept the warning first.
admin is the Kamiwaza application administrator. It is not the installer
identity and not the Keycloak management account.
Call the API. Exchange the same credentials for a token, and use it as a bearer token:
TOKEN=$(kubectl -n kamiwaza get secret kamiwaza-user-admin -o jsonpath='{.data.password}' \
| base64 -d | curl -sk https://kamiwaza.example.com/api/auth/token \
--data-urlencode username=admin --data-urlencode password@- | jq -r .access_token)
curl -sk -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $TOKEN" \
https://kamiwaza.example.com/api/auth/users/me # expect 200
Step 6: Deploy the default embedding model
Kamiwaza ships a default embedding model, all-MiniLM-L6-v2, inside the release
and adds it to the model catalog during the install. On OpenShift with
GreyMatter it does not deploy it for you. Until you do, embedding requests fail
with HTTP 503 and the code embedding_unavailable, and features that embed
text, such as semantic search in documents, are unavailable. Deploying it
needs no internet access.
- Log in as
adminand open Models. - In the
all-MiniLM-L6-v2row, click Deploy. - Select the
embeddingconfiguration and click Deploy.
Deploy it once. The model's first pod can start before GreyMatter has added its sidecar to the model; Kamiwaza replaces that pod with one that has the sidecar, usually within a minute.
Check that the embedding answers, using the TOKEN from step 5 (rerun the
step 5 command that sets it if the token has expired):
curl -sk -X POST -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"text":"embedding check"}' https://kamiwaza.example.com/api/embedding/generate \
| jq 'if .embedding then (.embedding | length) else . end'
# Expect 384. Otherwise the response is printed: the code embedding_deploying
# means the model is still starting; retry after a minute.
kubectl -n kamiwaza get pods \
-o custom-columns='POD:.metadata.name,INIT:.spec.initContainers[*].name' | grep '^llamacpp-'
# Expect one pod, with greymatter-sidecar listed. Within the first minute you
# can also see the replaced pod terminating.
The pod check applies with tenant-mode inference off, which is the default.
With it on, model pods are named tmi-<id>; see
Tenant-mode inference.
The platform is installed. Deploying models and apps is covered in the Quickstart.
Troubleshooting
The install stops before creating anything. The chart checks its values
before rendering. A missing EULA acceptance or license prints
INSTALLATION HALTED; a missing egressCidrs value names the value it needs.
Nothing was installed. Fix the values file and rerun step 4.
no Kamiwaza license is configured, with the license set. The license value
belongs under core:. Check that it is core.license.existingSecret, not
license.existingSecret. To see values Helm received but no chart reads, run
helm get values kamiwaza -n kamiwaza.
The domain gives a TLS error, or no response, after step 3. Check that the
Route's TLS termination is passthrough, that the edge pod is ready, and that
greymatter-edge-ingress holds tls.crt and tls.key for the domain. An
admitted Route with a ready edge still fails the handshake when the edge cannot
read its certificate.
No Route or edge pod appears after step 3. GreyMatter has not applied the
configuration. Ask the platform administrator to check that the namespace is
enrolled, that the tenant controllers (for example greymatter-tcm) are ready,
and that greymatter-tenant-repo names the repository, branch, and
_tenants/kamiwaza path you pushed to. If greymatter-tcm is running but its
log repeats Failed to connect to a NATS cluster, the edge certificate isn't
issued by the mesh CA; see
Edge certificate.
Pods do not start: container has runAsNonRoot and image has non-numeric user, or image will run as root. The pod was admitted under an SCC other
than restricted-v2. See which with
kubectl -n kamiwaza get pod <pod> -o jsonpath='{.metadata.annotations.openshift\.io/scc}',
and ask the platform administrator to remove the extra SCC from Kamiwaza's
service accounts, as in
Security context constraints.
Then delete the pod so it is recreated.
Pods stay in CreateContainerConfigError, naming greymatter-tenant-repo.
The Secret is missing or lacks a key. core-scheduler and core-raycluster need
url, branch, relative_path, http_username, and http_password.
The core-raycluster pods cannot pull the GreyMatter proxy image. Check
global.greymatter.proxy.image against the image you recorded, and that the pull
Secret exists: kubectl -n kamiwaza get secret greymatter-image-pull.
The install hangs for about 15 minutes, then fails on a post-install job.
Check core-scheduler first: kubectl -n kamiwaza logs deploy/core-scheduler.
A license problem shows here as License check failed; see
License File. If frontend or keycloak are not
ready, GreyMatter has not given them sidecars; check step 3. While the job runs,
every helm upgrade fails with
another operation (install/upgrade/rollback) is in progress.
Pods are Pending on their volumes. A StorageClass named in the values does
not exist or cannot provision. Compare kubectl get storageclass with
global.storage.stateful.*.
Images fail to pull with unauthorized. The pull Secret named in
global.imagePullSecrets does not exist in the namespace, or has the wrong
credentials. Compare it with kubectl -n kamiwaza get secret.
Apps install but cannot reach the Kubernetes API. egressCidrs is missing
the endpoint addresses. Add every address from
Kubernetes API addresses and rerun step 4.
An app returns "Page Not Found" or the platform's own page, or a model's
endpoint returns no healthy upstream. Kamiwaza adds each app's and model's
route to the tenant configuration repository, and GreyMatter then applies it.
Wait a minute or two after its pods are ready, or after the model shows as
deployed. If it stays unreachable, check the core-scheduler logs for Git
errors: the branch must accept direct pushes, and the namespace must reach the
Git server.
Embedding requests return HTTP 503 with embedding_unavailable. No
embedding model is deployed. Deploy the default one as in
step 6.
A GPU model deployment fails with HTTP 422 and Core cannot see a GPU on this install. Tenant-mode inference is off, which is the default. Prepare the GPU
nodes, then turn it on; see
Tenant-mode inference.
Next steps
- Quickstart: start using Kamiwaza.
- License File: check and rotate the license.
- Uninstall.