Skip to main content
Version: 1.3.3 (Latest)

Install on OpenShift with GreyMatter

This page installs Kamiwaza into one namespace of an OpenShift cluster whose service mesh is GreyMatter. You publish Kamiwaza's GreyMatter configuration, then run a single helm upgrade --install as a namespace-scoped identity. The chart creates only namespaced objects.

Before you start​

You needNotes
A prepared clusterEvery check in Platform Prerequisites passes, and you have the facts it told you to record.
The installer kubeconfigThe namespace-scoped identity from Installer identity. Never install with a cluster-admin kubeconfig.
oc or kubectl, and Helm 3.14 or laterHelm pulls the chart from an OCI registry. With Helm 4, add --server-side=false to every helm upgrade --install on this page; see step 4.
jq, curl, and gitUsed by the configuration and verification steps.
Push access to the tenant configuration repositoryThe repository and branch from Tenant configuration repository. If the platform administrator keeps that access, they run step 3.
A license filelicense.lic, issued by Kamiwaza. See License File.

No internet access? Do Install Without Internet Access first. It copies the release into your registry and gives you values to merge into your values file in step 2. In steps 3 and 4, use the chart from your registry in place of the release location.

Point kubectl and helm at the installer identity for every command on this page:

export KUBECONFIG=<path to the installer kubeconfig>

Kamiwaza installs into the kamiwaza namespace. The examples use kamiwaza.example.com as the domain; substitute your own.

Step 1: Add the license​

Create a Secret from the license file. The key must be license.lic:

kubectl -n kamiwaza create secret generic kamiwaza-license \
--from-file=license.lic=./license.lic --dry-run=client -o yaml \
| kubectl apply -f -

This also works on a reinstall: uninstalling keeps this Secret, and the command updates it in place.

The chart refuses to render until a license is configured and the EULA is accepted. You do both in the values file in the next step.

Step 2: Write your values file​

Your values file selects the OpenShift with GreyMatter platform and carries this cluster's facts and your acceptance of the EULA. Everything else is the chart's defaults for that platform. Keep the file somewhere safe and outside any shared repository; you will need it again for upgrades.

Required values​

ValueSet it to
global.platformosgm. This switches the chart to GreyMatter routing and OpenShift's security model.
global.eula.acceptedtrue, after reading the EULA. The render fails without it.
core.license.existingSecretkamiwaza-license, the Secret from step 1. The render fails without it.
global.domainThe Kamiwaza domain.
global.namespaces.platform, .extensions, .sandboxes, .observability, .ca, .registry, .systemkamiwaza, for each one. Everything runs in that one namespace.
global.storage.stateful.standard, .postgres, .etcd, .registry, .extensionsThe StorageClass you recorded, for each one.
core.extensionRuntime.kubernetesApi.egressCidrsEvery Kubernetes API Service address and every endpoint address you recorded, each as /32 (IPv4) or /128 (IPv6). Copy them exactly: a wrong address installs cleanly and then fails at runtime. The render fails without this value.
global.greymatter.proxy.imageThe GreyMatter proxy image you recorded, split into registry, repository, and tag. It must match the proxy version your GreyMatter platform runs.

Values that depend on your cluster​

ValueSet it when
global.greymatter.proxy.imagePullSecretGreyMatter's image pull Secret in the namespace is not named greymatter-image-pull.
global.imagePullSecretsKamiwaza images come from a registry that needs credentials. List the pull Secret you created, for example [{name: kamiwaza-pull}].
core.inferenceResources.enabled, .bundleConfigMap, .ownerPublicKey, .clusterBindingIdYou serve models on GPUs. Set them after the cluster owner publishes a signed profile, as in Tenant-mode inference. Without them, a GPU deployment is refused. Prepare the nodes first; see GPU nodes.

Optional values​

ValueNotes
registry.persistence.size, seaweedfs.persistence.sizeVolume sizes for the model registry cache (default 500Gi) and the built-in object store (default 100Gi). Lower them when storage is small; see Block storage. Set them before the first install; they can't be reduced later.
core.userDefaults.adminPasswordThe password for the admin user. Leave it unset and Kamiwaza generates one, which you read back in step 5. Setting it here stores it in the Helm release. It applies at first install only; changing it later does not rotate the password.

Example​

global:
platform: osgm
domain: kamiwaza.example.com
eula:
accepted: true
namespaces:
platform: kamiwaza
extensions: kamiwaza
sandboxes: kamiwaza
observability: kamiwaza
ca: kamiwaza
registry: kamiwaza
system: kamiwaza
storage:
stateful:
standard: gp3-csi
postgres: gp3-csi
etcd: gp3-csi
registry: gp3-csi
extensions: gp3-csi
greymatter:
proxy:
image: # the proxy image you recorded
registry: oci.download.greymatter.io
repository: greymatter-proxy
tag: 1.17.2

core:
license:
existingSecret: kamiwaza-license
extensionRuntime:
kubernetesApi:
egressCidrs:
- 172.30.0.1/32 # kubectl -n default get service kubernetes
- 10.0.1.10/32 # the kubernetes EndpointSlice addresses
- 10.0.1.11/32
- 10.0.1.12/32

Save it as kamiwaza-values.yaml.

Step 3: Publish the GreyMatter configuration​

GreyMatter needs Kamiwaza's configuration before the install. It defines the edge that serves the Kamiwaza domain, the routes to Kamiwaza's services, and the sidecars GreyMatter adds to them. The chart generates this configuration from your values file, so it always matches the release you install. You render it, commit it to the tenant configuration repository, and GreyMatter Sync applies it.

  1. Clone the tenant configuration repository on its branch, and go to the project the platform administrator created:

    git clone -b tenant/kamiwaza https://git.example.com/org/kamiwaza-gsl.git kamiwaza-gsl
    cd kamiwaza-gsl/_tenants/kamiwaza

    If you already have a clone, run git switch tenant/kamiwaza and git pull in it instead.

  2. Render the configuration and write its files into the project. This does not change the cluster.

    helm template kamiwaza \
    oci://ghcr.io/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza \
    --version 1.3.3 \
    --namespace kamiwaza \
    --values <absolute path to kamiwaza-values.yaml> \
    --show-only charts/network/templates/greymatter/gsl-bundle.yaml \
    | kubectl create --dry-run=client --validate=false -f - -o json > gsl-bundle.json

    jq -r '.data["file-list.txt"]' gsl-bundle.json | grep '^greymatter/' \
    | while read -r f; do
    case "$f" in greymatter/generated/*) [ -e "$f" ] && continue ;; esac
    mkdir -p "$(dirname "$f")"
    jq -r --arg k "${f//\//__}" '.data[$k]' gsl-bundle.json > "$f"
    done
    rm gsl-bundle.json

    This writes the greymatter/ directory. It leaves the .greymatter file and cue.mod/ directory that the GreyMatter CLI created unchanged. It creates the files under greymatter/generated/ only if they do not exist: Kamiwaza records its app and model routes there while it runs, so an upgrade must keep them.

  3. Check that the configuration names your domain, then commit and push:

    grep -rh 'api_endpoint' greymatter | sort -u
    # Expect https://kamiwaza.example.com, and paths under it
    git add greymatter
    git diff --cached --quiet || git commit -m "Kamiwaza 1.3.3 GreyMatter configuration"
    git pull --rebase
    git push

    On a reinstall or an upgrade that changes nothing here, there is nothing to commit, and the push is a no-op.

Check that GreyMatter applied it and the edge serves the domain:

kubectl -n kamiwaza get route \
-o custom-columns='NAME:.metadata.name,HOST:.spec.host,TLS:.spec.tls.termination'
# Expect a Route for kamiwaza.example.com with TLS "passthrough"
kubectl -n kamiwaza get pods -l app.kubernetes.io/name=edge
# Expect the edge pod Running and ready

curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/
# Any HTTP status from the edge proves the Route and edge certificate work.
# Before Kamiwaza is installed, expect 404 or 503.

GreyMatter can take a few minutes to sync. A connection that fails, or ends during the TLS handshake, is not a sync delay: see Troubleshooting.

Step 4: Install​

Record the start time first. The verification step uses it to ignore events from before this install.

INSTALL_START=$(date -u +%Y-%m-%dT%H:%M:%SZ)

helm upgrade --install kamiwaza \
oci://ghcr.io/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza \
--version 1.3.3 \
--namespace kamiwaza \
--values <absolute path to kamiwaza-values.yaml> \
--wait --wait-for-jobs --timeout 30m
  • Do not add --create-namespace. The namespace belongs to the platform administrator.
  • With Helm 4, add --server-side=false, on the first install and every upgrade. OpenShift adds its own pull Secret to each of Kamiwaza's service accounts, and Helm 4's default server-side apply then refuses every later upgrade with conflict with "openshift.io/image-registry-pull-secrets_service-account-controller". Helm 3 is not affected.
  • Helm prints WARNING: core cannot see GPUs on this install. until tenant-mode inference is on. It is expected; see Tenant-mode inference.
  • A typical install takes 5 to 15 minutes, most of it pulling images.

Install and upgrade with Helm, not rendered manifests​

Always install and upgrade with helm upgrade --install against the cluster. Rendering the chart with helm template or --dry-run and applying the output, for example through Argo CD or Flux, is not supported. On first install the chart generates the platform's encryption keys, signing keys, and database passwords, and on each upgrade it reads them back from the cluster to keep them. A rendered chart cannot read the cluster, so every render generates new ones. Applying them makes existing encrypted data and access tokens unreadable.

Step 3 renders only the GreyMatter configuration files, which contain no generated secrets, and applies nothing to the cluster.

To upgrade to a later release in the same line, repeat step 3 with the new version, then rerun the command in this step with the new --version and the same values file.

Step 5: Verify​

Check what the platform does, not only its status fields.

Pods are running and ready:

kubectl -n kamiwaza get pods --no-headers \
| awk '$3 != "Completed" { split($2, r, "/"); if ($3 != "Running" || r[1] != r[2]) print }'
# Expect no output

An Error pod whose Job later completed is a retried attempt, not a failure. Confirm with kubectl -n kamiwaza get jobs: every Job should show all completions.

GreyMatter sidecars are in place. The chart adds one to the core-raycluster pods; GreyMatter adds one to frontend and keycloak. A pod without it cannot receive traffic from the edge.

kubectl -n kamiwaza get pods \
-o custom-columns='POD:.metadata.name,INIT:.spec.initContainers[*].name' \
| grep -E '^(core-raycluster|frontend|keycloak)-' | grep -v '^keycloak-postgres'
# Expect greymatter-sidecar listed for each pod

Every pod runs under OpenShift's default security context constraint:

kubectl -n kamiwaza get pods \
-o jsonpath='{range .items[*]}{.metadata.name} {.metadata.annotations.openshift\.io/scc}{"\n"}{end}' \
| grep -v -E ' restricted-v2$' | grep -v -E '^greymatter-'
# Expect no output

GreyMatter's own controllers in the namespace (greymatter-*) use GreyMatter's SCCs and are left out.

No image pulls failed during this install. Older events in the namespace are ignored. Run this in the same shell as step 4, or set INSTALL_START to the time you started the install, in the same format:

kubectl -n kamiwaza get events -o json | jq --arg start "$INSTALL_START" '
[.items[]
| select((.lastTimestamp // .eventTime // .metadata.creationTimestamp) >= $start)
| select((.message // "") | test("ErrImagePull|ImagePullBackOff|Failed to pull image|Back-off pulling image"))
] | length'
# Expect 0

The platform answers at its domain:

curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/ # expect 200
curl -sk -o /dev/null -w '%{http_code}\n' https://kamiwaza.example.com/api/ # expect 401

Log in. If you did not set an admin password, read the generated one:

kubectl -n kamiwaza get secret kamiwaza-user-admin \
-o jsonpath='{.data.password}' | base64 -d; echo

Open https://kamiwaza.example.com/ and log in as admin. With a certificate your browser does not trust, accept the warning first.

admin is the Kamiwaza application administrator. It is not the installer identity and not the Keycloak management account.

Call the API. Exchange the same credentials for a token, and use it as a bearer token:

TOKEN=$(kubectl -n kamiwaza get secret kamiwaza-user-admin -o jsonpath='{.data.password}' \
| base64 -d | curl -sk https://kamiwaza.example.com/api/auth/token \
--data-urlencode username=admin --data-urlencode password@- | jq -r .access_token)
curl -sk -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $TOKEN" \
https://kamiwaza.example.com/api/auth/users/me # expect 200

Step 6: Deploy the default embedding model​

Kamiwaza ships a default embedding model, all-MiniLM-L6-v2, inside the release and adds it to the model catalog during the install. On OpenShift with GreyMatter it does not deploy it for you. Until you do, embedding requests fail with HTTP 503 and the code embedding_unavailable, and features that embed text, such as semantic search in documents, are unavailable. Deploying it needs no internet access.

  1. Log in as admin and open Models.
  2. In the all-MiniLM-L6-v2 row, click Deploy.
  3. Select the embedding configuration and click Deploy.

Deploy it once. The model's first pod can start before GreyMatter has added its sidecar to the model; Kamiwaza replaces that pod with one that has the sidecar, usually within a minute.

Check that the embedding answers, using the TOKEN from step 5 (rerun the step 5 command that sets it if the token has expired):

curl -sk -X POST -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"text":"embedding check"}' https://kamiwaza.example.com/api/embedding/generate \
| jq 'if .embedding then (.embedding | length) else . end'
# Expect 384. Otherwise the response is printed: the code embedding_deploying
# means the model is still starting; retry after a minute.

kubectl -n kamiwaza get pods \
-o custom-columns='POD:.metadata.name,INIT:.spec.initContainers[*].name' | grep '^llamacpp-'
# Expect one pod, with greymatter-sidecar listed. Within the first minute you
# can also see the replaced pod terminating.

The pod check applies with tenant-mode inference off, which is the default. With it on, model pods are named tmi-<id>; see Tenant-mode inference.

The platform is installed. Deploying models and apps is covered in the Quickstart.

Troubleshooting​

The install stops before creating anything. The chart checks its values before rendering. A missing EULA acceptance or license prints INSTALLATION HALTED; a missing egressCidrs value names the value it needs. Nothing was installed. Fix the values file and rerun step 4.

no Kamiwaza license is configured, with the license set. The license value belongs under core:. Check that it is core.license.existingSecret, not license.existingSecret. To see values Helm received but no chart reads, run helm get values kamiwaza -n kamiwaza.

The domain gives a TLS error, or no response, after step 3. Check that the Route's TLS termination is passthrough, that the edge pod is ready, and that greymatter-edge-ingress holds tls.crt and tls.key for the domain. An admitted Route with a ready edge still fails the handshake when the edge cannot read its certificate.

No Route or edge pod appears after step 3. GreyMatter has not applied the configuration. Ask the platform administrator to check that the namespace is enrolled, that the tenant controllers (for example greymatter-tcm) are ready, and that greymatter-tenant-repo names the repository, branch, and _tenants/kamiwaza path you pushed to. If greymatter-tcm is running but its log repeats Failed to connect to a NATS cluster, the edge certificate isn't issued by the mesh CA; see Edge certificate.

Pods do not start: container has runAsNonRoot and image has non-numeric user, or image will run as root. The pod was admitted under an SCC other than restricted-v2. See which with kubectl -n kamiwaza get pod <pod> -o jsonpath='{.metadata.annotations.openshift\.io/scc}', and ask the platform administrator to remove the extra SCC from Kamiwaza's service accounts, as in Security context constraints. Then delete the pod so it is recreated.

Pods stay in CreateContainerConfigError, naming greymatter-tenant-repo. The Secret is missing or lacks a key. core-scheduler and core-raycluster need url, branch, relative_path, http_username, and http_password.

The core-raycluster pods cannot pull the GreyMatter proxy image. Check global.greymatter.proxy.image against the image you recorded, and that the pull Secret exists: kubectl -n kamiwaza get secret greymatter-image-pull.

The install hangs for about 15 minutes, then fails on a post-install job. Check core-scheduler first: kubectl -n kamiwaza logs deploy/core-scheduler. A license problem shows here as License check failed; see License File. If frontend or keycloak are not ready, GreyMatter has not given them sidecars; check step 3. While the job runs, every helm upgrade fails with another operation (install/upgrade/rollback) is in progress.

Pods are Pending on their volumes. A StorageClass named in the values does not exist or cannot provision. Compare kubectl get storageclass with global.storage.stateful.*.

Images fail to pull with unauthorized. The pull Secret named in global.imagePullSecrets does not exist in the namespace, or has the wrong credentials. Compare it with kubectl -n kamiwaza get secret.

Apps install but cannot reach the Kubernetes API. egressCidrs is missing the endpoint addresses. Add every address from Kubernetes API addresses and rerun step 4.

An app returns "Page Not Found" or the platform's own page, or a model's endpoint returns no healthy upstream. Kamiwaza adds each app's and model's route to the tenant configuration repository, and GreyMatter then applies it. Wait a minute or two after its pods are ready, or after the model shows as deployed. If it stays unreachable, check the core-scheduler logs for Git errors: the branch must accept direct pushes, and the namespace must reach the Git server.

Embedding requests return HTTP 503 with embedding_unavailable. No embedding model is deployed. Deploy the default one as in step 6.

A GPU model deployment fails with HTTP 422 and Core cannot see a GPU on this install. Tenant-mode inference is off, which is the default. Prepare the GPU nodes, then turn it on; see Tenant-mode inference.

Next steps​