Skip to main content
Version: 1.3.3 (Latest)

Install Without Internet Access

An air-gapped install is an ordinary install that pulls from your own registry. This page covers the two things that differ: copying the release into a registry your cluster can reach, and the values that point Kamiwaza at it.

Do this page before you install. The checks in the prerequisites for your environment use two public images that step 3 copies, so on a cluster that already has no internet access, do steps 2 and 3 first (and step 4, if your registry needs credentials), then run those checks, then continue. How you move files across the gap is up to you and your site's transfer process; the commands below assume one machine that can reach both the release registry and yours, and they work the same through any intermediate hop that preserves image digests.

Before you start​

You needNotes
A container registry your cluster's nodes can pull fromAny OCI registry. It must accept repository paths several segments deep, for example kamiwaza-ai/releases/kamiwaza/images/core. On a registry that groups repositories into projects, such as Harbor, create a project named kamiwaza-ai.
About 20 GB free in that registryFor the images without GPU inference engines. The GPU engine images are several GB each; add them if you will serve models on GPUs.
A tool that copies images between registriesIt must copy every platform of a multi-platform image and keep each image's digest unchanged. The examples use skopeo with --all --preserve-digests; crane copy and oras cp -r also work.
yq 4 (mikefarah/yq), on the machine you copy fromTurns the release manifest into the image list in step 1. The Python tool also named yq takes different arguments.
Helm 3.14 or later, on the machine you copy fromChecks the copied chart and renders it against your registry before you install.
A cluster whose own images pull without internet accessKamiwaza's release covers only Kamiwaza. The cluster's own components, including your service mesh's sidecar image, must already pull from inside your network. On OpenShift with GreyMatter, you also copy the GreyMatter proxy image that Kamiwaza's Ray pods run; see step 3.

The examples use these names. Set them once in the shell you copy from:

SOURCE=ghcr.io/kamiwaza-ai/releases/kamiwaza # where Kamiwaza publishes the release
REGISTRY=<your registry host> # bare host[:port], no scheme or path

Step 1: Get the image list​

The release is one Helm chart plus the images it deploys, and the chart of the optional VRAM plugin that a cluster administrator installs beside it. Beside the charts, the release publishes a manifest that names every image the release can deploy, at its location in the release and with the tags it is published under: the platform, its databases, the default extensions, the model serving engines, and the default embedding model.

Fetch the manifest and turn it into a list with one image per line:

skopeo copy docker://$SOURCE/manifest:1.3.3 dir:./release-manifest
layer=$(yq -r '.layers[0].digest | sub("^sha256:"; "")' ./release-manifest/manifest.json)
yq '(.components[].artifacts[] | select((.tags // []) | length == 0)
| .repository + "@" + .digest),
(.components[].artifacts[] | select((.tags // []) | length > 0)
| .repository + ":" + .tags[] + "@" + .digest)' \
"./release-manifest/$layer" > images.txt

The manifest is an OCI artifact with one YAML file in it, which skopeo saves under its digest. With oras, oras pull $SOURCE/manifest:1.3.3 saves the same file as release-1.3.3.yaml.

Each line of images.txt is one image at the release location, pinned by digest. Most lines carry the tag the image is published under, for example <release location>/images/core:<tag>@sha256:<digest>, and an image published under two tags has a line for each. From 1.3.1, every line carries a tag: an image that Kamiwaza pulls only by digest, such as a model serving engine, is published under release-<version>- followed by the first 12 hexadecimal characters of its digest. Lists from earlier releases name such images without a tag: <release location>/images/<name>@sha256:<digest>. The copy loop in step 3 handles both.

To leave out the GPU inference engines on a cluster without GPUs, filter the list now:

grep -vE 'vllm|cuda|rocm|diffusion' images.txt > images.cpu.txt && mv images.cpu.txt images.txt

Check:

wc -l < images.txt
# Expect a count above 0
grep -vE "^$SOURCE/[a-z0-9._/-]+(:[A-Za-z0-9_][A-Za-z0-9._-]*)?@sha256:[0-9a-f]{64}\$" images.txt
# Expect no output: every line is an image at the release location, by digest

Step 2: Trust your registry's certificate​

If your registry uses a certificate from a public certificate authority, skip this step.

If it uses a certificate from your own certificate authority, everything that talks to the registry must trust that authority, or it fails with x509: certificate signed by unknown authority:

  • the machine you copy from, before step 3;
  • the machine you run helm from, if it is a different one;
  • every node in the cluster, for its container runtime. The platform administrator does this.

On most distributions, adding the authority to the operating system's trust store is enough:

# RHEL 9
sudo cp my-ca.crt /etc/pki/ca-trust/source/anchors/
sudo update-ca-trust

# Ubuntu
sudo cp my-ca.crt /usr/local/share/ca-certificates/my-ca.crt
sudo update-ca-certificates

On each node, then restart the container runtime so it reloads the trust store. Where the runtime is a separate service, restart it, for example sudo systemctl restart containerd. Where your Kubernetes distribution runs the runtime itself, restart that service instead, for example sudo systemctl restart k0scontroller (or k0sworker) on k0s. On a single control-plane node this briefly interrupts the Kubernetes API.

On OpenShift, do not change the nodes' trust store. Give the authority to the cluster's image configuration instead, as the platform administrator. Create a ConfigMap in openshift-config with one key per registry, named after the registry's host (write host:port as host..port), holding the authority's certificate. The image configuration names a single ConfigMap, so check first whether it already names one:

oc get image.config.openshift.io/cluster -o jsonpath='{.spec.additionalTrustedCA.name}{"\n"}'

If that prints a name, add your registry's key to that ConfigMap, so the registries it already trusts stay trusted:

# For registry.example.com:5000; for a registry without a port, the key is the
# host alone, for example registry.example.com
oc -n openshift-config set data configmap/<name it printed> \
--from-file=registry.example.com..5000=my-ca.crt

If it prints nothing, create the ConfigMap and point the image configuration at it:

oc -n openshift-config create configmap registry-ca \
--from-file=registry.example.com..5000=my-ca.crt
oc patch image.config.openshift.io/cluster --type merge \
-p '{"spec":{"additionalTrustedCA":{"name":"registry-ca"}}}'

The nodes pick it up within a few minutes, without restarting. On a managed OpenShift whose provider owns the image configuration, set the same authority through the provider's tooling instead.

Check on each of those machines:

curl -sS -o /dev/null -w '%{http_code}\n' https://<your registry host>/v2/
# Expect 200, or 401 if the registry requires credentials. A certificate error
# means this machine does not trust the authority yet.

On OpenShift, run this check only on the machines you copy from and run helm from. The image configuration gives the authority to the nodes' container runtime, not to their operating system, so curl on a node is not a valid check. The pull check in step 3 is the check for the nodes.

Step 3: Copy the release into your registry​

If your registry requires credentials, log in to it with your copy tool, and also with helm registry login: Helm keeps its own logins, and it pulls the chart from your registry at install. The release registry needs no login.

Copy each image to the same path in your registry, with only the registry host changed: $REGISTRY/kamiwaza-ai/releases/kamiwaza/.... Both the chart and Kamiwaza look for images there. The chart pulls the platform images from it, and when Kamiwaza deploys an App Garden extension it rewrites each of the extension's images to the same location, with the step 5 values. Then copy the charts to their original paths: the product chart, and the VRAM plugin's chart if a cluster administrator will install the plugin. A chart is an OCI artifact and copies the same way as an image.

Copy by digest, to the tags the list names. The chart pulls some images by digest and some by tag, and App Garden extensions pull by tag, so a copy that changes a digest or drops a tag fails to pull. skopeo does not accept a tag and a digest together, so the loop below copies each image from its digest to its tag. An image listed without a tag gets one derived from its digest, because a registry needs a tag to accept an image; the digest is unchanged.

: > copy-failures.txt
while read -r ref || [ -n "$ref" ]; do
repo=${ref%@*} digest=${ref##*@}
case ${repo##*/} in
*:*) tag=${repo##*:} repo=${repo%:*} ;;
*) tag=sha256-${digest#sha256:} ;;
esac
skopeo copy --all --preserve-digests --retry-times 3 \
"docker://$repo@$digest" \
"docker://$REGISTRY/kamiwaza-ai/releases/kamiwaza/${repo#"$SOURCE"/}:$tag" \
|| echo "$ref" >> copy-failures.txt
done < images.txt

skopeo copy --all --preserve-digests --retry-times 3 \
"docker://$SOURCE/charts/kamiwaza:1.3.3" \
"docker://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza:1.3.3" \
|| echo chart >> copy-failures.txt

# Only if a cluster administrator will install the VRAM plugin:
skopeo copy --all --preserve-digests --retry-times 3 \
"docker://$SOURCE/charts/vram-plugin:1.3.3" \
"docker://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/vram-plugin:1.3.3" \
|| echo vram-plugin chart >> copy-failures.txt

The checks in the prerequisites for your environment use two public images from Docker Hub that are not part of the release. Copy them too, keeping their Docker Hub paths. On a registry that uses projects, this needs projects named library and curlimages:

skopeo copy --all --preserve-digests \
docker://docker.io/library/busybox:1.37 "docker://$REGISTRY/library/busybox:1.37"
skopeo copy --all --preserve-digests \
docker://docker.io/curlimages/curl:latest "docker://$REGISTRY/curlimages/curl:latest"

Then, in those checks, use <your registry host>/library/busybox:1.37 in place of busybox:1.37, and <your registry host>/curlimages/curl:latest in place of curlimages/curl, everywhere each one appears. Where a check names its image both in --image and inside --overrides, as the Kubernetes with Istio DNS check does, change both: the pod runs the one in --overrides.

If your registry requires credentials, run the checks after step 4, in the kamiwaza namespace, where the pull Secret lives; a pod in another namespace cannot use it. Give each check pod the Secret: add "imagePullSecrets":[{"name":"kamiwaza-pull"}] to the spec in its --overrides (add --overrides='{"spec":{"imagePullSecrets":[{"name":"kamiwaza-pull"}]}}' where the command has none), or imagePullSecrets: under spec: in a check's manifest.

Check:

cat copy-failures.txt
# Expect no output. Rerun the copy for any line listed; copies that already
# succeeded are skipped quickly.

helm show chart oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza --version 1.3.3
# Expect "name: kamiwaza" and "version: 1.3.3"

helm show chart oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/vram-plugin --version 1.3.3
# If you copied it: expect "name: vram-plugin" and "version: 1.3.3"

Check that the cluster's nodes can pull from your registry, which is the only proof that step 2 reached their container runtime. Like the prerequisite checks, the platform administrator runs this, with REGISTRY set. If your registry requires credentials, run it after step 4 instead, in the kamiwaza namespace, and add --overrides='{"spec":{"imagePullSecrets":[{"name":"kamiwaza-pull"}]}}' to kubectl run. In the kamiwaza namespace, kubectl prints a PodSecurity restricted warning for this pod; the check still runs.

kubectl -n default run registry-pull-check --restart=Never \
--image=$REGISTRY/library/busybox:1.37 \
--annotations=sidecar.istio.io/inject=false -- true
kubectl -n default wait --for=jsonpath='{.status.phase}'=Succeeded \
pod/registry-pull-check --timeout=120s
# Expect "condition met". A pull error in
# `kubectl -n default describe pod registry-pull-check` names the cause:
# x509 means step 2 has not reached this node; "not found" means a missing copy.
kubectl -n default delete pod registry-pull-check

On OpenShift with GreyMatter, Kamiwaza's Ray pods also run the GreyMatter proxy image named in global.greymatter.proxy.image. That image is not part of the Kamiwaza release, and the step 5 values do not move it: copy it from the registry your GreyMatter platform pulls it from, e.g. to $REGISTRY/greymatter-proxy:1.17.2, and set global.greymatter.proxy.image to that copy.

A plain-HTTP registry. The commands on this page assume HTTPS. Against a registry that serves plain HTTP, turn off TLS in your copy tool (for skopeo, --dest-tls-verify=false on copy and --tls-verify=false on login and inspect), add --plain-http to every helm command that names your registry, including the install, and configure your cluster's container runtime to allow the registry. Registry credentials then cross the network unencrypted, so use plain HTTP only on a network you trust.

Step 4: Create a pull Secret, if your registry needs credentials​

If your nodes can pull from your registry without credentials, skip this step.

Otherwise, create the pull Secret as described in Registry access in the prerequisites for Kubernetes with Istio or OpenShift with GreyMatter, with your registry's host as <registry-host>. It must match $REGISTRY exactly, including any port. Kamiwaza passes this Secret to its own pods and to every extension it deploys.

If a Secret of that name already exists, e.g. on an install that pulled from the internet until now, delete it first with kubectl -n kamiwaza delete secret kamiwaza-pull, then create it as above. Running pods keep running; pods started afterwards pull with the new Secret.

Step 5: Know the values to add​

When you write your values file on your install page, merge these keys into it. Put them under the file's existing keys, and do not repeat any key at any level: not a second global: or core:, and not a second scheduler: under core:. When a key appears twice, Helm keeps only the last one and silently drops the other.

global:
imageRegistryPrefix: <your registry host>/kamiwaza-ai/releases/kamiwaza
# Only if you created a pull Secret in step 4:
# imagePullSecrets:
# - name: kamiwaza-pull

core:
scheduler:
extensionImageRelocateRegistry: <your registry host>/kamiwaza-ai/releases/kamiwaza
templates:
availability:
enabled: false
sync:
stage: LOCAL
localMode:
enabled: true
ValueWhat it doesIf you leave it out
global.imageRegistryPrefixMoves every image the chart deploys from the release location to the same path in your registry, <prefix>/images/<name>. It also moves the engine images of tenant-mode inference.Every pod tries to pull from the release location on the internet and fails.
core.scheduler.extensionImageRelocateRegistryRewrites each App Garden extension image, when the extension deploys, to this value followed by the last two segments of the image's path, which is where step 3 copied it.Each App Garden extension tries to pull its images from the internet when it deploys, and fails. The install itself succeeds.
core.templates.availability.enabled: falseTurns off the App Garden's online catalog lookup.The platform keeps trying to reach Kamiwaza's online catalog, which your network cannot reach.
core.templates.sync.stage: LOCALSelects the catalog that ships inside the release.The install fails to render: templates.sync.localMode.enabled=true requires templates.sync.stage=LOCAL.
core.templates.sync.localMode.enabled: trueLoads the App Garden from that catalog.The install fails to render: templates.sync.stage=LOCAL requires templates.sync.localMode.enabled=true.
global.imagePullSecretsNames the pull Secret for the platform and its extensions.Every pull from a registry that needs credentials fails with an authorization error.

The two registry values are the same: your registry host followed by the release path, e.g. registry.example.com:5000/kamiwaza-ai/releases/kamiwaza, with no scheme such as https://. The chart and Kamiwaza supply the rest of each image's path.

Step 6: Install​

Go to Installation Environments and open the install page for your environment. Follow it as written, with the step 5 values merged into your values file. In every command on that page that names the release location, including the install command, use the chart from your registry instead:

oci://<your registry host>/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza

Check your completed values file against the chart just before you run the install command. The render needs no cluster access. Run it from the directory that holds kamiwaza-values.yaml, with REGISTRY set as above:

helm template kamiwaza oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza \
--version 1.3.3 --namespace kamiwaza --values kamiwaza-values.yaml > rendered.yaml

grep -E '^\s+(- )?image: ' rendered.yaml | grep -v "image: \"\{0,1\}$REGISTRY/"
# Expect no output: every image the chart deploys comes from your registry

# Every image path contains /images/; this skips the prefix value itself.
grep -oE "$REGISTRY/[^\"' ,\\\\}]+" rendered.yaml | grep '/images/' \
| sed -E 's/:[^/@]+@/@/' | sort -u > rendered-images.txt

# App Garden extension images keep their original registry in the render and
# are relocated when an extension deploys, so check their relocated copies too.
# Extensions run images pinned as tag@digest and relocated by tag; a reference
# without a digest is catalog display data that nothing pulls, so this checks
# only tag@digest references.
grep -oE "ghcr\.io/[^\"' ,\\\\}]+@sha256:[0-9a-f]{64}" rendered.yaml | grep -E ':[^/@]+@' \
| sed -E "s#^ghcr\.io/(.*/)?([^/]+/[^/:@]+):([^/@]+).*#$REGISTRY/kamiwaza-ai/releases/kamiwaza/\2:\3#" \
| sort -u >> rendered-images.txt

while read -r image; do # with skopeo; any tool that reads a manifest works
skopeo inspect --raw "docker://$image" > /dev/null 2>&1 || echo "missing: $image"
done < rendered-images.txt
# Expect no output: every image the chart deploys, and every App Garden image
# the render names in full, is in your registry. If you left out the GPU
# engines in step 1, expect them, and only them, listed.

If the render fails, the error names the value to fix. Seeing ghcr.io inside the rendered App Garden catalog is expected; Kamiwaza rewrites those references when an extension deploys.

Engine images for tenant-mode inference​

Skip this section unless you will turn on tenant-mode inference, which serves models on GPUs and CPUs in a namespaced install.

Tenant-mode inference runs engine images that the core image's engine catalog pins by digest. The chart does not deploy them, so the step 6 check does not cover them, and Kamiwaza pulls each one by digest from global.imageRegistryPrefix followed by images/ and the last segment of the catalog's repository, such as $REGISTRY/kamiwaza-ai/releases/kamiwaza/images/vllm-cuda@sha256:.... The release manifest lists every engine image at that path, so step 3 has already copied them, by digest, to where Kamiwaza looks; there is nothing more to copy. If you left out the GPU inference engines in step 1, those engines are not in your registry, and a model that needs one fails to deploy.

If this install served models before you moved it off the internet, its running inference deployments, including the platform's embedding model, still name the registry they were planned with. After the upgrade, stop each one and deploy it again. While core.context.embedding.autoProvision is on (the default, except on global.platform: osgm), the embedding model's deployment needs only the stop: the next embedding request deploys it again, from your registry. That request returns HTTP 503 with the code embedding_deploying until the new deployment is ready; retry it. With it off, the platform does not redeploy the model, and embedding requests return HTTP 503 with the code embedding_unavailable until you deploy it again yourself. See Engine images. On later upgrades, keep the previous release's engine images in your registry until every inference deployment has been redeployed.

VRAM plugin​

Skip this section unless a cluster administrator will install the VRAM plugin. Its image, images/vram-plugin, is in the step 1 image list, and step 3 copies the image and the plugin's chart. Install it from your registry, after helm registry login as in step 3:

helm upgrade --install vram-plugin oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/vram-plugin \
--version 1.3.3 -n kamiwaza \
--set global.imageRegistryPrefix=$REGISTRY/kamiwaza-ai/releases/kamiwaza \
--set 'imagePullSecrets[0].name=kamiwaza-pull'

Leave out the imagePullSecrets line if you skipped step 4. Without global.imageRegistryPrefix, the chart names the release location, and the plugin's pods cannot pull their image.

Troubleshooting​

An App Garden extension fails to deploy, with pods in ImagePullBackOff. Describe one of the extension's pods and read the image it tries to pull:

kubectl -n kamiwaza get pods --field-selector=status.phase=Pending
kubectl -n kamiwaza describe pod <pod> | grep -E 'Image:|Failed'

If the image is <your registry host>/kamiwaza-ai/releases/kamiwaza/<path>:<tag>, that copy is missing: rerun the step 3 copy. If it still names ghcr.io, core.scheduler.extensionImageRelocateRegistry did not reach the platform: check your values file and rerun the install.

Pulls fail with x509: certificate signed by unknown authority. A node does not trust your registry's certificate authority. Repeat step 2 on that node, including the container runtime restart.

Pulls fail with unauthorized or authentication required. Check that the pull Secret exists in the kamiwaza namespace, that global.imagePullSecrets names it, and that its server matches your registry host exactly, including the port:

kubectl -n kamiwaza get secret kamiwaza-pull -o jsonpath='{.data.\.dockerconfigjson}' \
| base64 -d | jq -r '.auths | keys[]'
# Expect your registry host, exactly as in $REGISTRY

helm returns unauthorized from your registry, though your copy tool works. Helm keeps its own login. Run helm registry login as in step 3.

Pods fail with InvalidImageName, image paths contain the same segment twice, or the render fails with must include a repository path or must be a registry host[:port] with an optional lowercase /path. A registry value has the wrong shape. Set both global.imageRegistryPrefix and core.scheduler.extensionImageRelocateRegistry to your registry host followed by /kamiwaza-ai/releases/kamiwaza, without a scheme.

Pods still pull from the release location on the internet, and the step 6 check lists them. global.imageRegistryPrefix did not reach the install. It must be set in your values file under the file's only global: key; the release's chart sets it to the release location, and a value you set wins.

The App Garden is empty after install. Ask the platform for the catalog error directly, with an administrator token:

curl -sS -X POST -H "Authorization: Bearer <token>" \
https://<your Kamiwaza domain>/api/apps/remote/sync

An error that says the catalog targets a different Core version, or contains values unsupported by the catalog schema, means the chart and its catalog do not match. Check that you copied and installed the chart for the version you meant to install.