Install Without Internet Access
An air-gapped install is an ordinary install that pulls from your own registry. This page covers the two things that differ: copying the release into a registry your cluster can reach, and the values that point Kamiwaza at it.
Do this page before you install. The checks in the prerequisites for your environment use two public images that step 3 copies, so on a cluster that already has no internet access, do steps 2 and 3 first (and step 4, if your registry needs credentials), then run those checks, then continue. How you move files across the gap is up to you and your site's transfer process; the commands below assume one machine that can reach both the release registry and yours, and they work the same through any intermediate hop that preserves image digests.
Before you start
| You need | Notes |
|---|---|
| A container registry your cluster's nodes can pull from | Any OCI registry. It must accept repository paths several segments deep, for example kamiwaza-ai/releases/kamiwaza/images/core. On a registry that groups repositories into projects, such as Harbor, create a project named kamiwaza-ai. |
| About 20 GB free in that registry | For the images without GPU inference engines. The GPU engine images are several GB each; add them if you will serve models on GPUs. |
| A tool that copies images between registries | It must copy every platform of a multi-platform image and keep each image's digest unchanged. The examples use skopeo with --all --preserve-digests; crane copy and oras cp -r also work. |
yq 4 (mikefarah/yq), on the machine you copy from | Turns the release manifest into the image list in step 1. The Python tool also named yq takes different arguments. |
| Helm 3.14 or later, on the machine you copy from | Checks the copied chart and renders it against your registry before you install. |
| A cluster whose own images pull without internet access | Kamiwaza's release covers only Kamiwaza. The cluster's own components, including your service mesh's sidecar image, must already pull from inside your network. On OpenShift with GreyMatter, you also copy the GreyMatter proxy image that Kamiwaza's Ray pods run; see step 3. |
The examples use these names. Set them once in the shell you copy from:
SOURCE=ghcr.io/kamiwaza-ai/releases/kamiwaza # where Kamiwaza publishes the release
REGISTRY=<your registry host> # bare host[:port], no scheme or path
Step 1: Get the image list
The release is one Helm chart plus the images it deploys, and the chart of the optional VRAM plugin that a cluster administrator installs beside it. Beside the charts, the release publishes a manifest that names every image the release can deploy, at its location in the release and with the tags it is published under: the platform, its databases, the default extensions, the model serving engines, and the default embedding model.
Fetch the manifest and turn it into a list with one image per line:
skopeo copy docker://$SOURCE/manifest:1.3.3 dir:./release-manifest
layer=$(yq -r '.layers[0].digest | sub("^sha256:"; "")' ./release-manifest/manifest.json)
yq '(.components[].artifacts[] | select((.tags // []) | length == 0)
| .repository + "@" + .digest),
(.components[].artifacts[] | select((.tags // []) | length > 0)
| .repository + ":" + .tags[] + "@" + .digest)' \
"./release-manifest/$layer" > images.txt
The manifest is an OCI artifact with one YAML file in it, which skopeo saves
under its digest. With oras, oras pull $SOURCE/manifest:1.3.3
saves the same file as release-1.3.3.yaml.
Each line of images.txt is one image at the release location, pinned by
digest. Most lines carry the tag the image is published under, for example
<release location>/images/core:<tag>@sha256:<digest>, and an image published
under two tags has a line for each. From 1.3.1, every line carries a tag: an
image that Kamiwaza pulls only by digest, such as a model serving engine, is
published under release-<version>- followed by the first 12 hexadecimal
characters of its digest. Lists from earlier releases name such images without
a tag: <release location>/images/<name>@sha256:<digest>. The copy loop in
step 3 handles both.
To leave out the GPU inference engines on a cluster without GPUs, filter the list now:
grep -vE 'vllm|cuda|rocm|diffusion' images.txt > images.cpu.txt && mv images.cpu.txt images.txt
Check:
wc -l < images.txt
# Expect a count above 0
grep -vE "^$SOURCE/[a-z0-9._/-]+(:[A-Za-z0-9_][A-Za-z0-9._-]*)?@sha256:[0-9a-f]{64}\$" images.txt
# Expect no output: every line is an image at the release location, by digest
Step 2: Trust your registry's certificate
If your registry uses a certificate from a public certificate authority, skip this step.
If it uses a certificate from your own certificate authority, everything that
talks to the registry must trust that authority, or it fails with
x509: certificate signed by unknown authority:
- the machine you copy from, before step 3;
- the machine you run
helmfrom, if it is a different one; - every node in the cluster, for its container runtime. The platform administrator does this.
On most distributions, adding the authority to the operating system's trust store is enough:
# RHEL 9
sudo cp my-ca.crt /etc/pki/ca-trust/source/anchors/
sudo update-ca-trust
# Ubuntu
sudo cp my-ca.crt /usr/local/share/ca-certificates/my-ca.crt
sudo update-ca-certificates
On each node, then restart the container runtime so it reloads the trust store.
Where the runtime is a separate service, restart it, for example
sudo systemctl restart containerd. Where your Kubernetes distribution runs the
runtime itself, restart that service instead, for example
sudo systemctl restart k0scontroller (or k0sworker) on k0s. On a single
control-plane node this briefly interrupts the Kubernetes API.
On OpenShift, do not change the nodes' trust store. Give the authority to the
cluster's image configuration instead, as the platform administrator. Create a
ConfigMap in openshift-config with one key per registry, named after the
registry's host (write host:port as host..port), holding the authority's
certificate. The image configuration names a single ConfigMap, so check first
whether it already names one:
oc get image.config.openshift.io/cluster -o jsonpath='{.spec.additionalTrustedCA.name}{"\n"}'
If that prints a name, add your registry's key to that ConfigMap, so the registries it already trusts stay trusted:
# For registry.example.com:5000; for a registry without a port, the key is the
# host alone, for example registry.example.com
oc -n openshift-config set data configmap/<name it printed> \
--from-file=registry.example.com..5000=my-ca.crt
If it prints nothing, create the ConfigMap and point the image configuration at it:
oc -n openshift-config create configmap registry-ca \
--from-file=registry.example.com..5000=my-ca.crt
oc patch image.config.openshift.io/cluster --type merge \
-p '{"spec":{"additionalTrustedCA":{"name":"registry-ca"}}}'
The nodes pick it up within a few minutes, without restarting. On a managed OpenShift whose provider owns the image configuration, set the same authority through the provider's tooling instead.
Check on each of those machines:
curl -sS -o /dev/null -w '%{http_code}\n' https://<your registry host>/v2/
# Expect 200, or 401 if the registry requires credentials. A certificate error
# means this machine does not trust the authority yet.
On OpenShift, run this check only on the machines you copy from and run helm
from. The image configuration gives the authority to the nodes' container
runtime, not to their operating system, so curl on a node is not a valid
check. The pull check in step 3 is the check for the nodes.
Step 3: Copy the release into your registry
If your registry requires credentials, log in to it with your copy tool, and
also with helm registry login: Helm keeps its own logins, and it pulls the
chart from your registry at install. The release registry needs no login.
Copy each image to the same path in your registry, with only the registry host
changed: $REGISTRY/kamiwaza-ai/releases/kamiwaza/.... Both the chart and
Kamiwaza look for images there. The chart pulls the platform images from it, and
when Kamiwaza deploys an App Garden extension it rewrites each of the
extension's images to the same location, with the step 5 values. Then copy the
charts to their original paths: the product chart, and the VRAM plugin's chart
if a cluster administrator will install the plugin. A chart is an OCI artifact
and copies the same way as an image.
Copy by digest, to the tags the list names. The chart pulls some images by
digest and some by tag, and App Garden extensions pull by tag, so a copy that
changes a digest or drops a tag fails to pull. skopeo does not accept a tag
and a digest together, so the loop below copies each image from its digest to
its tag. An image listed without a tag gets one derived from its digest,
because a registry needs a tag to accept an image; the digest is unchanged.
: > copy-failures.txt
while read -r ref || [ -n "$ref" ]; do
repo=${ref%@*} digest=${ref##*@}
case ${repo##*/} in
*:*) tag=${repo##*:} repo=${repo%:*} ;;
*) tag=sha256-${digest#sha256:} ;;
esac
skopeo copy --all --preserve-digests --retry-times 3 \
"docker://$repo@$digest" \
"docker://$REGISTRY/kamiwaza-ai/releases/kamiwaza/${repo#"$SOURCE"/}:$tag" \
|| echo "$ref" >> copy-failures.txt
done < images.txt
skopeo copy --all --preserve-digests --retry-times 3 \
"docker://$SOURCE/charts/kamiwaza:1.3.3" \
"docker://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza:1.3.3" \
|| echo chart >> copy-failures.txt
# Only if a cluster administrator will install the VRAM plugin:
skopeo copy --all --preserve-digests --retry-times 3 \
"docker://$SOURCE/charts/vram-plugin:1.3.3" \
"docker://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/vram-plugin:1.3.3" \
|| echo vram-plugin chart >> copy-failures.txt
The checks in the prerequisites for your environment use two public
images from Docker Hub that are not part of the release. Copy them too, keeping
their Docker Hub paths. On a registry that uses projects, this needs projects
named library and curlimages:
skopeo copy --all --preserve-digests \
docker://docker.io/library/busybox:1.37 "docker://$REGISTRY/library/busybox:1.37"
skopeo copy --all --preserve-digests \
docker://docker.io/curlimages/curl:latest "docker://$REGISTRY/curlimages/curl:latest"
Then, in those checks, use <your registry host>/library/busybox:1.37 in place
of busybox:1.37, and <your registry host>/curlimages/curl:latest in place of
curlimages/curl, everywhere each one appears. Where a check names its image
both in --image and inside --overrides, as the Kubernetes with Istio DNS
check does, change both: the pod runs the one in --overrides.
If your registry requires credentials, run the checks after step 4, in the
kamiwaza namespace, where the pull Secret lives; a pod in another namespace
cannot use it. Give each check pod the Secret: add
"imagePullSecrets":[{"name":"kamiwaza-pull"}] to the spec in its
--overrides (add --overrides='{"spec":{"imagePullSecrets":[{"name":"kamiwaza-pull"}]}}'
where the command has none), or imagePullSecrets: under spec: in a check's
manifest.
Check:
cat copy-failures.txt
# Expect no output. Rerun the copy for any line listed; copies that already
# succeeded are skipped quickly.
helm show chart oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza --version 1.3.3
# Expect "name: kamiwaza" and "version: 1.3.3"
helm show chart oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/vram-plugin --version 1.3.3
# If you copied it: expect "name: vram-plugin" and "version: 1.3.3"
Check that the cluster's nodes can pull from your registry, which is the
only proof that step 2 reached their container runtime. Like the prerequisite
checks, the platform administrator runs this, with REGISTRY set.
If your registry requires credentials, run it after step 4 instead, in the
kamiwaza namespace, and add
--overrides='{"spec":{"imagePullSecrets":[{"name":"kamiwaza-pull"}]}}' to
kubectl run. In the kamiwaza namespace, kubectl prints a PodSecurity
restricted warning for this pod; the check still runs.
kubectl -n default run registry-pull-check --restart=Never \
--image=$REGISTRY/library/busybox:1.37 \
--annotations=sidecar.istio.io/inject=false -- true
kubectl -n default wait --for=jsonpath='{.status.phase}'=Succeeded \
pod/registry-pull-check --timeout=120s
# Expect "condition met". A pull error in
# `kubectl -n default describe pod registry-pull-check` names the cause:
# x509 means step 2 has not reached this node; "not found" means a missing copy.
kubectl -n default delete pod registry-pull-check
On OpenShift with GreyMatter, Kamiwaza's Ray pods also run the GreyMatter
proxy image named in global.greymatter.proxy.image. That image is not part of
the Kamiwaza release, and the step 5 values do not move it: copy it from the
registry your GreyMatter platform pulls it from, e.g. to
$REGISTRY/greymatter-proxy:1.17.2, and set global.greymatter.proxy.image to
that copy.
A plain-HTTP registry. The commands on this page assume HTTPS. Against a
registry that serves plain HTTP, turn off TLS in your copy tool (for skopeo,
--dest-tls-verify=false on copy and --tls-verify=false on login and
inspect), add --plain-http to every helm command that names your registry,
including the install, and configure your cluster's container runtime to allow
the registry. Registry credentials then cross the network unencrypted, so use
plain HTTP only on a network you trust.
Step 4: Create a pull Secret, if your registry needs credentials
If your nodes can pull from your registry without credentials, skip this step.
Otherwise, create the pull Secret as described in Registry access in the
prerequisites for
Kubernetes with Istio or
OpenShift with GreyMatter, with your
registry's host as <registry-host>. It must match $REGISTRY exactly,
including any port. Kamiwaza passes this Secret to its own pods and to every
extension it deploys.
If a Secret of that name already exists, e.g. on an install that pulled
from the internet until now, delete it first with
kubectl -n kamiwaza delete secret kamiwaza-pull, then create it as above.
Running pods keep running; pods started afterwards pull with the new Secret.
Step 5: Know the values to add
When you write your values file on your install page, merge these keys into it.
Put them under the file's existing keys, and do not repeat any key at any level:
not a second global: or core:, and not a second scheduler: under core:.
When a key appears twice, Helm keeps only the last one and silently drops the
other.
global:
imageRegistryPrefix: <your registry host>/kamiwaza-ai/releases/kamiwaza
# Only if you created a pull Secret in step 4:
# imagePullSecrets:
# - name: kamiwaza-pull
core:
scheduler:
extensionImageRelocateRegistry: <your registry host>/kamiwaza-ai/releases/kamiwaza
templates:
availability:
enabled: false
sync:
stage: LOCAL
localMode:
enabled: true
| Value | What it does | If you leave it out |
|---|---|---|
global.imageRegistryPrefix | Moves every image the chart deploys from the release location to the same path in your registry, <prefix>/images/<name>. It also moves the engine images of tenant-mode inference. | Every pod tries to pull from the release location on the internet and fails. |
core.scheduler.extensionImageRelocateRegistry | Rewrites each App Garden extension image, when the extension deploys, to this value followed by the last two segments of the image's path, which is where step 3 copied it. | Each App Garden extension tries to pull its images from the internet when it deploys, and fails. The install itself succeeds. |
core.templates.availability.enabled: false | Turns off the App Garden's online catalog lookup. | The platform keeps trying to reach Kamiwaza's online catalog, which your network cannot reach. |
core.templates.sync.stage: LOCAL | Selects the catalog that ships inside the release. | The install fails to render: templates.sync.localMode.enabled=true requires templates.sync.stage=LOCAL. |
core.templates.sync.localMode.enabled: true | Loads the App Garden from that catalog. | The install fails to render: templates.sync.stage=LOCAL requires templates.sync.localMode.enabled=true. |
global.imagePullSecrets | Names the pull Secret for the platform and its extensions. | Every pull from a registry that needs credentials fails with an authorization error. |
The two registry values are the same: your registry host followed by the
release path, e.g. registry.example.com:5000/kamiwaza-ai/releases/kamiwaza,
with no scheme such as https://. The chart and Kamiwaza supply the rest of
each image's path.
Step 6: Install
Go to Installation Environments and open the install page for your environment. Follow it as written, with the step 5 values merged into your values file. In every command on that page that names the release location, including the install command, use the chart from your registry instead:
oci://<your registry host>/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza
Check your completed values file against the chart just before you run the
install command. The render needs no cluster access. Run it from the directory
that holds kamiwaza-values.yaml, with REGISTRY set as above:
helm template kamiwaza oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/kamiwaza \
--version 1.3.3 --namespace kamiwaza --values kamiwaza-values.yaml > rendered.yaml
grep -E '^\s+(- )?image: ' rendered.yaml | grep -v "image: \"\{0,1\}$REGISTRY/"
# Expect no output: every image the chart deploys comes from your registry
# Every image path contains /images/; this skips the prefix value itself.
grep -oE "$REGISTRY/[^\"' ,\\\\}]+" rendered.yaml | grep '/images/' \
| sed -E 's/:[^/@]+@/@/' | sort -u > rendered-images.txt
# App Garden extension images keep their original registry in the render and
# are relocated when an extension deploys, so check their relocated copies too.
# Extensions run images pinned as tag@digest and relocated by tag; a reference
# without a digest is catalog display data that nothing pulls, so this checks
# only tag@digest references.
grep -oE "ghcr\.io/[^\"' ,\\\\}]+@sha256:[0-9a-f]{64}" rendered.yaml | grep -E ':[^/@]+@' \
| sed -E "s#^ghcr\.io/(.*/)?([^/]+/[^/:@]+):([^/@]+).*#$REGISTRY/kamiwaza-ai/releases/kamiwaza/\2:\3#" \
| sort -u >> rendered-images.txt
while read -r image; do # with skopeo; any tool that reads a manifest works
skopeo inspect --raw "docker://$image" > /dev/null 2>&1 || echo "missing: $image"
done < rendered-images.txt
# Expect no output: every image the chart deploys, and every App Garden image
# the render names in full, is in your registry. If you left out the GPU
# engines in step 1, expect them, and only them, listed.
If the render fails, the error names the value to fix. Seeing ghcr.io inside
the rendered App Garden catalog is expected; Kamiwaza rewrites those references
when an extension deploys.
Engine images for tenant-mode inference
Skip this section unless you will turn on tenant-mode inference, which serves models on GPUs and CPUs in a namespaced install.
Tenant-mode inference runs engine images that the core image's engine catalog
pins by digest. The chart does not deploy them, so the step 6 check does not
cover them, and Kamiwaza pulls each one by digest from
global.imageRegistryPrefix followed by images/ and the last segment of the
catalog's repository, such as
$REGISTRY/kamiwaza-ai/releases/kamiwaza/images/vllm-cuda@sha256:.... The
release manifest lists every engine image at that path, so step 3 has already
copied them, by digest, to where Kamiwaza looks; there is nothing more to copy.
If you left out the GPU inference engines in step 1, those engines are not in
your registry, and a model that needs one fails to deploy.
If this install served models before you moved it off the internet, its running
inference deployments, including the platform's embedding model, still name
the registry they were planned with. After the upgrade, stop each one and deploy
it again. While core.context.embedding.autoProvision is on (the default,
except on global.platform: osgm), the embedding model's deployment needs only
the stop: the next embedding request deploys it again, from your registry. That
request returns HTTP 503 with the code embedding_deploying until the new
deployment is ready; retry it. With it off, the platform does not redeploy the
model, and embedding requests return HTTP 503 with the code
embedding_unavailable until you deploy it again yourself. See
Engine images. On later
upgrades, keep the previous release's engine images in your registry until
every inference deployment has been redeployed.
VRAM plugin
Skip this section unless a cluster administrator will install the
VRAM plugin. Its image,
images/vram-plugin, is in the step 1 image list, and step 3 copies the image
and the plugin's chart. Install it from your registry, after
helm registry login as in step 3:
helm upgrade --install vram-plugin oci://$REGISTRY/kamiwaza-ai/releases/kamiwaza/charts/vram-plugin \
--version 1.3.3 -n kamiwaza \
--set global.imageRegistryPrefix=$REGISTRY/kamiwaza-ai/releases/kamiwaza \
--set 'imagePullSecrets[0].name=kamiwaza-pull'
Leave out the imagePullSecrets line if you skipped step 4. Without
global.imageRegistryPrefix, the chart names the release location, and the
plugin's pods cannot pull their image.
Troubleshooting
An App Garden extension fails to deploy, with pods in ImagePullBackOff.
Describe one of the extension's pods and read the image it tries to pull:
kubectl -n kamiwaza get pods --field-selector=status.phase=Pending
kubectl -n kamiwaza describe pod <pod> | grep -E 'Image:|Failed'
If the image is <your registry host>/kamiwaza-ai/releases/kamiwaza/<path>:<tag>,
that copy is missing: rerun the step 3 copy. If it still names ghcr.io,
core.scheduler.extensionImageRelocateRegistry did not reach the platform:
check your values file and rerun the install.
Pulls fail with x509: certificate signed by unknown authority. A node does
not trust your registry's certificate authority. Repeat step 2 on that node,
including the container runtime restart.
Pulls fail with unauthorized or authentication required. Check that the
pull Secret exists in the kamiwaza namespace, that global.imagePullSecrets
names it, and that its server matches your registry host exactly, including the
port:
kubectl -n kamiwaza get secret kamiwaza-pull -o jsonpath='{.data.\.dockerconfigjson}' \
| base64 -d | jq -r '.auths | keys[]'
# Expect your registry host, exactly as in $REGISTRY
helm returns unauthorized from your registry, though your copy tool works.
Helm keeps its own login. Run helm registry login as in step 3.
Pods fail with InvalidImageName, image paths contain the same segment twice,
or the render fails with must include a repository path or
must be a registry host[:port] with an optional lowercase /path. A registry
value has the wrong shape. Set both global.imageRegistryPrefix and
core.scheduler.extensionImageRelocateRegistry to your registry host followed
by /kamiwaza-ai/releases/kamiwaza, without a scheme.
Pods still pull from the release location on the internet, and the step 6
check lists them. global.imageRegistryPrefix did not reach the install. It
must be set in your values file under the file's only global: key; the
release's chart sets it to the release location, and a value you set wins.
The App Garden is empty after install. Ask the platform for the catalog error directly, with an administrator token:
curl -sS -X POST -H "Authorization: Bearer <token>" \
https://<your Kamiwaza domain>/api/apps/remote/sync
An error that says the catalog targets a different Core version, or contains values unsupported by the catalog schema, means the chart and its catalog do not match. Check that you copied and installed the chart for the version you meant to install.