• Get started

    • Getting started
    • Concepts
    • Tour
  • Using shpyrd

    • Deploying
    • shpyrd.yaml
    • Resources
    • Databases and caches
    • Domains and exposure
    • Sign-in for your app
    • Teams, roles and security
    • AI assistants (MCP)
    • Logs
    • Dashboard
    • CLI reference
  • Running it yourself

    • Installation
    • Oracle Cloud (OKE)
    • AWS (EKS)
    • Extensions and sign-in
    • Platform backups
    • Architecture guide
  • Project

    • Design principles
    • Roadmap
    • How to contribute
    • Support the project
  1. Running it yourself
  2. Architecture guide
Add toClaudeClaude CodeClaudeClaudeOpenAICodexCursorCursorVisual Studio CodeVS Code

Architecture guide

How the installer, the App controller and the server turn a Kubernetes cluster into a PaaS.

shpyrd is two binaries and a set of well-known open source components. This guide is for running shpyrd yourself, or for knowing how it works; on shpyrd cloud the platform is run for you: see Getting started. The CLI installs and operates the cluster from your machine; the server runs inside it as API, controller and dashboard.

Components

ComponentRole in shpyrd
kind (local profile)the Kubernetes cluster itself, driven as a Go library
cert-manager + trust-managercertificates for every project and the dashboard; on the local profile from a development CA generated by the CLI, distributed to all namespaces as a trust bundle
ingress-nginxHTTP entry point; publishes the per-host request and latency metrics behind the dashboard charts
registry (registry:3)in-cluster OCI registry for builder and application images, TLS from the platform CA, one platform credential, weekly garbage collection
Calico (policy only, oci profile)enforces the per-project NetworkPolicy where the provider's CNI does not
ExternalDNS + an OCI DNS-01 solver (oci profile, optional)publishes records for every hostname and issues the platform's wildcard certificate
kpack + Paketo buildpacksbuilds source into images inside the cluster; polls Git branches for new commits
kube-prometheus-stackPrometheus, Grafana, node-exporter, kube-state-metrics
shpyrd-serverAPI for the CLI upload and the dashboard, the App controller, source archive store, the two embedded applications (the console for the operator, the workspace for its members), served by host
shpyrd CLIcluster bootstrap and every project operation

Control-plane database. Teams, project grants, the workspace and the people it has seen sign in live in PostgreSQL, not in Kubernetes objects: the control-plane-db component runs a single instance with a persistent volume (SHPYRD_CONTROL_PLANE_DB_SIZE), or SHPYRD_DATABASE_URL names a managed database and the component is skipped. Workloads (App, Volume, Postgres, Redis, LogDrain, ObjectBucket) stay custom resources reconciled by controllers. Every identifier is a native uuid; the schema is versioned with golang-migrate (embedded up and down files in pkg/store/migrations, applied at start under an advisory lock, a dirty flag when a migration was interrupted), and the server imports the Team/ProjectMember objects of older installs once. Take a database backup before upgrading across a schema migration; shpyrd cluster backup carries the database's content, and pg_dump of the control-plane-db pod is the plain alternative.

The installer

shpyrd cluster init applies a profile: an ordered list of components, each a Helm chart (installed with the Helm SDK, values embedded in the binary), a Kustomize tree (rendered in-process and applied with server-side apply), or both, plus readiness conditions: Deployments available, CRDs established, an admission webhook accepting a dry-run request, a status condition true for the current generation.

Components are grouped in runlevels (rc0...rc4). A level is applied in parallel and waited for before the next starts, so CRDs exist before the resources that use them and webhooks are serving before objects they validate. Variables (${SHPYRD_DOMAIN}, ${SHPYRD_REGISTRY_HOST}...) are substituted in the manifests before Kustomize parses them. The result is recorded in a ConfigMap (shpyrd-system/shpyrd-install) that cluster status and the dashboard read. Re-runs upgrade in place; cluster export renders everything to disk for GitOps tools.

Cloud profiles (oci today) swap the pieces that differ - a cloud load balancer on a reserved address, Let's Encrypt, the provider's storage - and run cluster init in two phases: everything that needs no certificate first, then a wait for the load balancer address and for DNS, then the rest (Oracle Cloud (OKE)). Explicitly given settings are recorded as overrides and carried over by later runs.

Three decisions shaped the profiles:

  • Registry trust. The in-cluster registry sits on a fixed ClusterIP (10.96.0.50:5000) and serves TLS only, with a certificate from the platform CA whose SAN is that address, plus one generated credential. kpack has no setting for a private CA and every step of a build talks to the registry, so the server runs an admission webhook that mounts the cluster's trust bundle into each container of a build pod and sets SSL_CERT_FILE (every step is a Go binary); the kpack controller and BuildKit get the same bundle, and a DaemonSet writes the CA where each node's runtime reads it (containerd's certs.d, CRI-O's certs.d). TLS-only matters: go-containerregistry treats RFC 1918 addresses as insecure-eligible and races plain HTTP against HTTPS, and only a registry without a plaintext listener makes that race always end on HTTPS.
  • Certificates per host. Each hostname that needs a certificate has its own Certificate object owned by the App - custom domains always, the project hostname only when no platform wildcard serves it - so a domain whose DNS is not ready never blocks the others. With a DNS provider, one wildcard certificate is ingress-nginx's default certificate and project Ingresses carry none.
  • Multi-arch builders. The prebuilt Paketo builder images are amd64-only, which breaks on Apple Silicon. The builder is assembled from the individual Paketo buildpackages (Go, Node.js, Java, Python, Ruby, .NET, web servers, Procfile), which ship amd64 and arm64, so kpack resolves the right architecture per node.

The App controller

Projects are App custom resources living in their own namespace (app-<name>) with everything derived from them, which gives garbage collection and per-project RBAC for free. The controller (controller-runtime, kubebuilder layout) reconciles an App into:

  • a kpack Image for the source (Git or uploaded archive), with a build cache volume, or, for the dockerfile strategy, a rootless BuildKit Job per source revision (an init container fetches the archive or clones the revision, the build pushes to the registry with a layer cache; the digest travels in the container's termination message). Images go to apps/<workspace id>/<slug> in the registry, the workspace's id rendered in base36 (25 lowercase characters): ids never change and slugs are per workspace, so two workspaces with a project of the same name never share a repository, its tags, the build cache or the builder;
  • the claims and mounts of the project's Volume resources (single-instance volumes force one replica and a Recreate rollout) and the Secret <app>-bindings with the config vars of attached resources;
  • one Deployment per process type, sized (default 1 CPU / 512 MiB), web running the image entrypoint and other types /cnb/process/<type>, with PORT injected and the config var Secret mounted as environment;
  • a Service per process with a port and an Ingress with a cert-manager certificate for web.

Status is derived from the kpack Image or build Job (build progress and failures), the Deployments (instances updated and ready) and the pods behind them (instances that cannot start, with the container's reason). A release is recorded whenever the running image or the configuration hash changes; its config vars are snapshotted into a Secret, and a rollback request (annotation) restores that snapshot before pinning the release's build. Status updates use optimistic locking so a reconcile working from a stale cache never overwrites a newer state.

The server

shpyrd-server runs the controller manager (with leader election) and an HTTP API in one process. The API serves the dashboard (projects, builds, logs streamed as NDJSON with Heroku-style instance names, Prometheus-backed metrics per process and cluster capacity, write-only config vars, scale/rollback/deploy/destroy) and receives source archives from shpyrd deploy, which it stores by SHA-256 on a persistent volume and serves to kpack at a cluster-internal URL.

Authentication is a single admin token generated at install time, accepted as Authorization: Bearer or X-Shpyrd-Token (the API server's service proxy, through which the CLI uploads, strips Authorization).

The CLI

Apart from the upload, the CLI never talks to the server: it reads and writes App objects, config var Secrets and pod logs directly through your kubeconfig, like flux or kubectl argo rollouts. shpyrd deploy archives the committed tree of the current directory (git archive HEAD), applies shpyrd.yaml, then follows the kpack build step by step and the rollout until the release is running.

Repository layout

cmd/shpyrd            CLIcmd/shpyrd-server     server: API + App controller + the embedded applicationsapi/v1alpha1          App CRD types (kubebuilder layout; `make generate`)internal/controller   App reconcilerinternal/cli          CLI commandspkg/install           runlevel installer: embedded Kustomize + Helm SDK + server-side applypkg/api               HTTP API, source store, Prometheus clientpkg/kind              kind cluster provisioningpkg/localca           development root CApkg/configvars        config vars (names + metadata; values are write-only)deploy/               components and profiles embedded in the binarydesign/ui             the component library and its gallery, shared by the applications and the siteapps/                 the console and the workspace applications (Next, static files, embedded by pkg/ui)examples/hello        example project (Go, web + worker)rfcs/                 design documents

On this page

  • Components
  • The installer
  • The App controller
  • The server
  • The CLI
  • Repository layout
  • Docs

shpyrd is open source under MPL-2.0, and in beta.