oxid/ docs

// security

What you are actually running

The daemon mounts the Docker socket. Anything that can reach its API with a valid token can build and run containers on the host, which is the same reach as root. That is inherent to the job — a control plane that starts containers needs to start containers — and it is why every control below exists. This page explains what each one protects and why it defaults the way it does; PRODUCTION.md has the deployment checklist itself.

Where the API is published

The daemon always binds 0.0.0.0:8080 inside its container. What decides exposure is the publish, and the shipped compose publishes 127.0.0.1:8080:8080 — the control API is reachable from the host and nowhere else.

The shipped compose publishes on every interface, because a preview-environment server nobody can reach is not one: the team's developers point their CLI at it and the Git host has to deliver pushes to it. The port is not the security boundary — the credential is.

  • Every /api/v1/* route requires a bearer token.
  • The daemon refuses to start on a non-loopback bind with no OXID_API_TOKEN, overridable only by an explicit OXID_ALLOW_OPEN_API=1.
  • Credentials carry a role, a scope and an expiry.
  • The one pre-auth endpoint that hands over a token is disabled in that compose — see below.

What the open port does not give you is confidentiality: a bearer token over plain HTTP is readable by anything on the path. Terminate TLS in front of the daemon before it crosses a network you do not control, or narrow the publish back to 127.0.0.1:8080:8080 and reach it over SSH.

The bootstrap token endpoint

GET /api/v1/setup/token hands the auto-generated master token to a caller before authentication — it exists so the onboarding wizard can finish without asking an operator to read container logs. That is a real hole if it answers the wrong caller, so who it answers is a setting:

OXID_BOOTSTRAP_TOKEN_ACCESSHands the token to
loopback (default)Only a caller from the daemon's own host.
anyAny caller that can reach the port. Correct only when the port itself is private.
offNobody. Retrieve the token from the file or the logs instead.

The default withholds because the daemon cannot tell a safe deployment from an exposed one: containerized, it is always bound to 0.0.0.0 and every caller arrives from the bridge gateway, so loopback detection alone would answer the whole LAN. The operator knows what the publish looks like; the daemon does not.

The shipped docker-compose.yml sets off, because it publishes on every interface. loopback would not mean loopback there — a containerized daemon sees every caller arrive from the bridge gateway — and any would hand the master token to whoever reached the port first. The installer prints the token instead, which reaches exactly one person on the machine that ran it.

Giving people access

OXID_API_TOKEN is the master credential. Everything else is issued from it, is attributable in the audit trail, and says three things: what a person may do, where, and until when.

# a developer on one project, for a quarter:
oxid token create juan --project 1 --role developer --expires-in 90d

# read-only, for whoever just watches the previews:
oxid token create ana --project 1 --role viewer

# another operator, who can then issue access themselves:
oxid token create devops2 --role admin

oxid token list                # role, status, scope, expiry
oxid token suspend <id>        # reversible — someone on leave
oxid token resume <id>
oxid token revoke <id>         # permanent

The roles

RoleCanCannot
viewerSee projects, environments, logs, history.Change anything at all.
developerDeploy, roll back, pause, wake, destroy environments.Read or write secrets; change the project.
maintainerIts projects' secrets, settings and branch filter; delete the project.Anything node-wide.
adminThe node: stats, infra, backups, and issuing access to others.Rotate the master key or read the webhook secret — those stay with the master credential.

Roles are cumulative — each one can do everything the one below it can, and a test pins that, because a matrix where a middle role is missing something a lower one has is a bug nobody finds by reading it.

The line worth understanding is between developer and maintainer: shipping the code and reading the credentials that code is handed are different powers, and most of a team only needs the first.

Scope, and what a denial reveals

  • Another project's routes answer 404, not 403 — existence is not disclosed to a credential with no business knowing. Scope is checked before role for exactly this reason: answering “your role is too low” would confirm the project exists.
  • Everything else answers 403 with the reason — expired, suspended, or the role needed. A developer told “your access expired” opens a ticket; one told 404 files a bug against the daemon.
  • A scoped credential can never act node-wide, whatever its role. “Admin of project 3” is not an admin of the server, and a global secret — injected into every project's deploys — counts as node-wide.

Omitting --role reproduces the power a token had before roles existed: maintainer when scoped, admin when not. An upgrade never quietly removes a permission — which means least privilege is something you ask for. oxid token create prints the role it granted.

Webhook verification

Push webhooks are verified by HMAC-SHA256 against OXID_WEBHOOK_SECRET (GitHub, Gitea, Gogs) or by token echo (GitLab). Webhooks are rejected entirely while the secret is unset — there is no unauthenticated deploy path.

A verified push is matched to a project by exact repository path, never by substring, so you/app-staging cannot trigger a deploy of you/app. The pushing user is recorded as the audit operator.

A push is answered 202 queued and deployed off a persisted queue. Providers abandon a delivery in seconds — GitHub at 10, with no retry for push events — while a real first build takes far longer.

Secrets at rest, and rotating the key

Project and branch secrets are encrypted with AES-GCM under a master key at {data}/secret.key (0600), generated on first start unless OXID_MASTER_KEY supplies one. Values are never echoed back by the API or printed in logs.

oxid rotate-key   # re-encrypts every secret under a fresh key, zero downtime

Rotation runs in a BEGIN IMMEDIATE transaction. With a deferred one the write lock is only taken at the first write, leaving a window in which a secret written under the old key survives the swap and becomes undecryptable.

Back up {data} as a unit: the database and secret.key belong together, and a database restored without its key has unreadable secrets. oxid backup produces both in one archive.

Restore is opt-in, and never in place

POST /api/v1/backup/restore accepts an upload only when the daemon runs with OXID_ALLOW_RESTORE=1. An accepted archive is staged and applied on the next startup — the live database is never overwritten under running requests, and a restore that fails validation leaves the running system untouched.

Rate limiting and transport

ControlNotes
OXID_RATE_LIMIT_PER_SECOND + _BURSTPer-client-IP token bucket on protected routes. Both are required together; setting one alone does nothing. Public setup routes get their own bucket, so hammering them cannot exhaust an operator's.
OXID_TLS_CERT / OXID_TLS_KEYWhen both are set the daemon serves HTTPS itself. Otherwise terminate at Traefik or another proxy — do not publish plain HTTP beyond a trusted network.
Heartbeat endpointDeliberately unauthenticated: Traefik calls it on every request to every environment. It only writes a coalesced timestamp feeding idle detection, and exposes nothing.

The audit trail

Every deploy, pause, wake, destroy, secret write and token action is recorded with its operator, and a failed deploy is recorded too — a broken Dockerfile leaves a build_failed environment, an audit event and an ERROR line rather than nothing at all.

oxid audit --limit 50
oxid audit --branch feat/login
oxid audit --json | jq '.[] | select(.action == "destroy")'

Trusting a node

A fleet node exposes its Docker API over the network, and a Docker socket is root on that machine — not "root inside a container", root: anything that can reach it can mount the host filesystem and start a privileged container. mTLS bounds who can reach it. It does not bound what they can do once they have.

So the rule is short: an Oxid node must be a machine you would already trust the control plane on. Registering a remote endpoint with no TLS material is refused outright; OXID_ALLOW_INSECURE_NODES=1 is the explicit opt-in, for a network you control end to end.

Secrets never land on a node's disk. They stay encrypted on the control plane and are injected over the mTLS connection as container environment variables — which are readable with docker inspect on that node, exactly as they already are on a single machine. Worth saying rather than leaving to be discovered.

The node's certificate paths live in the database; the files do not. A restored backup brings back the rows and not the certificates, so back those up alongside secret.key.

Before you expose it

  • A long random OXID_API_TOKEN is set — and you are not using it for day-to-day work.
  • OXID_WEBHOOK_SECRET is set and matches what the Git host sends.
  • OXID_BOOTSTRAP_TOKEN_ACCESS matches your publish: any only while the port is private, otherwise loopback or off.
  • TLS terminates somewhere — the daemon or a proxy in front of it.
  • Rate limits are set if the port is reachable by anything untrusted.
  • Multi-node only: every node registered with TLS paths, on machines you already trust the control plane on.
  • OXID_ALLOW_RESTORE stays unset except while you are actually restoring.
  • Backups run, and you have restored one at least once.

Found a hole? SECURITY.md has the disclosure process and what is most interesting to report.