Two roots of trust: Authentik for people, Vault for machines

Isometric illustration of two different vault doors side by side, a key on a hook between them and a sealed envelope on the floor.

In short. Authentik is the system of record for humans and fronts every application at the richest protocol it supports. Vault issues what machines need: per-consumer AppRole credentials, TLS from an offline root, eight-hour SSH certificates. Login is passkey first. Nothing in either root may depend on something that root protects, and an offline go-bag exists for the day both are down.

Why two roots

A person joins, changes role, forgets a password, loses a phone, leaves. A service is deployed, scaled, rotated, redeployed, deleted. Forcing both through one directory produces either service accounts with birthdays or people with API tokens. Two systems, each with one job, and a thin bridge between them.

The principal model

The protocol ladder

Each application is connected at the richest level it supports, in this order:

  1. OIDC with a groups claim. Authorisation follows group membership automatically.
  2. SAML where OIDC is not offered.
  3. SCIM for lifecycle, so a disabled person is disabled in the application too.
  4. LDAP or RADIUS outposts for devices and mail servers that speak nothing newer.

Two rules apply at every rung:

The hypervisor, backup server, source-of-record, monitoring, SIEM, file sharing and GitOps tools all take OIDC. Per-zone ingress runs the identity provider's proxy outpost for forward-auth in front of anything that has no login of its own.

Login policy

The identity provider's configuration is Terraform: brand, role-* and platform-* groups, providers, applications and policy bindings. Scripts that drive the REST API directly are banned, after one created an OAuth provider with no grant types and broke every login until someone noticed.

Vault as the machine root

flowchart TB
  H["People<br/>passkey first, TOTP fallback, no SMS"]
  M["Machines and automation"]
  AK["Authentik<br/>system of record for humans"]
  V["Vault CE<br/>root of trust for machines"]
  H --> AK
  M --> V
  AK -- "OIDC with groups claim, keyed on sub" --> APPS["Proxmox, PBS, NetBox, monitoring, SIEM, Nextcloud, Argo CD"]
  AK -- "proxy outpost forward-auth" --> ING["nginx ingress per zone"]
  AK -- "LDAP and RADIUS outposts" --> DEV["Mail and network gear"]
  AK -. "OIDC as a second login path only" .-> V
  V -- "AppRole per consumer<br/>wrapped SecretID, CIDR-bound" --> SVC["Services"]
  V -- "PKI: offline root, 5y intermediate<br/>90-day host certs" --> TLS["TLS everywhere"]
  V -- "SSH CA: 8h human, 1h automation" --> SSH["Hosts and jumphost"]
  GB["Offline go-bag<br/>breaks the recovery dependency loop"] -.-> V

What never goes in Vault: customer passwords, and application-specific app passwords. The latter have their own credential plane; rotating a person's identity-provider password does not revoke their app passwords, so offboarding has a separate step for them.

Why not Keycloak

It was the first choice and was dropped. Its claim to serve LDAP natively did not hold up for the devices that needed it, and the outpost model on the replacement did. The decision record is one paragraph and has saved several re-litigations.

Recovery and the dependency loop

Vault must depend on nothing it protects. The first design violated that twice:

The fix is an offline go-bag: the DNS token, the mesh VPN break-glass, the recovery shares (five shares, threshold three, encrypted to named holders) and the procedure, held outside every system they would be used to recover. The init playbook refuses to run until a custody register names the five holders.

The honest trade-off: on a small estate the Raft voters currently span the data centre and two home sites, so quorum depends on home internet links. The target is voters inside one site with asynchronous DR; the current layout is a budget compromise and is recorded as one.

FAQ

Is Authentik too heavy for a small team? It runs in two containers. The weight is in the configuration discipline, not the software.

Why not use a cloud identity provider and be done? Several of the things it fronts are out-of-band controllers and switches that must work when the internet does not. An identity provider you cannot reach during an outage is a second outage.

How are agent sessions handled? Automation that commits code signs with its own signing-only key and authenticates with a short-lived certificate, the same as any other machine principal. It never borrows a human's identity.

Related