CephFS subvolumes under LXC: snapshots your backups can see

Isometric illustration of a small translucent box on a wide slab, a folder shape inside it, and a stack of thin sheets behind.

In short. Databases stay on RBD. Bulk user data (files, mail blobs) goes on a CephFS subvolume with a quota, mounted on every node and bind-mounted into an unprivileged container. Snapshots run on the subvolume path every six hours. A nightly file-level backup to PBS is taken from a snapshot. The container backup stays small on purpose.

The question that starts this

"Why is the Nextcloud container backup 400 MB when users have 300 GB of files in it?"

Because the files are not in the container. They are on a bind mount, and vzdump does not follow bind mounts. The container backup is the runtime. The data needs its own path to the backup server, and this post is that path.

The split

DataWhereWhy
Container root filesystemRBDDisposable; rebuilt from the template
Databases, key-value storesRBDNeed block semantics and their own consistency
User files, mail blobs, uploadsCephFS subvolumeShared, quota-managed, snapshot-friendly
ISOs and templatesA separate CephFS subvolumeIts own quota so it cannot eat the file budget

One CephFS per cluster, with one active metadata server and two standby. Isolation is by subvolume group per application and subvolume per tenant, each with a quota. Never a second filesystem: that doubles the metadata servers and the failure modes for no isolation you could not get from a subvolume.

Mounting on every node

The container can start on any node, so every node must have the subvolume mounted at the same path before guests start. Each subvolume gets a systemd mount unit with three properties:

  1. A path-scoped cephx client. The key can read and write that subvolume path and nothing else. Inside the container, df used to show the whole cluster; with a path-scoped mount it shows the quota.
  2. Ordered before the guest service. The mount unit is a dependency of guest startup, so a container never starts with an empty directory where its data should be.
  3. An immutable mountpoint. chattr +i on the empty mountpoint means that if the mount fails, writes fail. Without this, a failed mount gives the application a perfectly writable empty directory on the root disk, and you find out at the next restore.

Bind-mounting into an unprivileged container

The container gets mp0: /mnt/cephfs/<group>/<subvolume>,mp=/data and nothing else about Ceph. It does not hold a Ceph key, it does not know the monitors' addresses, and it cannot see other tenants' subvolumes. The UID map is the standard unprivileged one; the subvolume is owned by the mapped UID.

This is the part that makes the container disposable. Destroy it, create a new one from the template with the same bind mount, and the application comes back with its data.

flowchart TB
  subgraph Ceph["Ceph cluster"]
    FS["One CephFS per cluster<br/>1 active MDS, 2 standby"]
    SVG["Subvolume group per app"]
    SV1["Subvolume: user files, quota"]
    SV2["Subvolume: mail blobs, quota"]
    RBD["RBD pool<br/>container rootfs and databases"]
    FS --> SVG
    SVG --> SV1
    SVG --> SV2
  end
  subgraph Host["Every PVE node"]
    MNT["systemd mount per subvolume<br/>path-scoped cephx, chattr +i, before pve-guests"]
  end
  SV1 --> MNT
  SV2 --> MNT
  MNT -- "bind mount mp0, unprivileged" --> LXC["LXC: app runtime only<br/>disposable"]
  RBD --> LXC
  SV1 -- "snap-schedule on the subvolume path<br/>6h keep 3, daily keep 14" --> SNAP[".snap"]
  SNAP -- "nightly pxar from a snapshot" --> PBS["PBS, then cross-site replica"]
  LXC -- "vzdump: skips bind mounts" --> PBS

Snapshots

CephFS snapshots are cheap and instant. The schedule is:

Two things to get right:

A thirty-minute recovery point for files is a snapshot, not a backup. If the cluster dies, the snapshots die with it.

The backup that sees the data

A nightly job on the backup server takes a pxar archive of the most recent snapshot directory and stores it in the tenant's namespace next to the container backup. Taking it from a snapshot rather than the live tree means the archive is internally consistent even if users are writing during the run.

The archive then rides the normal cross-site sync, so the off-site copy includes the files. That sync has its own post.

The restore drill

Three restores, documented and rehearsed:

  1. One file. Browse the pxar archive in the PBS UI, download the file. Or cp it from .snap if the snapshot is still there, which is faster.
  2. One subvolume. Restore the pxar into a fresh subvolume, point a scratch container at it, check the application sees what you expect, then swap the bind mount.
  3. From the other site. Same as 2, from the replica namespace, on the other cluster.

The drill found a prune running during a restore and a verification script whose pipe swallowed a failure. Both fixed; both would have stayed hidden without the drill.

Why not the alternatives

FAQ

Is this safe for mail stores? The mail server here keeps its index on RBD and its blob store on CephFS. The blob store is append-mostly, so a snapshot is consistent enough; the index can be rebuilt from blobs.

What about Ceph's own cephfs-mirror? It mirrors snapshots to another CephFS. It is a good fit when both sites run Ceph of similar size; here the second site's cluster was smaller and PBS already handled the transport and verification.

Does the evidence-preservation runbook use the same primitive? Yes. Sealing a compromised guest ends with a CephFS snapshot of the exported artefacts and a protected PBS backup. Same tools, different reason.

Related