CephFS subvolumes under LXC: snapshots your backups can see

In short. Databases stay on RBD. Bulk user data (files, mail blobs) goes on a CephFS subvolume with a quota, mounted on every node and bind-mounted into an unprivileged container. Snapshots run on the subvolume path every six hours. A nightly file-level backup to PBS is taken from a snapshot. The container backup stays small on purpose.
The question that starts this
"Why is the Nextcloud container backup 400 MB when users have 300 GB of files in it?"
Because the files are not in the container. They are on a bind mount, and vzdump does not follow bind mounts. The container backup is the runtime. The data needs its own path to the backup server, and this post is that path.
The split
| Data | Where | Why |
|---|---|---|
| Container root filesystem | RBD | Disposable; rebuilt from the template |
| Databases, key-value stores | RBD | Need block semantics and their own consistency |
| User files, mail blobs, uploads | CephFS subvolume | Shared, quota-managed, snapshot-friendly |
| ISOs and templates | A separate CephFS subvolume | Its own quota so it cannot eat the file budget |
One CephFS per cluster, with one active metadata server and two standby. Isolation is by subvolume group per application and subvolume per tenant, each with a quota. Never a second filesystem: that doubles the metadata servers and the failure modes for no isolation you could not get from a subvolume.
Mounting on every node
The container can start on any node, so every node must have the subvolume mounted at the same path before guests start. Each subvolume gets a systemd mount unit with three properties:
- A path-scoped cephx client. The key can read and write that subvolume path and nothing else. Inside the container,
dfused to show the whole cluster; with a path-scoped mount it shows the quota. - Ordered before the guest service. The mount unit is a dependency of guest startup, so a container never starts with an empty directory where its data should be.
- An immutable mountpoint.
chattr +ion the empty mountpoint means that if the mount fails, writes fail. Without this, a failed mount gives the application a perfectly writable empty directory on the root disk, and you find out at the next restore.
Bind-mounting into an unprivileged container
The container gets mp0: /mnt/cephfs/<group>/<subvolume>,mp=/data and nothing else about Ceph. It does not hold a Ceph key, it does not know the monitors' addresses, and it cannot see other tenants' subvolumes. The UID map is the standard unprivileged one; the subvolume is owned by the mapped UID.
This is the part that makes the container disposable. Destroy it, create a new one from the template with the same bind mount, and the application comes back with its data.
flowchart TB
subgraph Ceph["Ceph cluster"]
FS["One CephFS per cluster<br/>1 active MDS, 2 standby"]
SVG["Subvolume group per app"]
SV1["Subvolume: user files, quota"]
SV2["Subvolume: mail blobs, quota"]
RBD["RBD pool<br/>container rootfs and databases"]
FS --> SVG
SVG --> SV1
SVG --> SV2
end
subgraph Host["Every PVE node"]
MNT["systemd mount per subvolume<br/>path-scoped cephx, chattr +i, before pve-guests"]
end
SV1 --> MNT
SV2 --> MNT
MNT -- "bind mount mp0, unprivileged" --> LXC["LXC: app runtime only<br/>disposable"]
RBD --> LXC
SV1 -- "snap-schedule on the subvolume path<br/>6h keep 3, daily keep 14" --> SNAP[".snap"]
SNAP -- "nightly pxar from a snapshot" --> PBS["PBS, then cross-site replica"]
LXC -- "vzdump: skips bind mounts" --> PBSSnapshots
CephFS snapshots are cheap and instant. The schedule is:
- every six hours, keep three
- daily, keep fourteen
Two things to get right:
- Schedule the subvolume path, not the inner directory. A subvolume is a directory containing a UUID-named directory where the data lives. Scheduling the inner UUID directory reports "active" and creates nothing. Schedule the subvolume itself and verify that
.snapfills up. - Verify that snapshots appear. The check is a timer that lists
.snapand alerts if the newest is older than the interval. Schedulers that report success while doing nothing are a recurring theme in this series.
A thirty-minute recovery point for files is a snapshot, not a backup. If the cluster dies, the snapshots die with it.
The backup that sees the data
A nightly job on the backup server takes a pxar archive of the most recent snapshot directory and stores it in the tenant's namespace next to the container backup. Taking it from a snapshot rather than the live tree means the archive is internally consistent even if users are writing during the run.
The archive then rides the normal cross-site sync, so the off-site copy includes the files. That sync has its own post.
The restore drill
Three restores, documented and rehearsed:
- One file. Browse the pxar archive in the PBS UI, download the file. Or
cpit from.snapif the snapshot is still there, which is faster. - One subvolume. Restore the pxar into a fresh subvolume, point a scratch container at it, check the application sees what you expect, then swap the bind mount.
- From the other site. Same as 2, from the replica namespace, on the other cluster.
The drill found a prune running during a restore and a verification script whose pipe swallowed a failure. Both fixed; both would have stayed hidden without the drill.
Why not the alternatives
- A Ceph client inside the container. It would need a key and monitor access, which is exactly the blast radius a tenant container must not have.
- RBD for files. Block devices cannot be mounted on more than one node at once without a cluster filesystem on top, and they cannot be quota-shared between applications.
- Backing up from inside the application. Application exports are slow, inconsistent under load, and unique per application. A filesystem snapshot is the same for every application.
FAQ
Is this safe for mail stores? The mail server here keeps its index on RBD and its blob store on CephFS. The blob store is append-mostly, so a snapshot is consistent enough; the index can be rebuilt from blobs.
What about Ceph's own cephfs-mirror?
It mirrors snapshots to another CephFS. It is a good fit when both sites run Ceph of similar size; here the second site's cluster was smaller and PBS already handled the transport and verification.
Does the evidence-preservation runbook use the same primitive? Yes. Sealing a compromised guest ends with a CephFS snapshot of the exported artefacts and a protected PBS backup. Same tools, different reason.