Proxmox: bind-mount CephFS into an unprivileged container
In short. One subvolume per tenant, under a subvolume group per application, with a quota. A cephx client that can only see that path. A systemd mount unit on every node, ordered before
pve-guests, on an immutable mount point. Thenpct set --mp0withshared=1, achownto the mapped UID, and the container sees/datawith the right size and nothing else about Ceph.
Goal: a container's data on CephFS without a Ceph client in the container
You run an unprivileged container (files, mail, uploads) and want its bulk data on CephFS so the container stays disposable and the data has snapshots. This is Proxmox VE 9.2 with Ceph Squid 19.2; the subvolume commands are the same on Reef.
What you need
- A CephFS named
cephfswith one active MDS and two standby, on a three-node cluster. - Root on every node and a way to put the same file on all three.
- A container ID (
101), an application name for the group (files) and a tenant name for the subvolume (tenant7). - The UID the application runs as inside the container (
www-data, UID 33, in this example).
Steps
- Create the subvolume group and the subvolume with its quota, then read back its real path. The path includes a UUID; the mount unit uses it.
ceph fs subvolumegroup create cephfs files
ceph fs subvolume create cephfs tenant7 --group_name files --size 322122547200
ceph fs subvolume getpath cephfs tenant7 --group_name files
# /volumes/files/tenant7/9a1c2d3e-0000-4000-8000-000000000001
- Create a cephx client scoped to that path. It can read and write under the subvolume and nothing above it.
ceph fs authorize cephfs client.files-tenant7 \
/volumes/files/tenant7/9a1c2d3e-0000-4000-8000-000000000001 rw
ceph auth get-key client.files-tenant7 > /etc/ceph/files-tenant7.secret
chmod 600 /etc/ceph/files-tenant7.secret
Copy the secret file to the other two nodes with the same mode. It is a secret; it goes through the vault, not the repository.
- On every node, make the mount point and freeze it.
mkdir -p /mnt/cephfs/files/tenant7
chattr +i /mnt/cephfs/files/tenant7
- On every node, the mount unit. The unit name is the path with dashes; the ordering against
pve-guests.serviceis what stops a container starting on an empty directory.
# /etc/systemd/system/mnt-cephfs-files-tenant7.mount
[Unit]
Description=CephFS subvolume files/tenant7
After=network-online.target
Wants=network-online.target
Before=pve-guests.service
[Mount]
What=10.1.5.51,10.1.5.52,10.1.5.53:/volumes/files/tenant7/9a1c2d3e-0000-4000-8000-000000000001
Where=/mnt/cephfs/files/tenant7
Type=ceph
Options=name=files-tenant7,secretfile=/etc/ceph/files-tenant7.secret,fs=cephfs,_netdev
[Install]
RequiredBy=pve-guests.service
systemctl daemon-reload
systemctl enable --now mnt-cephfs-files-tenant7.mount
df -h /mnt/cephfs/files/tenant7
df should show 300G, the quota, because the mount root sits inside the quota'd directory. If it shows the whole cluster, the mount path or the key scope is wrong.
- Fix ownership for the unprivileged UID map. Container UID 0 is host UID 100000, so
www-data(33) is host UID 100033.
chown 100033:100033 /mnt/cephfs/files/tenant7
- Attach it to the container with
shared=1, which is honest only because step 4 ran on every node.
pct set 101 --mp0 /mnt/cephfs/files/tenant7,mp=/data,shared=1
pct start 101
flowchart TB
FS["CephFS<br/>1 active MDS, 2 standby"] --> G["subvolume group: files"]
G --> SV["subvolume: tenant7<br/>quota 300G"]
SV -- "client.files-tenant7<br/>rw on this path only" --> M["every node<br/>/mnt/cephfs/files/tenant7<br/>mount unit, chattr +i"]
M -- "mp0 shared=1<br/>UID +100000" --> CT["CT 101 unprivileged<br/>/data, no Ceph key"]Verify it worked
From the host:
pct exec 101 -- df -h /data
pct exec 101 -- su -s /bin/sh www-data -c 'touch /data/.write-test && ls -l /data/.write-test'
pct exec 101 -- ls /etc/ceph 2>&1
Expect /data at 300G, a file owned by www-data, and no /etc/ceph inside the container. On a different node, ls -l /mnt/cephfs/files/tenant7/.write-test shows the file owned by UID 100033, which proves the path is shared and the mapping is right.
Gotchas
- Inside the container
dfshowed the whole cluster until the mount was path-scoped. A key withrwon/and a mount of a subpath reports the filesystem, not the quota. Scope the key and mount the subvolume path. chattr +ibefore the first mount, on the empty directory. After the mount the attribute is on the CephFS root of the subvolume, which is not what you want, andrmdiron the mount point will fail for the right reason.pct set --mp0replaces the entry. Putshared=1andmp=in the same command every time.- Unprivileged UID arithmetic is per container only if you have changed the default map in
/etc/pve/lxc/101.conf. With the default map, 100000 plus the in-container UID is always right. vzdumpskips bind mounts. The container backup does not contain/data. The data is backed up from a CephFS snapshot with the backup client, which the decision post covers.
FAQ
Why not give the container a Ceph key and mount CephFS inside it? A Ceph client needs the monitors reachable, a key on the container's root disk and, for an unprivileged container, capabilities it should not have. The bind mount gives the container a directory and nothing else. If the container is compromised, the attacker has a directory with a quota.
Can two containers share one subvolume?
Yes, if the application tolerates it. Both get the same mp0 and the UID mapping has to agree. Separate tenants never share a subvolume; that is what the quota and the path-scoped key are for.
What happens if the mount fails on one node?
The container cannot start there, because pve-guests requires the mount unit and the mount point is immutable. The HA manager moves on to a node where it did mount. That is a loud failure, and loud is the design.