Proxmox: bind-mount CephFS into an unprivileged container

In short. One subvolume per tenant, under a subvolume group per application, with a quota. A cephx client that can only see that path. A systemd mount unit on every node, ordered before pve-guests, on an immutable mount point. Then pct set --mp0 with shared=1, a chown to the mapped UID, and the container sees /data with the right size and nothing else about Ceph.

Goal: a container's data on CephFS without a Ceph client in the container

You run an unprivileged container (files, mail, uploads) and want its bulk data on CephFS so the container stays disposable and the data has snapshots. This is Proxmox VE 9.2 with Ceph Squid 19.2; the subvolume commands are the same on Reef.

What you need

Steps

  1. Create the subvolume group and the subvolume with its quota, then read back its real path. The path includes a UUID; the mount unit uses it.
ceph fs subvolumegroup create cephfs files
ceph fs subvolume create cephfs tenant7 --group_name files --size 322122547200
ceph fs subvolume getpath cephfs tenant7 --group_name files
# /volumes/files/tenant7/9a1c2d3e-0000-4000-8000-000000000001
  1. Create a cephx client scoped to that path. It can read and write under the subvolume and nothing above it.
ceph fs authorize cephfs client.files-tenant7 \
  /volumes/files/tenant7/9a1c2d3e-0000-4000-8000-000000000001 rw
ceph auth get-key client.files-tenant7 > /etc/ceph/files-tenant7.secret
chmod 600 /etc/ceph/files-tenant7.secret

Copy the secret file to the other two nodes with the same mode. It is a secret; it goes through the vault, not the repository.

  1. On every node, make the mount point and freeze it.
mkdir -p /mnt/cephfs/files/tenant7
chattr +i /mnt/cephfs/files/tenant7
  1. On every node, the mount unit. The unit name is the path with dashes; the ordering against pve-guests.service is what stops a container starting on an empty directory.
# /etc/systemd/system/mnt-cephfs-files-tenant7.mount
[Unit]
Description=CephFS subvolume files/tenant7
After=network-online.target
Wants=network-online.target
Before=pve-guests.service

[Mount]
What=10.1.5.51,10.1.5.52,10.1.5.53:/volumes/files/tenant7/9a1c2d3e-0000-4000-8000-000000000001
Where=/mnt/cephfs/files/tenant7
Type=ceph
Options=name=files-tenant7,secretfile=/etc/ceph/files-tenant7.secret,fs=cephfs,_netdev

[Install]
RequiredBy=pve-guests.service
systemctl daemon-reload
systemctl enable --now mnt-cephfs-files-tenant7.mount
df -h /mnt/cephfs/files/tenant7

df should show 300G, the quota, because the mount root sits inside the quota'd directory. If it shows the whole cluster, the mount path or the key scope is wrong.

  1. Fix ownership for the unprivileged UID map. Container UID 0 is host UID 100000, so www-data (33) is host UID 100033.
chown 100033:100033 /mnt/cephfs/files/tenant7
  1. Attach it to the container with shared=1, which is honest only because step 4 ran on every node.
pct set 101 --mp0 /mnt/cephfs/files/tenant7,mp=/data,shared=1
pct start 101
flowchart TB
  FS["CephFS<br/>1 active MDS, 2 standby"] --> G["subvolume group: files"]
  G --> SV["subvolume: tenant7<br/>quota 300G"]
  SV -- "client.files-tenant7<br/>rw on this path only" --> M["every node<br/>/mnt/cephfs/files/tenant7<br/>mount unit, chattr +i"]
  M -- "mp0 shared=1<br/>UID +100000" --> CT["CT 101 unprivileged<br/>/data, no Ceph key"]

Verify it worked

From the host:

pct exec 101 -- df -h /data
pct exec 101 -- su -s /bin/sh www-data -c 'touch /data/.write-test && ls -l /data/.write-test'
pct exec 101 -- ls /etc/ceph 2>&1

Expect /data at 300G, a file owned by www-data, and no /etc/ceph inside the container. On a different node, ls -l /mnt/cephfs/files/tenant7/.write-test shows the file owned by UID 100033, which proves the path is shared and the mapping is right.

Gotchas

FAQ

Why not give the container a Ceph key and mount CephFS inside it? A Ceph client needs the monitors reachable, a key on the container's root disk and, for an unprivileged container, capabilities it should not have. The bind mount gives the container a directory and nothing else. If the container is compromised, the attacker has a directory with a quota.

Can two containers share one subvolume? Yes, if the application tolerates it. Both get the same mp0 and the UID mapping has to agree. Separate tenants never share a subvolume; that is what the quota and the path-scoped key are for.

What happens if the mount fails on one node? The container cannot start there, because pve-guests requires the mount unit and the mount point is immutable. The HA manager moves on to a node where it did mount. That is a loud failure, and loud is the design.

Related