Proxmox: automated install with an answer file and PXE

In short. The official proxmox-auto-install-assistant turns the stock ISO into one that fetches a TOML answer file over HTTP and installs without a keyboard. The answer file is rendered from a template per node: ZFS mirror boot, static management address, filter on the first NIC's MAC. Boot it with iDRAC virtual media from an HTTP/1.1 server, or PXE. Then let configuration management own the host, because the next reinstall overwrites anything you fixed by hand.

Goal: install Proxmox VE on a bare Dell server without touching it

You have a rack of R630-class nodes at a site you do not visit often, an iDRAC on each, and a jumphost on the management VLAN. You want a reinstall to be a command, not a trip. This uses Proxmox VE 9.2 and the assistant from the same repository.

What you need

Steps

  1. Write the answer template. This is Jinja, rendered per node by a small script from the inventory. Keys are the assistant's; the country code must be lower case.
[global]
keyboard = "en-us"
country = "au"
fqdn = "{{ hostname }}.sitea.internal"
mailto = "hostmaster@example.net"
timezone = "Etc/UTC"
root_password_hashed = "{{ root_hash }}"
root_ssh_keys = ["{{ ops_pubkey }}"]
reboot_on_error = false

[network]
source = "from-answer"
cidr = "{{ mgmt_ip }}/24"
gateway = "10.1.1.254"
dns = "10.1.1.254"
filter.ID_NET_NAME_MAC = "*{{ mgmt_nic_mac_nocolons }}"

[disk-setup]
filesystem = "zfs"
zfs.raid = "raid1"
zfs.ashift = 12
zfs.arc_max = 2048
filter.ID_MODEL = "{{ boot_ssd_model }}*"

[first-boot]
source = "from-url"
url = "http://10.1.1.199/pve/first-boot.sh"
ordering = "network-online"
  1. Validate every rendered file before it goes anywhere near a server.
proxmox-auto-install-assistant validate-answer answers/pve01.toml
  1. Serve the answers. The installer sends an HTTP POST with the machine's system information (DMI serial, NIC MACs) and expects the TOML back. A forty-line handler on the jumphost matches the serial against the inventory and returns that node's file. Put it behind nginx on port 80.

  2. Prepare the ISO once; it is the same ISO for every node because the answer comes from the URL.

proxmox-auto-install-assistant prepare-iso proxmox-ve_9.2-1.iso \
  --fetch-from http \
  --url http://10.1.1.199/answer \
  --output /srv/http/iso/proxmox-ve_9.2-1-auto.iso
  1. Boot it. With iDRAC, attach the ISO as virtual media over HTTP, set boot-once to the virtual DVD and power cycle. The iDRAC password is prompted for; do not put it on the command line.
racadm -r 10.1.1.21 -u root remoteimage -c -l http://10.1.1.199/iso/proxmox-ve_9.2-1-auto.iso
racadm -r 10.1.1.21 -u root set iDRAC.VirtualMedia.BootOnce 1
racadm -r 10.1.1.21 -u root set iDRAC.ServerBoot.FirstBootDevice VCD-DVD
racadm -r 10.1.1.21 -u root serveraction powercycle

For PXE, the jumphost's DHCP points at a boot loader that loads the installer kernel and an initrd built from the prepared ISO (the Proxmox installer boots from an initrd that contains the whole ISO, so the initrd is large and ramdisk_size on the kernel line must allow for it). The prepared ISO's own boot entry shows the exact kernel arguments to copy; use those rather than guessing.

  1. Before the first boot, shut the switch ports for every NIC except the management one, or leave the cables out. With all NICs up, the installer's answer fetch went out an interface with no route and failed with ENETUNREACH, and the node sat at a prompt nobody was watching.

  2. After the reboot, the first-boot script enables the ops key and nothing else. Everything the host needs, bonds, VLAN interfaces, repositories, keepalived, the CephFS mount units, comes from the configuration baseline, run from the jumphost against the new node.

sequenceDiagram
  participant O as Operator
  participant J as Jumphost 10.1.1.199
  participant I as iDRAC .21
  participant N as Node pve01
  O->>J: render answers, prepare ISO
  O->>I: racadm remoteimage, boot once
  I->>J: GET iso (HTTP/1.1, ranges)
  I->>N: boot virtual DVD
  N->>J: POST system info to /answer
  J-->>N: pve01.toml
  N->>N: ZFS mirror install, reboot
  N->>J: GET first-boot.sh
  O->>N: configuration baseline over SSH

Verify it worked

ssh root@10.1.1.51 'pveversion; zpool status rpool | grep -A3 mirror; ip -br addr show'

Expect PVE 9.2, a mirror-0 with two SSDs ONLINE, and exactly the management address until the baseline adds the rest. Then run the baseline and check ip -br addr again for bond1.5 with the storage address.

Gotchas

FAQ

Can I bake the answer into the ISO instead of fetching it? Yes, --fetch-from iso --answer-file pve01.toml makes one ISO per node. It works and avoids the HTTP handler, but you are then managing a per-node ISO and the MAC filter is doing less for you. The URL method scales to the third site without new ISOs.

Where does the root password hash come from? mkpasswd -m sha-512 on the jumphost, with the result stored in the secrets vault and injected at render time. The rendered answer files are not committed anywhere.

Why a first-boot script at all if the baseline does everything? It installs the ops key so the baseline can connect without a password, and that is the whole job. Anything more and the script becomes a second, undocumented baseline.

Related