What a webshell taught a small hosting shop: rebuild, never clean

Isometric illustration of a film-wrapped server being lifted out of a rack by straps while a new server slides into the empty slot.

In short. Found by accident. Contained every listener, not just the web port. Preserved evidence before touching anything. Sealed it with a filesystem snapshot and a protected backup. Rebuilt on a new host with every credential treated as disclosed. Then changed the platform so the next one has nowhere to land.

Names, addresses and dates are left out on purpose. Durations are approximate. Nothing here identifies the site, the host or the people involved.

Timeline

  1. A routine request. Someone asked for editor accounts on a content-management site hosted on a shared control-panel VM. Looking at the site to add them, the entry file was not what it should have been.
  2. What was there. A search-engine cloaking injector in the site's front controller, a password-protected webshell disguised as the login page, and several droppers that could reinstall the others.
  3. How it got in. The best later account is a pre-authentication flaw in the platform, exploited within days of the fix being released. Patching was weekly. Weekly was too slow.
  4. How long it had been there. File timestamps suggested a year. Timestamps are set by whoever writes the file, and attackers back-date. Search-engine evidence showed about seven weeks of cloaked spam.
  5. The first rebuild was reinfected within hours. Two loaders had survived a sweep that reported clean.
  6. The second rebuild was on a new host, from a known-good source, with no file copied from the old one.

The three mistakes that made it worse

What containment means

Containment is not "take the website down". It is every listener on the host, including the control panel, the mail ports and SSH, because a webshell on a shared host is a foothold on all of them. Every credential on the host is treated as disclosed: database passwords, API keys, mail passwords, and the panel's own admin accounts. No new accounts are created on a host under investigation.

Access for the response itself was by an eight-hour SSH certificate from the certificate authority, so the response session was itself bounded and logged.

Evidence before anything else

The runbook that came out of this has five stages, and each refuses to start until the previous one is proven:

flowchart LR
  D["Detect<br/>found during a routine account request"] --> C["Contain<br/>every listener, every credential disclosed"]
  C --> E["Preserve<br/>export disks and RAM, byte-compare, checksums"]
  E --> S["Seal<br/>read-only files, CephFS snapshot, protected PBS backup"]
  S --> R["Rebuild<br/>new host, never clean in place"]
  R --> P["Prevent<br/>no internet-facing panels, web user owns no code, patch in a day, default-deny egress"]
  X["What made it worse<br/>read-only docroot pinned the malware<br/>scan said clean, two loaders survived<br/>panel stayed online during containment"] -.-> R
  1. Plan. Write down what will be exported and where it will go.
  2. Export. Every disk snapshot and the RAM state. The normal guest backup captures only the current disk, with no snapshots and no memory, so it is not evidence.
  3. Verify. Byte-compare each export against its source and write checksums.
  4. Seal. Make the exports read-only, take a filesystem snapshot of the export directory, and take a protected backup that cannot be pruned.
  5. Destroy. Only now is the original guest removed.

Why "clean" is the wrong goal

A scan that reports clean tells you what the scanner knows. Two loaders survived one here. The rule now is that a scan that cannot finish is a failure, and a clean result is not a reason to keep a host. Rebuild on a new host, from the repository and from data that was restored, not copied.

Controls that came out of it

FAQ

Would a managed WordPress host have prevented this? It would have patched faster, and patch speed was the root cause. It would not have provided the evidence trail or the per-tenant isolation that the replacement platform does.

Is a read-only document root ever right? Only when updates are deployed by replacing the whole tree from a build, so nothing ever needs to write inside it. Making a self-updating platform read-only is the worst of both.

What about notifying people? The people whose site it was were told the same day, with what was known and what was not. The seven-week cloaking window was reported to them when the evidence supported it, not before.

Related