Restart your machines
I had a small Ubuntu VM in Singapore with 214 days of uptime. I liked seeing that number in uptime -p, the way you like any long streak. Last month DNS on that box started acting strange. Lookups took seconds, sometimes they timed out, and I spent close to two hours reading resolver configs and container logs at an hour when I should have been asleep. I rebooted at midnight because I had run out of ideas. Everything came back clean in about ninety seconds.
The kernel on that machine was three versions behind. Ubuntu had been writing a reboot-required flag for weeks and I kept scrolling past it.
What piles up while a box stays on
Pending kernel updates are the obvious one. apt upgrade pulls down the new kernel image, writes it to disk, updates the bootloader, and then leaves you running whatever you booted into months ago. The patch notes you skimmed apply to a kernel that is not loaded. Livepatch covers part of this on machines that have it, and most of mine do not.
Stale file handles are quieter. When a shared library like glibc or openssl gets replaced, every process that already mapped the old copy keeps using it. Those processes are fine until they are not, and nothing in your monitoring will tell you about it. On Debian and Ubuntu, needrestart will list the services still holding deleted library files, which is a useful thing to run even when you have no intention of rebooting that day.
Memory is the slow one. Daemons leak a little, page cache fills up, a worker process gets stuck and nobody notices because the box has 8 GB and the leak is 40 MB a week. Six months later you are swapping and blaming the application.
Then there is config that only gets read at boot. Network settings, mount options in fstab, systemd units you edited and reloaded but never tested cold. The box is running on a configuration that exists only in memory, and the on-disk version is whatever you last typed. You find out which is which the next time the power flickers.
Where this has bitten me
A container on one of my nodes handles log shipping. It rotated its own logs, the old files got deleted, and the process kept the handles open. df reported the volume full while du showed plenty of space, which is a confusing twenty minutes if you have never seen it. Restarting the container released the handles and gave back about 6 GB.
A VM I use for small internal services got a network change earlier this year. I edited the config, applied it live, confirmed it worked, and moved on. In August the host went down briefly for a UPS test and that VM came back on the old settings, because I had edited one file and the boot path read another. The outage lasted four minutes. Finding out why took an hour, and it would have taken ten seconds if I had rebooted the VM the day I made the change.
The Proxmox host is where I am most careful. One of mine sat at over 300 days while I kept applying updates to it. A new pve-kernel had been installed three times over, so the running kernel and the installed kernel modules had drifted apart. A container refused to start with an error about an unsupported mount feature. The fix was a reboot into the kernel that matched the modules already on disk. The diagnosis was a Saturday morning.
What I check before I reboot
On Ubuntu guests I look at two things. Uptime tells me how long it has been, and the flag file tells me whether an update is waiting on a restart:
uptime -p
cat /var/run/reboot-required 2>/dev/null || echo "no reboot flagged"
apt list --upgradable 2>/dev/null | head -20
sudo needrestart -b
If something is flagged, I also want to know what will come back on its own. Anything that matters should be enabled:
systemctl list-unit-files --state=enabled --type=service
On Proxmox I check the same information from the web UI under Updates, then I move or stop the one VM that holds recordings before I touch the host. My setup is the multi-site lab I wrote about in how my homelab is wired, so there is always another node answering DNS while one is down. If you have a single node, the honest version of this routine is to pick an hour when nobody is home and accept the downtime.
The order I use
Small things first. Containers, then VMs, then the host if the update asks for it. One site at a time. I wait for Tailscale to reconnect and for the dashboard to load before I move to the next box, which is most of the time spent. A full round across three sites takes me about fifteen minutes, nearly all of it watching a login page.
From a Proxmox host shell:
pct reboot 116
qm reboot 102
I run one, confirm it is back on the tailnet, then run the next. Rebooting them in parallel saves four minutes and costs me the ability to tell which one broke.
Keeping risky work apart from routine work
For a long time I batched everything into one late-night session. Host package updates, guest updates, container image pulls, all of it, followed by a round of reboots. When something came back wrong I had four candidate causes and no way to separate them.
Now guest updates go in whenever during the week, with a reboot right after if the flag file is there. Those are low stakes, since a guest that fails to boot can be rolled back from a snapshot in under a minute. Host updates wait for a weekend morning when I can sit in front of the console, with IPMI or a physical screen reachable, and nothing else scheduled. I take a snapshot of the VMs I care about first, and I check that my backups ran that week before I type reboot on a hypervisor.
The guest side is easy enough to automate:
sudo apt update && sudo apt upgrade -y
[ -f /var/run/reboot-required ] && sudo reboot
I run that from a short script across the Ubuntu guests. The hosts stay manual, and I expect they always will.
Closing
Uptime is still fun to look at. I reboot when the flag is there.