Theo's Corner
dev / irl / thoughts
← Back
tech

Things that surprised me about running infrastructure

Theo|Jul 2026|~4 min read

You learn a lot of things about running actual infrastructure that you don't get from tutorials or docs. Here's what genuinely surprised me once I was actually doing it for real.

Most outages aren't dramatic

You picture server outages as big, obvious events — everything on fire, alarms going off. In reality most of them are quiet and almost boring. A disk fills up slowly over weeks. A cert renewal silently fails and nobody notices until it expires. A config drifted out of sync months ago and only just caused a problem. The dramatic outage is rare. The slow, quiet failure is common.

Documentation for yourself matters more than you'd think

I used to think documentation was mainly for other people. Then I came back to a server I'd set up six months earlier and had genuinely no memory of why I'd made certain decisions. Writing things down isn't for some imagined future colleague — it's for future you, who will absolutely forget.

Every time I've thought "I'll remember this, no need to write it down" — I did not remember it. Every single time.

The boring stuff is what actually matters

Backups, monitoring, firewall configs, log rotation — none of it is interesting to set up. All of it is what determines whether a bad day becomes a catastrophe or just an inconvenience. The exciting parts of infrastructure work are a small fraction of what actually keeps things running.

Small scale problems don't disappear at larger scale — they just change shape

I assumed running more servers would mean the same problems, just more of them. Actually the problems change character entirely. Coordination becomes the issue rather than any individual server. Managing consistency across multiple machines is a genuinely different problem to managing one machine well.

Users find bugs you'd never find yourself

No amount of testing I do myself surfaces the edge cases that real users find within days. Someone will do something you never anticipated and it'll break in a way that reveals an assumption you didn't know you were making. This happens constantly and it's honestly one of the most useful parts of running something real rather than a personal project nobody else touches.

You get calmer about outages over time, not more anxious

The first time something broke in production I was genuinely stressed. Now, after enough incidents, there's a much calmer process — check the logs, isolate the problem, fix it, document what happened. Experience doesn't make outages less likely. It makes them less frightening, which is arguably more valuable.