[HN Gopher] Ask HN: How do you keep system context from rotting ...
___________________________________________________________________
Ask HN: How do you keep system context from rotting over time?
Former SRE here, looking for advice. I know there are a lot of
tools focused on root cause analysis after things break. Cool, but
that's not what's wearing me down. What actually hurts is the
constant context switching while trying to understand how a system
fits together, what depends on what, and what changed recently. As
systems grow, this feels like it gets exponentially harder. Add
logs and now you've created a million new events to reason about.
Add another database and suddenly you're dealing with subnet
constraints or a DB choice that's expensive as hell, and no one
noticed until later. Everyone knows their slice, but the full
picture lives nowhere, so bit rot just keeps creeping in. This
feels even worse now that AI agents are pushing large amounts of
code and config changes quickly. Things move faster, but shared
understanding falls behind even faster. I'm honestly stuck on how
people handle this well in practice. For folks dealing with real
production systems, what's actually helped? Diagrams, docs, tribal
knowledge, tooling, something else? Where does it break down?
Author : kennethops
Score : 3 points
Date : 2026-01-20 16:40 UTC (6 hours ago)
___________________________________________________________________
(page generated 2026-01-20 23:01 UTC)