[HN Gopher] Route leak incident on January 22, 2026
___________________________________________________________________
Route leak incident on January 22, 2026
Author : nomaxx117
Score : 97 points
Date : 2026-01-23 17:54 UTC (5 hours ago)
(HTM) web link (blog.cloudflare.com)
(TXT) w3m dump (blog.cloudflare.com)
| dfajgljsldkjag wrote:
| We already have the tools to stop this from happening today. The
| problem is not the technology but the fact that companies do not
| want to work together to fix it. It is sad that we let the
| internet break because people are too slow to use the safety
| features we have.
| btown wrote:
| > we pushed a change via our policy automation platform to remove
| the BGP announcements from Miami
|
| Is there any way to test these changes against a simulation of
| real world routes? Including to ensure that traffic that
| shouldn't hit Cloudflare servers, continues to resolve routes
| that don't hit Cloudflare?
|
| I have to imagine there's academic research on how to simulate a
| fork of global BGP state, no? Surely there's a tensor
| representation of the BGP graph that can be simulated on GPU
| clusters?
|
| If there's a meta-rule I think of when these incidents occur,
| it's that configuration rules need change management, and change
| management is only as good as the level of automated testing.
| Just because code hasn't changed doesn't mean you shouldn't test
| the baseline system behavior. And here, that means testing that
| the Internet works.
| Analemma_ wrote:
| I assume it's not possible unless you know the in-memory state
| of all the other gateway routers on the internet, no? You can
| know what they advertise, but that's not the same thing as a
| full description of their internal state and how they will
| choose to update if a route gets withdrawn.
| erredois wrote:
| I think you could know the state of the peers and simulate
| what they advertise and receive and validate that. The test
| unit would need to be a simulated router that behaves exactly
| as the real one, I actually think its technically doable with
| tight version control for routers.
| hnuser123456 wrote:
| You can cross-reference RADB, the RIRs, and looking glass
| servers, and you'd find 3 different pictures of the internet.
| PunchyHamster wrote:
| > Is there any way to test these changes against a simulation
| of real world routes? Including to ensure that traffic that
| shouldn't hit Cloudflare servers, continues to resolve routes
| that don't hit Cloudflare?
|
| You can get access to view of routes from different parts of
| networks but you do not have access to those routers policies,
| so no
|
| > I have to imagine there's academic research on how to
| simulate a fork of global BGP state, no? Surely there's a
| tensor representation of the BGP graph that can be simulated on
| GPU clusters?
|
| Just simulating your peers and maybe layer after is most likely
| good enough. And you can probably do it with a bunch of cgroups
| and some actual routing software. There are also network sims
| like GNS3 that can even just run router images
| 0xy wrote:
| The string of recent incidents don't really make the new CTO look
| good. Too much focus on shipping, not enough on shipping
| correctly.
| iLoveOncall wrote:
| Welcome to the age of AI-assisted coding.
| SketchySeaBeast wrote:
| I could have sworn "move fast and break things" existed
| before AI.
| arter45 wrote:
| I've had to read the RCA a couple of times to (probably) get what
| happened, even if I'm reasonably familiar with BGP.
|
| Basically, my understanding (simplified) is:
|
| - they originally had a Miami router advertise Bogota prefixes
| (=subnets) to Cloudflare's peers. Essentially, Miami was handling
| Bogota's subnets. This is not an issue.
|
| - because you don't normally advertise arbitrary prefixes via
| BGP, policies were used. These policies are essentially if/then
| statements, carrying out certain actions (advertise or not, add
| some tags or remove them,...) if some conditions are matched.
| This is completely normal.
|
| - Juniper router configuration for this kind of policy is
| (simplifying):
|
| set <BGP POLICY NAME> from <CONDITION1>
|
| set <BGP POLICY NAME> from <CONDITION2>
|
| set <BGP POLICY NAME> then <ACTION1>
|
| set <BGP POLICY NAME> then <ACTION2>
|
| ...
|
| - prior to the incident, CF changed its network so that Miami
| didn't have to handle Bogota subnets (maybe Bogota does it on its
| own, maybe there's another router somewhere else)
|
| - the change aimed at removing the configurations on Miami which
| were advertising Bogota subnets
|
| - the change implementation essentially removed all lines from
| all policies containing "from IP in the list of Bogota prefixes".
| This is somewhat reasonable, because you could have the same
| policy handling both Bogota and, say, Quito prefixes, so you just
| want to remove the Bogota part.
|
| HOWEVER, there was at least one policy like this:
|
| (Before)
|
| set <BGP POLICY NAME> from is_internal(prefix) == True
|
| set <BGP POLICY NAME> from prefix in bogota_prefix_list
|
| set <BGP POLICY NAME> then advertise
|
| (After)
|
| set <BGP POLICY NAME> from is_internal(prefix) == True
|
| set <BGP POLICY NAME> then advertise
|
| Which basically means: if you have an internal prefix advertise
| it
|
| - an "internal prefix" is any prefix that was not received by
| another BGP entity (autonomous system)
|
| - BGP routers in Cloudflare exchange routes to one another. This
| is again pretty normal.
|
| - As a result of this change, all routes received by Miami
| through some other Cloudflare router were readvertised by Miami
|
| - the result is CF telling the Internet (more accurately, its
| peers) "hey, you know that subnet? Go ask my Miami router!"
|
| - obviously, this increases bandwidth utilization and latency for
| traffic crossing the Miami router.
| vlovich123 wrote:
| I'm a huge fan of flapping when it's really hard to do
| progressive rollouts. What this would mean here is you switch
| advertising the old and new routes back and forth automatically
| and this happens let's say for 1 minute max before the old config
| is restored. Then a human looks at various metrics before they
| push a button to really make the new config permanent. It gives
| you a cheap way to preflight what will happen when you make a
| globally impacting config change.
| arter45 wrote:
| I'm not sure this would be a good idea in this kind of change.
|
| Flapping is bad in the networking world.
|
| Flapping BGP routes, specifically, is bad because it can stress
| all BGP routers involved to the point where they can "go
| crazy". Routes are explicitly advertised, so if you keep
| changing the routes, you are tasking the router CPU to process
| new stuff, discard it and process new stuff. In fact, BGP route
| flaps are specifically the focus of an entire RFC:
| https://datatracker.ietf.org/doc/html/rfc2439
|
| More in general, a flapping link (on/off/on/off) can really
| mess with TCP.
|
| Flapping in the networking world is not something you want to
| do intentionally.
| PunchyHamster wrote:
| nice way to 100% the router CPUs for all your peers
| jacquesm wrote:
| That's like what, one major incident per month now, Nov 18, Dec
| 5, and now this one?
|
| I'll bet JGC can write his own ticket by now, but unretiring
| would be really bad optics. He's on the board though and still
| keeping a watchful eye. But a couple more of these and CFs
| reputation will be in the gutter.
| betaby wrote:
| Weak engineering. Both from the CloudFlare side and their peers.
| PunchyHamster wrote:
| Damn, I missed the fact Juniper was acquired by HPE, RIP
___________________________________________________________________
(page generated 2026-01-23 23:00 UTC)