[HN Gopher] Route leak incident on January 22, 2026
       ___________________________________________________________________
        
       Route leak incident on January 22, 2026
        
       Author : nomaxx117
       Score  : 97 points
       Date   : 2026-01-23 17:54 UTC (5 hours ago)
        
 (HTM) web link (blog.cloudflare.com)
 (TXT) w3m dump (blog.cloudflare.com)
        
       | dfajgljsldkjag wrote:
       | We already have the tools to stop this from happening today. The
       | problem is not the technology but the fact that companies do not
       | want to work together to fix it. It is sad that we let the
       | internet break because people are too slow to use the safety
       | features we have.
        
       | btown wrote:
       | > we pushed a change via our policy automation platform to remove
       | the BGP announcements from Miami
       | 
       | Is there any way to test these changes against a simulation of
       | real world routes? Including to ensure that traffic that
       | shouldn't hit Cloudflare servers, continues to resolve routes
       | that don't hit Cloudflare?
       | 
       | I have to imagine there's academic research on how to simulate a
       | fork of global BGP state, no? Surely there's a tensor
       | representation of the BGP graph that can be simulated on GPU
       | clusters?
       | 
       | If there's a meta-rule I think of when these incidents occur,
       | it's that configuration rules need change management, and change
       | management is only as good as the level of automated testing.
       | Just because code hasn't changed doesn't mean you shouldn't test
       | the baseline system behavior. And here, that means testing that
       | the Internet works.
        
         | Analemma_ wrote:
         | I assume it's not possible unless you know the in-memory state
         | of all the other gateway routers on the internet, no? You can
         | know what they advertise, but that's not the same thing as a
         | full description of their internal state and how they will
         | choose to update if a route gets withdrawn.
        
           | erredois wrote:
           | I think you could know the state of the peers and simulate
           | what they advertise and receive and validate that. The test
           | unit would need to be a simulated router that behaves exactly
           | as the real one, I actually think its technically doable with
           | tight version control for routers.
        
         | hnuser123456 wrote:
         | You can cross-reference RADB, the RIRs, and looking glass
         | servers, and you'd find 3 different pictures of the internet.
        
         | PunchyHamster wrote:
         | > Is there any way to test these changes against a simulation
         | of real world routes? Including to ensure that traffic that
         | shouldn't hit Cloudflare servers, continues to resolve routes
         | that don't hit Cloudflare?
         | 
         | You can get access to view of routes from different parts of
         | networks but you do not have access to those routers policies,
         | so no
         | 
         | > I have to imagine there's academic research on how to
         | simulate a fork of global BGP state, no? Surely there's a
         | tensor representation of the BGP graph that can be simulated on
         | GPU clusters?
         | 
         | Just simulating your peers and maybe layer after is most likely
         | good enough. And you can probably do it with a bunch of cgroups
         | and some actual routing software. There are also network sims
         | like GNS3 that can even just run router images
        
       | 0xy wrote:
       | The string of recent incidents don't really make the new CTO look
       | good. Too much focus on shipping, not enough on shipping
       | correctly.
        
         | iLoveOncall wrote:
         | Welcome to the age of AI-assisted coding.
        
           | SketchySeaBeast wrote:
           | I could have sworn "move fast and break things" existed
           | before AI.
        
       | arter45 wrote:
       | I've had to read the RCA a couple of times to (probably) get what
       | happened, even if I'm reasonably familiar with BGP.
       | 
       | Basically, my understanding (simplified) is:
       | 
       | - they originally had a Miami router advertise Bogota prefixes
       | (=subnets) to Cloudflare's peers. Essentially, Miami was handling
       | Bogota's subnets. This is not an issue.
       | 
       | - because you don't normally advertise arbitrary prefixes via
       | BGP, policies were used. These policies are essentially if/then
       | statements, carrying out certain actions (advertise or not, add
       | some tags or remove them,...) if some conditions are matched.
       | This is completely normal.
       | 
       | - Juniper router configuration for this kind of policy is
       | (simplifying):
       | 
       | set <BGP POLICY NAME> from <CONDITION1>
       | 
       | set <BGP POLICY NAME> from <CONDITION2>
       | 
       | set <BGP POLICY NAME> then <ACTION1>
       | 
       | set <BGP POLICY NAME> then <ACTION2>
       | 
       | ...
       | 
       | - prior to the incident, CF changed its network so that Miami
       | didn't have to handle Bogota subnets (maybe Bogota does it on its
       | own, maybe there's another router somewhere else)
       | 
       | - the change aimed at removing the configurations on Miami which
       | were advertising Bogota subnets
       | 
       | - the change implementation essentially removed all lines from
       | all policies containing "from IP in the list of Bogota prefixes".
       | This is somewhat reasonable, because you could have the same
       | policy handling both Bogota and, say, Quito prefixes, so you just
       | want to remove the Bogota part.
       | 
       | HOWEVER, there was at least one policy like this:
       | 
       | (Before)
       | 
       | set <BGP POLICY NAME> from is_internal(prefix) == True
       | 
       | set <BGP POLICY NAME> from prefix in bogota_prefix_list
       | 
       | set <BGP POLICY NAME> then advertise
       | 
       | (After)
       | 
       | set <BGP POLICY NAME> from is_internal(prefix) == True
       | 
       | set <BGP POLICY NAME> then advertise
       | 
       | Which basically means: if you have an internal prefix advertise
       | it
       | 
       | - an "internal prefix" is any prefix that was not received by
       | another BGP entity (autonomous system)
       | 
       | - BGP routers in Cloudflare exchange routes to one another. This
       | is again pretty normal.
       | 
       | - As a result of this change, all routes received by Miami
       | through some other Cloudflare router were readvertised by Miami
       | 
       | - the result is CF telling the Internet (more accurately, its
       | peers) "hey, you know that subnet? Go ask my Miami router!"
       | 
       | - obviously, this increases bandwidth utilization and latency for
       | traffic crossing the Miami router.
        
       | vlovich123 wrote:
       | I'm a huge fan of flapping when it's really hard to do
       | progressive rollouts. What this would mean here is you switch
       | advertising the old and new routes back and forth automatically
       | and this happens let's say for 1 minute max before the old config
       | is restored. Then a human looks at various metrics before they
       | push a button to really make the new config permanent. It gives
       | you a cheap way to preflight what will happen when you make a
       | globally impacting config change.
        
         | arter45 wrote:
         | I'm not sure this would be a good idea in this kind of change.
         | 
         | Flapping is bad in the networking world.
         | 
         | Flapping BGP routes, specifically, is bad because it can stress
         | all BGP routers involved to the point where they can "go
         | crazy". Routes are explicitly advertised, so if you keep
         | changing the routes, you are tasking the router CPU to process
         | new stuff, discard it and process new stuff. In fact, BGP route
         | flaps are specifically the focus of an entire RFC:
         | https://datatracker.ietf.org/doc/html/rfc2439
         | 
         | More in general, a flapping link (on/off/on/off) can really
         | mess with TCP.
         | 
         | Flapping in the networking world is not something you want to
         | do intentionally.
        
         | PunchyHamster wrote:
         | nice way to 100% the router CPUs for all your peers
        
       | jacquesm wrote:
       | That's like what, one major incident per month now, Nov 18, Dec
       | 5, and now this one?
       | 
       | I'll bet JGC can write his own ticket by now, but unretiring
       | would be really bad optics. He's on the board though and still
       | keeping a watchful eye. But a couple more of these and CFs
       | reputation will be in the gutter.
        
       | betaby wrote:
       | Weak engineering. Both from the CloudFlare side and their peers.
        
       | PunchyHamster wrote:
       | Damn, I missed the fact Juniper was acquired by HPE, RIP
        
       ___________________________________________________________________
       (page generated 2026-01-23 23:00 UTC)