Post B44Oia7fC0tpuW8Yd6 by bytex64@awesome.garden
 (DIR) More posts by bytex64@awesome.garden
 (DIR) Post #B44OBYBPNp0eL48d8K by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       got another few valid hits on my bitflip experiment:statistics:hits53731after filter19bot rate99.96%addrs filtered200UAs filtered8paths filtered86new hits:- one from an Android 9 device on Rogers (ipv6) using gmail webview- 4 from google-owned IPs(!): three tracking pixels from blogger domains, and one pagespeed proxy requestI am slightly intrigued by the google IPs - do they run a lot of gear without ECC memory?
       
 (DIR) Post #B44Oia7fC0tpuW8Yd6 by bytex64@awesome.garden
       0 likes, 0 repeats
       
       @astraleureka Just guessing, but Google’s hardware engineering MO for a long while was “buy the cheapest hardware and flog it to death”. Someone else would have to tell you how true that is nowadays, but I wouldn’t be surprised if they’ve spared the expense of ECC.
       
 (DIR) Post #B44RAPhxhtyVcwVxYm by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       @cursedsql @bytex64 yeah, I am effectively re-running these same experiments (although I can't afford the full set of domains quite yet) to see what's different in the modern era of smaller silicon processes (more likely bitflips), DDR5 ECC, and the like.I knew google went hard on the cheap hardware back in the day, but ECC RDIMMs aren't all that more expensive than consumer-grade, it's only the weird low-volume stuff like unregistered/buffered UDIMMs and SODIMMs that get pricy ime. but those small differences do add up at google scale I suppose
       
 (DIR) Post #B4H1nXPgvA7f46RW2y by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       update since the 8th:hits110397after filter27bot rate99.98%addrs filtered382UAs filtered13paths filtered141new hits:- two hits from distinct AWS IPv4s ~2 seconds apart, to a Gmail asset URL- two more hits to the exact same tracking pixel URL from before (same referrer as well), one from DigitalOcean and another from residential .VN ISP- one hit to the default Google profile picture (referrer accounts.google.com) from a possible proxy in .PK- one hit to a placeholder image used in the Google Photos app (com.google.android.apps.photos in the UA) from a residential IPv6 in .VN- one hit to a user's google profile picture from a residential IPv4 in .IN (referrer speedtypingonline.com)it's getting a little trickier to filter out all the weird noise, my regex rules are starting to get kinda cluttered and I didn't provide for any means of commenting/documenting the rules. I think I will pick up another batch of 15 domains next paycheck, as it looks like there is still a surprising amount of activity even with only my small sample set so far.next steps will also be to start logging all DNS queries - it seems like 99.9% of the garbage traffic is hitting the base domain, while all the interesting stuff is hitting well known subdomains. I can see this sort of analysis being a lot harder for non-CDN domains that don't have unique subdomains...i'll probably try to cobble together a custom DNS server for this and run the nodes on my anycast routers, perhaps there will be some interesting geographical bias in where corrupt requests come from once more data is available. I am also wondering if there are possibly other services outside of HTTP{,S} running on googleusercontent.com - does anybody know if that's a thing?
       
 (DIR) Post #B4cbh5eSu0UDtajifg by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       +11 day update:hits228081after filter51bot rate99.98%addrs filtered634UAs filtered14paths filtered208notable or interesting hits:- a hit from a facebook crawler- several hits for GCP block storage downloads from Nepal- several pagespeed hits for Horse Talk- a few hits from a chromecast dongle with a *lot* of flips in the URL, poor thing must be really suffering- a number of hits for user profile pictures from what I think is PUBG mobile? game=ShadowTrackerExtra, engine=UE4, version=4.18.1-0+++UE4+Release-4.18, platform=IOS, osver=26.3.1- some unknown unity app? UnityPlayer/2019.4.40f1, libcurl/7.80.0-DEV- classic Opera with the Presto rendering engine, on a 32 bit Linux machine in Egypt! getting paid next week and will pick up another batch of domains, which should hopefully increase the hitrate. like originally expected it mostly seems to be mobile devices, but there have been a few desktops and servers in the data so far.
       
 (DIR) Post #B4nh9rRHSbkMiufPcm by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       +5 day update:hits252444after filter75bot rate99.97%addrs filtered760UAs filtered15paths filtered243notable hits:- a whole lot more hits from Unity apps, all for the same URL, in short temporal succession from all over the world. I suspect a game server experienced a bitflip and served the bad URL to many clients- doc-0s-7g-apps-viewer.guc.c/viewer/secure/pdf/ via CF Warp, from an iPhone- several hits to different URLs from a ~2016-2018 Sony Bravia TV running Android 9 (cpu: MT5891) from a Japanese IPv4 address- several more profile photos w/ referrer accounts.g.cit's the end of the month and payday, which means I can add another ~10 domains to the experiment and collect more data! :D
       
 (DIR) Post #B4oCFZ8zUUcYIbJQNE by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       @nini google shows no interest in registering any of the single-bit flip variants on their CDN domain(s), so I am registering as many as I can limited by cost :((there are a sufficient amount of devices out there with flaky memory or other issues that have been able to successfully establish connections to my stand-in server, primarily mobile and embedded devices so far but surprisingly a handful of servers too!
       
 (DIR) Post #B4oCOhaBUAgF4qO8LA by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       @nini arguably yes I could be malicious and try to steal session cookies or register variants of a more sensitive domain (I am analyzing variants of googleusercontent[.]com) but that's boring, I'd rather see a sampling of more day-to-day traffic that ends up inadvertently taking a wrong turn on the internet
       
 (DIR) Post #B4oGPGdZYLKxc7Uwhk by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       just registered the next batch of domains, now at 36 of 81 total / 78 available (a few of the variants were already registered)remaining cost is $465.36.. maybe finish it up next month 😔
       
 (DIR) Post #B4oGUBbfgU2fSXxEaO by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       now the proud owner of coogle and goofle user content dot com
       
 (DIR) Post #B4oXMk73CpGhNlfUvo by ska@social.treehouse.systems
       0 likes, 0 repeats
       
       @astraleureka
       
 (DIR) Post #B4w97h0Be67ZkuQacC by astraleureka@social.treehouse.systems
       0 likes, 0 repeats
       
       +4 day update:statistics:hits330360after filter543bot rate99.84%addrs filtered1242UAs filtered17paths filtered369traffic has popped off significantly with this new batch of domains, plus a whole lot more scanner traffic that I think is mostly filtered now.notable hits:- hundreds of hits from the same Level3/Lumen IPv4 for /proxy/<encoded> endpoints- a handful of drive-thirdparty.guc.com hits for video players- one of the first hits from a semi-modern device: SM-A032M(Samsung Galaxy A03 Core) rel 2021, Android 13- GmsCore/261133035 which looks like an alternative Play Services implementationI should probably browse through the debug logs of blocked requests to make sure I'm not accidentally filtering anything legit, but it seems unlikely