Post B5hN0ftft1L65a2vR2 by aredridel@kolektiva.social
 (DIR) More posts by aredridel@kolektiva.social
 (DIR) Post #B5hN0ftft1L65a2vR2 by aredridel@kolektiva.social
       0 likes, 0 repeats
       
       One thing I don't see anyone talking about that we probably should is the proliferation of captcha-busting anubis-busting browser-as-a-service services. It's not that the big model companies are scraping the web and ignoring robots.txt. (Some are, almost certainly, but there are datasets to train on already and they're not scraping random sites so much)It's that agent _users_ and the people serving them have a very large demand to access information with semi-automated systems. And they're building whole armies of ways around blocking.
       
 (DIR) Post #B5hN0gQHvlQXijIyGm by IngaLovinde@embracing.space
       0 likes, 0 repeats
       
       @aredridel > It's that agent _users_ and the people serving them have a very large demand to access information with semi-automated systemsWhich information though? What kind of agent user could have demand to access information on git blame page of some obscure file in one of my repos on my obscure gitea insurance?> It's not that the big model companies are scraping the web and ignoring robots.txtI heard that big model companies outsource the scraping to the lowest bidders.