Post B6azAX5xwsd9dfidt2 by tiredbun@akko.wtf
 (DIR) More posts by tiredbun@akko.wtf
 (DIR) Post #B6ayNXxRx8Ho7nNSTI by niconiconi@mk.absturztau.be
       0 likes, 0 repeats
       
       LLM crawlers have the best archives of the 2020s Web today, likely better than the Wayback Machine. This is seriously an extremely valuable contribution to historians if they remain accessible decades later. Hopefully all the data would eventually be donated to a foundation, instead of silently going bankrupt with the raw data buried.
       
 (DIR) Post #B6az3437XgFRWlpKfg by rubinjoni@mastodon.social
       0 likes, 1 repeats
       
       @niconiconi like tearsinrain
       
 (DIR) Post #B6azAX5xwsd9dfidt2 by tiredbun@akko.wtf
       0 likes, 0 repeats
       
       @niconiconi The issue is they likely delete at least part of this data after training, partially to avoid liabilities in future...
       
 (DIR) Post #B6b4pGvZSpPN8i6Omm by jernej__s@infosec.exchange
       0 likes, 0 repeats
       
       @niconiconi Are you sure they even cache any data? It seems that they just re-fetch any time the data is needed…