[HN Gopher] More than 2M research papers have disappeared from t...
       ___________________________________________________________________
        
       More than 2M research papers have disappeared from the Internet
        
       Author : Brajeshwar
       Score  : 136 points
       Date   : 2024-03-04 12:30 UTC (10 hours ago)
        
 (HTM) web link (www.nature.com)
 (TXT) w3m dump (www.nature.com)
        
       | jtbayly wrote:
       | Scihub seems more and more beneficial every day.
        
         | kristofferR wrote:
         | Sci-hub is dead, hasn't been updates since 2021. Sad state of
         | affairs, hope Anna's Archive can take up the mantle.
        
           | nonrandomstring wrote:
           | It's moving onto overlay networks. As everything of human
           | value on the internet will eventually have to do.
        
           | bluish29 wrote:
           | Probably Alexandra hope too much that Delhi court will score
           | sci-hub a win that makr it legal in India. That would be a
           | precedent that she aims to be imitated in other places. She
           | is still complying with the request of the court to stop
           | uploading new papers.
        
           | araes wrote:
           | Until they manage to get the servers shut down, there's still
           | 88,000,000 articles for download.
        
           | esdf wrote:
           | A lawsuit is underway with a named defendant for Anna's
           | Archive https://torrentfreak.com/lawsuit-accuses-annas-
           | archive-of-ha...
        
             | kristofferR wrote:
             | Ah, major bummer, thought the lawsuit was just a useless
             | one against all John Does. Not good if they've found her :/
             | 
             | Another case to add to the follow case list, I guess.
             | https://www.courtlistener.com/docket/68157923/oclc-online-
             | co...
             | 
             | The other tech case I follow is the Yuzu vs Nintendo case.
             | Wish the best for both!
        
               | jtbayly wrote:
               | Looks like you can stop following the latter:
               | 
               | https://news.ycombinator.com/item?id=39593647
               | 
               | Yuzu loses.
        
               | kristofferR wrote:
               | Ah shit
        
       | perihelions wrote:
       | - _" It is also true that much content is "preserved," often
       | illegally, in shadow libraries/archives such as Library Genesis
       | and Sci-Hub (for more, see Bodo, 2018a, 2018b; Eve, 2022). These
       | archives are at great legal threat of shutdown but have also
       | proved surprisingly resilient. We do not, in this article, count
       | material stored in such archives, even though it constitutes an
       | additional storage location."_
       | 
       | A major limitation.
        
         | SiempreViernes wrote:
         | I don't know, this article is explicitly about the problem of
         | trying to link to something you find valuable but its
         | supposedly durable url has been taken offline.
         | 
         | The DOI is supposed to ensure a copy can be found, if it fails
         | to do that in 30% of the cases it's not a very useful,
         | regardless of if a copy exists somewhere online.
        
           | toomuchtodo wrote:
           | One of the few useful cases for IPFS tied to DOI.
        
           | jszymborski wrote:
           | I will say that DOI.org isn't the only dDOI resolver and that
           | sci-hub in many ways behaves like one.
        
           | kragen wrote:
           | while of course you should only do this if you are in a
           | country where it is legal, you can look up a doi in
           | libstc.cc, sci-hub.ru, or annas-archive.org
        
             | titanomachy wrote:
             | Are there jurisdictions where merely _accessing_ this
             | content is punishable by law?
             | 
             | EDIT: genuinely asking, I know very little about IP law.
        
               | kragen wrote:
               | there are jurisdictions where making a copy of it is,
               | even for personal use, but others where it is not
        
               | InSteady wrote:
               | I assume viewing a PDF in browser is generally safe
               | everywhere as long as you aren't downloading to HD or
               | cloud? Or is even this activity realistically dangerous
               | without a VPN in some places?
        
         | eviks wrote:
         | Indeed, "we can't find many books if we ignore the libraries!"
        
           | lcnPylGDnU4H9OF wrote:
           | More like "we can't find many publicly-available books if we
           | ignore home libraries". It seems prudent to wonder where so
           | many previously-publicly-available books have gone.
        
         | Brian_K_White wrote:
         | You don't think this alarm should be raised because there
         | exists illegal copies?
         | 
         | Otherwise what is the limitation? What is the failing that
         | makes it not that useful of a study?
        
       | advael wrote:
       | Arxiv feels like the most legitimate "journal" these days. It'd
       | be sci-hub except they seem to have managed to kill it finally
       | 
       | The business of IP has so hollowed out the credibility of its
       | brokers that the business part (like the stock price) has become
       | its only remaining value, and scientific journals are absolutely
       | included here if their publishing method doesn't _preserve the
       | goddamn body of evidence we claim science relies on_
        
         | pk-protect-ai wrote:
         | Do not forget that you are charged when you publish your work
         | in scientific journals. Would you not expect your published
         | work to be preserved?
        
           | advael wrote:
           | One might think this, but that you are guaranteed anything
           | (even if you are explicitly promised it) by a for-profit
           | corporation is a foolhardy and sometimes even dangerous thing
           | to take for granted
        
       | setgree wrote:
       | Prior: most published research is false [0].
       | 
       | If a paper is being forgotten by the internet, that's a weak
       | signal that it was not valued by its field. If you combine that
       | weak signal with the strong prior, you'd conclude that the paper
       | wasn't a meaningful contribution and can be safely forgotten.
       | 
       | I can see why this mass forgetting is a problem for
       | preservationists and historians of science.
       | 
       | [0]
       | https://journals.plos.org/plosmedicine/article?id=10.1371/jo...
       | 
       | [1] e.g. Gordon Allport's "The Nature of Prejudice" was published
       | 1954, has been cited >50,000 times, and is still in print:
       | https://www.amazon.com/Nature-Prejudice-25th-Anniversary/dp/...
        
         | marcosdumay wrote:
         | > If a paper is being forgotten by the internet, that's a weak
         | signal that it was not valued by its field.
         | 
         | That would be valid if the decision was made by popular vote.
         | 
         | It's completely invalid if a single entity is allowed to decide
         | whether it lives or dies.
        
         | plumeria wrote:
         | Some research is forgotten/dismissed by some generation only to
         | be revived by a later generation (see e.g. Geometric Algebra)
        
         | squigz wrote:
         | Science doesn't work like that; we can't look at data/research
         | at our current point in time and deduce its value forever.
         | Having this information years down the line can inspire some
         | more research ideas or contribute some groundwork to future
         | research
        
         | kragen wrote:
         | to tell whether a piece of research is false or not, you need
         | to be able to find out what it was. then you can conclusively
         | prove it false. lost research like tesla's weather control
         | notes, the alchemists' transmutation of lead into gold, or
         | water-powered cars can instead give rise to conspiracy theories
         | about how it was suppressed by the powerful and wasted
         | lifetimes that could have been dedicated to something
         | productive
        
         | eviks wrote:
         | The probability assessment is fine, but meaningless even
         | outside history. Baby and the bathwater
        
       | jjslocum3 wrote:
       | I imagine this would be of great benefit to plagiarists.
        
       | daniel31x13 wrote:
       | This is one of the main reasons I created Linkwarden - to combat
       | Link-Rot.
       | 
       | Linkwarden is an open-source collaborative bookmark manager to
       | collect, organize and preserve webpages:
       | 
       | https://linkwarden.app
        
         | Brian_K_White wrote:
         | Suggestion: Find your own name that isn't coat-tailing on
         | bitwarden.
         | 
         | Or don't, I'm not your mom, but the name tells me not to even
         | investigate, and I care about such things enough to run a YaCy
         | node for example.
        
           | severine wrote:
           | So, it has to be called Yet Another Something?
           | 
           | C'mon, let's not gatekeep or trademarkify normal descriptive
           | or allusive words. Linkwarden is fine!
        
       | nemoniac wrote:
       | Use DOIs, they said. URLs are just transient, they said. But DOIs
       | are permanent, they said.
       | 
       | Could it be that the whole DOI system was just a ruse by the
       | commercial academic publishers to re-intermediate themselves
       | between academic authors and readers in the age of the WWW?
        
         | jszymborski wrote:
         | I really don't see DOIs as the problem here. They are just
         | catalogue numbers, not an archive or repository.
         | 
         | The DOIs are still permanent, but their resolvers might not be.
         | Copyright problems definitely makes that harder to fix, but
         | again, I don't think the DOI is the problem here.
        
       | chaseadam17 wrote:
       | Arweave fixes this. It's pay once permanent decentralized
       | storage. One of the few web3 projects with fast-growing real
       | world usage.
        
       | bogtog wrote:
       | I wonder if this is mostly just the DOIs associated with weak or
       | predatory journals
        
       | renegade-otter wrote:
       | By the way, what's on the Internet is NOT forever. I expect that
       | Facebook and Twitter datasets, for example, will eventually go
       | offline, sold to AI companies, slowly reaching irrelevancy and
       | finally resting in some cold data storage or used for
       | training/educational purposes.
        
         | wepple wrote:
         | It's the worst of both worlds
         | 
         | You have to assume an unappealing photo of yourself you posted
         | in your youth _is_ stored somewhere forever
         | 
         | And, things you actually want to be stored forever and be
         | retrievable when useful, probably won't.
        
           | renegade-otter wrote:
           | Good point :)
        
           | InSteady wrote:
           | We shall henceforth refer to this as Wepple's paradox.
        
       | kkfx wrote:
       | That's why we need to push a desktop-centric internet, not a
       | thin-client/endpoints one. People MUST learn to store their data,
       | distributed storage must became the norm for public stuff, so
       | anything someone value will be preserved.
       | 
       | Just start with Zotero to have your papers LOCALLY (not on Zotero
       | cloud), it's a small simple step anyone have enough storage to
       | make.
        
       | ericra wrote:
       | Think of how trivial it would be, from a storage perspective, to
       | have a central repo of every (publicly-funded, at a minimum)
       | research paper ever written if copyright issues were not present.
       | 
       | It could be run by an international non-profit or as a
       | collaboration between governments. It wouldn't be particularly
       | expensive to run when weighed against the huge long-term
       | benefits.
       | 
       | It's really a failure of the U.S. and EU governments to not clamp
       | down on the Elseviers of the world who bring very few benefits at
       | a high cost to open science and society at large.
        
         | observationist wrote:
         | The Library of Congress has the digital infrastructure and
         | expertise to be able to handle something like this issue.
         | 
         | Unfortunately, scientific journals are a sacred cash cow and
         | infringing on any of their territory, real or imagined,
         | prevents any meaningful top down change within the system.
         | They've got money to pay lawyers to prevent any reform or
         | mandates or flexing of existing regulatory powers.
         | 
         | Pirate everything.
         | 
         | Publicly funded research gets lost because it negatively
         | affects profits to maintain unpopular, unread material in any
         | sort of diligent and effective way.
         | 
         | Journals have close ties to universities and academia, and big
         | commercial research outfits, and all of the social ties being
         | involved with those circles can bring. They've got lobbying
         | perfected to an art, and pay good, ruthless lawyers to protect
         | their interests.
         | 
         | The average voter won't ever care enough to make a popular
         | revolt, bottom up change possible; scientific publishing is too
         | dry and anemic when you contrast against the million other,
         | more outrageous, imminently threatening issues people care
         | about. When you look at the conflicted interests that benefit
         | from the status quo, such as companies that can pay Journals so
         | their studies and papers will appear alongside distinguished,
         | credentialed works, there doesn't appear to be any place where
         | effective leverage can be applied.
         | 
         | Pirating it all is the only ethical solution. Outlaws with
         | rogue copies is the only feasible way much of this data will
         | carry on into the future. None of the people who can change
         | things actually seem to want to, and the public has more
         | pressing matters capturing their attention.
         | 
         | Aaron Swartz had it right. Apathy, greed, and politics aren't a
         | problem with a solution that will come from this space. The
         | only winning move is not to play their game.
         | 
         | The same largely applies to any media content being gatekept by
         | entities requiring repeated, endless rent on their "property"
         | despite bringing no value to the market, existing simply to
         | expand, endlessly, mindlessly shoveling other people's money
         | into their shareholder's pockets.
         | 
         | We live in a world that is awfully stupid sometimes. So stupid
         | that important scientific literature is being faded into
         | oblivion simply because it can't be monetized under perverse
         | adtech incentive schemes. This has crucial implications,
         | because if the habit of lazy citation takes root, it creates a
         | kind of secular system of faith in scientism, with authors
         | being given the benefit of the doubt when their citations can't
         | be verified. That could be catastrophic if it affects
         | medicines, engineering, environmental management, urban
         | planning, forensics, or a myriad other categories where a
         | flawed scientific paper might pose a threat.
         | 
         | Knowing that flawed papers get through, what could an actual
         | malicious actor get away with?
         | 
         | Don't let stupid win.
         | 
         | Pirate everything.
        
       ___________________________________________________________________
       (page generated 2024-03-04 23:02 UTC)