[HN Gopher] More than 2M research papers have disappeared from t...
___________________________________________________________________
More than 2M research papers have disappeared from the Internet
Author : Brajeshwar
Score : 136 points
Date : 2024-03-04 12:30 UTC (10 hours ago)
(HTM) web link (www.nature.com)
(TXT) w3m dump (www.nature.com)
| jtbayly wrote:
| Scihub seems more and more beneficial every day.
| kristofferR wrote:
| Sci-hub is dead, hasn't been updates since 2021. Sad state of
| affairs, hope Anna's Archive can take up the mantle.
| nonrandomstring wrote:
| It's moving onto overlay networks. As everything of human
| value on the internet will eventually have to do.
| bluish29 wrote:
| Probably Alexandra hope too much that Delhi court will score
| sci-hub a win that makr it legal in India. That would be a
| precedent that she aims to be imitated in other places. She
| is still complying with the request of the court to stop
| uploading new papers.
| araes wrote:
| Until they manage to get the servers shut down, there's still
| 88,000,000 articles for download.
| esdf wrote:
| A lawsuit is underway with a named defendant for Anna's
| Archive https://torrentfreak.com/lawsuit-accuses-annas-
| archive-of-ha...
| kristofferR wrote:
| Ah, major bummer, thought the lawsuit was just a useless
| one against all John Does. Not good if they've found her :/
|
| Another case to add to the follow case list, I guess.
| https://www.courtlistener.com/docket/68157923/oclc-online-
| co...
|
| The other tech case I follow is the Yuzu vs Nintendo case.
| Wish the best for both!
| jtbayly wrote:
| Looks like you can stop following the latter:
|
| https://news.ycombinator.com/item?id=39593647
|
| Yuzu loses.
| kristofferR wrote:
| Ah shit
| perihelions wrote:
| - _" It is also true that much content is "preserved," often
| illegally, in shadow libraries/archives such as Library Genesis
| and Sci-Hub (for more, see Bodo, 2018a, 2018b; Eve, 2022). These
| archives are at great legal threat of shutdown but have also
| proved surprisingly resilient. We do not, in this article, count
| material stored in such archives, even though it constitutes an
| additional storage location."_
|
| A major limitation.
| SiempreViernes wrote:
| I don't know, this article is explicitly about the problem of
| trying to link to something you find valuable but its
| supposedly durable url has been taken offline.
|
| The DOI is supposed to ensure a copy can be found, if it fails
| to do that in 30% of the cases it's not a very useful,
| regardless of if a copy exists somewhere online.
| toomuchtodo wrote:
| One of the few useful cases for IPFS tied to DOI.
| jszymborski wrote:
| I will say that DOI.org isn't the only dDOI resolver and that
| sci-hub in many ways behaves like one.
| kragen wrote:
| while of course you should only do this if you are in a
| country where it is legal, you can look up a doi in
| libstc.cc, sci-hub.ru, or annas-archive.org
| titanomachy wrote:
| Are there jurisdictions where merely _accessing_ this
| content is punishable by law?
|
| EDIT: genuinely asking, I know very little about IP law.
| kragen wrote:
| there are jurisdictions where making a copy of it is,
| even for personal use, but others where it is not
| InSteady wrote:
| I assume viewing a PDF in browser is generally safe
| everywhere as long as you aren't downloading to HD or
| cloud? Or is even this activity realistically dangerous
| without a VPN in some places?
| eviks wrote:
| Indeed, "we can't find many books if we ignore the libraries!"
| lcnPylGDnU4H9OF wrote:
| More like "we can't find many publicly-available books if we
| ignore home libraries". It seems prudent to wonder where so
| many previously-publicly-available books have gone.
| Brian_K_White wrote:
| You don't think this alarm should be raised because there
| exists illegal copies?
|
| Otherwise what is the limitation? What is the failing that
| makes it not that useful of a study?
| advael wrote:
| Arxiv feels like the most legitimate "journal" these days. It'd
| be sci-hub except they seem to have managed to kill it finally
|
| The business of IP has so hollowed out the credibility of its
| brokers that the business part (like the stock price) has become
| its only remaining value, and scientific journals are absolutely
| included here if their publishing method doesn't _preserve the
| goddamn body of evidence we claim science relies on_
| pk-protect-ai wrote:
| Do not forget that you are charged when you publish your work
| in scientific journals. Would you not expect your published
| work to be preserved?
| advael wrote:
| One might think this, but that you are guaranteed anything
| (even if you are explicitly promised it) by a for-profit
| corporation is a foolhardy and sometimes even dangerous thing
| to take for granted
| setgree wrote:
| Prior: most published research is false [0].
|
| If a paper is being forgotten by the internet, that's a weak
| signal that it was not valued by its field. If you combine that
| weak signal with the strong prior, you'd conclude that the paper
| wasn't a meaningful contribution and can be safely forgotten.
|
| I can see why this mass forgetting is a problem for
| preservationists and historians of science.
|
| [0]
| https://journals.plos.org/plosmedicine/article?id=10.1371/jo...
|
| [1] e.g. Gordon Allport's "The Nature of Prejudice" was published
| 1954, has been cited >50,000 times, and is still in print:
| https://www.amazon.com/Nature-Prejudice-25th-Anniversary/dp/...
| marcosdumay wrote:
| > If a paper is being forgotten by the internet, that's a weak
| signal that it was not valued by its field.
|
| That would be valid if the decision was made by popular vote.
|
| It's completely invalid if a single entity is allowed to decide
| whether it lives or dies.
| plumeria wrote:
| Some research is forgotten/dismissed by some generation only to
| be revived by a later generation (see e.g. Geometric Algebra)
| squigz wrote:
| Science doesn't work like that; we can't look at data/research
| at our current point in time and deduce its value forever.
| Having this information years down the line can inspire some
| more research ideas or contribute some groundwork to future
| research
| kragen wrote:
| to tell whether a piece of research is false or not, you need
| to be able to find out what it was. then you can conclusively
| prove it false. lost research like tesla's weather control
| notes, the alchemists' transmutation of lead into gold, or
| water-powered cars can instead give rise to conspiracy theories
| about how it was suppressed by the powerful and wasted
| lifetimes that could have been dedicated to something
| productive
| eviks wrote:
| The probability assessment is fine, but meaningless even
| outside history. Baby and the bathwater
| jjslocum3 wrote:
| I imagine this would be of great benefit to plagiarists.
| daniel31x13 wrote:
| This is one of the main reasons I created Linkwarden - to combat
| Link-Rot.
|
| Linkwarden is an open-source collaborative bookmark manager to
| collect, organize and preserve webpages:
|
| https://linkwarden.app
| Brian_K_White wrote:
| Suggestion: Find your own name that isn't coat-tailing on
| bitwarden.
|
| Or don't, I'm not your mom, but the name tells me not to even
| investigate, and I care about such things enough to run a YaCy
| node for example.
| severine wrote:
| So, it has to be called Yet Another Something?
|
| C'mon, let's not gatekeep or trademarkify normal descriptive
| or allusive words. Linkwarden is fine!
| nemoniac wrote:
| Use DOIs, they said. URLs are just transient, they said. But DOIs
| are permanent, they said.
|
| Could it be that the whole DOI system was just a ruse by the
| commercial academic publishers to re-intermediate themselves
| between academic authors and readers in the age of the WWW?
| jszymborski wrote:
| I really don't see DOIs as the problem here. They are just
| catalogue numbers, not an archive or repository.
|
| The DOIs are still permanent, but their resolvers might not be.
| Copyright problems definitely makes that harder to fix, but
| again, I don't think the DOI is the problem here.
| chaseadam17 wrote:
| Arweave fixes this. It's pay once permanent decentralized
| storage. One of the few web3 projects with fast-growing real
| world usage.
| bogtog wrote:
| I wonder if this is mostly just the DOIs associated with weak or
| predatory journals
| renegade-otter wrote:
| By the way, what's on the Internet is NOT forever. I expect that
| Facebook and Twitter datasets, for example, will eventually go
| offline, sold to AI companies, slowly reaching irrelevancy and
| finally resting in some cold data storage or used for
| training/educational purposes.
| wepple wrote:
| It's the worst of both worlds
|
| You have to assume an unappealing photo of yourself you posted
| in your youth _is_ stored somewhere forever
|
| And, things you actually want to be stored forever and be
| retrievable when useful, probably won't.
| renegade-otter wrote:
| Good point :)
| InSteady wrote:
| We shall henceforth refer to this as Wepple's paradox.
| kkfx wrote:
| That's why we need to push a desktop-centric internet, not a
| thin-client/endpoints one. People MUST learn to store their data,
| distributed storage must became the norm for public stuff, so
| anything someone value will be preserved.
|
| Just start with Zotero to have your papers LOCALLY (not on Zotero
| cloud), it's a small simple step anyone have enough storage to
| make.
| ericra wrote:
| Think of how trivial it would be, from a storage perspective, to
| have a central repo of every (publicly-funded, at a minimum)
| research paper ever written if copyright issues were not present.
|
| It could be run by an international non-profit or as a
| collaboration between governments. It wouldn't be particularly
| expensive to run when weighed against the huge long-term
| benefits.
|
| It's really a failure of the U.S. and EU governments to not clamp
| down on the Elseviers of the world who bring very few benefits at
| a high cost to open science and society at large.
| observationist wrote:
| The Library of Congress has the digital infrastructure and
| expertise to be able to handle something like this issue.
|
| Unfortunately, scientific journals are a sacred cash cow and
| infringing on any of their territory, real or imagined,
| prevents any meaningful top down change within the system.
| They've got money to pay lawyers to prevent any reform or
| mandates or flexing of existing regulatory powers.
|
| Pirate everything.
|
| Publicly funded research gets lost because it negatively
| affects profits to maintain unpopular, unread material in any
| sort of diligent and effective way.
|
| Journals have close ties to universities and academia, and big
| commercial research outfits, and all of the social ties being
| involved with those circles can bring. They've got lobbying
| perfected to an art, and pay good, ruthless lawyers to protect
| their interests.
|
| The average voter won't ever care enough to make a popular
| revolt, bottom up change possible; scientific publishing is too
| dry and anemic when you contrast against the million other,
| more outrageous, imminently threatening issues people care
| about. When you look at the conflicted interests that benefit
| from the status quo, such as companies that can pay Journals so
| their studies and papers will appear alongside distinguished,
| credentialed works, there doesn't appear to be any place where
| effective leverage can be applied.
|
| Pirating it all is the only ethical solution. Outlaws with
| rogue copies is the only feasible way much of this data will
| carry on into the future. None of the people who can change
| things actually seem to want to, and the public has more
| pressing matters capturing their attention.
|
| Aaron Swartz had it right. Apathy, greed, and politics aren't a
| problem with a solution that will come from this space. The
| only winning move is not to play their game.
|
| The same largely applies to any media content being gatekept by
| entities requiring repeated, endless rent on their "property"
| despite bringing no value to the market, existing simply to
| expand, endlessly, mindlessly shoveling other people's money
| into their shareholder's pockets.
|
| We live in a world that is awfully stupid sometimes. So stupid
| that important scientific literature is being faded into
| oblivion simply because it can't be monetized under perverse
| adtech incentive schemes. This has crucial implications,
| because if the habit of lazy citation takes root, it creates a
| kind of secular system of faith in scientism, with authors
| being given the benefit of the doubt when their citations can't
| be verified. That could be catastrophic if it affects
| medicines, engineering, environmental management, urban
| planning, forensics, or a myriad other categories where a
| flawed scientific paper might pose a threat.
|
| Knowing that flawed papers get through, what could an actual
| malicious actor get away with?
|
| Don't let stupid win.
|
| Pirate everything.
___________________________________________________________________
(page generated 2024-03-04 23:02 UTC)