[HN Gopher] Visualizing All ISBNs
___________________________________________________________________
Visualizing All ISBNs
Author : RyanShook
Score : 315 points
Date : 2025-01-10 04:45 UTC (18 hours ago)
(HTM) web link (annas-archive.org)
(TXT) w3m dump (annas-archive.org)
| skrebbel wrote:
| I thought it was my color blindness that made me not able to
| distinguish between the red and green pixels as described (i only
| see red and black ones), but even with a browser extension that
| counters color blindness i can't distinguish more colors. Is this
| just me, or is the graph weird?
| superzamp wrote:
| The graph seems to be alright, there are indeed red and (some)
| green pixels, looks like an issue with your extension
| unfortunately.
| psychoslave wrote:
| No idea of were the issue might land, but I can see the
| difference in colors.
| rendx wrote:
| I see green dots and a few lines of green dots. Did you try
| zooming in?
| asfasdfasdfn wrote:
| The graphs are very easy to read, albeit depend on your ability
| to distinguish between red and green.
|
| Can you change the green channel to blue to better view it?
| saithound wrote:
| Fwiw (not color-blind) I can see red, green and black pixels.
| The graph doesn't look weird to the naked eye.
|
| Find the interactive visualiser by scrolling down, and switch
| it to "Files in Anna's Archive [md5]". This will highlight the
| location of the green pixels in grey.
| Finnucane wrote:
| I am also color blind and the graph is not good.
| Muehe wrote:
| If you have red-green blindness like me try this:
|
| - Right-click the image and select "Inspect".
|
| - Add a new CSS hue-rotate filter to the element:
| element { max-width: 100%; margin: 0
| auto; filter: hue-rotate(-90deg); }
|
| Usually I use "filter: saturate(100);", but that didn't really
| work well for this image. You might have to adjust the rotation
| degree though, -90 worked best for me.
| thaumasiotes wrote:
| I see red, green, and a bit of yellow. I assume the yellow is
| what happens when the red and green pixels come too close to
| each other.
| whataguy wrote:
| > Each pixel represents 2,500 ISBNs. If we have a file for an
| ISBN, we make that pixel more green.
|
| What do you mean by "more green"? I don't see any shaded green.
|
| And I presume the black pixels are unregistered ISBNs?
| lmm wrote:
| If you look closely there are definitely some brownish pixels
| and some dim greens.
| quink wrote:
| Kind of hard to tell what corresponds to what in these graphs,
| maybe if someone could point out Bookland (i.e. 978), it would be
| a bit easier to orient oneself?
| seszett wrote:
| Making it easier to visualise is the whole point of the bounty
| announced by this post.
| eporomaa wrote:
| Hm, I got:
|
| "...
|
| European sanctions
|
| The Council of Europe has decided that the websites of RT
| (formerly Russia Today) and Sputnik News may no longer be
| transmitted. The website you are trying to visit falls under this
| European sanction.
|
| ..."
| TonyTrapp wrote:
| Works fine here from a European IP.
| jaapz wrote:
| It's blocked at least in the Netherlands. Weirdly it mentions
| it being part of the sanctions against Russia, while from a
| cursory search I only found a judge ordering the site to be
| blocked because of copyright issues (thanks Brein). They
| probably just show the wrong error page?
| rchard2scout wrote:
| It's blocked by my corporate networking filter for me, in
| the category "Illegal downloads". So the Russian sanctions
| message is probably incorrect indeed.
| Cthulhu_ wrote:
| Must be ISP specific, I'm also in NL and can access it
| fine.
| rollulus wrote:
| I'm also in NL. Ziggo's DNS server blocks it:
| $ dig annas-archive.org @89.101.251.228 annas-
| archive.org. 360 IN CNAME unavailable.for.legal.reasons.
| unavailable.for.legal.reasons. 339 IN A 213.46.185.10
|
| 213.46.185.10 serves a generic page mentioning Russia Today
| and the Pirate Bay. Not sure which one applies here.
| Freak_NL wrote:
| Same for KPN:
|
| http://195.121.82.125/
|
| Would Tweak have blocked this? Most households in the
| Netherlands currently have the choice of Ziggo, KPN, and
| Odido. Long live VPNs...
| xp84 wrote:
| Is that _three broadband providers_ serving the same
| address?? You guys are so lucky you don't even know. In
| America we generally have a choice of _one_ if you aren't
| including Starlink or legacy slow satellite. And perhaps
| a joke of a 1-6Mbps DSL option in some parts.
| seszett wrote:
| > CNAME unavailable.for.legal.reasons.
|
| Not really standards compliant, but an interesting use of
| DNS.
| reddalo wrote:
| I think the website is censored at DNS level but they chose the
| wrong error page.
|
| In Italy it just errors out with a NS_ERROR_CONNECTION_REFUSED.
| flir wrote:
| You're just cleared up a minor mystery I never bothered to
| investigate (BT, UK). Thanks.
|
| Flipping DNS to 8.8.4.4 fixed it for now but I really need to
| move this connection to A&A.
| powerhugs wrote:
| Switch DNS to like 1.1.1.1 (Cloudflare) or 8.8.8.8 (Google)
| sebstefan wrote:
| >$10,000 bounty
|
| >There is much to explore here, so we're announcing a bounty for
| improving the visualization above. Unlike most of our bounties,
| this one is time-bound. You have to submit your open source code
| by 2025-01-31 (23:59 UTC).
|
| >The best submission will get $6,000, second place is $3,000, and
| third place is $1,000.
|
| >All bounties will be awarded using Monero (XMR).
|
| ? Why are they using crypto, and, weirdly enough, specifically
| the crypto people use for buying drugs, to award this?
|
| Is it some kind of scam?
| yawndex wrote:
| Because the efforts of Anna's Archive are unfortunately
| currently very much illegal, and XMR is one of the few
| cryptocurrencies that can actually offer some privacy to its
| users.
| sebstefan wrote:
| I've used XMR before. Just surprised seeing it to pay for
| legitimate & harmless visualization work.
|
| I see, that makes sense
| aprilnya wrote:
| So what you're saying is you think XMR is just for buying
| drugs, and you're also saying you've used XMR before.
|
| Hmmmmmm
|
| /s
| Klaus23 wrote:
| Because it is a book download site, which is illegal in every
| country that has copyright, and revealing one's identity with a
| bank transfer would be a stupid way to go to jail.
| akimbostrawman wrote:
| >Why are they using crypto, and, weirdly enough, specifically
| the crypto people use for buying drugs, to award this?
|
| You really have to ask why a illegal/grey site is using
| currency that is build to protect privacy and anonymity?
|
| is this some kind of sarcasm?
| fear-anger-hate wrote:
| They use monero because what they are doing (copyright
| infringement) will get you in to big trouble anywhere in the
| western world. Without cryptocurrencies much of the modern
| large scale archival efforts wouldn't be possible, or at the
| very least would significantly increase risks for the people
| participating in it. For me this alone is a good enough reason
| to admit that there are valid reasons for existence of privacy
| coins.
|
| The harm they may cause in the short term via tax avoidance or
| being used to buy drugs is minimal, but the possibility that
| because of them archivists are able to fund servers for data
| that future historians wouldn't have otherwise been able to get
| their hands on? Priceless.
| billpg wrote:
| Anyone else seeing this?
|
| "This server couldn't prove that it's annas-archive.org; its
| security certificate is from *.hs.llnwd.net. This may be caused
| by a misconfiguration or an attacker intercepting your
| connection."
| c0balt wrote:
| No, sounds like you are being mitm for them. Though the domain
| appears like a legitimate CDN.
| swores wrote:
| Same for me
| masfuerte wrote:
| Yes. A DNS request for annas-archive.org to my ISP (EE in the
| UK) returns an address for www.ukispcourtorders.co.uk, which
| also gives a security warning. If I click through the warning
| on either site I get an HTTP 400 error.
|
| According to Wikipedia, www.ukispcourtorders.co.uk used to list
| the blocked domains and the court orders responsible.
|
| https://en.wikipedia.org/wiki/List_of_websites_blocked_in_th...
| WillAdams wrote:
| The thing is, ISBNs aren't hierarchical --- they are bought in
| blocks (or even individually at an exorbitant markup, says the
| guy who bought one to reprint a single book), so this doesn't
| show anything really interesting/useful.
|
| A visualization using LoC or even Dewey Decimal would be far more
| useful, esp. if it also linked to public domain and copyright-
| free repositories/lists, say an interactive and visual version of
| John Mark Ockerbloom's:
|
| https://onlinebooks.library.upenn.edu/
| MarceColl wrote:
| It shows what they want to show, which is mostly how much of
| the world books they have. Hierarchical has nothing to do with
| it.
| Finnucane wrote:
| It only sort of shows that. ISBNs are issued by edition, not
| title, so many books would have more than one. And books
| published before 1970 or so might not be represented at all
| if they have no recent edition.
| NoMoreNicksLeft wrote:
| They can't even have a tiny fraction of the world's books.
| Each edition of the book gets a new ISBN... if a book is
| released as a paperback, hardback, kindle edition, pdf, and
| epub then there are supposed to be five ISBNs.
|
| The vast, vast majority have only been released as dead-tree
| versions. They have none of those. The books they scan may
| have an ISBN, but the scans do not have them. Like all
| Project Gutenberg books, their books have no ISBNs at all.
| From a strict point of view, they've released new editions of
| these books.
| nickelpro wrote:
| Worthless semantics in the context of the mission of the
| project.
|
| What you've described is that the archived content can be
| mapped to multiple ISBNs. It's clear the only element of
| concern here is the content itself. The failure to preserve
| a particular binding or printer's choice of typeface is
| irrelevant.
|
| Failing to recognize this requires an almost malicious
| level of pedantry
| jameshart wrote:
| A successful archival of one of those ISBNs will light
| up; four of those ISBNs remain dark. Yet they have that
| content archived. It means that lighting up the entire
| grid is not necessary to achieve their goal.
|
| Indeed a bigger problem is that it's much harder to know
| which areas of the grid are never going to light up
| because the ISBN has not been used.
| NoMoreNicksLeft wrote:
| >Worthless semantics in the context of the mission of the
| project.
|
| Hardly worthless... often times, the edition of the book
| matters as much as the title. Steven King wrote two books
| named _The Stand_ , and one isn't anything like the
| other. He pulled a Lucas pretty early on.
|
| He's hardly the only author to ever do this. But it's not
| just authors either. Editors, collectors, translators all
| make their mark, and give you works that though they
| might be slightly different to you, the differences
| actually matter to the rest of us. It's not that you're
| ignorant that offends me, it's the arrogance about a
| subject you seem to know so little about that makes it
| difficult to tolerate.
|
| There is no pedantry here, just a desire to actually
| preserve books and to organize them.
| mmooss wrote:
| > The books they scan may have an ISBN, but the scans do
| not have them. Like all Project Gutenberg books, their
| books have no ISBNs at all. From a strict point of view,
| they've released new editions of these books.
|
| Are you saying they actively remove ISBN numbers from
| scans? If I downloaded one of the books, it wouldn't have
| an ISBN?
|
| Why? That seems like a bunch of extra processing per book,
| makes it harder for users to specifically identify a book,
| and probably does nothing for legality. Also, can people
| search by ISBN?
| Tomte wrote:
| > Are you saying they actively remove ISBN numbers from
| scans?
|
| No, he's playing the pointless ,,well, actually a scan of
| a book is a different thing from the book itself" game.
| NoMoreNicksLeft wrote:
| No, I'm saying that the ISBN doesn't describe titles, it
| describes editions, and editions matter.
| est31 wrote:
| ISBN's are hierarchical, what do you mean? Like Gaul, ISBNs are
| divided into multiple parts, where one part is for the
| language, another is for the publisher, and the last is for the
| title. The last part is a checksum.
| https://en.wikipedia.org/wiki/ISBN#Overview
| WillAdams wrote:
| Yes, but this internal hierarchy for an issued number doesn't
| tell anything beyond those facts about a specific edition of
| a specific text.
|
| One can't use ISBNs alone to create a hierarchical listing of
| texts which is useful for anything beyond browsing by
| language/publisher/order in which the ISBN was generated.
|
| A visual and interactive representation of books by LoC or
| some other cataloging system would actually be useful.
| convolvatron wrote:
| totally agree, but thats not in the data. however, since
| blocks are assigned to agencies associated with countries
| and publishers, you might find some utility in showing
| coverage by likely language and/or country of origin and
| date.
| PaulHoule wrote:
| I got into an argument with the manager of South End Press
| back in '94 about whether 'Futuresplash' (soon to be
| Macromedia Flash) had a future, he thought it did and he
| was right.
|
| Years later I was working at the library and got a little
| bit steamed because South End Press was reusing ISBN's
| after books went out of print which was allowed but, I
| think, lame.
|
| One of my strategies for researching a topic is looking a
| few up in the OPAC, finding them in the stacks, and finding
| more books on the topic in those areas. (In the Library of
| Congress system, machine vision could be under QA56 with
| the rest of computer science or around TA1630, thus
| "areas".)
|
| From time to time I've thought about trying to replicate
| the feel of this with some kind of UI given that our
| library moved a lot of the collection into deep archives
| and we have a very fast 'Borrow Direct' service with other
| peers)
| omoikane wrote:
| One thing it shows is how ISBNs are allocated much faster than
| they are used, judging by the amount of black pixels.
|
| The image contains 1000*800 pixels at 2500 ISBNs per pixel, so
| it's visualizing 2e9 ISBNs. ISBN-13 contains 12 digits plus one
| check digit, so we might have expected the image to be 500
| times bigger/denser than the current image. The fact that it's
| at its current size suggests that only ISBNs with 978 and 979
| prefixes are included, and since the bottom half is more
| sparse, that probably corresponds to the new 979 range.
| greenie_beans wrote:
| is it illegal to download and use their isbn file? like what is
| wrong with having that information?
| karel-3d wrote:
| I don't think this page, which links to libgen and sci-hub, is
| that concerned about copyright.
| greenie_beans wrote:
| annoying non-answer to my question. i already know all about
| anna's archive. i'm asking if a person can download these
| isbns and use them to make data visualizations without fear
| of breaking a law? https://software.annas-
| archive.li/AnnaArchivist/annas-archiv...
| salomonk_mur wrote:
| They explicitly provide that data for you to do as you
| wish. They are in a grey area, not you. You can download it
| no problem.
| greenie_beans wrote:
| is there legal precedent for that?
|
| already asked LLMs so please don't copy/paste an LLM
| response.
| eemil wrote:
| Depends on your jurisdiction.
| karel-3d wrote:
| Sorry, I misunderstood your question.
| qingcharles wrote:
| Seeing as nobody has provided a real answer. The question
| is, maybe.
|
| Anna's Archive is getting sued currently for scraping vast
| amounts of essentially public metadata which was being
| gate-keeped by a single organisation.
|
| Here's the longer and more complicated answer for you:
|
| https://libraries.emory.edu/research/copyright/copyright-
| dat...
| greenie_beans wrote:
| feist is what comes up when i search around, too. the
| ISBNs might be poisoned if anna broke terms of service to
| get the ISBNs
| jdblair wrote:
| It appears that the IP of the server is blocked in the EU. I get
| this from my ISP (Ziggo, in the Netherlands):
|
| Deze website is geblokkeerd
|
| Europese sancties
|
| De Raad van Europa heeft besloten dat de websites van RT
| (voorheen Russia Today) en Sputnik News niet meer mogen worden
| doorgegeven. De website die je probeert te bezoeken, valt onder
| deze Europese sanctie.
|
| VodafoneZiggo is verplicht de sanctie uit te voeren en heeft de
| website geblokkeerd.
| voytec wrote:
| Works in Poland, but here you go:
|
| https://web.archive.org/web/20250106112552/https://annas-arc...
| hk__2 wrote:
| No issue here in France.
| manosyja wrote:
| Running your own recursive resolver has certain advantages...
| glimshe wrote:
| Anna's archive is one of the wonders of the world. If we almost
| destroyed our species but Anna's archive endured, there would be
| hope for a relatively expedient reconstruction.
| graypegg wrote:
| I see that bounty at the bottom, so tossing away my chances here,
| but this visualization is just asking to be mapped onto a Hilbert
| Curve. [0] When you "stripe" the data like this, points that are
| sorted close together could end up pretty far apart, since a
| distance in the Y axis skips an entire row of data as you move
| down, rather than a distance in the X axis which is 1-to-1 with
| the source data.
|
| If you map it onto a hilbert curve, the X and Y axis mean
| nothing, but visually points that are close together in the
| sorted list, will be visually close together in the output image.
|
| Since the first part of an ISBN is the country, then the second
| part is the publisher, and the third part is the title, with a
| check sum at the end, I would remove the checksum and sort them
| each as a big number. (no hyphens)
|
| You should end up with "islands", where you see big areas covered
| by big publishing countries, with these "islands" having bright
| spots for the publisher codes.
|
| Bonus points for labeling these areas!
|
| I set up something a while ago [1] for an interview that does
| this with weather data. It makes the seasons really obvious since
| they're all grouped together.
|
| [0] https://en.wikipedia.org/wiki/Hilbert_curve
|
| [1] https://graypegg.com/hilbert
| (https://github.com/graypegg/hilbertcurveplayground code if
| anyone wants to go for the prize using this! Please at least
| mention me if you decide to reuse this code, but I can't stop ya
| lol)
| abetusk wrote:
| And there's a generalized Hilbert curve, the Gilbert curve, for
| non powers of two rectangular regions [0] (online demo [1]).
|
| [0] https://github.com/jakubcerveny/gilbert
|
| [1] https://jakubcerveny.github.io/gilbert/demo/
| qingcharles wrote:
| Now do ISSNs, please.
| ge96 wrote:
| Ooh prize money, D3 those are fun, where you can map a million
| things/zoom into it
___________________________________________________________________
(page generated 2025-01-10 23:00 UTC)