[HN Gopher] Visualizing All ISBNs
       ___________________________________________________________________
        
       Visualizing All ISBNs
        
       Author : RyanShook
       Score  : 315 points
       Date   : 2025-01-10 04:45 UTC (18 hours ago)
        
 (HTM) web link (annas-archive.org)
 (TXT) w3m dump (annas-archive.org)
        
       | skrebbel wrote:
       | I thought it was my color blindness that made me not able to
       | distinguish between the red and green pixels as described (i only
       | see red and black ones), but even with a browser extension that
       | counters color blindness i can't distinguish more colors. Is this
       | just me, or is the graph weird?
        
         | superzamp wrote:
         | The graph seems to be alright, there are indeed red and (some)
         | green pixels, looks like an issue with your extension
         | unfortunately.
        
         | psychoslave wrote:
         | No idea of were the issue might land, but I can see the
         | difference in colors.
        
         | rendx wrote:
         | I see green dots and a few lines of green dots. Did you try
         | zooming in?
        
         | asfasdfasdfn wrote:
         | The graphs are very easy to read, albeit depend on your ability
         | to distinguish between red and green.
         | 
         | Can you change the green channel to blue to better view it?
        
         | saithound wrote:
         | Fwiw (not color-blind) I can see red, green and black pixels.
         | The graph doesn't look weird to the naked eye.
         | 
         | Find the interactive visualiser by scrolling down, and switch
         | it to "Files in Anna's Archive [md5]". This will highlight the
         | location of the green pixels in grey.
        
         | Finnucane wrote:
         | I am also color blind and the graph is not good.
        
         | Muehe wrote:
         | If you have red-green blindness like me try this:
         | 
         | - Right-click the image and select "Inspect".
         | 
         | - Add a new CSS hue-rotate filter to the element:
         | element {            max-width: 100%;            margin: 0
         | auto;            filter: hue-rotate(-90deg);         }
         | 
         | Usually I use "filter: saturate(100);", but that didn't really
         | work well for this image. You might have to adjust the rotation
         | degree though, -90 worked best for me.
        
         | thaumasiotes wrote:
         | I see red, green, and a bit of yellow. I assume the yellow is
         | what happens when the red and green pixels come too close to
         | each other.
        
       | whataguy wrote:
       | > Each pixel represents 2,500 ISBNs. If we have a file for an
       | ISBN, we make that pixel more green.
       | 
       | What do you mean by "more green"? I don't see any shaded green.
       | 
       | And I presume the black pixels are unregistered ISBNs?
        
         | lmm wrote:
         | If you look closely there are definitely some brownish pixels
         | and some dim greens.
        
       | quink wrote:
       | Kind of hard to tell what corresponds to what in these graphs,
       | maybe if someone could point out Bookland (i.e. 978), it would be
       | a bit easier to orient oneself?
        
         | seszett wrote:
         | Making it easier to visualise is the whole point of the bounty
         | announced by this post.
        
       | eporomaa wrote:
       | Hm, I got:
       | 
       | "...
       | 
       | European sanctions
       | 
       | The Council of Europe has decided that the websites of RT
       | (formerly Russia Today) and Sputnik News may no longer be
       | transmitted. The website you are trying to visit falls under this
       | European sanction.
       | 
       | ..."
        
         | TonyTrapp wrote:
         | Works fine here from a European IP.
        
           | jaapz wrote:
           | It's blocked at least in the Netherlands. Weirdly it mentions
           | it being part of the sanctions against Russia, while from a
           | cursory search I only found a judge ordering the site to be
           | blocked because of copyright issues (thanks Brein). They
           | probably just show the wrong error page?
        
             | rchard2scout wrote:
             | It's blocked by my corporate networking filter for me, in
             | the category "Illegal downloads". So the Russian sanctions
             | message is probably incorrect indeed.
        
             | Cthulhu_ wrote:
             | Must be ISP specific, I'm also in NL and can access it
             | fine.
        
             | rollulus wrote:
             | I'm also in NL. Ziggo's DNS server blocks it:
             | $ dig annas-archive.org @89.101.251.228       annas-
             | archive.org. 360 IN CNAME unavailable.for.legal.reasons.
             | unavailable.for.legal.reasons. 339 IN A 213.46.185.10
             | 
             | 213.46.185.10 serves a generic page mentioning Russia Today
             | and the Pirate Bay. Not sure which one applies here.
        
               | Freak_NL wrote:
               | Same for KPN:
               | 
               | http://195.121.82.125/
               | 
               | Would Tweak have blocked this? Most households in the
               | Netherlands currently have the choice of Ziggo, KPN, and
               | Odido. Long live VPNs...
        
               | xp84 wrote:
               | Is that _three broadband providers_ serving the same
               | address?? You guys are so lucky you don't even know. In
               | America we generally have a choice of _one_ if you aren't
               | including Starlink or legacy slow satellite. And perhaps
               | a joke of a 1-6Mbps DSL option in some parts.
        
               | seszett wrote:
               | > CNAME unavailable.for.legal.reasons.
               | 
               | Not really standards compliant, but an interesting use of
               | DNS.
        
         | reddalo wrote:
         | I think the website is censored at DNS level but they chose the
         | wrong error page.
         | 
         | In Italy it just errors out with a NS_ERROR_CONNECTION_REFUSED.
        
           | flir wrote:
           | You're just cleared up a minor mystery I never bothered to
           | investigate (BT, UK). Thanks.
           | 
           | Flipping DNS to 8.8.4.4 fixed it for now but I really need to
           | move this connection to A&A.
        
         | powerhugs wrote:
         | Switch DNS to like 1.1.1.1 (Cloudflare) or 8.8.8.8 (Google)
        
       | sebstefan wrote:
       | >$10,000 bounty
       | 
       | >There is much to explore here, so we're announcing a bounty for
       | improving the visualization above. Unlike most of our bounties,
       | this one is time-bound. You have to submit your open source code
       | by 2025-01-31 (23:59 UTC).
       | 
       | >The best submission will get $6,000, second place is $3,000, and
       | third place is $1,000.
       | 
       | >All bounties will be awarded using Monero (XMR).
       | 
       | ? Why are they using crypto, and, weirdly enough, specifically
       | the crypto people use for buying drugs, to award this?
       | 
       | Is it some kind of scam?
        
         | yawndex wrote:
         | Because the efforts of Anna's Archive are unfortunately
         | currently very much illegal, and XMR is one of the few
         | cryptocurrencies that can actually offer some privacy to its
         | users.
        
           | sebstefan wrote:
           | I've used XMR before. Just surprised seeing it to pay for
           | legitimate & harmless visualization work.
           | 
           | I see, that makes sense
        
             | aprilnya wrote:
             | So what you're saying is you think XMR is just for buying
             | drugs, and you're also saying you've used XMR before.
             | 
             | Hmmmmmm
             | 
             | /s
        
         | Klaus23 wrote:
         | Because it is a book download site, which is illegal in every
         | country that has copyright, and revealing one's identity with a
         | bank transfer would be a stupid way to go to jail.
        
         | akimbostrawman wrote:
         | >Why are they using crypto, and, weirdly enough, specifically
         | the crypto people use for buying drugs, to award this?
         | 
         | You really have to ask why a illegal/grey site is using
         | currency that is build to protect privacy and anonymity?
         | 
         | is this some kind of sarcasm?
        
         | fear-anger-hate wrote:
         | They use monero because what they are doing (copyright
         | infringement) will get you in to big trouble anywhere in the
         | western world. Without cryptocurrencies much of the modern
         | large scale archival efforts wouldn't be possible, or at the
         | very least would significantly increase risks for the people
         | participating in it. For me this alone is a good enough reason
         | to admit that there are valid reasons for existence of privacy
         | coins.
         | 
         | The harm they may cause in the short term via tax avoidance or
         | being used to buy drugs is minimal, but the possibility that
         | because of them archivists are able to fund servers for data
         | that future historians wouldn't have otherwise been able to get
         | their hands on? Priceless.
        
       | billpg wrote:
       | Anyone else seeing this?
       | 
       | "This server couldn't prove that it's annas-archive.org; its
       | security certificate is from *.hs.llnwd.net. This may be caused
       | by a misconfiguration or an attacker intercepting your
       | connection."
        
         | c0balt wrote:
         | No, sounds like you are being mitm for them. Though the domain
         | appears like a legitimate CDN.
        
         | swores wrote:
         | Same for me
        
         | masfuerte wrote:
         | Yes. A DNS request for annas-archive.org to my ISP (EE in the
         | UK) returns an address for www.ukispcourtorders.co.uk, which
         | also gives a security warning. If I click through the warning
         | on either site I get an HTTP 400 error.
         | 
         | According to Wikipedia, www.ukispcourtorders.co.uk used to list
         | the blocked domains and the court orders responsible.
         | 
         | https://en.wikipedia.org/wiki/List_of_websites_blocked_in_th...
        
       | WillAdams wrote:
       | The thing is, ISBNs aren't hierarchical --- they are bought in
       | blocks (or even individually at an exorbitant markup, says the
       | guy who bought one to reprint a single book), so this doesn't
       | show anything really interesting/useful.
       | 
       | A visualization using LoC or even Dewey Decimal would be far more
       | useful, esp. if it also linked to public domain and copyright-
       | free repositories/lists, say an interactive and visual version of
       | John Mark Ockerbloom's:
       | 
       | https://onlinebooks.library.upenn.edu/
        
         | MarceColl wrote:
         | It shows what they want to show, which is mostly how much of
         | the world books they have. Hierarchical has nothing to do with
         | it.
        
           | Finnucane wrote:
           | It only sort of shows that. ISBNs are issued by edition, not
           | title, so many books would have more than one. And books
           | published before 1970 or so might not be represented at all
           | if they have no recent edition.
        
           | NoMoreNicksLeft wrote:
           | They can't even have a tiny fraction of the world's books.
           | Each edition of the book gets a new ISBN... if a book is
           | released as a paperback, hardback, kindle edition, pdf, and
           | epub then there are supposed to be five ISBNs.
           | 
           | The vast, vast majority have only been released as dead-tree
           | versions. They have none of those. The books they scan may
           | have an ISBN, but the scans do not have them. Like all
           | Project Gutenberg books, their books have no ISBNs at all.
           | From a strict point of view, they've released new editions of
           | these books.
        
             | nickelpro wrote:
             | Worthless semantics in the context of the mission of the
             | project.
             | 
             | What you've described is that the archived content can be
             | mapped to multiple ISBNs. It's clear the only element of
             | concern here is the content itself. The failure to preserve
             | a particular binding or printer's choice of typeface is
             | irrelevant.
             | 
             | Failing to recognize this requires an almost malicious
             | level of pedantry
        
               | jameshart wrote:
               | A successful archival of one of those ISBNs will light
               | up; four of those ISBNs remain dark. Yet they have that
               | content archived. It means that lighting up the entire
               | grid is not necessary to achieve their goal.
               | 
               | Indeed a bigger problem is that it's much harder to know
               | which areas of the grid are never going to light up
               | because the ISBN has not been used.
        
               | NoMoreNicksLeft wrote:
               | >Worthless semantics in the context of the mission of the
               | project.
               | 
               | Hardly worthless... often times, the edition of the book
               | matters as much as the title. Steven King wrote two books
               | named _The Stand_ , and one isn't anything like the
               | other. He pulled a Lucas pretty early on.
               | 
               | He's hardly the only author to ever do this. But it's not
               | just authors either. Editors, collectors, translators all
               | make their mark, and give you works that though they
               | might be slightly different to you, the differences
               | actually matter to the rest of us. It's not that you're
               | ignorant that offends me, it's the arrogance about a
               | subject you seem to know so little about that makes it
               | difficult to tolerate.
               | 
               | There is no pedantry here, just a desire to actually
               | preserve books and to organize them.
        
             | mmooss wrote:
             | > The books they scan may have an ISBN, but the scans do
             | not have them. Like all Project Gutenberg books, their
             | books have no ISBNs at all. From a strict point of view,
             | they've released new editions of these books.
             | 
             | Are you saying they actively remove ISBN numbers from
             | scans? If I downloaded one of the books, it wouldn't have
             | an ISBN?
             | 
             | Why? That seems like a bunch of extra processing per book,
             | makes it harder for users to specifically identify a book,
             | and probably does nothing for legality. Also, can people
             | search by ISBN?
        
               | Tomte wrote:
               | > Are you saying they actively remove ISBN numbers from
               | scans?
               | 
               | No, he's playing the pointless ,,well, actually a scan of
               | a book is a different thing from the book itself" game.
        
               | NoMoreNicksLeft wrote:
               | No, I'm saying that the ISBN doesn't describe titles, it
               | describes editions, and editions matter.
        
         | est31 wrote:
         | ISBN's are hierarchical, what do you mean? Like Gaul, ISBNs are
         | divided into multiple parts, where one part is for the
         | language, another is for the publisher, and the last is for the
         | title. The last part is a checksum.
         | https://en.wikipedia.org/wiki/ISBN#Overview
        
           | WillAdams wrote:
           | Yes, but this internal hierarchy for an issued number doesn't
           | tell anything beyond those facts about a specific edition of
           | a specific text.
           | 
           | One can't use ISBNs alone to create a hierarchical listing of
           | texts which is useful for anything beyond browsing by
           | language/publisher/order in which the ISBN was generated.
           | 
           | A visual and interactive representation of books by LoC or
           | some other cataloging system would actually be useful.
        
             | convolvatron wrote:
             | totally agree, but thats not in the data. however, since
             | blocks are assigned to agencies associated with countries
             | and publishers, you might find some utility in showing
             | coverage by likely language and/or country of origin and
             | date.
        
             | PaulHoule wrote:
             | I got into an argument with the manager of South End Press
             | back in '94 about whether 'Futuresplash' (soon to be
             | Macromedia Flash) had a future, he thought it did and he
             | was right.
             | 
             | Years later I was working at the library and got a little
             | bit steamed because South End Press was reusing ISBN's
             | after books went out of print which was allowed but, I
             | think, lame.
             | 
             | One of my strategies for researching a topic is looking a
             | few up in the OPAC, finding them in the stacks, and finding
             | more books on the topic in those areas. (In the Library of
             | Congress system, machine vision could be under QA56 with
             | the rest of computer science or around TA1630, thus
             | "areas".)
             | 
             | From time to time I've thought about trying to replicate
             | the feel of this with some kind of UI given that our
             | library moved a lot of the collection into deep archives
             | and we have a very fast 'Borrow Direct' service with other
             | peers)
        
         | omoikane wrote:
         | One thing it shows is how ISBNs are allocated much faster than
         | they are used, judging by the amount of black pixels.
         | 
         | The image contains 1000*800 pixels at 2500 ISBNs per pixel, so
         | it's visualizing 2e9 ISBNs. ISBN-13 contains 12 digits plus one
         | check digit, so we might have expected the image to be 500
         | times bigger/denser than the current image. The fact that it's
         | at its current size suggests that only ISBNs with 978 and 979
         | prefixes are included, and since the bottom half is more
         | sparse, that probably corresponds to the new 979 range.
        
       | greenie_beans wrote:
       | is it illegal to download and use their isbn file? like what is
       | wrong with having that information?
        
         | karel-3d wrote:
         | I don't think this page, which links to libgen and sci-hub, is
         | that concerned about copyright.
        
           | greenie_beans wrote:
           | annoying non-answer to my question. i already know all about
           | anna's archive. i'm asking if a person can download these
           | isbns and use them to make data visualizations without fear
           | of breaking a law? https://software.annas-
           | archive.li/AnnaArchivist/annas-archiv...
        
             | salomonk_mur wrote:
             | They explicitly provide that data for you to do as you
             | wish. They are in a grey area, not you. You can download it
             | no problem.
        
               | greenie_beans wrote:
               | is there legal precedent for that?
               | 
               | already asked LLMs so please don't copy/paste an LLM
               | response.
        
               | eemil wrote:
               | Depends on your jurisdiction.
        
             | karel-3d wrote:
             | Sorry, I misunderstood your question.
        
             | qingcharles wrote:
             | Seeing as nobody has provided a real answer. The question
             | is, maybe.
             | 
             | Anna's Archive is getting sued currently for scraping vast
             | amounts of essentially public metadata which was being
             | gate-keeped by a single organisation.
             | 
             | Here's the longer and more complicated answer for you:
             | 
             | https://libraries.emory.edu/research/copyright/copyright-
             | dat...
        
               | greenie_beans wrote:
               | feist is what comes up when i search around, too. the
               | ISBNs might be poisoned if anna broke terms of service to
               | get the ISBNs
        
       | jdblair wrote:
       | It appears that the IP of the server is blocked in the EU. I get
       | this from my ISP (Ziggo, in the Netherlands):
       | 
       | Deze website is geblokkeerd
       | 
       | Europese sancties
       | 
       | De Raad van Europa heeft besloten dat de websites van RT
       | (voorheen Russia Today) en Sputnik News niet meer mogen worden
       | doorgegeven. De website die je probeert te bezoeken, valt onder
       | deze Europese sanctie.
       | 
       | VodafoneZiggo is verplicht de sanctie uit te voeren en heeft de
       | website geblokkeerd.
        
         | voytec wrote:
         | Works in Poland, but here you go:
         | 
         | https://web.archive.org/web/20250106112552/https://annas-arc...
        
         | hk__2 wrote:
         | No issue here in France.
        
         | manosyja wrote:
         | Running your own recursive resolver has certain advantages...
        
       | glimshe wrote:
       | Anna's archive is one of the wonders of the world. If we almost
       | destroyed our species but Anna's archive endured, there would be
       | hope for a relatively expedient reconstruction.
        
       | graypegg wrote:
       | I see that bounty at the bottom, so tossing away my chances here,
       | but this visualization is just asking to be mapped onto a Hilbert
       | Curve. [0] When you "stripe" the data like this, points that are
       | sorted close together could end up pretty far apart, since a
       | distance in the Y axis skips an entire row of data as you move
       | down, rather than a distance in the X axis which is 1-to-1 with
       | the source data.
       | 
       | If you map it onto a hilbert curve, the X and Y axis mean
       | nothing, but visually points that are close together in the
       | sorted list, will be visually close together in the output image.
       | 
       | Since the first part of an ISBN is the country, then the second
       | part is the publisher, and the third part is the title, with a
       | check sum at the end, I would remove the checksum and sort them
       | each as a big number. (no hyphens)
       | 
       | You should end up with "islands", where you see big areas covered
       | by big publishing countries, with these "islands" having bright
       | spots for the publisher codes.
       | 
       | Bonus points for labeling these areas!
       | 
       | I set up something a while ago [1] for an interview that does
       | this with weather data. It makes the seasons really obvious since
       | they're all grouped together.
       | 
       | [0] https://en.wikipedia.org/wiki/Hilbert_curve
       | 
       | [1] https://graypegg.com/hilbert
       | (https://github.com/graypegg/hilbertcurveplayground code if
       | anyone wants to go for the prize using this! Please at least
       | mention me if you decide to reuse this code, but I can't stop ya
       | lol)
        
         | abetusk wrote:
         | And there's a generalized Hilbert curve, the Gilbert curve, for
         | non powers of two rectangular regions [0] (online demo [1]).
         | 
         | [0] https://github.com/jakubcerveny/gilbert
         | 
         | [1] https://jakubcerveny.github.io/gilbert/demo/
        
       | qingcharles wrote:
       | Now do ISSNs, please.
        
       | ge96 wrote:
       | Ooh prize money, D3 those are fun, where you can map a million
       | things/zoom into it
        
       ___________________________________________________________________
       (page generated 2025-01-10 23:00 UTC)