RECOVERIES 004 What the archive does not hold ------------------------------------------------------------------ CORRECTION, 11 AUGUST 2026. When first published this post said 882 addresses were unusable and 1,151 entries were at issue. Both were too high. 614 of them were news: and mailto: addresses, which are good URIs that simply have no hostname, because that is how those schemes work. My test looked for a hostname, did not find one, and called them broken. The real figures are 268 unusable, 537 at issue, 1.5 per cent, floor 35,253. A post about not trusting a count got its own count wrong, in the same direction it was warning about. Every entry merged from comp.infosystems.www.announce and The Scout Report carried a tag that said Unverified. The tag was honest. It meant only that no one had ever checked. So I checked. I put 6,204 addresses to the Wayback Machine, one at a time, six seconds apart, over most of a day. Addresses asked of the archive 6,204 Excluded - the lookup itself errored 25 MEASURED - the denominator 6,179 Exact address archived, with a capture year 5,224 Page gone, but the host was reached 806 Host empty, but gopher/FTP/telnet/IP/odd port 126 NO ARCHIVE HOLDS THE HOST AT ALL 23 An error is an unresolved question, not a zero, so those twenty- five sit outside every figure above. The years land where you would expect. 1,034 first seen in 1996, 1,279 in 1997, 1,151 in 1998, 1,167 in 1999. Then the count collapses: 453 in 2000, 68 in 2001, 32 in 2002. The Wayback Machine was not watching the early web as it happened. It was watching it as it ended. The interesting number is twenty-three. Twenty-three addresses were published by an author or an editor, on a dated day, and no archive anywhere holds them. Not the page. Not the host. Nothing. That number was very nearly one hundred and forty-nine. I had written it down and I was ready to say it. Then I looked at what the 149 actually were, and 109 of them were gopher addresses. The Wayback Machine barely indexes gopher. A few more were bare IP addresses, which are never crawled under a hostname, and a handful were FTP, telnet, or an odd port. An empty result for those says something about the archive's coverage and nothing whatever about the site. Calling them absent would have been a finding the data does not support. So the number is twenty-three, and twenty-three that I can defend is worth more than one hundred and forty-nine that I cannot. The caveat has to be carried properly. An empty result means never crawled, or excluded by robots.txt, or removed on request. From outside, those three are indistinguishable. "The archive holds nothing" is not the same sentence as "It never existed". THE COUNT ON MY OWN FRONT PAGE IS WRONG While doing this, I measured the corpus itself, and I do not like what it says. The front page claims 35,790 catalogued sites. Of those, 268 have an address that cannot be used, and 269 are the same address recorded twice. That is 537 entries, or 1.5 per cent. The defensible floor is 35,253. Worse, the note where I recorded this problem was itself wrong. It said 1,127 duplicates. The real figure is 269. The canonicaliser returns nothing when it cannot parse an address, so every unparseable entry collapsed into the same bucket and was counted as a duplicate of the others. An error was counted as a finding. That is the exact mistake this project keeps making, and this time it was sitting in the very file meant to catch it. THE 232 ARE NOT RUBBISH 232 of those addresses are not addresses at all. They read like this: mail almanac@acenet.auburn.edu (in body of letter...) That is not a broken address. That is how you reached a resource before you could link to one. Somebody wrote down an access method, and the field it went into only knows how to hold addresses. 98 are instructions to send mail, 34 are finger, 23 gopher, 15 ftp, 14 telnet, two WAIS, and the rest are prose. Twelve more are simply mistyped, http// with the colon missing, and those I can repair. I am not going to delete them. They are a record of a way of using the internet that stopped existing, sitting in a catalogue with no field to describe it. They need a field, not a broom. The other 614 needed nothing at all. They were news: newsgroups and mailto: addresses, and they had been correct the whole time. What was broken was the test I measured them with. The catalogue is smaller than I said it was. Twenty-three addresses have no archive anywhere. And the file I keep to protect me from bad numbers had a bad number in it. I would rather publish all three than a round figure I cannot stand behind. ------------------------------------------------------------------ Khalid Alshaikh, 7Z1FP https://ge97.com/ .