[HN Gopher] Internet search tips
       ___________________________________________________________________
        
       Internet search tips
        
       Author : herbertl
       Score  : 199 points
       Date   : 2023-07-22 03:22 UTC (19 hours ago)
        
 (HTM) web link (gwern.net)
 (TXT) w3m dump (gwern.net)
        
       | skimdesk wrote:
       | Another useful Google search trick not mentioned in the article
       | is numeric range queries.
       | 
       | You can use two numbers separated by two dots to represent all
       | numbers in the range.
       | 
       | For example, a search for                 taki 100..200
       | 
       | gives me results for taki 183.
       | 
       | This is useful when you can't remember an exact year or number.
        
       | Jakob wrote:
       | +1 for his tip of "Search The Internet Archive".
       | 
       | The other day I searched for the Tom & Jerry full episodes on the
       | web to no avail (streaming platforms and video platforms like
       | youtube only have rubbish cuts).
       | 
       | The internet archive has every episode starting from the first
       | one in 1940, in an easily accessible player without any ads or
       | recommendations: https://archive.org/details/tom-and-jerry-
       | all-114-episodes
        
         | nomilk wrote:
         | TIL: IA is more than Wayback Machine.
         | 
         | Something I'd love to learn to do better is search WBM. I use
         | WBM only a couple of times per month, but when I need it it's
         | the only tool that can do the job and is therefore very
         | valuable. Trouble is, I don't really know how to search it
         | unless I have a record of the exact URL I want, which isn't
         | always possible.
        
         | andrepd wrote:
         | Your comment just reminded me of how great the Internet Archive
         | is, so I guess it's time to donate another 100 bucks to IA
         | again.
        
           | thakoppno wrote:
           | it's within the realm of a few apple stores to serve the
           | internet archive?
        
           | WaffleIronMaker wrote:
           | Just as a quick link for others, it's possible to set up a
           | single or monthly donation to the internet archive at this
           | link: https://archive.org/donate/
        
       | tzfld wrote:
       | Wondering why we don't have web search engineers
        
         | cal85 wrote:
         | We do have 'researchers' though
        
         | ImaCake wrote:
         | Mostly people don't think it's possible to scam your way into a
         | job with that title.
         | 
         | But that is a _great_ reposte for  "prompt engineer"; my
         | coworker and I joke about being prompt engineers, but really we
         | are technical people putting new tools to use.
        
       | saltysalt wrote:
       | Great article, but can't help feel it just highlights how broken
       | Google search is from a UX perspective.
        
       | 38 wrote:
       | my problem is some operators just dont work any more, either on
       | purpose or because of crappy quality control. for example, you
       | used to be able to do:                   allintitle:Neil Diamond
       | If You Go Away
       | 
       | on YouTube, and get exactly what you would think, results with
       | all those words in the title. but now, you dont:
       | 
       | https://youtube.com/results?search_query=allintitle:Neil+Dia...
       | 
       | now, I get crap like this:                   Neil Diamond &
       | Shirley Bassey - Play Me - "high quality"         Barbra
       | Streisand - If You Go Away (Ne Me Quitte Pas)
       | 
       | how is that what I searched for? also, what is this:
       | 
       | > A search for [site:nytimes.com] will work, but
       | [site:nytimes.com] won't.
       | 
       | https://support.google.com/websearch/answer/2466433
       | 
       | did I just have a stroke? those two searches are exactly the
       | same. I try to be understanding, but I am constantly tripping
       | over big companies glaring software and/or documentation issues,
       | it gets old.
        
         | hddqsb wrote:
         | > https://support.google.com/websearch/answer/2466433
         | 
         | > those two searches are exactly the same
         | 
         | Yes, that appears to be a recently-introduced typo -- the
         | archived version from April does have a space:
         | https://web.archive.org/web/20230412181331/https://support.g...
         | 
         | I submitted a feedback comment, hopefully they'll fix the typo.
         | 
         | (For future reference, here is a snapshot of the current
         | version without a space: https://web.archive.org/web/2023072208
         | 5337/https://support.g...)
        
         | Daub wrote:
         | Agree. This even applies to simple Booleans, which are
         | sometimes completely ignored.
        
           | asteroidz wrote:
           | I sometimes wonder about PMs who sign off on decisions like
           | these, and the seeming lack of protest the developers put up.
        
         | bdn_ wrote:
         | Has anyone else experienced DuckDuckGo ignoring the exclusion
         | operator? For example, searching `kiwi -fruit`, with no space
         | between the hyphen and second word, used to bring up results
         | that did not include the word "fruit". This no longer seems to
         | be the case.
        
           | mdp2021 wrote:
           | > _duckduckgo ... exclusion operator_
           | 
           | Removed a few weeks ago.
           | 
           | Somebody posted in these pages the github diff showing the
           | removal of the options.
        
           | lloydatkinson wrote:
           | Someone needs to build a meta search engine...
        
             | ghaff wrote:
             | At least one existed in the dot com era. Forget the name.
        
               | fuzztester wrote:
               | metacrawler, IIRC.
        
         | pdanpdan wrote:
         | The one not working has a space after the colon. It's even in
         | the section taking about this :)
        
           | 38 wrote:
           | no, it doesn't.
        
         | chkal wrote:
         | > did I just have a stroke? those two searches are exactly the
         | same.
         | 
         | I just had the exact same though while reading the page.
        
           | nomilk wrote:
           | The fact Google's own documentation is this poorly kept is a
           | astonishing.
           | 
           | Makes me a firm believer companies should have their
           | documentation on GitHub (or similar) so anyone can make a PR
           | to tidy these things up.
        
       | self_awareness wrote:
       | Years ago the topic of Internet search was one of the core
       | philosophies of Search Lores:
       | 
       | https://web.archive.org/web/20191201105759/http://search.lor...
       | 
       | Gwern's site somehow reminds me of this Fravia's site.
        
       | Daub wrote:
       | If I had my way, I would introduce internet search techniques as
       | a core module in all university programs.
       | 
       | Too often I see in my students work evidence of lazy searching.
       | It is as if they expect Google to be able to read their minds or
       | even foresee their future intentions. Lack of variety of search
       | terms is a key shortcoming. Also, lack of exploration of terms
       | which are tangentially related to their search topic.
       | 
       | An internet search should be playful and exploratory. Above all,
       | it should be understood that the internet is beyond simple linear
       | indexing.
        
         | hattmall wrote:
         | So I actually did have this as a class, or at least part of
         | one.
         | 
         | I think there is / was an official standard or name, I can't
         | remember it though. It never worked with Google really, but it
         | worked on the search engines for our university and other
         | academic sites.
        
         | userbinator wrote:
         | _It is as if they expect Google to be able to read their minds
         | or even foresee their future intentions._
         | 
         | That's exactly what Google and other interests want ---
         | gullible, uncritical sheeple that can be exploited to extract
         | $$$ and worse.
        
         | aardshark wrote:
         | But search techniques change over time. You can see that in the
         | responses in this thread. What used to work doesn't anymore.
        
       | Thoeu388 wrote:
       | I found it useful to screenrecord all my computer activity. It is
       | OCR searchable. This way if I vaguely remember some bits, I can
       | reconstruct my sources. It also preserves original text, in case
       | sources get edited.
        
         | cole-k wrote:
         | I've thought about this or something similar (e.g. ArchiveBox).
         | 
         | It seems like 95-99% of the content you archive will be junk
         | that you never look at again, but that 1-4% really might be
         | worth it. Especially if it gets taken down.
         | 
         | What do you do for storage, though? I haven't invested in a NAS
         | or personal storage over 1TB.
         | 
         | I think I understand now why some of these archive utilities
         | offer you the option to upload to the internet archive in
         | addition to storing locally. You can build up a local cache and
         | then start removing the oldest page snapshots (but keep the
         | link trail) once you start to run out of space.
        
           | Thoeu388 wrote:
           | I buy 4TB external HDDs and use them as archival tapes.
        
             | cole-k wrote:
             | How often do you need to buy a new one?
        
         | orliesaurus wrote:
         | What tool do you use?
        
         | BMc2020 wrote:
         | I tried something similar, I had a text file called mylog.txt
         | and whenever I hit alt + m it would append whatever was in the
         | clipboard to the text file along with the date & time. But it
         | turned out I was terrible at predicting what I would want to
         | find again and it turned out to be not very useful to me.
        
       | dang wrote:
       | Related:
       | 
       |  _Internet Search Tips_ -
       | https://news.ycombinator.com/item?id=26847596 - April 2021 (77
       | comments)
       | 
       |  _Internet Search Tips_ -
       | https://news.ycombinator.com/item?id=18666574 - Dec 2018 (28
       | comments)
        
       | Thoeu388 wrote:
       | For complement I would recommend Russian yandex, and vk.com
       | (russian facebook). It is like snapshot of internet from 2018,
       | before everything was hit by massive censorship. It has some very
       | niche communities that are no longer on internet. For example
       | Ancient Egypt history, western net is filled with alien theories,
       | Russians are doing experimental archaeology. Also repair,
       | electronics...
       | 
       | Also Telegram is like second internet, nobody talks about!
        
         | vGPU wrote:
         | Yandex hits me with captchas every search and every 2-3 pages
         | on Firefox.
        
         | mxmlnkn wrote:
         | Telegram? How so? I'm still only using it for chatting. How
         | would you even find anything on Telegram? The simple global
         | search? I wouldn't like to join random shady groups found
         | through there.
        
           | mrweasel wrote:
           | Not sure either, but trying to follow the war in Ukraine has
           | taught me both Ukrainians and Russians get a surprising
           | amount of their news from various Telegrams channels.
        
         | moneywoes wrote:
         | How do you find google Telegrams? For me most seem bot spam
        
         | rwmj wrote:
         | Yandex image search is actually useful, especially after Google
         | destroyed their image search.
        
           | [deleted]
        
         | flyinghamster wrote:
         | The irony of someone recommending Russian platforms to evade
         | "censorship" is astounding.
        
         | jo_beef wrote:
         | > Also Telegram is like second internet, nobody talks about!
         | 
         | Do you have any groups/channels that you can recommend?
        
       | krmblg wrote:
       | There are a couple of comments on how search engine x dropped
       | feature y on date z (i time periods after competitor j did).
       | 
       | Can anyone recommend a site that tracks those changes?
       | 
       | I find myself getting annoyed by serps of a given vendor and I
       | might even be adopting to changes but ultimately jumping ship
       | towards the next best thing if I can't figure out a very low-
       | effort way of influencing the results in my favor using flags
       | that suddenly stop working - all I notice is a significant
       | degradation of result quality.
        
       | imchillyb wrote:
       | Soon, there will be an LLM that can utilize indexes and large
       | scale distributed databases, and that will replace every search
       | engine on earth.
       | 
       | Google should be quaking, and pissing themselves. It won't be
       | their bardum, and they won't have the financial power to purchase
       | said LLM.
       | 
       | Poof. Goodbye useless SEO spam, advertisement hell. We hated your
       | guts.
        
         | gnicholas wrote:
         | > _Goodbye useless SEO spam_
         | 
         | You think LLMs won't also be used to generate more SEO spam? I
         | think this will be an arms race, or a game of LLM-cat and LLM-
         | mouse.
        
         | computator wrote:
         | > _Goodbye useless SEO spam, advertisement hell. We hated your
         | guts._
         | 
         | That sounds great, but are you sure that there won't be legions
         | of LLMs generating SEO spam and advertisement hell in warfare
         | with the LLMs that do search?
         | 
         | Just as an aside: Sci-fi seems to have one superintelligence
         | (one Skynet) fighting humankind. If superintelligence emerges,
         | I think it's likely that we'll have thousands, millions, or
         | billions of superintelligences fighting each other, each with
         | its own agenda.
        
           | andirk wrote:
           | The searching of data doesn't need to be controlled by
           | insanely powerful insanely rich entities. Ex: Wikipedia,
           | Craigslist, polio vaccine. If this LLM arms race leads to an
           | organic solution that anyone can grow their own, it might not
           | have as many "This 1 Trick Your Grocer Doesn't Want You To
           | Know" ads in it.
           | 
           | And for superintelligence, let's clarify that this
           | "Artificial Intelligence" moniker has almost nothing to do
           | with actual intelligence, so we're probably a few years off
           | from that.
        
         | Yemeth wrote:
         | I'm not sure how LLMs are going to be immune to SEO spam and
         | advertising. As if human nature would magically transform and
         | people would stop buying stupid stuff.
         | 
         | LLMs are already being used to make the web even less useful,
         | by shitting out vast ammounts of meaningless and even outright
         | wrong text for SEO purposes. And don't forget the systems are
         | being trained on the web in the first place... using a LLM that
         | is able to utilize a web search already thinks listicles are
         | useful information and not just a way to place affiliate links.
         | 
         | When the hype is over and the VC money dried out, companies
         | will find ways to make the LLM interfaces and outputs an
         | 'advertising friendly' affair.
        
           | winwang wrote:
           | I think the idea is that, if _we_ (some of us) can figure out
           | when something is SEO spam or rather, generally low quality,
           | an LLM should be able to but faster and more quantitatively
           | why.
        
             | tempodox wrote:
             | And who, pray tell, would have an interest in giving us
             | that? The internet overlords are advertising companies.
             | Would you pay for such an LLM? Those companies have gotten
             | so powerful because everyone wants stuff for "free".
        
       | 1vuio0pswjnm7 wrote:
       | Who remembers +fravia's searchlores.org
       | 
       | https://web.archive.org/web/20000620171446if_/http://www.sea...
        
         | marttt wrote:
         | a mirror without archive.org:
         | http://biostatisticien.eu/www.searchlores.org/indexo.htm
         | 
         | Thanks very much for this; I'm very much interested in
         | searching techniques, but wasn't aware of this site.
        
       | icinganalysis wrote:
       | Learned a lot of new tips today.
       | 
       | I have found a few obscure publications at the [Defense Technical
       | Information Center](https://discover.dtic.mil).
        
       | 8bitsrule wrote:
       | Many good scarce tips in there (once past Google Scholar, unless
       | you're monopoly-oriented). 'Dealing with paywalls' for example.
       | 
       | He did mention IA searches, but I didn't spot a mention of their
       | https://scholar.archive.org/ with "over 25 million research
       | articles and other scholarly documents preserved in the Internet
       | Archive."
       | 
       | For book metadata, I find https://openlibrary.org/ search has a
       | lot of 'MARC' type data. Esp useful for books with many editions.
       | 
       | Some services are getting more restrictive. e.g. WorldCat
       | recently got harder to use, rejecting many searches. But if you
       | can find a book's OCLC ###, then
       | https://www.worldcat.org/oclc/### works every time. Useful for
       | finding local hardcopy if you've got your location turned on.
       | (With an ISBN , WPedia will also do this for you at:
       | https://en.wikipedia.org/wiki/Special:BookSources/ )
        
         | jcmeyrignac wrote:
         | I also recommend https://www.semanticscholar.org/ because you
         | can easily retrieve some of the PDFs.
        
       | costco wrote:
       | Really good list. Here are some others I've discovered:
       | 
       | Some libgen clone sites like z-lib have fulltext search on books
       | with support for exact matches: https://zlibrary-
       | asia.se/fulltext/?q=%22frank+sinatra%22&typ...
       | 
       | Even if you are going to purchase book on a subject, this finds
       | so much stuff that is not in Google because of copyright
       | delisting and is sometimes useful in knowing which books to
       | consider.
       | 
       | Yandex results especially the non English ones can be good if you
       | are willing to use a translator.
        
       | orphea wrote:
       | I feel like Google became absolutely unusable as a search engine.
       | You search for "keyword1 keyword2", and at least half of the
       | results are either "Missing: keyword1" or "Missing: keyword2".
        
         | cma wrote:
         | They moved over to really dumb word vector stuff, they mention
         | it in the TPUv4 paper and it is pretty surprising but probably
         | monetized better somehow.
        
         | ImaCake wrote:
         | Until a year ago using DuckDuckGo instead of google felt like
         | only an equal or inferior option. But at some recent point I
         | have found that DDG has slightly improved and Google has gotten
         | much, much worse. With DDG and Bing Chat, google's future is
         | looking very Internet Explorer 6.
        
           | flyinghamster wrote:
           | DDG is following Google's path, with removal of the exclusion
           | operator. Nothing like searching for "foo -quux" and finding
           | "quux" in ALL of my search results, on DDG and Google alike.
        
         | discobean wrote:
         | Make your tool so idiots can use it and only idiots will use it
        
         | mrweasel wrote:
         | Google is especially bad that this. Often what will be remove
         | is the most important keyword, I assume because that yields
         | more result.
        
       | coldblues wrote:
       | Gwern is the person I'd become if I could properly monetize my
       | ADHD and rabbit hole seeking behavior. I absolutely love his
       | website and everything he publishes, and I am always delighted to
       | look over his new stuff. Wish there are more people like him to
       | read from.
        
         | jterrys wrote:
         | Are you perchance able to provide some kind of RSS feed to his
         | website? I'm having a hard time finding his newest stuff. You
         | can subscribe to his newsletters but he stopped doing those two
         | years ago
        
           | Tenoke wrote:
           | There is the firehose from his patreon though that might not
           | be exactly what you are looking for. Personally, I just check
           | 'newest' on his frontpage every now and then.
           | 
           | https://www.patreon.com/gwern
        
         | RGBCube wrote:
         | His website is the best designed website I have genuinely ever
         | visited, so fast, easy to navigate and very little blank space
         | which makes it information dense. I also love the inline link
         | opening, I definitely will be implementing that when I make my
         | own website in the future.
        
           | wintermutestwin wrote:
           | >His website is the best designed website
           | 
           | The forced R and L margins suck for people with oldster eyes
           | who have to increase the text size.
           | 
           | Kids these days!
        
       ___________________________________________________________________
       (page generated 2023-07-22 23:02 UTC)