[HN Gopher] Internet search tips
___________________________________________________________________
Internet search tips
Author : herbertl
Score : 199 points
Date : 2023-07-22 03:22 UTC (19 hours ago)
(HTM) web link (gwern.net)
(TXT) w3m dump (gwern.net)
| skimdesk wrote:
| Another useful Google search trick not mentioned in the article
| is numeric range queries.
|
| You can use two numbers separated by two dots to represent all
| numbers in the range.
|
| For example, a search for taki 100..200
|
| gives me results for taki 183.
|
| This is useful when you can't remember an exact year or number.
| Jakob wrote:
| +1 for his tip of "Search The Internet Archive".
|
| The other day I searched for the Tom & Jerry full episodes on the
| web to no avail (streaming platforms and video platforms like
| youtube only have rubbish cuts).
|
| The internet archive has every episode starting from the first
| one in 1940, in an easily accessible player without any ads or
| recommendations: https://archive.org/details/tom-and-jerry-
| all-114-episodes
| nomilk wrote:
| TIL: IA is more than Wayback Machine.
|
| Something I'd love to learn to do better is search WBM. I use
| WBM only a couple of times per month, but when I need it it's
| the only tool that can do the job and is therefore very
| valuable. Trouble is, I don't really know how to search it
| unless I have a record of the exact URL I want, which isn't
| always possible.
| andrepd wrote:
| Your comment just reminded me of how great the Internet Archive
| is, so I guess it's time to donate another 100 bucks to IA
| again.
| thakoppno wrote:
| it's within the realm of a few apple stores to serve the
| internet archive?
| WaffleIronMaker wrote:
| Just as a quick link for others, it's possible to set up a
| single or monthly donation to the internet archive at this
| link: https://archive.org/donate/
| tzfld wrote:
| Wondering why we don't have web search engineers
| cal85 wrote:
| We do have 'researchers' though
| ImaCake wrote:
| Mostly people don't think it's possible to scam your way into a
| job with that title.
|
| But that is a _great_ reposte for "prompt engineer"; my
| coworker and I joke about being prompt engineers, but really we
| are technical people putting new tools to use.
| saltysalt wrote:
| Great article, but can't help feel it just highlights how broken
| Google search is from a UX perspective.
| 38 wrote:
| my problem is some operators just dont work any more, either on
| purpose or because of crappy quality control. for example, you
| used to be able to do: allintitle:Neil Diamond
| If You Go Away
|
| on YouTube, and get exactly what you would think, results with
| all those words in the title. but now, you dont:
|
| https://youtube.com/results?search_query=allintitle:Neil+Dia...
|
| now, I get crap like this: Neil Diamond &
| Shirley Bassey - Play Me - "high quality" Barbra
| Streisand - If You Go Away (Ne Me Quitte Pas)
|
| how is that what I searched for? also, what is this:
|
| > A search for [site:nytimes.com] will work, but
| [site:nytimes.com] won't.
|
| https://support.google.com/websearch/answer/2466433
|
| did I just have a stroke? those two searches are exactly the
| same. I try to be understanding, but I am constantly tripping
| over big companies glaring software and/or documentation issues,
| it gets old.
| hddqsb wrote:
| > https://support.google.com/websearch/answer/2466433
|
| > those two searches are exactly the same
|
| Yes, that appears to be a recently-introduced typo -- the
| archived version from April does have a space:
| https://web.archive.org/web/20230412181331/https://support.g...
|
| I submitted a feedback comment, hopefully they'll fix the typo.
|
| (For future reference, here is a snapshot of the current
| version without a space: https://web.archive.org/web/2023072208
| 5337/https://support.g...)
| Daub wrote:
| Agree. This even applies to simple Booleans, which are
| sometimes completely ignored.
| asteroidz wrote:
| I sometimes wonder about PMs who sign off on decisions like
| these, and the seeming lack of protest the developers put up.
| bdn_ wrote:
| Has anyone else experienced DuckDuckGo ignoring the exclusion
| operator? For example, searching `kiwi -fruit`, with no space
| between the hyphen and second word, used to bring up results
| that did not include the word "fruit". This no longer seems to
| be the case.
| mdp2021 wrote:
| > _duckduckgo ... exclusion operator_
|
| Removed a few weeks ago.
|
| Somebody posted in these pages the github diff showing the
| removal of the options.
| lloydatkinson wrote:
| Someone needs to build a meta search engine...
| ghaff wrote:
| At least one existed in the dot com era. Forget the name.
| fuzztester wrote:
| metacrawler, IIRC.
| pdanpdan wrote:
| The one not working has a space after the colon. It's even in
| the section taking about this :)
| 38 wrote:
| no, it doesn't.
| chkal wrote:
| > did I just have a stroke? those two searches are exactly the
| same.
|
| I just had the exact same though while reading the page.
| nomilk wrote:
| The fact Google's own documentation is this poorly kept is a
| astonishing.
|
| Makes me a firm believer companies should have their
| documentation on GitHub (or similar) so anyone can make a PR
| to tidy these things up.
| self_awareness wrote:
| Years ago the topic of Internet search was one of the core
| philosophies of Search Lores:
|
| https://web.archive.org/web/20191201105759/http://search.lor...
|
| Gwern's site somehow reminds me of this Fravia's site.
| Daub wrote:
| If I had my way, I would introduce internet search techniques as
| a core module in all university programs.
|
| Too often I see in my students work evidence of lazy searching.
| It is as if they expect Google to be able to read their minds or
| even foresee their future intentions. Lack of variety of search
| terms is a key shortcoming. Also, lack of exploration of terms
| which are tangentially related to their search topic.
|
| An internet search should be playful and exploratory. Above all,
| it should be understood that the internet is beyond simple linear
| indexing.
| hattmall wrote:
| So I actually did have this as a class, or at least part of
| one.
|
| I think there is / was an official standard or name, I can't
| remember it though. It never worked with Google really, but it
| worked on the search engines for our university and other
| academic sites.
| userbinator wrote:
| _It is as if they expect Google to be able to read their minds
| or even foresee their future intentions._
|
| That's exactly what Google and other interests want ---
| gullible, uncritical sheeple that can be exploited to extract
| $$$ and worse.
| aardshark wrote:
| But search techniques change over time. You can see that in the
| responses in this thread. What used to work doesn't anymore.
| Thoeu388 wrote:
| I found it useful to screenrecord all my computer activity. It is
| OCR searchable. This way if I vaguely remember some bits, I can
| reconstruct my sources. It also preserves original text, in case
| sources get edited.
| cole-k wrote:
| I've thought about this or something similar (e.g. ArchiveBox).
|
| It seems like 95-99% of the content you archive will be junk
| that you never look at again, but that 1-4% really might be
| worth it. Especially if it gets taken down.
|
| What do you do for storage, though? I haven't invested in a NAS
| or personal storage over 1TB.
|
| I think I understand now why some of these archive utilities
| offer you the option to upload to the internet archive in
| addition to storing locally. You can build up a local cache and
| then start removing the oldest page snapshots (but keep the
| link trail) once you start to run out of space.
| Thoeu388 wrote:
| I buy 4TB external HDDs and use them as archival tapes.
| cole-k wrote:
| How often do you need to buy a new one?
| orliesaurus wrote:
| What tool do you use?
| BMc2020 wrote:
| I tried something similar, I had a text file called mylog.txt
| and whenever I hit alt + m it would append whatever was in the
| clipboard to the text file along with the date & time. But it
| turned out I was terrible at predicting what I would want to
| find again and it turned out to be not very useful to me.
| dang wrote:
| Related:
|
| _Internet Search Tips_ -
| https://news.ycombinator.com/item?id=26847596 - April 2021 (77
| comments)
|
| _Internet Search Tips_ -
| https://news.ycombinator.com/item?id=18666574 - Dec 2018 (28
| comments)
| Thoeu388 wrote:
| For complement I would recommend Russian yandex, and vk.com
| (russian facebook). It is like snapshot of internet from 2018,
| before everything was hit by massive censorship. It has some very
| niche communities that are no longer on internet. For example
| Ancient Egypt history, western net is filled with alien theories,
| Russians are doing experimental archaeology. Also repair,
| electronics...
|
| Also Telegram is like second internet, nobody talks about!
| vGPU wrote:
| Yandex hits me with captchas every search and every 2-3 pages
| on Firefox.
| mxmlnkn wrote:
| Telegram? How so? I'm still only using it for chatting. How
| would you even find anything on Telegram? The simple global
| search? I wouldn't like to join random shady groups found
| through there.
| mrweasel wrote:
| Not sure either, but trying to follow the war in Ukraine has
| taught me both Ukrainians and Russians get a surprising
| amount of their news from various Telegrams channels.
| moneywoes wrote:
| How do you find google Telegrams? For me most seem bot spam
| rwmj wrote:
| Yandex image search is actually useful, especially after Google
| destroyed their image search.
| [deleted]
| flyinghamster wrote:
| The irony of someone recommending Russian platforms to evade
| "censorship" is astounding.
| jo_beef wrote:
| > Also Telegram is like second internet, nobody talks about!
|
| Do you have any groups/channels that you can recommend?
| krmblg wrote:
| There are a couple of comments on how search engine x dropped
| feature y on date z (i time periods after competitor j did).
|
| Can anyone recommend a site that tracks those changes?
|
| I find myself getting annoyed by serps of a given vendor and I
| might even be adopting to changes but ultimately jumping ship
| towards the next best thing if I can't figure out a very low-
| effort way of influencing the results in my favor using flags
| that suddenly stop working - all I notice is a significant
| degradation of result quality.
| imchillyb wrote:
| Soon, there will be an LLM that can utilize indexes and large
| scale distributed databases, and that will replace every search
| engine on earth.
|
| Google should be quaking, and pissing themselves. It won't be
| their bardum, and they won't have the financial power to purchase
| said LLM.
|
| Poof. Goodbye useless SEO spam, advertisement hell. We hated your
| guts.
| gnicholas wrote:
| > _Goodbye useless SEO spam_
|
| You think LLMs won't also be used to generate more SEO spam? I
| think this will be an arms race, or a game of LLM-cat and LLM-
| mouse.
| computator wrote:
| > _Goodbye useless SEO spam, advertisement hell. We hated your
| guts._
|
| That sounds great, but are you sure that there won't be legions
| of LLMs generating SEO spam and advertisement hell in warfare
| with the LLMs that do search?
|
| Just as an aside: Sci-fi seems to have one superintelligence
| (one Skynet) fighting humankind. If superintelligence emerges,
| I think it's likely that we'll have thousands, millions, or
| billions of superintelligences fighting each other, each with
| its own agenda.
| andirk wrote:
| The searching of data doesn't need to be controlled by
| insanely powerful insanely rich entities. Ex: Wikipedia,
| Craigslist, polio vaccine. If this LLM arms race leads to an
| organic solution that anyone can grow their own, it might not
| have as many "This 1 Trick Your Grocer Doesn't Want You To
| Know" ads in it.
|
| And for superintelligence, let's clarify that this
| "Artificial Intelligence" moniker has almost nothing to do
| with actual intelligence, so we're probably a few years off
| from that.
| Yemeth wrote:
| I'm not sure how LLMs are going to be immune to SEO spam and
| advertising. As if human nature would magically transform and
| people would stop buying stupid stuff.
|
| LLMs are already being used to make the web even less useful,
| by shitting out vast ammounts of meaningless and even outright
| wrong text for SEO purposes. And don't forget the systems are
| being trained on the web in the first place... using a LLM that
| is able to utilize a web search already thinks listicles are
| useful information and not just a way to place affiliate links.
|
| When the hype is over and the VC money dried out, companies
| will find ways to make the LLM interfaces and outputs an
| 'advertising friendly' affair.
| winwang wrote:
| I think the idea is that, if _we_ (some of us) can figure out
| when something is SEO spam or rather, generally low quality,
| an LLM should be able to but faster and more quantitatively
| why.
| tempodox wrote:
| And who, pray tell, would have an interest in giving us
| that? The internet overlords are advertising companies.
| Would you pay for such an LLM? Those companies have gotten
| so powerful because everyone wants stuff for "free".
| 1vuio0pswjnm7 wrote:
| Who remembers +fravia's searchlores.org
|
| https://web.archive.org/web/20000620171446if_/http://www.sea...
| marttt wrote:
| a mirror without archive.org:
| http://biostatisticien.eu/www.searchlores.org/indexo.htm
|
| Thanks very much for this; I'm very much interested in
| searching techniques, but wasn't aware of this site.
| icinganalysis wrote:
| Learned a lot of new tips today.
|
| I have found a few obscure publications at the [Defense Technical
| Information Center](https://discover.dtic.mil).
| 8bitsrule wrote:
| Many good scarce tips in there (once past Google Scholar, unless
| you're monopoly-oriented). 'Dealing with paywalls' for example.
|
| He did mention IA searches, but I didn't spot a mention of their
| https://scholar.archive.org/ with "over 25 million research
| articles and other scholarly documents preserved in the Internet
| Archive."
|
| For book metadata, I find https://openlibrary.org/ search has a
| lot of 'MARC' type data. Esp useful for books with many editions.
|
| Some services are getting more restrictive. e.g. WorldCat
| recently got harder to use, rejecting many searches. But if you
| can find a book's OCLC ###, then
| https://www.worldcat.org/oclc/### works every time. Useful for
| finding local hardcopy if you've got your location turned on.
| (With an ISBN , WPedia will also do this for you at:
| https://en.wikipedia.org/wiki/Special:BookSources/ )
| jcmeyrignac wrote:
| I also recommend https://www.semanticscholar.org/ because you
| can easily retrieve some of the PDFs.
| costco wrote:
| Really good list. Here are some others I've discovered:
|
| Some libgen clone sites like z-lib have fulltext search on books
| with support for exact matches: https://zlibrary-
| asia.se/fulltext/?q=%22frank+sinatra%22&typ...
|
| Even if you are going to purchase book on a subject, this finds
| so much stuff that is not in Google because of copyright
| delisting and is sometimes useful in knowing which books to
| consider.
|
| Yandex results especially the non English ones can be good if you
| are willing to use a translator.
| orphea wrote:
| I feel like Google became absolutely unusable as a search engine.
| You search for "keyword1 keyword2", and at least half of the
| results are either "Missing: keyword1" or "Missing: keyword2".
| cma wrote:
| They moved over to really dumb word vector stuff, they mention
| it in the TPUv4 paper and it is pretty surprising but probably
| monetized better somehow.
| ImaCake wrote:
| Until a year ago using DuckDuckGo instead of google felt like
| only an equal or inferior option. But at some recent point I
| have found that DDG has slightly improved and Google has gotten
| much, much worse. With DDG and Bing Chat, google's future is
| looking very Internet Explorer 6.
| flyinghamster wrote:
| DDG is following Google's path, with removal of the exclusion
| operator. Nothing like searching for "foo -quux" and finding
| "quux" in ALL of my search results, on DDG and Google alike.
| discobean wrote:
| Make your tool so idiots can use it and only idiots will use it
| mrweasel wrote:
| Google is especially bad that this. Often what will be remove
| is the most important keyword, I assume because that yields
| more result.
| coldblues wrote:
| Gwern is the person I'd become if I could properly monetize my
| ADHD and rabbit hole seeking behavior. I absolutely love his
| website and everything he publishes, and I am always delighted to
| look over his new stuff. Wish there are more people like him to
| read from.
| jterrys wrote:
| Are you perchance able to provide some kind of RSS feed to his
| website? I'm having a hard time finding his newest stuff. You
| can subscribe to his newsletters but he stopped doing those two
| years ago
| Tenoke wrote:
| There is the firehose from his patreon though that might not
| be exactly what you are looking for. Personally, I just check
| 'newest' on his frontpage every now and then.
|
| https://www.patreon.com/gwern
| RGBCube wrote:
| His website is the best designed website I have genuinely ever
| visited, so fast, easy to navigate and very little blank space
| which makes it information dense. I also love the inline link
| opening, I definitely will be implementing that when I make my
| own website in the future.
| wintermutestwin wrote:
| >His website is the best designed website
|
| The forced R and L margins suck for people with oldster eyes
| who have to increase the text size.
|
| Kids these days!
___________________________________________________________________
(page generated 2023-07-22 23:02 UTC)