[HN Gopher] Show HN: Ichido, search engine that tags sites using...
___________________________________________________________________
Show HN: Ichido, search engine that tags sites using Google and
Cloudflare
Hello HN, In my spare time I work on an experimental search engine
named Ichido. Search is fascinating, there are so many features you
can add to a search engine, but I find that the existing search
engines are a bit limited in the features they have to offer. So I
decided to work on my own search engine to test out different
features, searching algorithms, and front ends in order to improve
my (and hopefully others) searching experience. Ichido includes a
tagging system that provides more info on search results. For
example, if a site links to Google services or uses Cloudflare, a
tag is shown with the search result that let's the user know about
that site's use of those services. Ichido also includes links to
RSS feeds in search results, making it much easier to find RSS
feeds. This search engine is free to use, but if you like the
service and want to support continued development please consider
making a donation (Ichido currently supports donations through
Libera Pay).
Author : anthonyhn
Score : 88 points
Date : 2023-02-26 15:12 UTC (7 hours ago)
(HTM) web link (ichi.do)
(TXT) w3m dump (ichi.do)
| berry_sortoro wrote:
| [dead]
| superasn wrote:
| I think the tags can be grouped like Extereme trackers, Moderate
| trackers, etc and clicking on them expands the full list.
|
| Also one really useful tag would be "Affiliate links" if there is
| a way to identify a page contains affiliate links like amazon
| affiliate, etc. Those pages are always almost crap.
|
| Also a tag for "Modal popups", those are too often just marketing
| related websites and definitely want to skip it if I know prior
| to visiting.
| simultsop wrote:
| What would be the issue of being hosted on CF? I believe it is a
| better option than the rest of the shared hosting industry.. If
| nothing critical whats the intention of tagging?
| danuker wrote:
| http://crimeflare.eu.org
|
| CloudFlare is a MitMaaS. Traffic is seen by them because they
| are in control of the HTTPS certificates, and you have to take
| them at their word that they do not log content (and even if
| they're not lying/under a gag order, just metadata is enough
| for a lot of evil things).
| Tijdreiziger wrote:
| So are AWS and Azure also MitMaaS?
|
| If yes, what's the endgame? Everyone goes back to managing
| their own servers?
|
| If no, why is Cloudflare the only hosting provider that gets
| singled out?
| danuker wrote:
| Good point. Yes, they are, as are all cloud services. But
| with AWS and Azure it's more explicit, in that the server
| is running on their machines.
|
| With Cloudflare, you may not realize this if you don't know
| how HTTPS works.
|
| The end game is not necessarily to avoid all use of cloud
| services, but to be aware of their functioning and their
| trade-offs, and avoid them where third party spying is
| crucial to prevent.
| simultsop wrote:
| Again better option than the rest of shared hosting
| industry where a company of 50 or 500 can lurk into user
| space and for courtesy take a look what is going on. I
| see CF promoting ZeroTrust services and I believe they
| use them for themselves first hand. Any digital
| information is always prone to compromise, whether it
| remains in my pocket or in the bunker of Area51...
| binarymax wrote:
| CF will force you to recaptcha if you try to remain anonymous
| callahad wrote:
| Is that universally true, or just when domains explicitly opt
| into specific traffic screening measures? Asking as I'm
| thinking of moving some stuff to Cloudflare Pages.
| flas9sd wrote:
| I see you offer an opensearch.xml already - if you embed it as
| link node with the appropriate type it will be straightforward to
| add it to the browser as (default) search engine:
| https://developer.mozilla.org/en-US/docs/Web/OpenSearch#auto...
|
| also: happy to give this a try, more knobs for power users
| daoudc wrote:
| This is really cool! Please consider joining forces with us at
| mwmbl.org, would love to incorporate some of these ideas.
| 1vuio0pswjnm7 wrote:
| The pagination keep increasing past the point where Bing will
| provide no more results. Testing a popular search term, for which
| there are no doubt millions of results, it was only possible to
| get new results up to page 45. Yet the website will keep
| incrementing the page number and result numbers as if new results
| are being returned.
|
| Then tried same search with popularity set to 500000 and could
| not even get a single full page of 10 results. It's laughable to
| assume from this "search" that only, say, 500004 out of the
| millions of websites in existence include this term. Not that I
| want to browse a full list, but at least I want to know how many
| hits I got. Then I can add more terms and try to reduce that
| number.
| danuker wrote:
| Thank you! I think any competition is welcome for search engines,
| with Google going down the monetization path.
|
| A piece of feedback: When I select "Remove top ...." and click
| Submit, then click Next, the popularity filter is gone.
|
| Edit: looks like the file type filter is dropped as well. Do add
| the arguments to the pagination links.
| anthonyhn wrote:
| >Edit: looks like the file type filter is dropped as well. Do
| add the arguments to the pagination links.
|
| Thank you, great feedback! You're right, I forgot to include
| some of the params in the pagination, will have to include
| those in the next update.
| jesprenj wrote:
| An interesting search proxy is also SearX. Written in Python, it
| supports many backend engines and can be self hosted.
|
| And here's a lightweight frontend/proxy I wrote in C for using
| Google search on low-end phones that can't render bloated HTML
| (SearX was too complicated to install):
|
| http://searc.4a.si:7327/search?q=news
|
| It's also nice that the structured never constantly changing HTML
| it produces makes it ideal to programatically query Google.
| Although you still run into captchas which it cannot solve if
| queries get too suspicious.
| b1ue64 wrote:
| Also see SearXNG https://github.com/searxng/searxng/
| mg wrote:
| I run this search engine comparison tool:
|
| https://www.gnod.com/search/
|
| Just added Ichido.
|
| Click on "more engines" to activate it.
| daoudc wrote:
| Nice! Please consider adding https://mwmbl.org
|
| Thanks!
| marban wrote:
| Can you add https://biztoc.com/search ? (Real-time
| business/finance News) Zero Tracking/Cookies.
| FireInsight wrote:
| You could add other AI search such as perplexity.ai and
| phind.com
| anthonyhn wrote:
| >Just added Ichido.
|
| Thanks, much appreciated
| culi wrote:
| I made a post with a bunch of suggestions from my list and then
| my browser extension that limits my time on HN lost my whole
| comment including all the little explanations I had for each
| one. So here's my raw list instead haha meta
| https://www.gnod.com/search/
| https://github.com/searx/searx categories
| independent https://www.crawlson.com/
| https://search.marginalia.nu/ https://wiby.me/
| https://searchmysite.net/ international
| https://bonzamate.com.au/ australia
| https://www.baidu.com/ china https://yandex.com/
| russia code https://searchcode.com/
| https://codesearch.ai/ http://symbolhound.com/
| https://publicwww.com/ https://search.feep.dev/
| http://codesearch.debian.net/
| https://codesearch.isocpp.org/
| https://www.programcreek.com/python/
| https://livegrep.com/search/linux https://grep.app/
| ai https://consensus.app/ scientific consensus
| https://github.com/jokenox/Goopt procedurally generated
| https://same.energy/ image similarity products
| https://www.looria.com/ https://knifist.com/ knives
| https://attic.city/ home and fashion from indie stores
| topical https://biztoc.com/search business news
| premium https://kagi.com/ other
| https://metager.org/ privacy centric engine that combines
| results of several engines https://thangs.com/ 3d
| models https://filmot.com/ youtube subtitles
| lists https://seirdy.one/posts/2021/03/10/search-
| engines-with-own-indexes/ https://web.archive.org/web/2
| 0200710091019/http://www.jaruzel.com/textfiles/Old%20Web%20Info
| /Internet%20Search%20Engines%20v2.61.txt
|
| Hope its useful still
| bastawhiz wrote:
| What's the use case for this? If I don't want Google scripts, I
| block them. I'll use a user agent that doesn't download or run
| them. If I don't want cookies, I'll instruct my browser not to
| save cookies. What situation would I be in where knowing whether
| a site uses these things is a search result I want to visit?
| brucethemoose2 wrote:
| I find the extra information useful, as I dont have to visit
| the site to find out.
| KomoD wrote:
| Too many tags, and if a site has something, like scripts, why do
| you say "may"?
|
| If a site has scripts then it's not "This site may be using
| Javascript", it's for sure that the site uses it...?
|
| And popularity filter doesn't work, the results are empty and if
| you try going to any of the other pages it removes the filter
| partyguy wrote:
| Nice project! However, when trying to search for my site
| (https://spacehey.com), it shows multiple tags, with most of them
| being false (Cloudflare, UTM Tracking, WEBP Images). I used
| Cloudflare at one point in the past, but don't anymore.
| Additionally, there has never been UTM tracking or anything like
| that nor WEBP images... Where do you get such data from?
|
| Apart from that, awesome project!
| anthonyhn wrote:
| Since spacehey includes user-submitted content, it's possible
| that:
|
| * Someone uploaded a WEBP image to the site.
|
| * Someone pasted a link with a utm_* param.
|
| * The page was crawled when cloudflare was used.
|
| Will look into it and see if I can find the pages that
| generated the tags. Search results are generally tagged by
| domain name (necessary since not all pages can be crawled, and
| even if the page the user connects to doesn't have, for example
| google trackers, a user would likely want to know if the site
| is using trackers elsewhere).
|
| Also love the spacehey project, really captures the feel of
| Myspace!
| anthonyhn wrote:
| EDIT: I found some of the pages with links that include UTM
| tracking params. Let me know if you want me to send you the
| pages with those links, can send them through email (my email
| is on the contact page of the site).
| return_to_monke wrote:
| what's wrong with webp?
| anthonyhn wrote:
| > what's wrong with webp?
|
| Nothing wrong with the format in particular. However some
| may prefer formats such as PNG and JPEG since:
|
| * A lot more software supports PNG and JPEG (backwards
| compatibility, better integration with one's existing
| system and tools).
|
| * You can often get the same file size, visual quality, and
| performance with PNG and JPEG as you can with WEBP with
| optimization.
| partyguy wrote:
| Oh, I see - that makes perfect sense! Thank you for the
| clarification!
|
| Glad you like SpaceHey :)
|
| Keep up the great work!
| ocdtrekkie wrote:
| This looks great, I am really glad to see things making it more
| obvious how pervasive malicious Google scripts are.
|
| I find the webp flag interesting, as I don't think webp itself is
| inherently harmful, except for being an image spec that solely
| exists because Google NIHs everything and wants to write their
| own everything. (Long live JPEG-XL!)
|
| I'm curious why you chose to tag it explicitly though.
| brucethemoose2 wrote:
| I love that tag, as (to me) it indicates a site is trying to be
| bandwidth efficient instead of just defaulting to JPEG.
|
| JXL is pretty much dead thanks to Google... and avif is still
| mostly suited to thumbnails.
| TekMol wrote:
| In your about page, I see you are using Bing's API. I didn't even
| know Bing has a search API that everyone can use!
|
| How much do you have to pay them for this?
| anthonyhn wrote:
| It's $4/1000 queries, but the rate is increasing in May to
| $18/1000 queries. The Bing API is available through Azure.
| danuker wrote:
| I hope the author knows this, and won't be surprised by a
| bill more than 4x larger.
| riidom wrote:
| That comment was written by the author.
| [deleted]
| danuker wrote:
| Oh! I am scatterbrained.
| gdcbe wrote:
| Pretty sure that only Google and Microsoft have the money and
| resources to crawl the entire internet. Or perhaps the only
| that can AND are willing to.
|
| Correct me if I'm wrong though, but I'm pretty certain that all
| other search engines in the same category use one of these as
| their backend. Eg I'm pretty certain that counts for duckduckgo
| as well.
| Santosh83 wrote:
| I think Yandex have their own index and some others too like
| Marginalia, but the latter couldn't be called "in the same
| category" as the other three.
| kkielhofner wrote:
| Yandex also suffered a security breach and their source
| code is available[0] although utilizing it in any way is
| ethically and legally dubious (at best).
|
| [0] - https://arstechnica.com/information-
| technology/2023/01/massi...
| quectophoton wrote:
| > Pretty sure that only Google and Microsoft have the money
| and resources to crawl the entire internet. Or perhaps the
| only that can AND are willing to.
|
| Money and resources _and_ a dominant-enough position so that
| your crawlers are not blocked by websites.
|
| Unfortunately.
| rezonant wrote:
| There's definitely gatekeeping on websites but having done
| a bunch of crawler work I can say you'd be surprised how
| rarely a site will outright block you if you just do the
| right things: have an identifiable user agent with a
| working URL that explains what your crawler does, respect
| robots.txt, implement polite crawling. As for actually
| being able to crawl the whole thing though, yeah it's
| stupidly expensive :-/
| antonok wrote:
| Brave Search has its own independent index too -
| https://brave.com/brave-search-beta/
| daoudc wrote:
| Mwmbl has its own index but it's orders of magnitude smaller
| than commercial search engines.
| jacooper wrote:
| Brave goggles also do something similar, allowing to filter
| search the way to you want.
| coolspot wrote:
| I would prefer more logical tags like "top 1k", "aggregator",
| "user-generated content" than technical like "utm" and
| "obfuscated scripts". Also, I would prefer tags grouped together
| into expandable lists and not shown all by default. Every site
| uses javascript, I don't want to see it over and over again
| unless specifically queried for that.
___________________________________________________________________
(page generated 2023-02-26 23:01 UTC)