[HN Gopher] Updates to our web search products and Programmable...
___________________________________________________________________
Updates to our web search products and Programmable Search Engine
capabilities
Author : 01jonny01
Score : 204 points
Date : 2026-01-23 09:38 UTC (13 hours ago)
(HTM) web link (programmablesearchengine.googleblog.com)
(TXT) w3m dump (programmablesearchengine.googleblog.com)
| 01jonny01 wrote:
| Google quietly announced that Programmable Search (ex-Custom
| Search) won't allow new engines to "search the entire web"
| anymore. New engines are capped at searching up to 50 domains,
| and existing full-web engines have until Jan 1, 2027 to
| transition.
|
| If you actually need whole-web search, Google now points you to
| an "interest form" for enterprise solutions (Vertex AI Search
| etc.), with no public pricing and no guarantee they'll even
| reply.
|
| This seems like it effectively ends the era of indie / niche
| search engines being able to build on Google's index. Anything
| that looks like general web search is getting pushed behind
| enterprise gates.
|
| I haven't seen much discussion about this yet, but for anyone who
| built a small search product on Programmable Search, this feels
| like a pretty big shift.
|
| Curious if others here are affected or already planning
| alternatives.
|
| UPDATE: I logged into Programmable Search and the message is even
| more explicit: Full web search via the "Search the entire web"
| feature will be discontinued within the next year. Please update
| your search engine to specify specific sites to search. With this
| link: https://support.google.com/programmable-
| search/answer/123971...
| throwaway_20357 wrote:
| What are some of the niche search engines build on Google's
| index affected by this?
| doublerabbit wrote:
| Kagi
| marginalia_nu wrote:
| They published this the other day:
|
| https://blog.kagi.com/waiting-dawn-search
|
| Which saw some discussion on HN.
| embedding-shape wrote:
| > some discussion
|
| ~450 score, ~247 comments and still on /best ("Most-
| upvoted stories of the last 48 hours"):
|
| https://news.ycombinator.com/item?id=46708678 - "Waiting
| for dawn in search: Search index, Google rulings and
| impact on Kagi"
| pell wrote:
| I think Kagi buys search engine results from SERP vendors
| who typically scrape Google's results and offer an API
| experience on top of it.
| echelon wrote:
| No wonder Kagi is angry.
|
| Google is a monopoly across several broad categories.
| They're also a taxation enterprise.
|
| Google Search took over as the URL bar for 91% of all web
| users across all devices.
|
| Since this intercepts trademarks and brand names, Google
| gets to tax all businesses unfairly.
|
| Tell your legislators in the US and the EU that Google
| shouldn't be able to sell ads against registered trademarks
| (+/- some edit distance). They re-engineered the web to be
| a taxation system for all businesses across all categories.
|
| Searching for Claude -> Ads in first place
|
| Searching for ChatGPT -> Ads in first place
|
| Searching for iPhone -> Ads in first place
|
| This is inexcusable.
|
| Only searches for "ChatGPT versus", "iPhone reviews", or
| "Nintendo game comparison" should allow ads. And one could
| argue that the "URL Bar" shouldn't auto suggest these
| either when a trademark is in the URL bar.
|
| If Google won't play fair, we have to kill 50% of their
| search revenue for being egregiously evil.
|
| If you own a trademark, Google shouldn't be able to sell
| ads against you.
|
| --
|
| Google's really bad. Ideally we'd get an antitrust breakup.
| They're worse than Ma Bell. I wouldn't even split Google
| into multiple companies by division - I'd force them to be
| multiple copies of the same exact entity that then have to
| compete with each other:
|
| Bell Systems -> {BellSouth, Bell Atlantic, Southwestern
| Bell, ...}
|
| Google -> {GoogleA, GoogleB, GoogleC, ...}
|
| They'd each have cloud, search, browser, and YouTube. But
| new brand names for new parent companies. That would create
| all-out war and lead to incredible consumer wins.
| marginalia_nu wrote:
| Could probably argued that search access is an essential
| facility[1], though it doesn't appear antitrust law has
| anywhere near the same sort of enforcement it did in the
| past.
|
| [1] https://en.wikipedia.org/wiki/Essential_facilities_do
| ctrine
| throwaway290 wrote:
| what stops Kagi from indexing internet and makes them pay
| some guys to scrape search results from Google? one guy
| at Marginalia can do it and entire dev team at a PAID
| search engine can't?
| DangitBobby wrote:
| As we've seen here on HN on the AI boom, it's not
| wonderful when a bunch of companies all use bots to
| scrape the entire web. Many sites only allow Google
| scrapers in robots.txt and the public will fight you hard
| if you scrape them without permission. It's just one of
| those things where it would be better for everyone if
| search engines could pay for access to the work that's
| done only once.
| echelon wrote:
| > Many sites only allow Google scrapers in robots.txt and
| the public will fight you hard if you scrape them without
| permission.
|
| This just lets a monopoly replace the website instead of
| distributing power and fostering open source. The same
| monopoly that was already bleeding off the web's utility
| and taxing it.
| mrweasel wrote:
| I don't know about others, but we have special rules for
| Google, Bing, and a few others, rate-limiting them less
| than some random bot.
|
| The problem is scrapers (mostly AI scrapers from what we
| can tell). They will pound a site into the ground and not
| care and they are becoming increasingly good at hiding
| their tracks. The only reasonable way to deal with them
| is to rate-limit every IP by default and then lifting
| some of those restrictions on known, well behaving bots.
| Now we will lift those restrictions if asked, and
| frequently look at statistics to lift the restrictions
| from search engines we might have missed, but it's an up
| hill battle if you're new and unknown.
| thewebguyd wrote:
| > If you own a trademark, Google shouldn't be able to
| sell ads against you.
|
| This is frustrating even from a consumer perspective.
| Before I ran adblock everywhere, I couldn't stand that
| typing in a specific company I was looking for would just
| serve ads from any number of related brands that I wasn't
| looking for that were competitors.
| nemosaltat wrote:
| > Kagi This seems to be true, but more indirectly. From
| Kagi's blog [0] which is a follow up to a Kagi blog post
| from last year [1].
|
| [0]> Google: Google does not offer a public search API. The
| only available path is an ad-syndication bundle with no
| changes to result presentation - the model Startpage uses.
| Ad syndication is a non-starter for Kagi's ad-free
| subscription model.[^1]
|
| [0]> The current interim approach (current as of Jan 21,
| 2026)
|
| [0]> Because direct licensing isn't available to us on
| compatible terms, we - like many others - use third-party
| API providers for SERP-style results (SERP meaning search
| engine results page). These providers serve major
| enterprises (according to their websites) including Nvidia,
| Adobe, Samsung, Stanford, DeepMind, Uber, and the United
| Nations.
|
| I'm an avid Kagi user, and it seems like Kagi and some
| other notable interested parties have _already_ been unable
| to do get what they want/need with Google's index.
|
| [0]> The fact that we - and companies like Stanford,
| Nvidia, Adobe, and the United Nations - have had to rely on
| third-party vendors is a symptom of the closed ecosystem,
| not a preference.
|
| Hopefully someone here can clarify for me, or enumerate
| some of these "third-party vendors" who seem like they
| will/might/could be directly affected by this.
|
| [0] antibabelic > relevant https://blog.kagi.com/waiting-
| dawn-search [1] https://blog.kagi.com/dawn-new-era-search >
| [^1]: A note on Google's existing APIs: Google offers PSE,
| designed for adding search boxes to websites. It can return
| web results, but with reduced scope and terms tailored for
| that narrow use case. More recently, Google offers
| Grounding with Google Search through Vertex AI, intended
| for grounding LLM responses. Neither is general-purpose
| index access. Programmable Search Engine is not designed
| for building competitive search. Grounding with Google
| Search is priced at $35 per 1,000 requests - economically
| unviable for search at scale, and structured as an AI add-
| on rather than standalone index syndication. These are not
| the FRAND terms the market needs
| tpetry wrote:
| I believe they try to indirectly say they are using
| SerpApi or a similar product that scrapes Google search
| results to use them. And other big ones use it too so it
| must be ok...
|
| That must be the reason why they limit the searches you
| can do in the starter plan. Every SerpApi call costs
| money.
| sixhobbits wrote:
| Google is also suing SerpAPI
|
| And I can't prove correlation but they refused to index
| one of my domains and I think it _might_ be because we
| had some content on there about how to use SerpAPI
| monooso wrote:
| Kagi _does not_ use Google 's search index. From their post
| which made the front page of HN yesterday [1]:
|
| > Google does not offer a public search API. The only
| available path is an ad-syndication bundle with no changes
| to result presentation - the model Startpage uses. Ad
| syndication is a non-starter for Kagi's ad-free
| subscription model.
|
| [1]: https://news.ycombinator.com/item?id=46708678
| jsnell wrote:
| They then go on to say that they pay a 3rd party company
| to scrape Google results (and serve those scraped results
| to their users). So their search engine is indeed based
| on unauthorized and uncompensated use of Google's index.
|
| But since they're not using/paying for a supported API
| but just taking what they want, they indeed are unlikely
| to be impacted by this API turndown.
| DangitBobby wrote:
| Congrats on saying that in the most one-sided way
| possible. Google makes it literally impossible for them
| to pay for access to search results to make the product
| they want (customizable subscription search with no ads),
| and Google also is the de-facto globally sanctioned
| crawler because they are the only search engine anyone
| gives a shit about, and also sites need to be indexed by
| them to survive. In short, Google owns the river and
| sells the boats, and the public built a wall around it.
| Google is in a monopoly position in search.
| nova22033 wrote:
| >In short, Google owns the river and sells the boats, and
| the public built a wall around it.
|
| That would be a monopoly if there was only 1 river in the
| whole world.
| DangitBobby wrote:
| Yeah I mean think whatever you need to for the metaphor
| to work.
| Ferret7446 wrote:
| They have a monopoly on their own search results. There's
| nothing stopping anyone from making their own (hell, a
| poster did so in the comments above). God forbid we
| aren't entitled access to the fruits of their labor; the
| reason you want it isn't because you can't make it
| (again, see above). It's because making it good is hard,
| and you want the good results without yourself putting in
| the effort to make it
| ipaddr wrote:
| They get results from another provider who has authorized
| access. Google doesn't provide search results to
| unauthorized requests as many on tor have experienced.
| direwolf20 wrote:
| Residential proxies are also cheaper than you might
| realize.
| saltysalt wrote:
| I built my own web search index on bare metal, index now up to
| 34m docs: https://greppr.org/
|
| People rely too much on other people's infra and services,
| which can be decommissioned anytime. The Google Graveyard is
| real.
| orf wrote:
| Number of docs isn't the limiting factor.
|
| I just searched for "stackoverflow" and the first result was
| this: https://www.perl.com/tags/stackoverflow/
|
| The actual Stackoverflow site was ranked way down, below some
| weird twitter accounts.
| saltysalt wrote:
| I don't weight home pages in any way yet to bump them up,
| it's just raw search on keyword relevance.
| orf wrote:
| Sure, but the point is results are not relevant at all?
|
| It's cool though, and really fast
| saltysalt wrote:
| I'll work on that adjustment, it's fair feedback thanks!
| direwolf20 wrote:
| Unfortunately this is the bulk of search engine work.
| Recursive scraping is easy in comparison, even with
| CAPTCHA bypassing. You either limit the index to only
| highly relevant sites (as Marginalia does) or you must
| work very hard to separate the spam from the ham. And
| spam in one search may be ham in another.
| saltysalt wrote:
| I limit it to highly relevant curated seed sites, and
| don't allow public submissions. I'd rather have a small
| high-quality index.
|
| You are absolutely right, it is the hardest part!
| globular-toast wrote:
| What do you mean they're not relevant? The top result you
| linked contained the word stackoverflow didn't it? It's
| showing you exactly what you searched for. Why would you
| need a search engine at all if you already know the name
| of the thing? Just type stackoverflow.com into your
| address bar.
|
| I feel like Google-style "search" has made people really
| dumb and unable to help themselves.
| orf wrote:
| the query is just to highlight that relevance is a
| complex topic. few people would consider "perl blog posts
| from 2016 that have the stack overflow tag" as the _most_
| relevant result for that query.
| dredmorbius wrote:
| Google's entire (initial) claim-to-fame was "PageRank",
| referring both to the ranking of pages and co-founder
| Larry Page, which strongly prioritised a _relevance
| attribute_ over raw keyword findings (which then-popular
| alternatives such as Alta Vista, Yahoo, AskJeeves, Lycos,
| Infoseek, HotBot, etc., relied on, or the rather more
| notorious paid-rankings schemes in which SERP order was
| effectively sold). When it was first introduced, Google
| Web Search was absolutely worlds ahead of any
| competition. I remember this well having used them
| previously and adopted Google quite early (1998 /99).
|
| Even with PageRank result prioritisation is highly
| subject to gaming. Raw keyword search is far more so
| (keyword stuffing and other shenanigans), moreso as any
| given search engine begins to become popular and catch
| the attention of publishers.
|
| Google now applies other additional ordering factors as
| well. And of course has come to dominate SERP results
| with paid, advertised, listings, which are all but
| impossible to discern from "organic" search results.
|
| (I've not used Google Web Search as my primary tool for
| well over a decade, and probably only run a few searches
| per month. DDG is my primary, though I'll look at a few
| others including Kagi and Marginalia, though those
| rarely.)
|
| <https://en.wikipedia.org/wiki/PageRank>
|
| "The anatomy of a large-scale hypertextual Web search
| engine" (1998)
| <http://infolab.stanford.edu/pub/papers/google.pdf> (PDF)
|
| Early (1990s) search engines: <https://en.wikipedia.org/w
| iki/Search_engine#1990s:_Birth_of_...>.
| saltysalt wrote:
| PageRank was an innovative idea in the early days of the
| Internet when trust was high, but yes it's absolutely
| gamed now and I would be surprised if Google still relies
| on it.
|
| Fair play to them though, it enabled them to build a
| massive business.
| marginalia_nu wrote:
| Anchor text information is arguably a better source for
| relevance ranking in my experience.
|
| I publish exports of the ones Marginalia is aware of[1]
| if you want to play with integrating them.
|
| [1] https://downloads.marginalia.nu/exports/ grab
| 'atags-25-04-20.parquet'
| saltysalt wrote:
| Very interesting, and it is very kind of you to share
| your data like that. Will review!
| dredmorbius wrote:
| Though I'd think that you'd want to weight _unaffiliated_
| sites ' anchor text to a given URL much higher than an
| affiliated site.
|
| "Affiliation" is a tricky term itself. Content farms were
| popular in the aughts (though they seem to have largely
| subsided), firms such as Claria and Gator. There are
| chumboxes (Outbrain, Taboola), and of course affiliate
| links (e.g., to Amazon or other shopping sites). SEO
| manipulation is its own whole universe.
|
| (I'm sure you know far more about this than I do, I'm
| mostly talking at other readers, and maybe hoping to
| glean some more wisdom from you ;-)
| marginalia_nu wrote:
| Oh yeah, there's definitely room for improvement in that
| general direction. Indexing anchor texts is much better
| than page rank, but in isolation, it's not sufficient.
|
| I've also seen some benefit fingerpinting the network
| traffic the websites make using a headless browser, to
| identify which ad networks they load. Very few spam sites
| have no ads, since there wouldn't be any economy in that.
|
| e.g. https://marginalia-
| search.com/site/www.salon.com?view=traffi...
|
| The full data set of DOM samples + recorded network
| traffic are in an enormous sqlite file (400GB+), and I
| haven't yet worked out any way of distributing the data
| yet. Though it's in the back of my mind as something I'd
| like to solve.
| dredmorbius wrote:
| Oh, _that_ is clever!
|
| I'd also suspect that there are networks / links which
| are more likely signs of low-value content than others.
| Off the top of my head, crypto, MLM, known scam/fraud
| sites, and perhaps share links to certain social networks
| might be negative indicators.
| marginalia_nu wrote:
| You can actually identify clusters of websites based on
| the cosine similarity of their outbound links. Pretty
| useful for identifying content farms spanning multiple
| websites.
|
| Have a lil' data explorer for this:
| https://explore2.marginalia.nu/
|
| Quite a lot of dead links in the dataset, but it's still
| useful.
| snowwrestler wrote:
| Google's biggest search signal now is aggregate
| behavioral data reported from Chrome. That pervasive
| behavioral surveillance is the main reason Apple has
| never allowed a native Chrome app on iOS.
|
| It's also why it is so hard to compete with Google. You
| guys are talking about techniques for analyzing the
| corpus of the search index. Google does that _and_ has a
| direct view into how millions of people interact with it.
| xnx wrote:
| > That pervasive behavioral surveillance is the main
| reason Apple has never allowed a native Chrome app on iOS
|
| The Chrome iOS app still knows every url visited,
| duration, scroll depth, etc.
| saltysalt wrote:
| Yes indeed, they have an impossibly deep moat and deeper
| pockets. I'm certainly not trying to compete with them
| with my little side project, it's just for fun!
| pjc50 wrote:
| Confluence search does this, for our intranet. As a
| result it's barely usable.
|
| Indexing is a nice compact CS problem; not completely
| simple for huge datasets like the entire internet, but
| well-formed. Ranking is the thing that makes a search
| engine valuable. Especially when faced with people trying
| to game it with SEO.
| renegat0x0 wrote:
| I made also something for my own search needs. It's just an
| SQLite table of domains, and places. I have your search
| engine there also ;-)
|
| https://github.com/rumca-js/Internet-Places-Database
|
| Demo for most important ones https://rumca-
| js.github.io/search
| saltysalt wrote:
| Thank you, will check it out!
| johnofthesea wrote:
| I tested it using a local keyword, as I normally do, and it
| took me to a Wikipedia page I didn't know existed. So thanks
| for that.
| saltysalt wrote:
| It will throw up weird and interesting results sometimes
| ;-)
| lolive wrote:
| Lol, a GooglePlus URL was mentionned on a webpage i browsed
| this week.#blastFromThePast
| saltysalt wrote:
| I still remember their circles interface ;-)
| jfindley wrote:
| Unfortunately the index is the easy part. Transforming user
| input into a series of tokens which get used to rank possible
| matches and return the top N, based on likely relevence, is
| the hard part and I'm afraid this doesn't appear to do an
| acceptable job with any of the queries I tested.
|
| There's a reason Google became so popular as quickly as it
| did. It's even harder to compete in this space nowadays, as
| the volume of junk and SEO spam is many orders of magnitude
| worse as a percentage of the corpus than it was back then.
| saltysalt wrote:
| I am definitely not trying to complete with Google, instead
| I am offering an old-school "just search" engine with no
| tracking, personalization filtering, or AI.
|
| It's driven by my own personal nostalgia for the early
| Internet, and to find interesting hidden corners of the
| Internet that are becoming increasingly hard to find on
| Google after you wade through all of the sponsored results
| and spam in the first few pages...
| prophesi wrote:
| There may be a free CS course out there that teaches how
| to implement a simplified version of Google's PageRank.
| It's essentially just the recursive idea that a page is
| important if important pages link to it. The original
| paper for it is a good read, too. Curiously, it took me
| forever to find the unaltered version of the paper that
| includes Appendix A: Advertising and Mixed Motives,
| explaining how any search engine with an ad-based
| business model will inherently be biased against the
| needs of their users[0]
|
| [0] https://www.site.uottawa.ca/~stan/csi5389/readings/go
| ogle.pd...
| saltysalt wrote:
| Nice find, will review!
| 1718627440 wrote:
| The input on the results page doesn't work, you always need
| to return to the start page on which the browser history is
| disabled. That's just confusing behaviour.
| saltysalt wrote:
| I guess you used the return key instead of clicking on the
| search icon? Seems to be a bug with the return key, I'll
| fix that this weekend sorry.
| 1718627440 wrote:
| True, didn't occur to me, that I should click on the icon
| instead. Once I have clicked on the search icon once,
| enter also works. When I input a short query (single
| letter) it sometimes just shows a blank page, but maybe
| that is just HNs hug of death. Consider putting the query
| term more prominently in the front of the URL, so users
| can edit it. Also from the startpage, the URL in the
| URLbar isn't updated. As I already wrote, the browser
| shows completion for the searchbar on the result page,
| but does not for the one one the startpage. For my taste
| I would prefer less JS trickery, which would maybe
| already get rid of some of these issues.
| saltysalt wrote:
| Appreciate the detailed feedback! A lot of the JS
| trickery and URL shenanigans I'm doing is to prevent bot
| spam attempts, which was a real problem in the beginning.
| 1718627440 wrote:
| Sad state the web is in.
|
| It is intended, that the page currently shows a link to
| the wordpress login?
| saltysalt wrote:
| It does not use WordPress.
| 1718627440 wrote:
| I'm sorry, I am dumb and visited http://grepper.org/ .
| Where does your name come from I guess from grep for the
| WWW?
| tosti wrote:
| This is pretty cool. Don't let the naysayers stop you. Taking
| a stab at beating Google at their core product is bravery in
| my book. The best of luck to you!
| saltysalt wrote:
| Thank you kindly! It's just for fun.
| toofy wrote:
| > it's just for fun.
|
| amazing, for real.
|
| everything i've read and heard about the good internet is
| that it was good because sooooo many of the people did
| stuff for exactly that, fun.
|
| i've spent some time reading through some of the old
| email lists from earlier internet folks, they predicted
| exactly what weve turned this into. reading the
| resistance against early adoption of cookies is
| incredible to see how prescient some of those people
| were. truly incredible.
|
| keep having fun with it, i think it's our only way out of
| whatever this thing is we have now.
| saltysalt wrote:
| Couldn't agree more! The early pioneers of the Internet
| were hackers and tinkers, I've tried to maintain the same
| ethos.
| 1718627440 wrote:
| You should consider filtering by input language. Showing the
| same Wikipedia article in different languages is not helpful
| when I am searching in English. Also you may unify by entries
| by URL, it shows the same URL, just with different publish
| dates, which is interesting and might be useful, but should
| maybe be behind a toggle, as it is confusing at first.
| saltysalt wrote:
| Great feedback, agree I need to filter here. Some website
| localization is very hard to work around, because they will
| try to geo-locate the IP address of your bot and redirect
| it accordingly to a given language...
| 1718627440 wrote:
| The issue I was having was with the query
| "term+wikipedia" it then shows the wikipedia article in
| Czech, Hungarian, Russian, some kind of Arab and other
| before finally showing the English version. Then also a
| lot of that occur 2,3,4+ times with the same URL, just
| differing in crawltime by a few minutes.
| saltysalt wrote:
| It's a difficult problem to fix, you can set an Accept-
| Language header on crawl requests but his only works if
| the target website uses "Content Negotiation." Some sites
| ignore headers and determine language based on the IP
| address (Geo-IP) or the URL structure (e.g., /es/ vs
| /en/), basically a mess...
| 1718627440 wrote:
| I don't get the problem you claim. You crawl something
| and get a document in whatever language the site delivers
| you. You know the language of that document with the
| lang=... attribute of the document. What results you show
| for a given language is under your control and not
| influenced by what the crawled site chose to serve to the
| crawler.
| dust-jacket wrote:
| This is mad but cool. Keep at it.
| saltysalt wrote:
| Thanks, mad is fun for me! It costs me nothing if it fails.
| Tenemo wrote:
| That's super cool! Do you have any plans to commercialize it
| or it's just a pet project?
| saltysalt wrote:
| Pet project just for fun, thanks!
| bflesch wrote:
| Thanks for sharing, this is really impressive.
|
| Can you talk a bit about your stack? The about page mentions
| grep but I'd assume it's a bit more complex than having a
| large volume and running grep over it ;)
|
| Is it some sort of custom database or did you keep it simple?
| Do you also run a crawler?
| salawat wrote:
| It's been clear for the last decade that we have to wean
| ourselves off of centralized search indexes if only to
| innoculate the Net against censorship/politically motivated
| black holing.
|
| I can only weep at this point, as the heroes that were the
| Silent and Greatest generations (in the U.S.), who fought hard
| to pass on as much institutional knowledge as possible through
| hardcore organization and distribution via public and
| University library, have had that legacy shit on by these ad
| obsessed cretins. The entirety of human published
| understanding; and we make it nigh impossible for all but the
| most determined to actually avail themselves of it.
| raincole wrote:
| > "search the entire web"
|
| TIL they allowed that before. It sounds a bit crazy. Like
| Google is inviting people to repackage google search itself and
| sell it / serve with their own ads.
| MrGilbert wrote:
| You know, back in the days, the web used to be more open.
| Also - just because you CAN do something, doesn't mean you
| HAVE to.
| shevy-java wrote:
| It basically means that Google is now transitioning into a
| private web.
|
| Others have to replace Google. We need access to public
| information. States can not allow corporations to hold us
| here hostage.
| whs wrote:
| I tried it and contributed to searx. It didn't give the same
| result as Google, and it also have 10k request rate limit
| (per month I believe). More than that you'll have to "contact
| us"
| zitterbewegung wrote:
| I know that duckduckgo uses Microsoft Bing Custom search and
| honestly it is a much more robust system since you don't have
| to worry about Google axing it. https://www.customsearch.ai
| embedding-shape wrote:
| Instead you worry about Microsoft axing it? Sure, it might
| take 3 years instead of 6 months, and the shutdown period
| would be 1 year instead of 1 month, but hardly either are
| long-term solutions.
| aylmao wrote:
| > it might take 3 years instead of 6 months, and the
| shutdown period would be 1 year instead of 1 month
|
| This matters much more than people (and evidently those
| within Google) realize
| embedding-shape wrote:
| It matters more the shorter your future planning is.
| Neither works if you're looking forward 3-4 years, for
| example.
| sumtechguy wrote:
| Having been downsizing my horde of computer junk I have
| several large boxes of full MSDN disc sets. There is a
| graveyard of MS stuff that is no longer supported. The only
| thing with MS is they seem to give you a better off ramp
| usually than 'oh well sucks to be you'.
| thayne wrote:
| Bing Custom Search was discontinued last year. Although
| duckduckgo probably has some kind of special contract with
| Microsoft.
| 1vuio0pswjnm7 wrote:
| If its Motion for a Partial Stay is denied, or if it loses on
| appeal, then under this Final Judgement Google will be forced
| to offer syndicated "full web" search to Qualified Competitors
|
| https://dn710204.ca.archive.org/0/items/gov.uscourts.dcd.223...
| Antibabelic wrote:
| Relevant: Waiting for dawn in search: Search index, Google
| rulings and impact on Kagi
| https://news.ycombinator.com/item?id=46708678
| mrweasel wrote:
| This might be me reading it wrong, but isn't shutting down the
| full-web search going against the ruling mentioned in the Kagi
| post?
|
| > Google must provide Web Search Index data (URLs, crawl
| metadata, spam scores) at marginal cost.
|
| Maybe they're shutting down the good integration and then Kagi,
| Ecosia and others can buy index data in an inconvenient way
| going forward?
| Hackbraten wrote:
| If I understand Kagi's blog post correctly, then here's what
| happened, chronologically:
|
| Kagi makes deals with many search engines so they can have
| raw search results in exchange for money.
|
| Google says: no, you can't have raw search results because
| only whales can get those. Only thing we can offer you is
| search results riddled with ads and we won't allow you to
| reorder or filter them.
|
| Kagi thinks Google's offer is unacceptable, so Kagi goes to a
| third party SERP API, which scrapes Google at scale and sells
| the raw search results to Kagi and others.
|
| August 2024: Court says Google is breaking the law by selling
| raw search results only to whales.
|
| December 2025: Court orders that for the next six years, 1.
| Google must no longer exclude non-whales from buying raw
| search results, 2. Google must offer the raw search results
| for a reasonable price, and 3. Google can no longer force
| partners to bundle the results with ads.
|
| December 2025: Google sues the third-party scraping
| companies.
|
| January 2026: Google says "hey, the old search offering is
| going to go away, there's going to be a new API by 2027, stay
| tuned."
| mrweasel wrote:
| I don't really see any mentioning of a new API, beyond
| their Vertex AI thing, and I don't know how comparable that
| might be. Also it is capped at 50 domains (by default).
|
| It is perhaps a clever legal workaround. They must sell
| access to their index, but the verdict didn't state how
| much of it you can buy access to at any one time. So they
| put a limit of 50 domains, because that accommodates
| everyone who's not a search engine, but effectively blocks
| Kagi and Ecosia, while not exactly refusing to sell to
| them.
| DangitBobby wrote:
| It's only a "clever" workaround in a captured legal
| system that isn't interested in anti-monopoly outcomes.
| Any competent legal system would slap that shit down.
| Just the thought that they could "hack this ruling with
| one weird trick" is infuriating.
| Hackbraten wrote:
| > I don't really see any mentioning of a new API, beyond
| their Vertex AI thing
|
| I was referring to the following statement about full web
| search where they don't mention a 50-domains limit:
|
| > if your use case necessitates full web search, contact
| us to express your interest in and get more information
| about our full web search solution. Your transition to an
| alternative solution needs to be completed by January 1,
| 2027.
| agosta wrote:
| Among the various rulings, Google is supposed to provide
| access at market rates... which they are. At least for what
| is published: $5 per thousand queries is market rate for a
| product like this - see Brave's Search API pricing
| https://api-dashboard.search.brave.com/app/plans?tab=ai.
|
| Granted, that is scoped to 50 domains. But we don't know if
| the enterprise package, which allows full web search, isn't
| roughly market rate.
| HPsquared wrote:
| I had misread the title as "Google is ending (full-web search)
| for [aka in favour of] (niche search engines)"
|
| The correct parsing is: "Google is ending (full-web search for
| niche search engines)"
| dredmorbius wrote:
| "Google will discontinue third-party niche search engine access
| to full-web search" would be far clearer.
|
| Given that the title supplied is effectively editorialised, and
| the original article's title is effectively content-free
| ("Updates to our Web Search Products & Programmable Search
| Engine Capabilities"), my rewording would be at least as fair.
|
| HN's policy is to try to use text from the article itself where
| the article _title_ is clickbait, sensational, vague, etc.,
| however. I suspect Google 's blog authors are aware of this,
| and they've carefully avoided any readily-extracted clear
| statements, though I'll take a stab...
|
| Here's the most direct 'graph from TFA:
|
| _Custom Search JSON API: Vertex AI Search is a favorable
| alternative for up to 50 domains. Alternatively, if your use
| case necessitates full web search, contact us to express your
| interest in and get more information about our full web search
| solution. Your transition to an alternative solution needs to
| be completed by January 1, 2027._
|
| We can get a clearer, 80-character head that's somewhat
| faithful to that with:
|
| "Google Search API alternative Vertex AI Search limited to 50
| domains" (70 chars).
|
| That's still pretty loosely adherent, though it (mostly) uses
| words from the original article. I'm suggesting it to mods via
| email at hn@ycominator.com; others may wish to suggest their
| own formulations.
| chromehearts wrote:
| Is this about the little Google Search Bar that is present on
| some websites? Or am I mistaking something
| 01jonny01 wrote:
| Kind of, however the Google Search Bar present on website is
| usually there to search across their domain, the search results
| are limited to their domain e.g example.com/page1,
| example.com/page2. Google will carry on supporting this.
|
| What they are ending is their support for websites to search
| across the entire web. The websites that search across the
| entire web are usually niche search engine websites.
| chromehearts wrote:
| Ahh; so that's the difference. Thanks!
| vaylian wrote:
| Meanwhile in Europe: Qwant and Ecosia team up to build their own
| search index: https://blog.ecosia.org/eusp/
| tweetle_beetle wrote:
| It's a noble effort, but they're so late to the game that it's
| hard to see them making a significant dent. I hope I'm wrong.
|
| They were:
|
| > aiming to serve 30% of French search queries [by end of 2025]
|
| https://blog.ecosia.org/launching-our-european-search-index/
| Gigachad wrote:
| I feel like soon there won't even be a point having a search
| engine since almost the entire internet will be useless AI
| slop.
| altairprime wrote:
| It's as though full-text search of websites you've never
| heard of was a mistake :)
|
| PageRank wouldn't exist without webrings, directories, and
| forums you could only search individually, and we _thrived_
| on that Internet.
|
| Welcome back, ye olde Internet.
| baubino wrote:
| The old internet is still there. It hasn't gone away; it's
| just undiscoverable with ad-based search. The more slop
| there is, the more necessary it is to have good search
| engines.
| zelphirkalt wrote:
| Recently, I set up a fresh system on a laptop. Ahahahaaa,
| how utterly crap Google search results now are! It fills
| me with some stress and disgust to use that. Now one of
| the first things I do, right after emergency using
| duckduckgo to search for uBlock Origin and NoScript, is
| to get Kagi search installed as default search. Then I
| can continue setting things up more calmly.
| direwolf20 wrote:
| If you haven't tried Marginalia Search yet, do so. It's a
| small web search.
| direwolf20 wrote:
| A search engine doesn't have to search the entire internet.
| Most of them are extremely opinionated about what they
| index.
| anal_reactor wrote:
| Seriously though. Five years ago Google already became
| unusable without "site:reddit.com" which is actually
| hilarious for a search engine that's supposed to search the
| entire internet. Nowadays reddit is also shit, which means
| that the only use case for me to use Google or any search
| engine is to find products that for some reason I don't
| want to buy on Amazon.
|
| Internet isn't a global village, it's a global ghetto, and
| it's becoming increasingly true that the only way not to
| lose is not to play.
| johnofthesea wrote:
| Better late than never.
|
| > The French index is at an advanced stage of completion, we
| have started creating the German language index, and the
| English one should start shortly. All progress is quickly
| integrated into the Qwant STAAN API.
|
| https://noc.social/@327ppm/115934198650900394
| blell wrote:
| They can build whatever they want with lots of #hashtags
| and public money, but that doesn't mean they'll get 30% of
| French people to use it.
|
| But of course they managed to cut themselves a nice salary
| with EU funds, paid in part by me and you, so that's all
| that matters.
| philistine wrote:
| Oh no, someone is making money outside of the big
| American monopolists. Quick, the vapors!
| 1718627440 wrote:
| Which honestly no user cares about. They only care about
| whether it is good enough that they can use it.
| Marketshare only matters if you fear the vendor might
| shut it down, or if you are running ads.
| johnofthesea wrote:
| > with lots of #hashtags
|
| I missed this one. What was it about?
| blell wrote:
| Click the link.
| alexgieg wrote:
| The French government managed to rein in Amazon so
| traditional French stores, both online and brick and
| mortar ones, don't go bankrupt due to Amazon's unending
| pockets.
|
| If they deem it necessary to rein in Google, they _will_
| rein in Google. There 's no lack of tools for this,
| ranging from obliging phones sold in French territory to
| offer the French search engine as the default, to forcing
| every Google search result to promote the local search
| engine prominently, to campaigns about how it's important
| for national security not to rely on an adversary/enemy
| country's services, to everything in between and beyond.
| whatarethembits wrote:
| The landscape has completely shifted now; now more than
| ever, there's a real need for Europe to pivot from
| relying too much on an adversary.
| londons_explore wrote:
| What examples are there of people using this?
| 01jonny01 wrote:
| There is literally thousands of independent search engines that
| use Programmable search to search the entire web. Many ISP
| providers use it on their homepage, kids-based search engines
| like wackysafe.com use it, also search engines that focus on
| privacy like gprivate.com etc
| TeMPOraL wrote:
| Also LLM tools. Programmable Search Engine API was a way to
| give third-party LLM frontends the ability to give LLMs a web
| search tool. Notably, this was a common practice long before
| any of the major LLM providers added search capabilities to
| their frontents.
| 01jonny01 wrote:
| Exactly, Google want every one depended on Gemini.
| bovermyer wrote:
| I'm curious about what it would take to build my own "toy" search
| engine with its own index. Anyone ever tried this?
| Gigachad wrote:
| Might find YaCy interesting. It's meant to be a decentralised
| search engine where users scrape the internet and can search
| other users indexes in a kind of torrent like way.
|
| I found it didn't really work as a real search engine but it
| was interesting.
| reddalo wrote:
| Good luck scraping websites without being blocked, if you're
| not Google.
| marginalia_nu wrote:
| Well you'll get blocked some places but it's not too big of a
| deal. If you're running an above board operation, you can
| surprisingly often successfully just email the admin
| explaining what you're doing, and ask to be unblocked.
| BolsunBacset wrote:
| Sounds very time consuming. Glad you're able to sustain
| yourself to be able to do it full time.
| marginalia_nu wrote:
| Yeah that's where I started out in 2021. Been at it for almost
| 5 years now, last three of which full time. I'm indexing about
| 1.1 billion documents now off a single server.
|
| Hard part is doing it at any sort of scale and producing useful
| results. It's easy to build something that indexes a few
| million documents. Pushing into billions is a bigger challenge,
| as you start needing a lot of increasingly intricate bespoke
| solutions.
|
| Devlog here:
|
| https://www.marginalia.nu/tags/search-engine/
|
| And search engine itself:
|
| https://marginalia-search.com/
|
| (... though it operates a bit sub-optimally now as I'm using a
| ton of CPU cores to migrate the index to use postings lists
| compression, will take about 4-5 days I think).
| rickette wrote:
| Curious on what (how much) hardware your running this.
| marginalia_nu wrote:
| Currently running off
|
| AMD EPYC 7543 x2 for 64 cores/128 threads
|
| 512 GB RAM
|
| ~ 90 TB of PM9A3 SSDs across 12 physical devices
|
| Storage is not very full though. I'm probably using about a
| third of it at this point.
| riku_iki wrote:
| I assume you able to monetize it since you work on it full
| time?
| marginalia_nu wrote:
| I'm mostly living off grants and donations at this point,
| but the plan down the line is to polish it up well enough
| to make some money off providing an API like the one Google
| is making it a hassle to access with this change :-)
| lighthouse1212 wrote:
| The 'Google Graveyard is real' sentiment captures something
| important: every dependency on a large platform is a loan that
| can be called in. The 34-million-document indie index project
| someone mentioned is the right response - own your core
| infrastructure. Easier said than done for whole-web search, but
| the same principle applies everywhere.
| 01jonny01 wrote:
| Much easier said than done, especially if you are serving users
| on scale.
| 1718627440 wrote:
| Since the issue here is self-hosting and "core
| infrastructure", that isn't a problem, but everyone has their
| own search index, isn't credible either.
| jamesbelchamber wrote:
| Are competing search indexes (Bing, Ecosia/Qwant, etc)
| objectively worse in significant ways, or is Google just so
| entrenched that people don't want to "risk it" with another
| provider (and/or preferences and/or inertia).
|
| I suppose I'm asking whether this is actually a _good thing_ in
| that it will stimulate competition in the space, or if it's just
| a case that Google's index is now too good for anyone to
| reasonably catch up at this point.
| 01jonny01 wrote:
| The beauty about Google Programmable Search across the entire
| web is that it's free and users can make money by linking it
| their Adsense account.
|
| Bing charge per query for the average user. Ecosia and Qwant
| use Bing to power their results, probably under some type of
| license, which results in them paying much less per query than
| a normal user.
| thayne wrote:
| Bing recently shut down their API product, which was already
| very expensive.
|
| If you want programmatic access to search results there aren't
| really many options left.
| SirHumphrey wrote:
| I can manage fine with other search indexes for English
| language searches; weather that is because others got better or
| google got worse i cannot tell, though I suspect the latter.
|
| But for searching in more niche languages google is usually the
| only decent option and I have little hope that others will ever
| reach the scale where they could compete.
| Antibabelic wrote:
| Bing's index is smaller than Google's, and anecdotally I get
| fewer relevant results when using it, particularly from sites
| like Reddit that have exclusive search deals with Google.
| carlosjobim wrote:
| Yes, for non English queries they are all rubbish. And that's
| billions of users.
| sreekanth850 wrote:
| Never build a product with core feature depending on a third-
| party, you will eventually get fucked up for sure. always have a
| 70:30 rule for revenue where 70% is core independent features.
| halapro wrote:
| Soon you'll find that you cannot exist on the web without
| relying on third parties. Sometimes you'll even have trouble
| getting paid thanks to the painful existence of payment
| processors.
| sreekanth850 wrote:
| True, you can't exist without 3rd parties. But you shouldn't
| let them be your core moat/USP. Jasper is a great example,
| they depended too much on LLM access, then ChatGPT launched
| and ate the value. Using third party APIs is fine, but
| building a product whose core depends on them is suicide.
| direwolf20 wrote:
| That's why I eschew HTTPS.
| cubefox wrote:
| Is this perhaps to prevent ChatGPT, Claude and Grok to use Google
| Search? It would make sense for Google to keep that ability for
| Gemini.
| 01jonny01 wrote:
| I suspect its going to hurt the indie developers and small
| start-ups who do not have special licensing agreements.
| direwolf20 wrote:
| They'll go adversarial interop through SerpAPI, just like Kagi
| does. SerpAPI will get the money instead of Google getting it.
| cubefox wrote:
| "Why we're taking legal action against SerpApi's unlawful
| scraping" https://blog.google/innovation-and-
| ai/technology/safety-secu...
| snackbroken wrote:
| What's their angle here? Courts have been over the "Is
| scraping websites that don't want to be scraped OK?"
| question plenty of times. From
|
| > SerpApi deceptively takes content that Google licenses
| from others (like images that appear in Knowledge Panels,
| real-time data in Search features and much more), and then
| resells it for a fee. In doing so, it willfully disregards
| the rights and directives of websites and providers whose
| content appears in Search.
|
| it sounds like they are somehow suing on behalf of whoever
| they are licensing content from, but does that even give
| Google standing?
|
| I guess I'm asking if they actually are hoping to win or
| just going for a "the process is the punishment"+"we have
| more money and lawyers than you" approach.
| solarkraft wrote:
| This will significantly impact (quite possibly kill) Startpage
| and Ecosia, who are effectively white-label Google, right?
|
| What alternatives are there besides Bing? Is it really so hard
| that it's not considered worth doing? Some of the AI companies
| (Perplexity, Anthropic) seem to have managed to get their own
| indexing up and running.
| ColinHayhurst wrote:
| Excuse the self-promotion but Mojeek offers a web search API
| (>9 billion pages): https://www.mojeek.com/services/search/web-
| search-api/
| direwolf20 wrote:
| Nice! I hope Kagi is licensing access from you?
| jonplackett wrote:
| Are search engines like Kagi completely screwed by this or is
| there a way for them to keep operating?
| direwolf20 wrote:
| Kagi doesn't have a partnership with Google - they work under
| adversarial interoperability, stealing results from Google
| against their will, and paying some third-party to enable this.
| They'd like to simply pay Google, but Google doesn't want their
| money.
| TiredOfLife wrote:
| Kagi is backed by russia so they will be fine.
| contagiousflow wrote:
| What do you mean "backed by"
| Hackbraten wrote:
| You're probably referring to the fact that their search
| results include entries from Yandex. That's something
| entirely different from being "backed by Russia." If
| anything, they pay Russia, not the other way around.
| direwolf20 wrote:
| And Yandex has better results than Google, so I support
| this move. The USA does more wars than Russia does anyway.
| zoobab wrote:
| Antitrust do not work against large companies.
|
| Just dissolve them in acid.
| marginalia_nu wrote:
| This is the type of monopoly abuse these laws were designed to
| target, and antitrust laws actually do work against large
| companies.
|
| If you actually enforce them.
|
| Unfortunately, during the Reagan administration, political
| sentiment toward monopolies shifted and since then antitrust
| law has been a paper tiger at best.
| zoobab wrote:
| I heard when Bush came to power, the antitrust complaint
| against Microsoft monopoly driven by the government was
| dropped.
| shevy-java wrote:
| Google has consistently ruined its search engine in the last
| (almost) 10 years. You can find numerous articles about this, as
| well as videos on youtube (which is also controlled by google).
|
| Not long ago they ruined ublock origin (for chrome; ublock origin
| lite is nowhere near as good and effective, from my own
| experience here).
|
| Now Google is also committing towards more evil and trying to
| ruin things for more - people, competitors, you name it. We can
| not allow Google to continue on its wiched path here. It'll just
| further erode the quality. There is a reason why "killed by
| google" is more than a mere meme - a graveyard of things killed
| by google.
|
| We need alternatives, viable ones, for ALL Google services. Let's
| all work to make this world better - a place without Google.
| philistine wrote:
| To me there are two eras of the Google Graveyard(tm). First,
| there's the _we 're a university research group with an ad
| company footing the bill_ era. That's the early Google era, and
| it was a consequence of its corporate structure. They valued
| new projects, market fit, profitability, and maintenance be
| damned.
|
| We're in the second era. The era of _the MBAs are shutting down
| the last remnants of openness the company ever had_.
| consumer451 wrote:
| Dumb question:
|
| I keep seeing posts about how ~"the volume of AI scrapers is
| making hosting untenable."
|
| There must a ton of new full-web datasets out there, right?
|
| What are the major hurdles that prevent the owners of these
| datasets from providing them to third parties via API? Is it the
| quality of SERP, or staleness? Otherwise, this seems like a
| potentially lucrative pivot/side hustle?
| Terretta wrote:
| > _the volume of AI scrapers is making hosting untenable_
|
| Aside from that potential, it's also not true.
|
| A Pentium Pro or PIII SSE with circa 1998-99 Apache happily
| delivers a billion hits a month w/o breaking a sweat unless you
| think generating pages for every visit is better than
| generating pages when they change.
| Tenemo wrote:
| I think it is true that it is a real problem (EDIT: but
| doesn't necessarily make "hosting untenable"), but you are
| correct to point out that modern pages tend to be horribly
| optimized (and that's the source of the problem). Even
| "dynamic" pages using React/Next.js etc. could be pre-
| rendered and/or cached and/or distributed via CDNs. A simple
| cache or a CDN should be enough to handle pretty much any
| scrapping traffic unless you need to do some crazy logic on
| _every_ page visit - which should almost never be the case on
| public-facing sites. As an example, my personal site is
| technically written in React, but it 's fully pre-rendered
| and doesn't even serve JS - it can handle huge amounts of
| bot/scrapping traffic via its CDN.
| consumer451 wrote:
| OK, I agree with both of you. I am an old who is aware of
| NGINX and C10k. However, my question is: what are the
| economic or technical difficulties that prevent one of
| these new web-scale crawlers from releasing og-pagerank-
| api.com? We all love to complain about modern Google SERP,
| but what actually prevents that original Google experience
| from happening, in 2026? Is it not possible?
|
| Or, is that what orgs like Perplexity are doing, but with
| an LLM API? Meaning that they have their own indexes, but
| the original q= SERP API concept is a dead end in the
| market?
|
| _Tone: I am asking genuine questions here, not trying to
| be snarky._
| consumer451 wrote:
| Please see my reply to the other child comment. That is my
| actual question, apologies for not being more clear.
| nairboon wrote:
| Regarding alternate search engines: I consider the idea of YaCy
| kind of interesting: a P2P search engine: https://yacy.net/
|
| Although, it needs some more work and peers to be usable as a
| general-purpose search engine.
| mark_l_watson wrote:
| Not directly covered by this blog, but for low cost and good
| performance the combination of gemini-3-flash with search
| grounding is hard to beat, at least for the many small
| experiments I use it for.
|
| One thing touched upon in comments here: I never understood how
| it was proper for 3rd parties to scrape Google search results and
| reuse/resell them.
|
| Really off topic, sorry, but I am surprised that more companies
| don't build local search indices for just the few hundred web
| domains that are important to their businesses. I have tried this
| in combination with local (small and fast) LLMs and I think this
| is unappreciated tech: fast, cheap, and local.
| bennydog224 wrote:
| I built many products on Google PSE (Custom Search). Results were
| nowhere near as good as regular Google, but still useful. I
| usually needed to use another library to get the DOM content
| anyway. But it still was solid for grounding/checking data.
|
| RIP, another one to the Google Graveyard.
| motoboi wrote:
| This and agressive anti-bot at YouTube is Alphabet closing the AI
| data leaking
| thayne wrote:
| Does this mean the !g bang will stop working in DuckDuckGo?
| direwolf20 wrote:
| Doesn't it just redirect you to Google? So it will still work.
| jpalepu33 wrote:
| This is a clear example of why building on proprietary APIs is
| risky for indie devs and small startups. I've seen similar
| patterns with Twitter's API restrictions and other platforms
| gradually closing down their ecosystems.
|
| For anyone affected: consider this a forcing function to either:
| 1. Build your own lightweight search infrastructure (tools like
| Meilisearch, Typesense make this more accessible now) 2. Use
| adversarial interop via services like SerpAPI (though Google is
| already taking legal action there) 3. Pivot to specialized
| vertical search where you control the data sources
|
| The real lesson here is the importance of owning your core value
| proposition. If your product's moat depends entirely on a third-
| party API that can be yanked away with 12 months notice, you
| don't really have a sustainable business.
|
| Google is essentially saying: indie search is dead, pay
| enterprise prices or leave. This will probably accelerate the
| trend toward specialized, domain-specific search engines that
| don't rely on Google's index at all.
| joelboersma wrote:
| I've been occasionally working on a toy project that's basically
| "Google search in a TUI" that used this API. I was already
| planning on adding Brave Search as an option for a different
| backend, and I was heavily considering making it the default just
| because it's much easier to set up on the user's end. This is the
| straw that broke the camel's back.
___________________________________________________________________
(page generated 2026-01-23 23:01 UTC)