[HN Gopher] Introducing deep research
       ___________________________________________________________________
        
       Introducing deep research
        
       Author : mfiguiere
       Score  : 553 points
       Date   : 2025-02-03 00:06 UTC (22 hours ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | tucnak wrote:
       | Look, who's copying who now. They added _the_ button!
        
         | blackeyeblitzar wrote:
         | I'm not sure I understand what you mean by "the button". If
         | you're comparing this to DeepSeek's copying, it's not really
         | the same thing right? DeepSeek essentially stole intellectual
         | property by violating OpenAI's terms of service. As I
         | understand it, this is a copy of Google's Deep Research
        
           | lompad wrote:
           | Deepseek proved that there is no moat. Thus no path to
           | profitability for openai, anthropic & co.
           | 
           | Stealing from thieves is fine by me. Sama was the one
           | claiming that all information could be used to train LLMs,
           | without permisdion of the copyright holders.
           | 
           | Now the same is being done to openai. Well, too bad.
        
             | blackeyeblitzar wrote:
             | > Stealing from thieves is fine by me. Sama was the one
             | claiming that all information could be used to train LLMs,
             | without permisdion of the copyright holders.
             | 
             | OpenAI and other LLMs scraping the internet is probably
             | covered under fair use. DeepSeek's violation of OpenAI's
             | terms is pretty clearly a violation of their terms and not
             | legal.
        
               | simion314 wrote:
               | >DeepSeek's violation of OpenAI's terms is pretty clearly
               | a violation of their terms and not legal.
               | 
               | Here is a new thing you learn today, ToS are not laws,
               | you can ignore any ToS and at worst the company might
               | close your account.
        
           | figers wrote:
           | Didn't OpenAI steal everyone's data they could consume from
           | the internet? Actively being sued by the NY Times and others
           | for this...
        
             | blackeyeblitzar wrote:
             | Yes those cases will be interesting. By default a lot of
             | copyrighted content may be legal to use for training (in
             | the US but also many other places) under what's called fair
             | use. The cases you're referring to will likely reinforce
             | this, but it isn't known yet. Note that it's not just
             | OpenAI on that side of the argument but also other (non
             | tech) organizations that believe protecting fair use here
             | is current law and essential.
        
           | therealpygon wrote:
           | Care to explain how something that cannot be copyrighted and
           | was not generated by a human is "intellectual property"? Or
           | are you just parroting a narrative?
        
             | blackeyeblitzar wrote:
             | Trade secrets are protected by the law. It doesn't require
             | copyright.
        
               | therealpygon wrote:
               | Explain the trade secrets contained in non-copyrightable
               | AI outputs and the reasonable efforts OpenAI takes to
               | keep its AI output "secret". Or are you confused about
               | what a "trade secret" actually is?
        
           | caspper69 wrote:
           | I chuckle every time I see this. Poor OpenAI.
           | 
           | Meanwhile, their entire training corpus was the result of
           | scraping the intellectual property and copyrighted materials
           | of THE ENTIRE PUBLIC INTERNET.
           | 
           | Woe is them to be sure.
        
             | blackeyeblitzar wrote:
             | OpenAI's scraping will likely be ruled as fair use.
        
       | tomrod wrote:
       | I'm not sure if this is worth a subscription. DSPy and DeepseekR1
       | can already move this direction, if I understand right.
        
         | apstls wrote:
         | What is the current state of DSPy optimizers? When I originally
         | checked it out it appeared to just be optimizing the set of
         | examples used for n-shot prompting.
        
         | tmnvdb wrote:
         | You understand wrong.
        
       | kenjackson wrote:
       | If it has access to play by play data for all sports this could
       | be an absolute playground for amateur sports statisticians. The
       | possibilities...
        
       | rvz wrote:
       | It appears that OpenAI is in panic mode after the release of
       | DeepSeek. Before they were confident in competing against Google
       | on any AI model they release.
       | 
       | Now they are scrambling against open-source after their
       | disastrous operator demonstration and using this deep research
       | demo as cover. Nothing that Google or Perplexity could not
       | already do themselves.
       | 
       | By the end of them month, this feature is going be added by a
       | bunch of other open-source projects and this feature won't be as
       | interesting very quickly.
        
         | blackeyeblitzar wrote:
         | I don't think you're comparing the right things here. This
         | feature is more like Google's Deep Research, which basically
         | goes off and does a whole lot of search and compute to produce
         | something more like a full research report. This has nothing to
         | do with open weight models like DeepSeek (note: DeepSeek,
         | Llama, etc are NOT open source). This feature doesn't just
         | require the research on the model but also enormous compute.
         | Plus anyone using such a feature for real work is not going to
         | be using DeepSeek or whatever, but a product with trustworthy
         | practices and guarantees.
        
       | PartiallyTyped wrote:
       | I feel that a lot of this can already be achieved via aider (not
       | affiliated), and any of the top models.
        
         | tmnvdb wrote:
         | Do you have any benchmarks to back up your 'feelings'?
        
           | PartiallyTyped wrote:
           | I really don't like the snarky tone of the parent comment.
           | 
           | Nonetheless, I don't think this is even something that can
           | easily be benchmarked. I'd recommend you take a look at aider
           | [1], and consider how I drew similarities between it and
           | what's presented here.
           | 
           | Has ClosedAI presented any benchmarks / evaluation protocols?
           | 
           | [1] https://aider.chat/
        
             | tmnvdb wrote:
             | Yes, they show benchmarks in the article linked here. Did
             | you not read it?
        
               | PartiallyTyped wrote:
               | I don't think you actually read it. The benchmarks are in
               | reference to the model that's underlying deep-research,
               | and not deep-research itself. For the latter, they have
               | anecdata from scientists.
        
       | DigitalSea wrote:
       | Not sure if people picked up on it, but this is being powered by
       | the unreleased o3 model. Which might explain why it leaps ahead
       | in benchmarks considerably and aligns with the claims o3 is too
       | expensive to release publicly. Seems to be quite an impressive
       | model and the leading out of Google, DeepSeek and Perplexity.
        
         | xbmcuser wrote:
         | It was expensive as they wanted to charge more for it but
         | deepseek has forced their hand
        
           | Sparkyte wrote:
           | Rightfully so, some models are getting super efficient.
        
           | willy_k wrote:
           | They've only released o3-mini, which is a powerful model but
           | not the full o3 that is being claimed as too expensive to
           | release. That being said, DeepSeek for sure forced their hand
           | to release o3-mini to the public.
        
             | shawabawa3 wrote:
             | o3 mini was previewed in December. Deepseek maybe made them
             | release it a few weeks early but it was already on its way
        
               | sdesol wrote:
               | I guess the question is, did DeepSeek force them to
               | rethink pricing? It's crazy how much cheaper it (v3 and
               | R1) is, but considering they (Deepseek) can't keep up
               | with demand, the price is kind of moot right now. I
               | really do hope they get the hardware to support the API
               | again. The v3 and R1 models that are hosted by others are
               | still cheap compared to the incumbents, but nothing can
               | compete with DeepSeek on price and performance.
        
             | kandesbunzler wrote:
             | no they didn't, this was literally all announced in
             | December with a release date for January
        
         | bbor wrote:
         | Interesting, thanks for highlighting! Did not pick up on that.
         | Re:"leading", tho:
         | 
         | Effectiveness in this task environment is well beyond the
         | specific model involved, no? Plus they'd be fools (IMHO) to
         | only use one size of model for each step in a research task --
         | sure, o3 might be an advantage when synthesizing a final answer
         | or choosing between conflicting sources, but there are many,
         | many steps required to get to that point.
        
           | xendipity wrote:
           | I don't believe we have any indication that the big offerings
           | (claude.ai, Gemini, operator, tasks, canvas, chatgpt) use
           | multiple models in one call (other than for different
           | modalities like having Gemini create an image). It seems to
           | actually be very difficult technically and I'm curious as to
           | why.
           | 
           | I wonder how much of an impact our being still so early in
           | the productization phase of this all is. Like it takes a ton
           | of work and training and coordination to get multiple models
           | synced up into an offering and I think the companies are
           | still optimizing for getting new ideas out there rather truly
           | optimizing them.
        
             | someothherguyy wrote:
             | ...or its all a farce, for now.
        
         | ai-christianson wrote:
         | Has anyone here tried it out yet?
        
           | maroonblazer wrote:
           | Per the below, seems it's not available to many yet.
           | 
           | https://news.ycombinator.com/item?id=42913575
        
           | nycdatasci wrote:
           | Pro user. No access like everyone else.
           | 
           | OpenAI is very much in an existential crisis and their poor
           | execution is not helping their cause. Operator or "deep
           | research" should be able to assume the role of a Pro user,
           | run a quick test, and reliably report on whether this is
           | working before the press release right?
        
         | mistercheph wrote:
         | I'm sure o3 will be a generation ahead of whatever deepseek,
         | google and meta are doing today when it launches in 10 months,
         | super impressive stuff.
        
           | petesergeant wrote:
           | I'm not sure if you're implying this subtly in your comment
           | or not, as it's early here, but it does of course need to be
           | a generation ahead of what 10 months of their competitors
           | moving forward have done too. Nobody is standing still
        
             | bruce511 wrote:
             | I read a fair amount of sarcasm in the parent comment ;)
        
         | lordofgibbons wrote:
         | > Which might explain why it leaps ahead in benchmarks
         | considerably and aligns with the claims o3 is too expensive to
         | release publicly
         | 
         | It's the only tool/system (I won't call it an LLM) in their
         | released benchmarks that has access to tools and the web. So,
         | I'd wager the performance gains are strictly due to that.
         | 
         | If an LLM (o3) is too expensive to be released to the public,
         | why would you use it in a tool that has to make hundreds of
         | inference calls to it to answer a single question? You'd use a
         | much cheaper model. Most likely o3-mini or o1-mini combined
         | with o4-mini for some tasks.
        
           | og_kalu wrote:
           | >why would you use it in a tool that has to make hundreds of
           | inference calls to it to answer a single question? You'd use
           | a much cheaper model.
           | 
           | The same reason a lot of people switched to GPT-4 when it
           | came out even though it was much more expensive than 3 -
           | doesn't matter how cheap it is if it isn't good enough/much
           | worse.
        
         | bitshiftfaced wrote:
         | > but this is being powered by the unreleased o3 model
         | 
         | What makes you believe that?
        
           | _bin_ wrote:
           | they explicitly stated it in the launch
        
             | bitshiftfaced wrote:
             | The linked article says,
             | 
             | > Powered by a version of the upcoming OpenAI o3 model
             | that's optimized for web browsing and data analysis, it
             | leverages reasoning to search, interpret, and analyze
             | massive amounts of text, images, and PDFs on the internet,
             | pivoting as needed in reaction to information it
             | encounters.
             | 
             | If that's what you're referring to, then it doesn't seem
             | that "explicit" to me. For example, how do we know that it
             | doesn't use _less_ thinking than o3-mini? Google 's version
             | of deep research uses their "not cutting edge version" 1.5
             | model, after all. Are you referring to something else?
        
               | golol wrote:
               | o3-mini is not really "a version of the o3 model", it is
               | a different model (less parameters). So their language
               | strongly suggests, imo, that Deep Research is powered by
               | a model with the same number of parameters as o3.
        
       | chrismarlow9 wrote:
       | This smells like when Google released Gemini to have a product in
       | the space.
        
         | lysace wrote:
         | Eh, not really. Google failed to launch first out of internal
         | political dysfunction and then made a crash effort to launch
         | _something_ to counter the first ChatGPT release.
         | 
         | I _highly_ doubt that the concerns of internal political
         | commissars were holding up this particular openai release.
        
           | chrismarlow9 wrote:
           | That's some fancy words friend.
        
         | xnx wrote:
         | I agree that OpenAI is trying to stay relevant by announcing a
         | lot of have baked products with little to no availability.
         | 
         | > when Google released Gemini to have a product in the space.
         | 
         | Bard preceded Gemini.
        
       | febin wrote:
       | Is this "deep research" tool exploiting open knowledge creators,
       | using their work without compensation?
        
         | johnneville wrote:
         | would it be open knowledge if it required payment to access ?
        
         | scarab92 wrote:
         | Are you exploiting open knowledge creators, using their work
         | without compensation?
        
           | febin wrote:
           | The creators are aware that a human is using this, can we say
           | the same for AI, does it have their consent?
        
             | handfuloflight wrote:
             | Then consent is granted by transitive property because
             | these AI are yielded by humans.
        
               | hnisoss wrote:
               | Yea but guy paying closedai to get "insights" that
               | basically copy-pasted content from my blog is definitely
               | violating my blogs copyright, and in the end no coin
               | comes to me either. What about that?
        
               | handfuloflight wrote:
               | Could you provide an example where OpenAI outputting
               | verbatim quotes actually constitutes the copyright
               | violation? Because mechanically retrieving relevant
               | quotes seems analogous to grep/search - the copyright
               | status would depend on how downstream users transform and
               | use that content. Like how quoting your blog in a
               | technical analysis or critique is fair use, but wholesale
               | republishing isn't. This suggests the violation occurs at
               | usage time, not retrieval time.
        
         | tmnvdb wrote:
         | How is using public information "exploitation"? A human
         | researcher with Google would do the same.
        
           | hnisoss wrote:
           | So its fine for OpenAI to effectively sell your CC BY-NC
           | content to others?
        
         | protocolture wrote:
         | You exploited my eyes by making me read this comment. Wheres my
         | compensation.
        
         | hnisoss wrote:
         | Of course. It's a child's play for SamA et al.
        
         | febin wrote:
         | I see many are offended, but I am genuinely asking a question.
         | 
         | I want to understand does this mean it's ethical for anyone to
         | create a research AI tool that will go through arXiv and
         | related GitHub repo and use it to solve problems, implement
         | ideas like cursor.
        
         | rapjr9 wrote:
         | It is also an agent, so it is using you without compensation
         | for your work.
        
       | usaar333 wrote:
       | Overall impressive.
       | 
       | Though, the jump for Gaia relative to SOTA is relatively not that
       | high. Especially given that this is o3
        
       | ldjkfkdsjnv wrote:
       | Say whatever you want about openAI, they are shipping more than
       | any other company on the planet.
        
         | kortilla wrote:
         | What does that even mean? Treating each iterative model as a
         | new product is not any different than Google changing its
         | search or youtube recommendation algorithm.
         | 
         | Different pre-cooked prompts and filters don't really amount to
         | new products either, despite them being marketed as such. It's
         | like adobe treating each tool in photoshop as its own product.
        
           | tmnvdb wrote:
           | Have you even watched the video? This is a new capability and
           | not a trivial one.
        
         | viraptor wrote:
         | How do you even compare different companies? I'd say the
         | massive farms ship every year more than OpenAI ever did.
        
         | navigaid wrote:
         | Name one single open source model released by OpenAI since 2020
        
       | YmiYugy wrote:
       | If I understood the graphs correctly, it only achieves 20% pass
       | rate on their internal tests. So I have to wait 30min and pay a
       | lot of money just to sift through walls of most likely incorrect
       | text? Unless the possibility of hallucinations is negligible,
       | this is just way too much content to review at once. The process
       | probably needs to be a lot more iterative.
        
         | tmnvdb wrote:
         | Only if you are asking questions at the level of a cutting edge
         | benchmark
        
           | rvnx wrote:
           | This is one of the actual questions:
           | 
           | > In Greek mythology, who was Jason's maternal great-
           | grandfather?
           | 
           | https://www.google.com/search?q=In+Greek+mythology%2C+who+wa.
           | ..
        
             | tmnvdb wrote:
             | This is a hard question for language models since it
             | targets one of their known weaknesses.
        
               | layer8 wrote:
               | Users don't care about how hard something is for LLMs if
               | they receive incorrect output.
        
               | andyg_blog wrote:
               | Greek mythology? But seriously please elaborate for my
               | less educated self.
        
               | tmnvdb wrote:
               | LLMs often don't do well on tasks that require
               | composition into smaller subtasks. In this case there is
               | a chain of relations that depend on the previous result.
        
               | _bin_ wrote:
               | it tests syllogistic reasoning: Jason's mother was Tyro,
               | whose father was Poesidon, whose father was Kronos. it
               | also tests whether it "eagerly" rather than
               | comprehensively considers something: a maternal great-
               | grandfather could be the father of either one's maternal
               | grandmother or maternal grandfather. so the answer could
               | also be king Aeolus of the Etruscans.
               | 
               | ideally a model would be able to answer this accurately
               | and completely.
        
               | nimithryn wrote:
               | I think there are more possible answers? Jason's mother
               | differs depending on the author...
               | 
               | For example, Jason's mother was Philonis, daughter of
               | Mestra, daughter of Daedalion, son of Hesporos. So
               | Jason's maternal great-grandfather was Hesporos.
        
               | 11101010001100 wrote:
               | It's categorically more than a weakness.
        
             | elicksaur wrote:
             | In Greek mythology, Jason's maternal great-grandfather was
             | Einstein.
        
             | pama wrote:
             | No it is not an actual question on this exam. From the
             | paper: "To ensure question quality and integrity, we
             | enforce strict submission criteria. Questions should be
             | precise, unambiguous, solvable, and _non-searchable_ ,
             | ensuring models cannot rely on memorization or simple
             | retrieval methods. All submissions must be original work or
             | non-trivial syntheses of published information, though
             | contributions from unpublished research are acceptable.
             | Questions typically require graduate-level expertise or
             | test knowledge of highly specific topics (e.g., precise
             | historical details, trivia, local customs) and have
             | specific, unambiguous answers...". (Emphasis mine)
        
               | Neynt wrote:
               | It's example #7 on https://lastexam.ai/
        
               | pama wrote:
               | This is an example of the submitted questions. Because it
               | is possible to search it on the web, it is not an example
               | of the accepted questions.
        
               | freehorse wrote:
               | I am selling a bridge, it is a great bargain.
        
             | johnfn wrote:
             | Did you intentionally flip through all the questions to
             | find the one that seemed the easiest? If so, why? That's
             | question #7, and all other 7 questions in the sample set
             | seem ridiculously difficult to me.
        
         | brokensegue wrote:
         | 26.6% on humanity's last exam is actually impressive.
         | 
         | pass rate really only matters in context of the difficulty of
         | the tasks
        
         | spyckie2 wrote:
         | I mean you want it to grill your steak and eat it for you too?
         | 
         | I mean I too can complain that my iPhone doesn't automatically
         | screen out spammers and send my mom flowers on Mother's Day.
        
           | scarab92 wrote:
           | Why doesn't the iPhone screen spammers yet? Pixel has had
           | this feature for a decade.
        
             | senordevnyc wrote:
             | Pixel hasn't even been around for a decade.
        
               | scarab92 wrote:
               | The Pixel branding is 12 years old, and IIRC this feature
               | also existed in Nexus before that.
        
               | senordevnyc wrote:
               | Haha, are you referring to the Chromebook Pixel? How is
               | that relevant to stopping spam calls?
               | 
               | Pixel phone launched in 2016.
        
         | roenxi wrote:
         | Maybe. Not enough data to say. Say it does a days worth of work
         | in a query. It is sensible to use if it takes less than a day
         | to review ~5 days worth of work. I don't know if we're near
         | that threshold yet but conceptually this would work well for
         | actual research where the amount of preparation is large
         | compared to the amount of output written.
         | 
         | And eyeballing the benchmarks, it'll probably reach a >50% rate
         | per query by the end of the year. Seems to double every model
         | or two.
        
         | itkovian_ wrote:
         | Here's an example of the type of question it is acheiving 20%
         | on;
         | 
         | The set of natural transformations between two functors F,G
         | [?]:C-DF,G:C-D can be expressed as the end
         | Nat(F,G)[?][?]AHomD(F(A),G(A)). Nat(F,G)[?][?]A HomD
         | (F(A),G(A)).
         | 
         | Define set of natural cotransformations from FF to GG to be the
         | coend CoNat(F,G)[?][?]AHomD(F(A),G(A)). CoNat(F,G)[?][?]AHomD
         | (F(A),G(A)).
         | 
         | Let: - F=B[?](S4)*/F=B[?] (S4 )*/ be the under [?][?]-category
         | of the nerve of the delooping of the symmetric group S4S4 on 4
         | letters under the unique 00-simplex ** of B[?]S4B[?] S4 . -
         | G=B[?](S7)*/G=B[?] (S7 )*/ be the under [?][?]-category nerve
         | of the delooping of the symmetric group S7S7 on 7 letters under
         | the unique 00-simplex ** of B[?]S7B[?] S7 .
         | 
         | How many natural cotransformations are there between FF and GG?
        
           | Davidzheng wrote:
           | btw isn't this question at least really badly worded (and
           | maybe incorrect?) the definitions they give for F and G are
           | categories not functors... (and both categories are in fact
           | one object with contractible space of morphisms...)
        
           | perching_aix wrote:
           | It's very interesting to think about what kind of "mental
           | model" might it have, if it's capable of "understanding" all
           | this (to me) gibberish, but is then unable to actually work
           | the problem.
        
           | slaterbug wrote:
           | As someone who doesn't understand anything beyond the word
           | 'set' in that question, can anyone give an indication of how
           | hard of a problem that actually is (within that domain)?
           | 
           | Also I'm curious as to what percentage of the questions in
           | this benchmark are of this type / difficulty, vs the
           | seemingly much easier example of "In Greek mythology, who was
           | Jason's maternal great-grandfather?".
           | 
           | I'd imagine the latter is much easier for an LLM, and almost
           | trivial for any LLM with access to external sources (such as
           | deep research).
        
           | baal80spam wrote:
           | That's easy Dave: 42.
        
           | YmiYugy wrote:
           | Do we actually know whether it got this specific example
           | right? It got 20% on HLE, but I think a few questions are
           | quite a bit easier.
        
         | random_cynic wrote:
         | The difference is that it takes few minutes to an hour at most
         | so it can be run multiple times a day, using the results of
         | previous runs to further refine the search and reasoning
         | process to get better outcomes. Pretty much how any human
         | research works but much faster and with potentially vastly more
         | world-knowledge and reasoning capability than average humans.
         | And these capabilities will rapidly improve with further RL.
        
         | dyauspitr wrote:
         | On questions even specialists in that field can't answer
         | correctly.
        
         | throwaway123lol wrote:
         | Yeah it can be more iterative. Just use individual queries and
         | build on it yourself. This is all this is doing. It's a trick,
         | and OpenAI is a PR hype company at this stage.
        
       | thefourthchime wrote:
       | OpenAI has a deep bench. I bet they pushed this out to change the
       | narrative about deepseek
        
         | btown wrote:
         | Also named specifically to muddle the SEO for the term "deep."
         | Nothing that OpenAI does is unintentional.
        
           | bbor wrote:
           | I have never believed a conspiracy theory more instantly.
           | Deep Search vs. DeepSeek is way more than enough to confuse
           | the average layman! Especially when you're googling something
           | you heard about at work a few hours ago, or on Bloomberg TV
        
             | bonoboTP wrote:
             | You might as well say that DeepSeek wanted to cause
             | confusion with DeepMind. Deep isn't such a distinguishing
             | name, deep learning has been a buzzword since 2012.
        
               | viraptor wrote:
               | Deepmind is not a consumer product. Gemini is part of it
               | but nobody calls it deepmind.
        
               | bonoboTP wrote:
               | The point is, "deep" is an extremely generic word in the
               | AI space.
        
           | leonheld wrote:
           | Oh God, this is such an astute observation. I think it worked
           | so well on me that I didn't even think about the "deep"
           | portion initially. Goes to show how effective these things
           | are psychologically.
        
           | kevlened wrote:
           | It's more likely this is a response to Gemini Deep Research
           | released in December
           | 
           | https://blog.google/products/gemini/google-gemini-deep-
           | resea...
        
             | nicce wrote:
             | Two birds with one stone: timing for Deepseek and feature
             | for Gemini
        
             | petra wrote:
             | That Google product isnt that good, it can't really replace
             | research done by a person.
        
               | nicce wrote:
               | Just one tool in the toolbox. It helps to see if some
               | sources have been missed.
        
               | alvah wrote:
               | It absolutely can replace the research done by one
               | person, for my use case at least. It's also available on
               | their $20/month subscription, unlike OpenAI's $200/month.
        
               | sadeshmukh wrote:
               | Nobody was going to hire a researcher for a quick
               | question.
        
           | dougb5 wrote:
           | Does the naming scheme they've used for models so far suggest
           | that they care about SEO?
        
           | xnx wrote:
           | Google publicly announced a model named "Deep Research" on
           | December 11th: https://blog.google/products/gemini/google-
           | gemini-deep-resea...
        
       | adriand wrote:
       | Feels like only a matter of time before these crawlers are
       | blocked from large swathes of the internet. I understand that
       | they're already prohibited from Reddit and YouTube. If that
       | spreads, this approach might be in trouble.
        
         | drcode wrote:
         | I suppose there is an equilibrium, where sites that penalize
         | these types of crawlers will also get less traffic from people
         | reading ai citations, so for many sites the upsides of allowing
         | it will be greater than the downsides.
        
         | bbor wrote:
         | TBF OpenAI in particular bought access to Reddit. Otherwise
         | yeah this is my main confusion with all of these products,
         | Perplexity being the biggest -- how do you get around the
         | status-quo of refusing access to bots? Just to start off with,
         | there is no Google Search API, and they work hard to make sure
         | headless browsers can't access the normal service.
         | 
         | They do say "Currently, deep research can access the open
         | web...", so maybe "open" there implies something significant.
         | Like, "websites that have agreements with OpenAI and/or do not
         | enforce norobot policies".
        
           | wahnfrieden wrote:
           | Client-side browsers that crawl for users (and prompt for
           | logins or captcha as needed) won't be as easily blockable
        
         | scarab92 wrote:
         | I doubt those crawler rules will be honoured for long.
         | 
         | I wouldn't even be surprised if a law is passed requiring sites
         | to provide equal access to humans whether accessed directly or
         | via these models.
         | 
         | It's too important an innovation to stall, especially
         | considering the US's competitors (China) won't respect
         | robots.txt either.
        
         | cj wrote:
         | Anyone selling anything would want to remain crawlable if
         | people use this to research something that could lead to a
         | purchase.
        
           | reaperman wrote:
           | Not necessarily. Southwest airlines doesnt allow itself on
           | price comparison sites or Google Flights.
           | 
           | Amazon listings are blocked from google shopping and other
           | price comparison sites.
        
             | shlomo_z wrote:
             | Your point is completely valid, but... Southwest now has an
             | arrangement with Google Flights to allow their listings
             | there.
        
             | rsanek wrote:
             | they finally bowed mid last year
             | https://www.nerdwallet.com/article/travel/southwest-
             | google-f...
        
             | yencabulator wrote:
             | > Amazon listings are blocked from google shopping
             | 
             | I see Amazon results there all the time. 3 of the visible 8
             | sponsored results are Amazon, in the non-sponsored results
             | an Amazon listing is either first or second in every
             | category.
        
         | crazylogger wrote:
         | This is trivially bypassed by OpenAI asking the user to take
         | control of their computer (or a sandboxed browser within it,)
         | then for all intents and purposes it's the user themselves
         | accessing your site (with some productivity/accessibility aid
         | from OAI.)
        
         | optimalsolver wrote:
         | Big Tech Podcast listener?
        
         | felindev wrote:
         | While people might attempt that, it's going to be an arms race,
         | just like ads vs adblocks. There's already multiple crawlers
         | that present fake user-agent when their original one is
         | blocked. Temptation of more data is just to irresistible to
         | them
        
         | sumedh wrote:
         | > Feels like only a matter of time before these crawlers are
         | blocked from large swathes of the internet.
         | 
         | How would you know its a crawler?
        
       | reader9274 wrote:
       | I think we're all reaching AI fatigue. Fewer and fewer people
       | care anymore
        
         | rvnx wrote:
         | Especially this is not a breakthrough justifying a 340B USD
         | valuation, but rather the work that junior developers can do;
         | implement a loop of Bing Searches connected to an LLM.
        
           | tmnvdb wrote:
           | Peak HN comment
        
             | rvnx wrote:
             | Doesn't make it untrue.
             | 
             | Agents that can search the internet exist for a while now
             | and have been essentially solved and happily used in
             | platforms like Perplexity.
             | 
             | It's really "meh", very far from revolutionary.
             | 
             | Keep in mind this company is trying to convince everybody
             | they need 500B USD now (through the Stargate project).
        
               | spyckie2 wrote:
               | To go from partially automated to fully automated is
               | thousands of non trivial edge cases and unforeseen
               | decision points that must be tamed.
               | 
               | To say this is trivial is like saying the one shot ai
               | prompted twitter clone is the same thing as twitter.
               | 
               | Peak HN indeed.
        
               | CamperBob2 wrote:
               | Let us know when your Bing-bot scores over 20% on the HLE
               | benchmark.
        
               | rvnx wrote:
               | It's literally a browsing agent that searches the
               | internet and they know the questions in advance when
               | preparing the agent
               | 
               | Without internet: 10%
               | 
               | With internet: 23%
               | 
               | In addition:
               | 
               | > We found that the ground-truth answers for one dataset
               | were widely leaked online
               | 
               | in very small letters, and they blocked these URLs at
               | runtime but not training time.
               | 
               | It's not bad, but not revolutionary at all compared to
               | the leap that was GPT-2 from GPT-3, or GPT-4o to
               | DeepSeek-R1
        
               | CamperBob2 wrote:
               | If they "knew the questions in advance," why'd they need
               | Internet access at all? The ability to use the same data
               | sources humans would use is not the insult you seem to
               | think it is.
               | 
               | Again: the assertion was yours, so let us know the
               | results of your own work.
        
               | alvah wrote:
               | I haven't tried the OpenAI version yet, as I'm on their
               | peasant-level $20 plan, but the Google equivalent is way
               | superior to Perplexity (I use both extensively). The web
               | search Perplexity carries out is superficial compared to
               | the Google product; it misses a large percentage of what
               | Gemini Deep Research finds, and for a particular task in
               | my business this makes a huge difference.
        
         | bonoboTP wrote:
         | Sure if you're viewing this as some kind of spectator thing, or
         | entertainment, maybe it's less interesting. But it doesn't
         | really matter whether "people care". What matters is whether
         | it's useful and has impact. It's enough if the small number of
         | people use it for whom it is useful. It doesn't matter if the
         | average Joe on the street is excited by it.
         | 
         | Few people care or even know about various advances in various
         | specialized fields. It's enough if AI simply seeps into various
         | applications in boring and non-flashy ways for it to have
         | significant effects that will affect a wider range of people,
         | whether they get hyped by the news announcements or not. Jobs
         | etc.
         | 
         | An analogy: the Internet as such is not very exciting nowadays,
         | certainly not in the way it was exciting in the 90s with all
         | the news segments about surfing the information superhighway or
         | whatever. There was a lot of buzz around the web, but then it
         | got normalized. It didn't disappear, it just got taken for
         | granted. No average person got excited around HTML5 or IPv6. It
         | just chugs along in the background. AI will similarly simply
         | build into the fabric of how things get done. Sometimes visibly
         | to the average person, sometimes just behind the scenes.
        
           | InkCanon wrote:
           | Not sure if it's just me, but it looks like all SOTA
           | companies are doubling down to chase the new benchmark, which
           | beyond hype, doesn't seem to translate into real world uses.
           | Why don't these companies just plug it into a popular git
           | repo and say, hey our AI fixed these 100 issues! Or something
           | real? The only people who seem to be doing something real is
           | DeepMind.
        
         | khazhoux wrote:
         | Incorrect. We are not all reaching AI fatigue.
        
       | gwerbret wrote:
       | To anyone who's tried it: how does it handle captchas? I can't
       | imagine that OpenAI's IP addresses are anyone's favorites for
       | unfettered access to web properties these days.
        
         | layer8 wrote:
         | And is it smart enough to use archive.today for paywalled
         | articles. ;)
        
         | feznyng wrote:
         | You can buy residential proxies to pretend you're a regular
         | person IIRC, some of the browser automation companies do that
         | to bypass rate limiting, captchas, etc.
        
       | xt00 wrote:
       | "will find, analyze, and synthesize hundreds of online sources"
       | 
       | Synthesize? Seems like the wrong word -- I think they would want
       | to say something like, "analyze, and synthesize useful outputs
       | from hundreds of online sources"..
        
         | nicce wrote:
         | On the other hand, accurate if it is prone for hallucination...
        
         | tmnvdb wrote:
         | You can synthesize the parts to get the whole. Both uses are
         | correct AFAIK
        
         | pjot wrote:
         | From New Oxford dictionary:                 > combine (a number
         | of things) into a coherent whole: pupils should synthesize the
         | data they have gathered | Darwinian theory has been synthesized
         | with modern genetics.
        
       | tmnvdb wrote:
       | Eating popcorn while the scaling doubters scramble to move the
       | goalposts for the nth time.
        
         | elicksaur wrote:
         | Its number for one of the benchmark has:
         | 
         | **with browsing + python tools
         | 
         | Maybe we have different definitions of scaling?
        
           | tmnvdb wrote:
           | I would consider unsupervised tool usage an achievement
        
             | elicksaur wrote:
             | But it's not simply scaling. Who is moving the goalposts
             | exactly?
        
         | khazhoux wrote:
         | Trying to parse this. What are you saying?
        
       | ADeerAppeared wrote:
       | I'm sorry but what the fuck is this product pitch?
       | 
       | Anyone who's done any kind of substantial document research knows
       | that it's a _NIGHTMARE_ of chasing loose ends  & citogenesis.
       | 
       | Trusting an LLM to critically evaluate every source and to be
       | deeply suspect of any unproven claim is a ridiculous thing to do.
       | These are not hard reasoning systems, they are probabilistic
       | language models.
        
         | timsh wrote:
         | this is so precise. I guess we'll need a global version of
         | https://datacolada.org/ quite soon to not get hit by a bus in
         | every scientific field
        
         | panarky wrote:
         | _> they are probabilistic language models_
         | 
         | This is like arguing an Airbus cannot possibly fly because it
         | is 165 tonnes of aluminum, steel and plastic.
         | 
         | The proof is in the fact that it flies, not what it is
         | constructed from.
        
           | ADeerAppeared wrote:
           | > The proof is in the fact that it flies, not what it is
           | constructed from.
           | 
           | And LLMs do not.
           | 
           | > "But it looks like reasoning to me"
           | 
           | My condolences. You should go see a doctor about your
           | inability to count the number of 'R's in a word.
        
             | CamperBob2 wrote:
             | OK, what's your next move, now that letter-counting has
             | been solved by the current generation of frontier models?
             | 
             | CoT reasoning _is_ reasoning, whether you like it or not.
             | If you don 't understand that, it means the models are
             | already smarter than you.
        
             | panarky wrote:
             | "Even though that Airbus looks like it's flying it's really
             | not because my personal definition of 'flying' requires
             | feathers and flapping wings."
        
         | lukeschlather wrote:
         | o1 and o3 are definitely not your run of the mill LLM. I've had
         | o1 correct my logic, and it had correct math to back up why I
         | was wrong. I'm very skeptical, but I do think at some point AI
         | is going to be able to do this sort of thing.
        
       | cye131 wrote:
       | Does anyone actually have access to this? It says available for
       | pro users on the website today - I have pro via my employer but
       | see no "deep research" option in the message composer.
        
         | snewman wrote:
         | Two different people I know with pro subscriptions report not
         | having access yet.
        
         | greatpostman wrote:
         | Have pro, can't see it yet
        
         | fosterfriends wrote:
         | I have pro, in US, not seeing yet
        
           | _bin_ wrote:
           | what about a full refresh of the page or perhaps jump into
           | the dev tools and check "disable cache"
           | 
           | could also be aggressive caching from cloudflare. could be
           | they're just trying to announce more stuff to maintain cachet
           | and can't yet support all users forking over 200/month.
        
             | energy123 wrote:
             | I relogged, disabled cache and reloaded the page with
             | Ctrl+Shift+R but it doesn't show up.
        
               | nijaar wrote:
               | same here. pro in the US and still no access. i even
               | logged in using my phone and a different browser
        
         | fizx wrote:
         | same same
        
         | chachamatcha wrote:
         | also US based, have pro and still no access.
        
         | labanimalster wrote:
         | same here
        
         | nycdatasci wrote:
         | Pro user. No access like everyone else.
         | 
         | OpenAI is very much in an existential crisis and their poor
         | execution is not helping their cause. Operator or "deep
         | research" should be able to assume the role of a Pro user, run
         | a quick test, and reliably report on whether this is working
         | before the press release right?
        
           | kandesbunzler wrote:
           | How many times are you going to post this exact same comment
           | here? Are you a Chinese bot or something?
        
         | dimitri-vs wrote:
         | I have access as of ~3 hours ago. Using the Win desktop app
         | too, which is behind on some features (Operator, tasks). I open
         | up any of the models and it shows up as a `(Deep research)` tag
         | on the input field next to the web search option. Didn't clear
         | cache or anything.
        
       | hi_hi wrote:
       | This is terrifying. Even though they acknowledge the issues with
       | hallucinations/errors, that is going to be completely overlooked
       | by everyone using this, and then injecting the outputs into their
       | own powerpoints.
       | 
       | Management Consulting was bad enough before the ability to mass
       | produce these graphs and stats on a whim. At least there was some
       | understanding behind the scenes of where the numbers came from,
       | and sources would/could be provided.
       | 
       | The more powerful these tools become, the more prevelant this
       | effect of seepage will become.
        
         | tmnvdb wrote:
         | > At least there was some understanding behind the scenes of
         | where the numbers came from, and sources would/could be
         | provided.
         | 
         | Oh Sweet summer child.
        
           | opdahl wrote:
           | Hi tmnvdb, since you seem to love these super smart LLMs I
           | thought it would be fun to have openais o3-mini-high analyze
           | your recent comments in contrast to the Hacker News Comment
           | Guidelines. Here is the output it gave me, hope it helps you:
           | 
           | ------
           | 
           | Hey, I've noticed a few things in your style that are both
           | strengths and opportunities for improvement:
           | 
           |  _Strengths:_
           | 
           | - You clearly have deep knowledge and back up your points
           | with solid data and examples.
           | 
           | - Your confidence and detailed analysis make your arguments
           | compelling.
           | 
           |  _Opportunities:_
           | 
           | - At times, your tone can feel a bit combative, which might
           | shut down conversation.
           | 
           | - Focusing on critiquing ideas rather than questioning
           | someone's honesty can help keep the discussion constructive.
           | 
           | - A clearer structure in longer posts could make your points
           | even more accessible.
           | 
           | Overall, your passion and expertise shine through--tweaking
           | the tone a bit might help foster even more productive
           | debates.
           | 
           | ------
           | 
           |  _Just reply here if you want the full 500+ words analysis
           | that goes into more detail._
        
         | autoconfig wrote:
         | Either you care about being correct or you don't. If you don't
         | care then it doesn't matter whether you made it up or the AI
         | did. If you care then you'll fact check before publishing. I
         | don't see why this changes.
        
           | spaceywilly wrote:
           | I think a lot about how differentiating facts and quality
           | content is like differentiating signal from noise in
           | electronics. The signal to noise ratio on many online
           | platforms was already quite low. Tools like this will
           | absolutely add more noise, and arguably the nature of the
           | tools themselves make it harder to separate the noise.
           | 
           | I think this is a real problem for these AI tools. If you
           | can't separate the signal from the noise, it doesn't provide
           | any real value, like an out of range FM radio station.
        
             | WOTERMEON wrote:
             | Not only that: by publishing noise, you're lowering the
             | signal/noise ratio.
        
           | layer8 wrote:
           | People are much less scrupulous using LLM output than making
           | up stuff themselves, because then they can blame the LLM.
        
           | hi_hi wrote:
           | Because maybe you want to, but you have a boss breathing down
           | your neck and KPIs to meet and you haven't slept properly in
           | days and just need a win, so you get the AI to put together
           | some impressive looking graphs and stats that will look
           | impressive in that client showcase thats due in a few hours.
           | 
           | Things aren't quite so black and white in reality.
        
             | dauhak wrote:
             | I mean those same conditions already just lead the human to
             | cutting corners and making stuff up themselves. You're
             | describing the problem where bad incentives/conditions lead
             | to sloppy work, that happens with or without AI
             | 
             | Catching errors/validating work is obviously a different
             | process when they're coming from an AI vs a human, but I
             | don't see how it's fundamentally that different here. If
             | the outputs are heavily cited then that might go someway
             | into being able to more easily catch and correct slip-ups
        
               | hi_hi wrote:
               | Yep, I agree with this to some extent, but I think the
               | difference in the future is all that stress will be
               | bypassed and people will reach for the AI from the start.
               | 
               | Previously there was alot of stress/pressure which might
               | or might not have led to sloppy work (some consultants
               | are of a high quality). With this, there will be no
               | stress which will (always?) lead to sloppy work. Perhaps
               | there's an argument for the high quality consultants
               | using the tools to produce accurate and high quality
               | work. There will obviously be a sliding scale here. Time
               | will tell.
               | 
               | I'd wager the end result will be sloppy work, at scale
               | :-)
        
               | tikhonj wrote:
               | Making it easier and cheaper to cut corners and make
               | stuff up will result in more cut corners and more made up
               | stuff. That's not good.
               | 
               | Same problem I have with code models, honestly. We
               | already have way too much boilerplate and bad code;
               | machines to generate _more_ boilerplate and bad code aren
               | 't going to help.
        
               | mquander wrote:
               | The technology also makes it easier and cheaper to make
               | good things, so the direction of the outcome isn't
               | guaranteed.
        
           | sbarre wrote:
           | How _hard_ it is to produce credible-looking bullshit makes a
           | really big difference in these scenarios.
           | 
           | Consultants aren't the ones doing the fact-checking, that
           | falls to the client, who ironically tend to assume the
           | consultants did it.
        
           | azinman2 wrote:
           | When things are easy, you're going to take the easy path even
           | if it means quality goes down. It's about trade offs. If you
           | had to do it yourself, perhaps quality would have been higher
           | because you had no other choice.
           | 
           | Lots of kids don't want to do homework. That said, previously
           | many would because there wasn't another choice. But now they
           | can just ask ChatGPT for the answers they'll write that down
           | verbatim with zero learning taking place.
           | 
           | Caring isn't a binary thing or works in isolation.
        
             | simonw wrote:
             | "Lots of kids don't want to do homework"
             | 
             | Sure, but if you're a professional you have to care about
             | your reputation. Presenting hallucinated cases from ChatGPT
             | didn't go very well for that lawyer:
             | https://www.nytimes.com/2023/05/27/nyregion/avianca-
             | airline-...
        
               | PeterStuer wrote:
               | That's a lawyer in an adverserial situation. Business
               | consultants tell their clients what they _want_ to
               | believe, the facts be dammed.
        
               | rsanek wrote:
               | it sounds like ai doesn't really change that situation
        
               | asimpletune wrote:
               | But the point is it does if you count making it worse
               | changing the situation.
        
             | financypants wrote:
             | what about tests?
        
             | jstummbillig wrote:
             | I don't think it follows that taken an easier path _would_
             | mean quality goes down.
        
           | ADeerAppeared wrote:
           | > If you care then you'll fact check before publishing.
           | 
           | Doing a proper fact check is as much work as doing the entire
           | research by hand, and therefore, this system is useless to
           | anyone who cares about the result being correct.
           | 
           | > I don't see why this changes.
           | 
           | And because of the above this system should not exist.
        
           | RainyDayTmrw wrote:
           | It's possible that you care, but the person next to you
           | doesn't, and external pressures force you to keep up with the
           | person who's willing to shovel AI slop. Most of us don't have
           | a complete luxury of the moral high ground at our jobs.
        
             | doomroot wrote:
             | It looks like the moral high just came more in demand.
        
             | navigate8310 wrote:
             | It's the high reps fault then of not caring about quality.
             | Either you assimilate in that low quality lower management
             | using AI slop or change job.
        
           | n4r9 wrote:
           | It's a bit like saying "my kids are going to hit themselves
           | anyway, so it doesn't matter if I give them foam rods or
           | metal rods".
        
             | ctoth wrote:
             | Maybe this would make sense if you saw the whole world as
             | "kids" that you had to protect. As an adult who lives in an
             | adult world, I would like adults to have access to metal
             | tools and not just foam ones.
        
               | n4r9 wrote:
               | I guess I can replace "kid" with "toddler" and add
               | "unsupervised" at the end.
        
           | mlsu wrote:
           | If 20% of people don't care about being correct, the rest of
           | everyone can deal with that. If 80% of people don't care
           | about being correct, the rest of us will not be able to deal
           | with that.
           | 
           | Same thing as misinformation. A sufficient quantitative
           | difference becomes a qualitative difference at some point.
        
         | scarab92 wrote:
         | Think of it like a vaccine.
         | 
         | The majority of human written consultant reports are already
         | complete rubbish. Low accuracy, low signal-to-noise, generic
         | platitudes in a quantity-over-quality format.
         | 
         | LLMs are innoculating people to this kind of low information
         | value content.
         | 
         | People who produce LLM quality output, are now being accused of
         | using LLMs, and can no longer pretend to be adding value.
         | 
         | The result of this is going to be higher quality expectations
         | from consultants and a shaking out of people who produce word
         | vommit rather than accurate, insightful, contextually relevent
         | information.
        
           | layer8 wrote:
           | This has been downvoted, but I think there's actually a
           | chance it might become true (until AGI comes along at least).
        
           | DrSiemer wrote:
           | Exactly what will happen with art. The tolerance for low
           | quality output will decrease.
        
           | randcraw wrote:
           | I don't think so. Instead of SEO, I think we'll soon see
           | 'LLMO' dominating such uses, where LLM summaries are reshaped
           | by vendors and etailers to misrepresent facts in ways that
           | favor them over others.
           | 
           | I suspect this can be done simply by poisoning a query with
           | supplemental suggestions of sources to use in a RAG, many of
           | which don't even have to be publicly available but are made
           | accessible to the LLM (perhaps by submitting hidden URLs that
           | mislead the summary along with the query).
           | 
           | But even after such a practice is uncovered and roundly
           | maligned, that won't stop the infinite supply of net con men
           | from continuing to inject their poisons into the background
           | that drives deep research, so long as the LLM maker doesn't
           | _actively_ oppose this practice actively and publicly --
           | which none of them have been willing to do with any other LLM
           | operational details so far.
           | 
           | In fact, I predict that if a LLM summary like DR's does NOT
           | soon provide references to the sources of the facts it relies
           | on, in no time users will disregard such summaries to be yet
           | more uselessly unreliable pfaff from yet another net
           | disreputable -- as we do with search engine summaries now.
        
         | _bin_ wrote:
         | let's be real for a sec, i've done consulting and have a lot of
         | friends who still do. three times in four, your mckinsey report
         | isn't super well-founded in reality and involves a lot of
         | guesstimation.
        
         | n144q wrote:
         | I think that ship has sailed many years ago since Facebook
         | allowed false information to spread freely on their site (if
         | not earlier).
        
         | anthonyshort wrote:
         | Then the hallucinated research is published in an article which
         | is then cited by other AI research, continuing the push the
         | false information until it's hard to know where the lie
         | started.
        
       | VerdisQuo5678 wrote:
       | The accuracy of this tool does not matter. This is exclusively
       | designed for box ticking "reports" that nobody reads and a
       | produced for the sake of itself.
        
         | arbywhy wrote:
         | 99% of corpo upper management slide deck work. ai only makes
         | more of this useless pencil-neck board of directors slop.
        
           | reaperman wrote:
           | "Pencil-neck" is a strange insult to use here. How are
           | software developers, or hardware design engineers, or finance
           | workers any less "pencil-neck" than "board of directors"?
        
         | tomrod wrote:
         | The new term for this is "AI Loopidity", highlighting the
         | unintelligent ouroboros nature of one side using AI to generate
         | content and then another side to consume content.
        
           | sockaddr wrote:
           | Similar to "Bullshit jobs"
           | 
           | All the AI commercials are designed to appeal to people that
           | don't produce any actual value but haven't been detected by
           | the system yet.
           | 
           | Need to send email to boss? Press magic button! Job well
           | done, idiot.
           | 
           | Someone send you big scary email? Press magic button! Good
           | job dummy!
           | 
           | Someone wants to go eat some Italian with you, push magic
           | button for totally not-ad result. Enjoy your Olive Garden,
           | moron.
        
             | rsanek wrote:
             | i think the apple ads are the poster child here. hopefully
             | we can see more inventive ones than just serving lazy
             | people.
        
       | ldjkfkdsjnv wrote:
       | So much cynicism and hate in these comments, especially as we are
       | likely witnessing AGI come to life. Its still early, but it might
       | be coming. Where is the excitement? This is an interesting time
       | to be alive.
       | 
       | HN has a huge cultural problem that makes this website almost
       | irrelevant. All the interesting takes have moved to X/twitter
        
         | bonoboTP wrote:
         | HN is and has always been quite negative/pessimistic/cynical in
         | general. That Dropbox comment was quite a long time ago
         | already.
        
         | roenxi wrote:
         | We're looking at trends that may well obliterate the economic
         | value of a well trained human mind sitting behind a keyboard
         | all day. That is a bit of a threat to most people on HN if the
         | trending continues at the current rate and direction.
        
         | rvz wrote:
         | > "So much cynicism and hate in these comments, especially as
         | we are likely witnessing AGI come to life. Its still early, but
         | it might be coming. Where is the excitement? This is an
         | interesting time to be alive."
         | 
         | Maybe you can define what "AGI" really means and what the end-
         | game and the economic implications are when 'AGI" is some-what
         | achieved? OpenAI somehow believes that they haven't achieved
         | "AGI" yet, which they continue to do this on purpose for
         | obvious reasons.
         | 
         | The first hint I will give you is that it certainly won't be a
         | utopia.
        
         | layer8 wrote:
         | "May you live in interesting times" is usually taken as a
         | curse. ;)
         | 
         | More seriously, it's unclear why one should be excited by the
         | prospect of AGI, especially when instrumentalized by
         | corporations and authoritarian governments.
        
         | rpcope1 wrote:
         | > especially as we are likely witnessing AGI come to life
         | 
         | Man, I've got a great deal on some oceanfront property in
         | Wyoming for you.
        
         | qgin wrote:
         | Never underestimate HN's capacity to be cynical about literally
         | everything
        
         | dutchbookmaker wrote:
         | I would be more excited if it wasn't $200 a month to try.
         | 
         | I don't feel like OpenAI does a good job of getting me excited
         | either.
         | 
         | Find the perfect snowboard? How can that idea get pitched and
         | make the final cut for a $200 a month service? The NFL kicker
         | example is also completely ridiculous.
         | 
         | The business and UX example seems interesting. Would love to
         | see more.
        
         | crvdgc wrote:
         | AGI aside, sometimes HN critics/cynicism indeed points out the
         | exact reason why something wouldn't work and is vindicated
         | after the fact, e.g. Apple Vision Pro. I guess it's just hard
         | to predict the future and for me, it's interesting to listen to
         | even pure contrarians.
        
       | wilg wrote:
       | I think this looks cool. Apparently unlike everyone else on this
       | website?
        
         | tmnvdb wrote:
         | HN is full of people who want to feel smart by complaining.
        
       | therealmarv wrote:
       | I don't know. OpenAI is so bad in naming... the average person on
       | the street will confuse Deepseek with Deep Research. Also not to
       | forget o1, o3 ... 4o
        
         | szvsw wrote:
         | > the average person on the street will confuse Deepseek with
         | Deep Research.
         | 
         | That's probably a feature not a bug (from OpenAI's
         | perspective...).
        
         | hipadev23 wrote:
         | Yes.
        
         | tmnvdb wrote:
         | You're not wrong but it feels like bikeshedding at this point.
        
       | Havoc wrote:
       | The descriptions of the product sounded substantially more
       | impressive than the actual samples tbh.
       | 
       | Still I think there is a big market for this sort of ,,go away
       | for 30 mins and figure this out" style agent
        
         | TechDebtDevin wrote:
         | This is 5-10 years out. What OpenAI is displaying here I've
         | been able to do with relatively little code, a bit of scraping
         | and far less capable models for a year. I really don't see what
         | is novel or useful here.
        
           | sumedh wrote:
           | Probably the accuracy.
        
       | layer8 wrote:
       | From the demo: "Use bullets and tables where necessary for
       | clarity." It's weird that it would be necessary to specify that.
       | I suppose they want to showcase that you can influence the output
       | style, but it's strange that you'd have to explicitly specify the
       | use of something that is "necessary for clarity". It comes across
       | as either a flaw in the default execution, or as a merely
       | performative incantation.
        
       | prng2021 wrote:
       | "Deep research was trained using end-to-end reinforcement
       | learning"
       | 
       | Does this mean they skipped supervised fine tuning like DeepSeek
       | did with R1?
        
         | OutOfHere wrote:
         | No, it just suggests that RL was used over a base SFT model,
         | and moreover that RL here was tuned to this research task.
         | Personally I don't think that RL is strictly necessary for this
         | task at all, but perhaps it helps.
        
       | jasonjmcghee wrote:
       | Surprised more comments aren't mentioning deepseek has this
       | feature (for free) already. Assuming this is why OpenAI scrambled
       | to release it.
       | 
       | The examples they have on the page work well on chat.deepseek.com
       | with r1 and search options both enabled.
       | 
       | Do I blindly trust the accuracy of either though? Absolutely not.
       | I'm pretty concerned about these models falling into gaming SEO
       | and finding inaccurate facts and presenting them as fact. (How
       | easy is it to fool / prompt inject these models?)
       | 
       | But has utility if held right.
        
         | nicce wrote:
         | I wish Kagi would work with similar performance. Their lenses
         | feature is perfect for this and they already filter out most of
         | the SEO spam based on trackers and other typical red flags.
        
         | starchild3001 wrote:
         | Not really accurate. The "Search" functionality you're
         | describing in DeepSeek is comparable to OpenAI's existing
         | "Search GPT." OpenAI's recent announcement refers to a more
         | advanced capability, similar to Gemini's existing "deep
         | research" feature. DeepSeek's current offerings are
         | significantly more limited in scope.
        
           | jasonjmcghee wrote:
           | Doesn't seem like access is available to try "deep research"
           | yet on OpenAI, so I can only speak to what I tried, which was
           | their examples on the blog post (using DeepSeek w/ R1 +
           | Search) and results were pretty similar.
           | 
           | AFAIK OpenAI's current offering uses 4o, and it does a web
           | search and then pipes it into 4o. I'm guessing adding CoT +
           | other R1/o3 like stuff is one of the key effective
           | differences. But time will tell how different it is. Maybe
           | it's a dramatic improvement.
        
           | WiSaGaN wrote:
           | SearchGPT is bad because its underlying model is not a
           | reasoning one. Deepseek one mentioned above is closer to deep
           | research than searchgpt.
        
           | TechDebtDevin wrote:
           | Are you unaware that there is a "Deepthink (R1)" button right
           | next to the "Search" button on DeepSeek's Chat app. Its been
           | there for some time, even before all the hype regarding R1.
        
             | starchild3001 wrote:
             | I'm well aware of that. That is not what openai calls "deep
             | research".
        
       | pjs_ wrote:
       | McKinsey mode
        
         | sockaddr wrote:
         | Heh
        
         | TechDebtDevin wrote:
         | More like high school intern mode.
        
         | rsanek wrote:
         | don't be hyperbolic. deep research would need to help cause an
         | opioid crisis for to get to that level.
        
       | spyckie2 wrote:
       | Why is HN not creating policy against moral prigotry? There is no
       | useful discussion here anymore.
       | 
       | Seriously begging the mods to take a closer look, or at least PG
       | to not abandon his curated internet space.
        
         | esafak wrote:
         | priggery or bigotry? And what are you referring to?
        
       | esafak wrote:
       | Is there a benchmark we can compare this against You.com's
       | research mode? It looks like R1 forced them to release o3
       | prematurely and give it Internet access. And they didn't want to
       | say they released o3 so they called it 'Deep Research'.
        
       | spyckie2 wrote:
       | Is this ability really a prerequisite to AGI and ASI?
       | 
       | Reasoning, problem solving, research validation - at the
       | fundamental outset it is all refinement thinking.
       | 
       | Research is one of those areas where I remain skeptical it is
       | that important because the only valid proof is in the execution
       | outcome, not the compiled answer.
       | 
       | For instance you can research all you want about the best vacuum
       | on the internet but until you try it out yourself you are going
       | to be caught in between marketing, fake reviews, influencers,
       | etc. maybe the science fields are shielded from this (by being
       | boring) but imagine medical pharmas realizing that they can get
       | whatever paper to say whatever by flooding the internet with
       | their curated blog articles containing advanced medical "research
       | findings". At some point you cannot trust the internet at all and
       | I imagine that might be soon.
       | 
       | I worry especially with the rapidly changing landscape of the
       | amount of generated text in the internet that research will lose
       | a lot of value due to massive amounts of information garbage.
       | 
       | It will be a thing we used to do when the internet was still
       | "real".
        
         | observationist wrote:
         | It's a direction in a vast landscape, not a feature of itself -
         | being better at different tasks, like search generally, and
         | research in conjunction with reasoning, gets the model closer
         | to AGI. An AGI will be able to do these tasks - so the point of
         | the research is to have more Venn diagrams of capabilities like
         | these to help narrow down the view on things that might
         | actually be fundamental mechanisms involved in AGI.
         | 
         | Moravec detailed the idea of a landscape of human capabilities
         | slowly being submerged by AI capabilities, and the point at
         | which AI can do anything a human can, in practice or in
         | principle, we'll know for certain we've reached truly general
         | AI. This idea includes things like feeling pain and pleasure,
         | planning, complex social, oral, and ethical dynamics, and
         | anything else you can possibly think of as relevant to human
         | intelligence. Deep Research is just another island being slowly
         | submerged by the relentless and relentlessly accelerating
         | flood.
        
           | numba888 wrote:
           | > hings like feeling pain and pleasure
           | 
           | can machine feel? without that there is no AGI according to
           | definition above.
           | 
           | and the second question: are animals "GI"? they don't have
           | language and don't solve math problems, never heard of np-
           | complete.
        
             | xwolfi wrote:
             | Are we not machines anyway ? Ofc a machine can feel, just
             | need to have priorities that are aligned to itself, and use
             | strong feedback when that self is either in danger or on
             | the right path to preservation...
             | 
             | Feelings are nothing very special you know...
        
         | simonw wrote:
         | > _Is this ability really a prerequisite to AGI and ASI?_
         | 
         | That depends entirely on how you choose to define "AGI".
        
         | BeetleB wrote:
         | > For instance you can research all you want about the best
         | vacuum on the internet but until you try it out yourself you
         | are going to be caught in between marketing, fake reviews,
         | influencers, etc.
         | 
         | So you wouldn't use this tool for those types of use cases.
         | 
         | But still, a valid point. I recall I once wanted to compare
         | Hydroflask, Klean Kanteen and Thermos to see how they perform
         | for hot/cold drinks. I was looking specifically for
         | articles/posts where people had performed _actual
         | measurements_. But those were very hard to find, with almost
         | all Google hits being generic comparisons with no hard data.
         | That didn 't stop them from ranking ("Hydroflask is better for
         | warm drinks!")
         | 
         | Would I be able to get this to ignore all of those and use
         | _only_ ones where actual experiments were performed. And
         | moreover, filter out duplicates (e.g. one guy does an
         | experiment, and several other bloggers link to his post and
         | repeat his findings in their own posts - it 's one experiment
         | but with many search results).
        
       | ejang0 wrote:
       | Can anyone confirm if this is available in Canada and other
       | countries? This site says "We are still working on bringing
       | access to users in the United Kingdom, Switzerland, and the
       | European Economic Area." But I'm not sure about other countries.
       | I don't have Pro currently, only Plus.
        
         | carbocation wrote:
         | I don't even see it in the US right now.
        
           | carbocation wrote:
           | (Update: it's visible for me now.)
        
       | 6gvONxR4sf7o wrote:
       | There are some people in the blogosphere who are known experts in
       | their niche or even niche-famous because they write popular
       | useful stuff. And there are a ton more people who write useful
       | stuff because they want that 'exposure.' At least, they do in the
       | very broadest sense of writing it for another human to read it. I
       | wonder if these people will keep writing when their readership is
       | all bots. Dead internet here we come.
        
         | seanmcdirmid wrote:
         | I'm all for writing just for the bots, if I can figure it out.
         | A lot of academic papers aren't really read anyways, just
         | briefly glanced at so they can be cited together, large
         | publications like journal pubs or dissertations even less so.
         | But the ability to add to a world of knowledge that is very
         | easy to access by people who want to use it...that is very
         | appealing to me as an author. No more trudging through a bunch
         | of papers with titles that might be relevant to what I want to
         | know about...and no more trudging through my papers, I'm OK
         | with that.
        
         | lmm wrote:
         | Of course they will. Loads of people go around taking hundreds
         | of photos with the biggest camera they can afford even though
         | no-one else will ever willingly look at them.
        
       | Bjorkbat wrote:
       | Actually sounds pretty cool, but the graph on expert level tasks
       | is confusing my expectations. Saying it has a pass rate of less
       | than 20% sounds a lot like saying this thing is wrong most of the
       | time.
       | 
       | Granted, these strike me as difficult tasks and I'd likely ask it
       | to do far simpler things, but I'm not really sure what to expect
       | from looking at these graphs.
       | 
       | Ah, but the fact that it bothers to cite its sources is a huge
       | plus. Between that and its search abilities it sounds valuable to
       | me
        
         | random_cynic wrote:
         | I think that's mostly because of the access to information it
         | has. Much of the highly useful information is not on the public
         | internet or shows up on search engines, only domain experts
         | know about them. Also, the websites may be paywalled or gated
         | by login. So a better comparison would be if the models had the
         | same level of access as an expert.
        
       | EcommerceFlow wrote:
       | Can't even get Sunday nights off trying to keep up fml.
        
       | jaco6 wrote:
       | I see lots of warranted skepticism about the capabilities of this
       | tool, but the reality is that this is an incremental step toward
       | full automation of white collar labor. No, it will not make all
       | analysts jobless overnight. But it may reduce hiring of said
       | people by 5 or 10 percent. And as people get better at using the
       | tool and the tool itself gets better, those numbers will grow.
       | Remember that it took decades for the giant pool of typing
       | secretaries in Mad Men to disappear, but they did disappear. Gone
       | forever. Interestingly, anger about the diminishment of
       | secretarial male white collar work in Germany due to the spread
       | of the typewriter a few decades earlier was one of the drivers of
       | the Nazi Party's popularity (see Evans, the Rise of the Third
       | Reich).
       | 
       | AI's triumph in the white collar workplace will be gradual, not
       | instantaneous. And it will be grimly quiet, because no one likes
       | white collar workers the way they like blue collar workers, for
       | some odd reason, and there's no tradition of solidarity among
       | white collar workers. Everyone will just look up one day and find
       | that the local Big Corp headquarters is...empty.
        
       | michaelgiba wrote:
       | Gemini has had this for a month or two, also named "Deep
       | Research" https://blog.google/products/gemini/google-gemini-deep-
       | resea...
       | 
       | Meta question: what's with all of the naming overlap in the AI
       | world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI)
       | all come to mind
        
         | james_promoted wrote:
         | I've always thought the Triton situation was intentional since
         | the name isn't generic and because the companies are stepping
         | on each others toes here (Nvidia's Triton simplifying owning
         | your inference; OpenAI's Triton eroding the need for
         | familiarity with CUDA). I couldn't figure out who publicly used
         | the name first though.
        
         | stonogo wrote:
         | It's a sort of unofficial trade association where they coalesce
         | on specific redefinitions of terms to meet their sales and PR
         | efforts. First they came for "intelligence," then "open
         | source," then "reason," and it will continue. Any word which
         | the PR wants but they can't achieve gets redefined -- "grok" is
         | a perfect example, since in the original sci-fi book it meant
         | "total understanding." The mythological Triton ruled the deeps,
         | so the "deep learning" sales copy immediately co-opted it.
        
           | albert_e wrote:
           | Also "accuracy" as a measure of model's performance used to
           | mean something objective in the traditional ML world.
           | 
           | Now with LLMs it is what human evaluators feel about the LLM
           | output?
        
             | yorwba wrote:
             | Traditional ML is no stranger to measuring accuracy in
             | terms of agreement with human evaluators.
        
               | albert_e wrote:
               | A customer churn model or revenue forecast did have hard
               | objective data (ground truth) to compare against - isn't
               | it?
        
               | ptrrrrrrppr wrote:
               | you seem to think all classical ML models were
               | supervised, which isn't true. and we have metrics for
               | unsupervised approaches as well
        
         | samplatt wrote:
         | >Meta question
         | 
         | I think you have to prefix the query with "@Meta AI", hope this
         | helps
        
         | toomim wrote:
         | John Stewart had something to say about this:
         | https://youtu.be/Byg8VZdKK88?si=pX1WbtRwZCBGpwHS&t=141
        
           | justaj wrote:
           | Without the tracking bits: https://youtu.be/Byg8VZdKK88#t=141
        
         | shihab wrote:
         | From the creator of Triton (OpenAI)-
         | 
         | "PS: The name Triton was coined in mid-2019 when I released my
         | PhD paper on the subject. I chose not to rename the project
         | when the "TensorRT Inference Server" was rebranded as "Triton
         | Inference Server" a year later since it's the only thing that
         | ties my helpful PhD advisors to the project."
        
         | chabes wrote:
         | > what's with all of the naming overlap in the AI world? Triton
         | (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to
         | mind
         | 
         | They seem to be ok with outsourcing any and all creativity to a
         | language model, so it's not surprising that they can't come up
         | with unique names themselves.
        
           | dncbfwa wrote:
           | lol
        
         | kavalerov wrote:
         | I am afraid Gemini's version is not really very "deep" - it
         | surfaces a lot of information, but on a quite superficial
         | level. OAIs version seems to make that one step forward to
         | proper depth.
         | 
         | We found in our experience it is pretty hard to force LLM to do
         | something in proper depth, and OAI's deep research definitely
         | feels like one of the first examples from big labs on how this
         | can be done. What we typically see is that it is not even the
         | "agent" part that is hard to do, but how to force model to not
         | "forget" to go deep...
        
         | svara wrote:
         | > Gemini has had this for a month or two,
         | 
         | Would have loved to try it when they released it, but I'm
         | apparently in the wrong country. I think it's not available
         | outside the US (?). OpenAI and DeepSeek have no such issues.
         | It's a bummer really, I'm happy paying for this but they don't
         | want me to.
        
       | pazimzadeh wrote:
       | > In Nature journal's Scientific Reports conference proceedings
       | from 2012, in the article that did not mention plasmons or
       | plasmonics, what nano-compound is studied?
       | 
       | Aren't there more than one articles that did not mention plasmons
       | or plasmonics in Scientific Reports in 2012?
       | 
       | Also, did they pay for access to all journal contents? that would
       | be useful
        
         | nicce wrote:
         | Maybe that is the only one with open access
        
       | getnormality wrote:
       | The demo on global e-commerce trends seems less useful than a
       | Google search, where the AI answer will at least give you links
       | to the claimed information.
        
       | jmount wrote:
       | I had no idea there was a market for "Compile a research report
       | on how the retail industry has changed in the last 3 years. Use
       | bullets and tables where necessary for clarity." I imagine
       | reading such a result is pure torture.
        
       | airstrike wrote:
       | "Deep research" is now somehow synonymous to searching online for
       | stats and pulling stuff from Statista? And when I want to make
       | changes to that report, do I have to tweak my prompt and get an
       | entirely different document?
       | 
       | Not sure if I'm too tired and can't see it but the lack of
       | images/examples of the resulting report in this announcement
       | doesn't inspire a lot of confidence just yet.
        
       | highfrequency wrote:
       | Can it compile and run (non-Python) code as part of its tool use?
       | Compile-run steps always seemed like they would be a huge value
       | add during reasoning loops - it feels very silly to get output
       | from ChatGPT, try to run it in terminal, get an error and paste
       | the error to have ChatGPT immediately fix it. Surely it should be
       | able to run code during the reasoning loop itself?
        
         | simonw wrote:
         | It sounds like it can run Python, which means it has access to
         | Code Interpreter, which means it can run various other
         | languages as well if you can convince it to do so.
         | 
         | I've used Code Interpreter to compile and run C code -
         | https://simonwillison.net/2024/Mar/23/building-c-extensions-...
         | - and I've managed to get it to run JavaScript (by uploading a
         | Deno binary) and even Lua and PHP in the past as well:
         | https://til.simonwillison.net/llms/code-interpreter-expansio...
        
       | RayVR wrote:
       | Each release from openAI gives me less hope for them and this
       | whole AI boom. They should be leading the charge of highlighting
       | how the current generation of LLMs fail, not churning out half-
       | baked overhyped products.
       | 
       | Yes, they can do some cool tricks, and tool calling is fun. No
       | one should trust the output of these models, though. The
       | hallucinations are bad, and my experience with the "reasoning"
       | models is that as soon as they fuck up (they always do) they go
       | off the rails worse than the base LLMs.
        
       | nycdatasci wrote:
       | Pro user. No access like everyone else.
       | 
       | OpenAI is very much in an existential crisis and their poor
       | execution is not helping their cause. Operator or "deep research"
       | should be able to assume the role of a Pro user, run a quick
       | test, and reliably report on whether this is working before the
       | press release right?
        
         | samplatt wrote:
         | That's the third time in this thread you've stated "OpenAI is
         | in an existential crisis". It looks very suspicious.
        
         | _bin_ wrote:
         | man you work for high flyer or something? i know that's not
         | really a fair question but oai still seems to lead the pack. i
         | know it's a hype-y area but responding to one (1) model that's
         | comparable to o4 but cheaper with "guys it's so over for
         | openai" is excessive.
        
       | resters wrote:
       | Still not seeing access on my account.
        
         | _bin_ wrote:
         | they're not giving it to us lowly $20/month users yet :( gotta
         | take out a second mortgage and throw them 200/month if you want
         | it now
        
           | resters wrote:
           | I have the $200/month version. Deep Research arrived this
           | morning.
           | 
           | So far I tried it on one problem and it seems limited by the
           | "front end" being 4o-mini. It ignored most of my initial
           | prompt and also ignored the previous research it asked for
           | which I provided. The final output was high quality and
           | definitely was enriched by the web searching it did, but it
           | left out a crucially important dimension of the problem
           | because it was unable to ingest the background info I
           | provided adequately.
           | 
           | I'd like to see a version of it where the front end model is
           | o1-pro
        
       | lolpanda wrote:
       | "synthesize large amounts of online information" does it heavily
       | depend on the search engine performance and relevance of the
       | search results? I don't see any mention of Google or Bing. Is
       | this using their internal search engine then?
        
       | RandomWorker wrote:
       | I'm a researcher and honestly not worried. 1. Developing the
       | right question has always been the largest barrier to great
       | research. Not sure OpenAI can develop the right question without
       | the Human experience. The second biggest part of my role is
       | influencing people that my questions are the right questions.
       | Which is made easier when you have a thorough understanding of
       | the first. That being said, I'm sure there will be many people
       | here that will tell me that algorithms already influence people,
       | and ai can think through much of any issues there are.
       | 
       | I do use these systems from time to time, but it just never
       | renders any specific information that would make it great
       | research.
        
         | RayVR wrote:
         | 100% agree.
         | 
         | These systems serve best at augmenting information discovery.
         | When I'm tackling a new area or looking for the right
         | terminology, these models provide a quick shortcut because they
         | have good probabilistic "understanding" of my naive, jargon-
         | free description. This allows me to pull in all of the jargon
         | for the area of research I'm interested in, and move on to
         | actually useful resources, whether that be journal articles,
         | textbooks, or - rarely - online posts/blogs/videos.
         | 
         | the current "meta" is probably something like Elicit +
         | notebookLM + Claude for accelerating understanding of complex
         | topics and extracting useful parts. But, again, each step
         | requires that I am closely involved, from selecting the
         | "correct" papers, to carefully aggregating and grooming the
         | information pulled in from notebookLM, to judging the the
         | usefulness of Claude's attempts to extract what I have asked
         | for
        
         | GeoAtreides wrote:
         | > Developing the right question has always been the largest
         | barrier to great research.
         | 
         | I thought funding was the biggest barrier to great research
        
       | elashri wrote:
       | It is actually interesting for people working in academia. I
       | would like to test it but no way I can afford $200/m right now.
       | 
       | Can someone test it with this prompt.
       | 
       | "As a research assistant with comprehensive knowledge of particle
       | physics, please provide a detailed analysis of next-generation
       | particle collider projects currently under consideration by the
       | international physics community.
       | 
       | The analysis should encompass the major proposed projects,
       | including the Future Circular Collider (FCC) at CERN,
       | International Linear Collider (ILC), Compact Linear Collider
       | (CLIC), various Muon Collider proposals, and any other
       | significant projects as of 2024.
       | 
       | For each proposal, examine the planned energy ranges and
       | collision types, estimated timeline for construction and
       | operation, technical advantages and challenges, approximate
       | costs, and key physics goals. Include information about current
       | technical design reports, feasibility studies, and the level of
       | international support and collaboration.
       | 
       | Present a thorough comparative analysis that addresses technical
       | feasibility, cost-benefit considerations, scientific potential
       | for new physics discoveries, timeline to first data collection,
       | infrastructure requirements, and environmental impact. The
       | projects should be compared in terms of their relative strengths,
       | weaknesses, and potential contributions to advancing our
       | understanding of fundamental physics.
       | 
       | Please format the response as a structured technical summary
       | suitable for presentation at a topical meeting of particle
       | physicists. Where appropriate, incorporate relevant figures and
       | tables to facilitate clear comparisons between proposals. Base
       | your analysis on information from peer-reviewed sources and
       | official design reports, focusing on the most current available
       | data and design specifications.
       | 
       | Consider the long-term implications of each proposal, including
       | potential upgrade paths, flexibility for future modifications,
       | and integration with existing research infrastructure."
        
         | sagarpatil wrote:
         | I'll do it. AFL right now.
        
         | fosterfriends wrote:
         | Ran it for you using o3-high! Here's a link to the results:
         | https://chatgpt.com/share/67a0b227-8ee4-800f-a8ed-882e7bab97...
         | 
         | Hopefully helpful, happy to test others for you :)
         | 
         | -- Raw output --
         | 
         | Understood. I will conduct a detailed technical analysis of
         | next-generation particle collider projects, focusing on the
         | Future Circular Collider (FCC), International Linear Collider
         | (ILC), Compact Linear Collider (CLIC), Muon Collider proposals,
         | and other significant initiatives.
         | 
         | The report will include: - Planned energy ranges and collision
         | types - Estimated construction and operation timelines -
         | Technical advantages and challenges - Approximate costs - Key
         | physics goals - Review of technical design reports and
         | feasibility studies - International collaboration and support -
         | Comparative analysis on feasibility, cost-benefit, scientific
         | potential, infrastructure needs, and environmental impact -
         | Long-term considerations, including upgrade paths and
         | integration with existing facilities
         | 
         | The analysis will be formatted as a structured technical
         | summary, suitable for presentation at a topical meeting of
         | particle physicists, with comparative tables and figures where
         | relevant.
         | 
         | I will update you when the report is ready.
         | 
         | # Next-Generation Particle Collider Projects: A Comparative
         | Technical Analysis
         | 
         | ## Introduction High-energy physics is charting plans for new
         | colliders to extend beyond the LHC and its High-Luminosity
         | upgrade (HL-LHC). The discovery of the 125 GeV Higgs boson has
         | motivated proposals for a dedicated "Higgs factory" lepton
         | collider to study its properties with precision, as well as
         | future energy-frontier machines to search for new physics.
         | Several next-generation collider initiatives are under
         | consideration internationally as of 2024. These include CERN's
         | proposed *Future Circular Collider (FCC)*, Japan's
         | *International Linear Collider (ILC)*, CERN's *Compact Linear
         | Collider (CLIC)*, various designs for a *Muon Collider*,
         | China's *Circular Electron-Positron Collider (CEPC)* and its
         | successor *Super Proton-Proton Collider (SppC)*, among others.
         | Each proposal differs in collision type (electron-positron,
         | proton-proton, muon-muon, etc.), energy scale, technology,
         | timeline, cost, and physics focus. This summary reviews each
         | project's key parameters - *planned energy ranges, collision
         | types, timeline, technical advantages/challenges, cost, and
         | physics goals* - based on technical design reports and
         | feasibility studies. A comparative analysis then contrasts
         | their *technical feasibility, cost-benefit, scientific
         | potential for discoveries, timeline to first data,
         | infrastructure needs, and environmental impact*, highlighting
         | the relative strengths and weaknesses of each approach. We also
         | discuss long-term implications such as upgrade paths,
         | flexibility for future modifications, and integration with
         | existing infrastructure.
         | 
         |  _(Citations refer to official reports and peer-reviewed
         | sources using the format [(source+lines)] .)_
         | 
         | ## Future Circular Collider (FCC) - CERN - *Type and Energy:*
         | The FCC is a *proposed 100 km circular collider* at CERN that
         | would be realized in stages. The first stage, *FCC-ee*, is an
         | electron-positron ($e^+e^-$) collider with center-of-mass
         | energy tunable from ~90 GeV up to 350-365 GeV, covering the Z
         | boson pole, WW threshold, Higgs production (240 GeV), and top-
         | quark pair threshold (~350 GeV). A second stage, *FCC-hh*,
         | would use the same tunnel for a proton-proton collider at up to
         | *100 TeV* center-of-mass energy (an order of magnitude above
         | the LHC's 14 TeV). Heavy-ion collisions (e.g. Pb-Pb) are also
         | envisioned. An *FCC-eh* option (electron-hadron collisions) is
         | considered by adding a high-energy electron injector to collide
         | with the proton beam. This integrated FCC program thus spans
         | both *precision lepton* collisions and *energy-frontier hadron*
         | collisions.
         | 
         | - *Timeline:* The conceptual schedule foresees *FCC-ee
         | construction in the 2030s* and a start of operations by around
         | *2040* (as the LHC/HL-LHC program winds down). According to the
         | FCC Conceptual Design Report, an $e^+e^-$ Higgs factory could
         | begin delivering physics in _~2040_ , running for 15-20 years.
         | The *hadron collider FCC-hh* would be constructed subsequently
         | (using the same tunnel and upgraded infrastructure), aiming for
         | *first proton-proton collisions in the late 2050s)] . This
         | staged approach (lepton collider first, hadron later) mirrors
         | the successful *LEP-LHC sequence*, leveraging the $e^+e^-$
         | machine to produce great precision data (and to build
         | infrastructure) before pushing to the highest energies with the
         | hadron machine. ...
         | 
         | (Too long for HN to write more)
        
           | fosterfriends wrote:
           | Honestly, these are the smartest and overall best LLM outputs
           | I've ever seen to date. Loving Deep Research, feels like
           | another level up in the race
        
             | elashri wrote:
             | Thank you very much for doing that. It is actually somehow
             | impressive. It got a lot of big picture comparison and
             | points correct. There are problem with some details but
             | overall it does save some work for initial search process.
             | 
             | What I like is that it asked you before clarifying
             | questions before but I wonder if it just generic. Because
             | the prompt mentioned that this would be for "presentation
             | at a topical meeting of particle physicists" but still
             | asked its last question about
             | 
             | > Intended Audience: Should the analysis assume a general
             | physics audience or a more specialized group of particle
             | physicists?
             | 
             | Also probably expected but it didn't include or reference
             | graphs/plots.
        
       | rajnathani wrote:
       | I remember about 10-15 years ago that Ray Kurzweil (who still
       | works at Google) or someone at Google had this idea for what
       | Google should be able to do: About doing deep research by itself
       | with a simple search query. I can't find the source. Obviously it
       | didn't pan out without transformers.
        
       | corentin88 wrote:
       | Curious about the use cases here. Building AI Agents? But which
       | one?
        
       | picografix wrote:
       | I think deep research as a service could be a really strong use
       | case for enterprises, as long as they have access to non-public
       | data. I assume that most of this guarded data is high quality,
       | and seeing progress in these areas might end up being even more
       | impressive than it is now.
        
       | anon373839 wrote:
       | Setting aside how well it works, I think this is a pretty nice
       | demonstration of how to do UX for an agentic RAG app. I like that
       | the intermediate steps have been pushed out to a sidebar, with
       | updates that both provide some transparency about the process and
       | make the high latency more palatable.
        
       | auggierose wrote:
       | The flow reminds me a bit of undermind.ai.
        
       | rob_c wrote:
       | Feels more and more like openAI doesn't have "that next big
       | thing".
       | 
       | To be clear I'm constantly impressed with what they have and what
       | I get as a customer, but the delivery since 4 hasn't exactly been
       | in line with Altman's Musk-tier vapoware promises...
        
       | gorgoiler wrote:
       | For "deep research" I'm also reading "getting the answers right".
       | 
       | Most people I talk to are at the point now where getting
       | completely incorrect answers 10% of the time -- either obviously
       | wrong from common sense, or because the answers are self
       | contradictory -- undermines a lot of trust in any kind of
       | interaction. Other than double checking something you already
       | know, language models aren't large enough to actually _know_
       | everything. They can only sound like they do.
       | 
       | What I'm looking for is therefore not just the correct answer,
       | but the correct answer in an amount of time that's faster than it
       | would take me to research the answer myself, _and also faster
       | than it takes me to verify the answer given by the machine_.
       | 
       | It's one thing to ask a pupil to answer an exam paper to which
       | you know the answers. It's a whole next level to have it answer
       | questions to which you don't know the answers, and on whose
       | answers you are relying to be correct.
        
         | sandos wrote:
         | I mean this all falls down due to the need of verification:
         | 
         | "Limitations Deep research unlocks significant new
         | capabilities, but it's still early and has limitations. It can
         | sometimes hallucinate facts in responses or make incorrect
         | inferences"
         | 
         | How do I know which parts are false? It will take as long to
         | verify as to research!
        
         | igleria wrote:
         | It's really worrying to me, even as a self proclaimed "LLM <->
         | AI" skeptic, to see what kind of stuff people pretend to get
         | out from an LLM. Typewriter monkeys as a service almost.
         | 
         | Still useful for the odd task here and there, but not as useful
         | as all the money being invested in this (except for the
         | companies getting that money, that is).
         | 
         | edit: actual example of something I'd expect a real AI to be
         | able to solve by itself, but currently LLMs fail miserably
         | https://x.com/RadishHarmers/status/1885884032220643587
        
           | mdp2021 wrote:
           | > _Typewriter monkeys as a service almost. // Still useful
           | for the odd task here and there_
           | 
           | 1) Paramount task: searching in naturally structured
           | language, as opposed to keywords. Odd tasks: oh yes, several
           | tasks of fuzzy sophisticated text processing previously
           | unsolved.
           | 
           | 2) They translate NN encodings in natural language! The issue
           | remains about the quality of /what/ they translate in natural
           | language, but one important chunk of the problem* is in a way
           | solved...
           | 
           | Now, I've been probably one of the most vocal here, shouting
           | "That's the opposite of intelligence!" - even in the past 24
           | hours -, but be objective: there are also progresses ...
           | 
           | (* Around five years ago we were still stuck with Hinton's
           | problem of interpreting pronouns as pointers in "the item
           | won't fit in the case: it's too big" vs "the item won't fit
           | in the case: it's too small" - look at it now...)
        
             | igleria wrote:
             | Of course I see progress, but I feel like the bridge of
             | "thinks" versus "regurgitates" is still far off, if it is
             | still in the horizon with the current approach. IMHO.
             | 
             | edit: furthermore, LLMs probably tackle very little "real
             | state" in the "make machines THINK" land. But a crucial
             | piece on the overall puzzle.
        
         | chombier wrote:
         | > and also faster than it takes me to verify the answer given
         | by the machine.
         | 
         | I always thought there was a kind of NP-flavor to the problems
         | for which LLMs-like AI are helpful in practice, in the sense
         | that solving the problem may be hard but checking the solution
         | must be fast.
         | 
         | Unless the domain can accomodate errors/hallucination, checking
         | the solution (by a human) should be exponentially faster than
         | finding it (by some AI) otherwise there's little practical
         | gain.
        
         | jeswin wrote:
         | > Most people I talk to are at the point now where getting
         | completely incorrect answers 10% of the time
         | 
         | A year back that number was 30%, and a couple of years back it
         | was 60%. There will be a point where it'll be good enough.
         | There are also better and better ways to verify answers these
         | days.
         | 
         | It'll never be a solution for everything, but that's similar to
         | many engineering problems we have: for example, ORMs aren't
         | great for all types of queries, but they're sufficient for a
         | good part of them.
        
           | dimitri-vs wrote:
           | It contributes little to discuss a hypothetical future. Maybe
           | we'll have fusion energy, delivery drones, everyone using VR,
           | etc. Maybe we will go into a deep recession due to trade
           | wars, or maybe not.
           | 
           | The meaningful discussion is about how they perform NOW and
           | the edge cases that have persisted since GPT-2 which no one
           | has yet found a good solution for.
        
             | infecto wrote:
             | We already have delivery drones though.
             | 
             | I disagree though, it is useful as this problem has been
             | whittled down and I think there is expectation that there
             | will be continued effort. Its of course worth discussing
             | but I find that for my workflows, I rarely encounter issues
             | with hallucinations, they certainly exist but its gotten to
             | a point that I don't have major issue with it.
        
               | skywhopper wrote:
               | At best, a proof of concept of experimental delivery
               | drones exist, but only for small, lightweight items, and
               | only in a few places, only in the right weather, and only
               | if you place a target on your driveway and are there to
               | receive the item in person, and all at the cost of a very
               | high noise level. That's not exactly a real service.
        
         | herculity275 wrote:
         | My worry is that all these recent capabilities attempt to
         | minimize hallucinations by relying on extensive web search,
         | however web itself is being actively degraded by unfiltered LLM
         | output. After a certain point running your research agent
         | against a ~5-year-old snapshot of the web will be strictly more
         | accurate (for non-current affairs queries) than querying live
         | web.
        
         | shakes_mcjunkie wrote:
         | > What I'm looking for is therefore not just the correct
         | answer, but the correct answer in an amount of time that's
         | faster than it would take me to research the answer myself, and
         | also faster than it takes me to verify the answer given by the
         | machine.
         | 
         | This is why I haven't found AI tools very useful. I find my
         | self spending more time verifying and fixing it's answers than
         | I would have just doing or learning the darn thing myself.
        
           | 7thpower wrote:
           | It is added cognitive load, but there is a lot of value in
           | async tasks _if_ you can trust the output or if the
           | opportunity cost of validating is low.
           | 
           | The challenge with something like this for research, in its
           | current state, is you'll need to go double check it because
           | you don't trust it and it will end up effectively being a
           | list of links.
           | 
           | It's progress though and evidently good enough to find a
           | sweet NSX in Japan, which is all some really need.
        
         | squigz wrote:
         | > They can only sound like they do.
         | 
         | More importantly, I think, is that they are _incapable of not
         | doing so_. Have we figured out how to make an LLM realize and
         | answer that it doesn 't know an answer?
        
         | HarHarVeryFunny wrote:
         | > language models aren't large enough to actually know
         | everything
         | 
         | I'd say they don't know _anything_.
         | 
         | An LLM base model, before it is post-trained with RL, just has
         | access to a sliced and diced corpus of human output. Take the
         | contents of 4chan and WikiPedia, put in blender and mix and
         | chop into "training sample" sized bites, then learn the
         | statistical regularities of this blended mess. It is what it is
         | - not exactly what I'd call a knowledge base, even though there
         | are bits of knowledge in there.
         | 
         | When you add RL-based post-training for reasoning, all you are
         | doing is trying get the model to be more selective when you are
         | sampling from it - encouraging it to suppress some statistics,
         | and emphazise others, such that when you sample from it the
         | output looks more like valid reasoning steps and/or
         | conclusions, per the verified reasoning examples you train it
         | on.
         | 
         | I'm well aware of how useful RL-tuned models (whatever the
         | goal) can be, but at the end of the day all they are doing is
         | taking a statistical babbler and saying "try to output patterns
         | more like this". It's not exactly a recipe for factuality or
         | rationality - we've just gone from hallucination-prone base
         | models, to gaslighting-prone RL-tuned "reasoning" models that
         | output stuff that _sounds like_ coherent reasoning.
         | 
         | What missing from all of this - what makes it different from
         | how animals learn - it that the model has no experience of it's
         | own, no autonomy or motivation to explore, learn and verify,
         | and hence no episodic memories of how it learnt something
         | (tried it and ran controlled experiments, or just overheard it
         | on the bus), and what that implies about it's trustworthiness.
         | 
         | It's amazing that LLMs work as well as they do, a reflection of
         | how much of what we do can be accomplished by reactive pattern
         | matching, but if you want to go beyond that to something that
         | can learn and discern the truth for itself, this seems the
         | wrong paradigm altogether.
        
       | taran_narat wrote:
       | isn't this just perplexity?
        
       | gigatexal wrote:
       | Ok so I do this as a noob in some field. How do I know or trust
       | the research conclusions? How do I know it's not hallucinated its
       | conclusions? I'll likely have to do my own research to just
       | verify it and then if I did I might as well have done the
       | research myself.
        
       | freehorse wrote:
       | I love that when "open"ai releases things last year or so, they
       | do not actually release them. So we get the chance in the
       | meantime to all enjoy a bunch of speculative, shilling comments
       | here about this next great thing being miles ahead of
       | competitors/close to AGI/the tool that will actually do X thing
       | that others complain so far llms are failing to do.
        
       | dazzaji wrote:
       | Late Sunday night, I gained access to OpenAI's newly launched
       | Deep Research and immediately tested it on a draft blog post
       | about Uniform Electronic Transactions Act (UETA) compliance and
       | AI-agent error handling [1]. Here's what I found:
       | 
       | Within minutes, it generated a detailed, well-cited research
       | report that significantly expanded my original analysis,
       | covering: * Legal precedents & case law interpretations
       | (including a nuanced breakdown of UETA Section 10). * Comparative
       | international frameworks (EU, UK, Canada). * Real-world technical
       | implementations (Stripe's AI-driven transaction handling). *
       | Industry perspectives & business impact (trust, risk allocation,
       | compliance). * Emerging regulatory standards (EU AI Act, FTC
       | oversight, ISO/NIST AI governance).
       | 
       | What stood out most was its ability to: - Synthesize complex
       | legal, business, and technical concepts into clear, actionable
       | insights. - Connect legal frameworks, industry trends, and real-
       | world case studies. - Maintain a business-first focus,
       | emphasizing practical benefits. - Integrate 2024 developments
       | with historical context for a deeper analysis.
       | 
       | The depth and coherence of the output were comparable to what I
       | would expect from a team of domain experts--but delivered in a
       | fraction of the time.
       | 
       | From the announcement: Deep Research leverages OpenAI's next-
       | generation model, optimized for multi-step research, reasoning,
       | and synthesis. It has already set new performance benchmarks,
       | achieving 26.6% accuracy on Humanity's Last Exam (the highest of
       | any OpenAI model) and a 72.57% average accuracy on the GAIA
       | Benchmark, demonstrating advanced reasoning and research
       | capabilities.
       | 
       | Currently available to Pro users (with up to 100 queries per
       | month), it will soon expand to Plus and Team users. While OpenAI
       | acknowledges limitations--such as occasional hallucinations and
       | challenges in source verification--its iterative deployment
       | strategy and continuous refinement approach are promising.
       | 
       | My key takeaway: This LLM agent-based tool has the potential to
       | save hours of manual research while delivering high-quality,
       | well-documented outputs. Automating tasks that traditionally
       | require expert-level investigation, it can complete complex
       | research in 5-30 minutes (just 6 minutes for my task), with
       | citations and structured reasoning.
       | 
       | I don't see any other comments yet from people who have actually
       | used it, but it's only been a few hours.I'd love to hear how it's
       | performing for others. What use cases have you explored? How did
       | it do?
       | 
       | (Note: This review is based on a single use case. I'll provide
       | further updates as I conduct broader testing.)
       | 
       | [1] https://www.dazzagreenwood.com/p/ueta-and-llm-agents-a-
       | deep-...
        
         | timabdulla wrote:
         | I tried it on a few things I was familiar with just to assess
         | its reliability.
         | 
         | The first was on a topic with which I am deeply familiar --
         | myself -- and it made three factual errors in a 500-word
         | report: https://news.ycombinator.com/item?id=42916899
         | 
         | The second was a task to do an industry analysis on a space in
         | which I worked for about ten years. I think its overall
         | synthesis was good (it accorded with my understanding of the
         | space), but there were a number of errors in the statistics and
         | supporting evidence it compiled, based upon my random review of
         | the source material.
         | 
         | I think the product is cool and will definitely be helpful, but
         | I would still recommend verifying its outputs. I think the
         | process of verification is less time-consuming than the process
         | of researching and writing, so that is likely an acceptable
         | compromise in many cases.
        
       | Xuban wrote:
       | This make sense, I often use the normal search feature to
       | research a very large ammount of information and it mostly does
       | not work well. If the new search feature increases the number of
       | websites scrapped and the pertinence of the websites, I'm all in.
        
       | timabdulla wrote:
       | I just gave it a whirl. Pretty neat, but definitely watch out for
       | hallucinations. For instance, I asked it to compile a report on
       | myself (vain, I know.) In this 500-word report (ok, I'm not that
       | important, I guess), it made at least three errors.
       | 
       | It stated that I had 47,000 reputation points on Stack Overflow
       | -- quite a surprise to me, given my minimal activity on Stack
       | Overflow over the years. I popped over to the link it had cited
       | (my profile on Stack Overflow) and it seems it confused my number
       | of people reached (47k) with my reputation, a sadly paltry 525.
       | 
       | Then it cited an answer I gave on Stack Overflow on the topic of
       | monkey-patching in PHP, using this as evidence for my technical
       | expertise. Turns out that about 15 years ago, I _asked_ a
       | question on this topic, but the answer was submitted by someone
       | else. Looks like I don't have much expertise, after all.
       | 
       | Finally, it found a gem of a quote from an interview I gave. Or
       | wait, that was my brother! Confusingly, we founded a company
       | together, and we were both mentioned in the same article, but he
       | was the interviewee, not I.
       | 
       | I would say it's decent enough for a springboard, but you should
       | definitely treat the output with caution and follow the links
       | provided to make sure everything is accurate.
        
         | giarc wrote:
         | What's faster? Writing a 500 word report "from scratch" by
         | researching the topic yourself, vs. having AI write it then
         | having to fact check every answer and correct each piece
         | manually?
         | 
         | This is why I don't use AI for anything that requires a
         | "correct" answer. I use it to re-write paragraphs or sentences
         | to improve readability etc, but I stop short of trusting any
         | piece of info that comes out from AI.
        
         | mdp2021 wrote:
         | > _Then it cited an answer I gave on Stack Overflow [...] using
         | this as evidence for my technical expertise. Turns out that
         | about 15 years ago, I _asked_ a question on this topic, but the
         | answer was submitted by someone else_
         | 
         | Artificial dementia...
         | 
         | Some parties are releasing products much earlier than the
         | ability to ship well working products (I am not sure that their
         | legal cover will be so solid), but database aided outputs
         | should and could become a strong limit to that phenomenon of
         | remembering badly. Very linearly, like humans: get an idea,
         | then compare it to the data - it is due diligence and part of
         | the verification process in reasoning. It is as if some moves
         | outside linear pure product progress reasoning are swaying the
         | RnD towards directions outside the primary concerns. It's a
         | form of procrastination.
        
         | RobinL wrote:
         | Interesting
         | 
         | You might find it amusing to compare it to: https://hn-
         | wrapped.kadoa.com/timabdulla
         | 
         | (Ref:https://news.ycombinator.com/item?id=42857604)
        
           | wholinator2 wrote:
           | This is... very uncomfortable. An (expanded) AI summary of my
           | HN and reddit usage would appear to be a pretty complete
           | representation of my "online" identity/character. I remember
           | when people would browse your entire comment history just to
           | find something to discredit you on reddit, and that behavior
           | was _heavily_ discouraged. Now, we can just run an AI model
           | to follow you and sentence you to a hell of being permanently
           | discredited online. Give it a bunch of accounts to rotate
           | through, send some voting power behind it (reddit or hn), and
           | just pick apart every value you hold. You could obliterate
           | someone's will to discuss anything online. You could
           | effectively silence all but the most stubborn, and those
           | people you would probably drive insane.
           | 
           | It's a very interesting usecase though, filter through
           | billions of comments and give everyone a score on which real
           | life person they probably are. I wonder if say, Ted Cruz
           | hides behind a username somewhere.
        
             | throwaway519 wrote:
             | throwaway/anonymous.
             | 
             | not just for when discussion of the content not the
             | personality behind it is important.
        
           | dlivingston wrote:
           | I put my profile in [0] and it's mostly silly; a few comments
           | extracted and turned into jokes. No deep insights into me,
           | and my "Top 3 Technologies" are hilariously wrong (I've never
           | written a single line of TypeScript!)
           | 
           | [0]: https://hn-wrapped.kadoa.com/dlivingston
        
           | ComputerGuru wrote:
           | That.. seems to just take a few (three or four) random
           | comments that received some attention and then extrapolate an
           | entire profile based on (incorrectly) interpreting their
           | contents?
           | 
           | https://hn-wrapped.kadoa.com/ComputerGuru
        
         | prof-dr-ir wrote:
         | > Pretty neat, but definitely watch out for hallucinations.
         | 
         | That would be exactly my verdict of any product based on LLMs
         | in the past few years.
        
         | toasteros wrote:
         | "Pretty neat, but definitely watch out for hallucinations."
         | 
         | We'd never hire someone who just makes stuff up (or at least
         | keep them employed for long). Why are we okay with calling "AI"
         | tools like this anything other than curious research projects?
         | 
         | Can't we just send LLMs back to the drawing board until they
         | have some semblance of reliability?
        
           | oldstrangers wrote:
           | > Can't we just send LLMs back to the drawing board until
           | they have some semblance of reliability?
           | 
           | Well at this point they've certainly proven a net gain for
           | everyone regardless of the occasional nonsense they spew.
        
             | DanHulton wrote:
             | That is... debatable. You may be entirely inside the
             | bubble, there.
        
               | taikahessu wrote:
               | Not sure if this was posted as humour, but I don't feel
               | that way. In today's world, where I certainly would
               | consider taking the blue pill, I'm having a blast with
               | LLMs!
               | 
               | It has helped me learn stuff incredibly faster.
               | Especially I find them useful for filling the gaps of
               | knowledge and exploring new topics in my own way and
               | language, without needing to wait an answer from a human
               | (that could also be wrong).
               | 
               | Why does it feel, that "we are entirely inside the
               | bubble" for you?
        
               | dingnuts wrote:
               | >It has helped me learn stuff incredibly faster.
               | Especially I find them useful for filling the gaps of
               | knowledge and exploring new topics in my own way and
               | language
               | 
               | and then you verify every single fact it tells you via
               | traditional methods by confirming them in human-written
               | documents, right?
               | 
               | Otherwise, how do you use the LLM for learning? If you
               | don't know the answer to what you're asking, you can't
               | tell if it's lying. It also can't tell if it's lying, so
               | you can't ask it.
               | 
               | If you have to look up every fact it outputs after it
               | does, using traditional methods, why not skip to just
               | looking things up the old fashioned way and save time?
               | 
               | Occasionally an LLM helps me surface unknown keywords
               | that make traditional searches easier, but they can't
               | teach anything because they don't know anything. They can
               | imagine things you might be able to learn from a real
               | authority, but that's it. That can be useful! But it's
               | not useful for learning alone.
               | 
               | And if you're not verifying literally everything an LLM
               | tells you.. are you sure you're learning anything real?
        
               | kardos wrote:
               | The Gell-Mann amnesia effect applies to LLMs as well!
               | 
               | https://en.m.wikipedia.org/wiki/Gell-Mann_amnesia_effect
        
               | taikahessu wrote:
               | I guess it all depends on the topic and levels of trust.
               | How can I be certain that I have a brain? I just have to
               | take something for granted, don't I? Of course I will
               | "verify" the "important stuff", but what is important?
               | How can I tell? Most of the time only thing I need is a
               | pointer in the right direction. Wrong advice? I know when
               | I get there I suppose.
               | 
               | I can remember numerous things I was told while growing
               | up, that aren't actually true. Either by plain lies and
               | rumours or because of the long list of our cognitive
               | biases.
               | 
               | > If you have to look up every fact it outputs after it
               | does, using traditional methods, why not skip to just
               | looking things up the old fashioned way and save time?
               | 
               | What is the old fashioned way? I mean people learn
               | "truths" these days from Tiktok and Youtube. Some of the
               | stuff is actually very good, you just have to distill it
               | based on the stuff I was being taught at school. Nonody
               | has yet declared LLMs as a subtitute for schools, maybe
               | they soon will, but neither "guarantees" us anything. We
               | could as well be taught political agendas.
               | 
               | I could order a book about construction, but I wouldn't
               | build a house without asking a "verified" expert. Some
               | people build anyway and we get some catastrofic results.
               | 
               | Levels of trust, it's all games and play until it gets
               | serious, like what to eat or doing something that
               | involves life threatening physics. I take it as playing
               | with a toy. Surely something great have come up from only
               | a few piece of legos?
               | 
               | > And if you're not verifying literally everything an LLM
               | tells you.. are you sure you're learning anything real?
               | 
               | I guess you shouldn't do it that way. But really, so far
               | the topics I've rigorously explored with ChatGPT for
               | example, have been better than your average journalism.
               | What is real?
        
               | dingnuts wrote:
               | > What is the old fashioned way?
               | 
               | Looking in a resource written by someone with sufficient
               | ethos that they can be considered trustworthy .
               | 
               | > What is real?
               | 
               | I'm not arguing ontology about systems that can't do
               | arithmetic. you're not arguing in good faith at all
        
               | dialup_sounds wrote:
               | Saying you need to verify "literally everything" both
               | overestimates the frequency of hallucinations and
               | underestimates the amount of wrong found in human-written
               | sources. e.g. the infamous case of Google's AI
               | recommending Elmer's glue on pizza was literally a human-
               | written suggestion first: https://www.reddit.com/r/Pizza/
               | comments/1a19s0/my_cheese_sli...
        
               | squigz wrote:
               | > without needing to wait an answer from a human (that
               | could also be wrong).
               | 
               | The difference is you have some reassurances that the
               | human is not wrong - their expertise and experience.
               | 
               | The problem with LLMs, as demonstrated by the top-level
               | comment here, is that they constantly make stuff up.
               | While you may think you're learning things quickly, how
               | do you know you're learning them "correctly", for lack of
               | a better word?
               | 
               | Until an LLM can say "I don't know", I really don't think
               | people should be relying on them as a first-class method
               | of learning.
        
               | toasteros wrote:
               | Are you sure it's helped you learn?
               | 
               | In the early days of ChatGPT where it seemed like this
               | fun new thing, I used it to "learn" C. I don't remember
               | anything it told me, and none of the answers it gave me
               | were anything that I couldn't find elsewhere in different
               | forms - heck I could have flipped open Kernighan &
               | Ritchie to the right page and got the answer.
               | 
               | I had a conversation with an AI/Bitcoin enthusiast
               | recently. Maybe that already tells you everything you
               | need to know about this person, but to the hammer the
               | point home, they made a claim to similar to you: "I learn
               | much more and much better with AI". They also said they
               | "fact check" things it "tells" them. Some moments later
               | they told me "Bitcoin has its roots in Occupy Wall
               | Street".
               | 
               | A simple web search tells you that Bitcoin is conceived a
               | full 2 years before Occupy. How can they be related?
               | 
               | It's a simple error that can be fact checked simply. It's
               | a pretty innocuous falsity in this particular case - but
               | how many more falsehoods have they collected? How do
               | those falsehoods influence them on a day-by-day basis?
               | 
               | How many falsehoods influence you?
               | 
               | A very well meaning activist posted a "comprehensive"
               | list of all the programs that were to be halted by the
               | grants and loans freezes last week. Some of the entries
               | on the list weren't real, or not related to the freeze.
               | They revealed they used ChatGPT to help compile the list
               | and then went down one-by-one to verify each one.
               | 
               | With such meticulous attention to detail, incorrect
               | information still filtered through.
               | 
               | Are you sure you are learning?
        
               | panarky wrote:
               | When your bitcoiner friend told you something that's not
               | true, that's a human who hallucinated, not an LLM.
               | 
               | Maybe we're already at AGI and just don't know it because
               | we overestimate the capabilities of most humans.
        
               | toasteros wrote:
               | The assertion is that they "learned" that Bitcoin came
               | from Occupy from an AI.
               | 
               | If AI is teaching you, you are going to collect a
               | thousand papercuts of lies.
        
               | taikahessu wrote:
               | I guess the real learning happens outside the AI, here in
               | real life. Does the code run? Sure, it's on my local and
               | not in production, but I would've never have the patience
               | to get "that new thing working" without AI as assistant.
               | 
               | Does the food taste good? Oops, there's a bit too much
               | vegetables here, they are never gonna fit in this pan of
               | mine. Not a big deal, next time I'll be wiser.
               | 
               | AI is like a hypothesis machine. You're gonna have to
               | figure out if the output is true. Few years ago, just
               | testing any machine's "intelligence" was pretty quickly
               | done and machine failed miserably. Now, the accuracy is
               | astounishing in comparison.
               | 
               | > How many falsehoods influence you?
               | 
               | That is a great question. The answer is definitely not
               | zero. I try to live by with a hacker mentality and I'm an
               | engineer by trade. I read news and comments, which I'm
               | not sure is good for me. But you also need some
               | compassion towards oneself. It's not like ripping
               | everything open will lead to salvation. I believe the
               | truth does set you free, eventually. But all in one's
               | time...
               | 
               | Anyway, AI is a tool like any other. Someone will hammer
               | their fingers with it. I just don't understand the hate.
               | It's not like we're drinking any AI koolaids here. It's
               | just like it was 30 years ago (in my personal journey),
               | you had a keyboard and a machine, you asked it things and
               | got gibberish. Now the conversation with it just started
               | to get interesting. Peace.
        
               | orangepanda wrote:
               | You overestimate the importance of being correct
        
             | aiono wrote:
             | No, from the research around it the findings are mixed.
             | There is no consensus that it's net gain.
        
             | kees99 wrote:
             | "Occasional nonsense" doesn't sound great, but would be
             | tolerable.
             | 
             | Problem is - LLMs pull answers from their behind, just like
             | a lazy student on the exam. "Halucinations" is the word
             | people use to describe this.
             | 
             | Those are extremely hard to spot - unless you happen to
             | know the right answer already, at which point - why ask?
             | And those are _everywhere_.
             | 
             | One example - recently there was quite a discussion about
             | llm being able to understand (and answer) base16 (aka
             | "hex") encoding on the fly, so I went on to try base64,
             | gzipped base64, zstd-compressed base64, etc...
             | 
             | To my surprise, LLM got most of those encoding/compressions
             | right, decoded/uncompressed the question, and answered it
             | flawlessly.
             | 
             | But with few encodings, LLM detected base64 correctly, got
             | compression algorithm correctly, and then... instead of
             | decompressing, made up a completely different payload, and
             | proceeded to answer that. Without any hint of anything
             | sinister going.
             | 
             | We really need LLMs to reliably calculate and express
             | confidence. Otherwise they will remain mere toys.
        
               | oldstrangers wrote:
               | Yeah, what you said represents a 'net gain' over not
               | having any of that at all.
        
             | majormajor wrote:
             | I think as these things get more integrated into customer
             | service workflows - especially for things like insurance
             | claims - there's gonna start being a lot more buyer's
             | remorse on everyone's part.
             | 
             | We've tried for decades to turn people into reliable
             | robots, now many companies are running to replace people
             | robots with (maybe less reliable?) robot-robots. What could
             | go wrong? What are the escalation paths going to be? Who's
             | going to be watching them?
        
             | hawaiianbrah wrote:
             | A net gain for everyone? Tell that to the artists its
             | screwing over!
        
           | throwing_away wrote:
           | > We'd never hire someone who just makes stuff up (or at
           | least keep them employed for long).
           | 
           | This is contrary to my experience.
        
             | kadushka wrote:
             | Our president begs to differ! Or pretty much any elected
             | official for that matter.
        
           | cdblades wrote:
           | > Why are we okay with calling "AI" tools like this anything
           | other than curious research projects?
           | 
           | Because they are a way to launder liability while reducing
           | costs to produce a service.
           | 
           | Look at the AI-based startups y-combinator has been funding.
           | They match that description.
        
           | deeviant wrote:
           | Yeah, I used to hire people, but then one of them made a
           | mistake, now I'm done with them forever, they are useless. It
           | is not I, who is directing the workers, who cannot create a
           | process that is resistant to errors, it's definitely the fact
           | that all people are worthless until they make no errors as
           | there truly is no other way of doing things other than
           | telling your intern to do a task then having them send it
           | directly to the production line.
        
           | ramon156 wrote:
           | 3k a month vs ~500 dollars a month. That's all u need to
           | know. Not saying its as good, but its all some managers care
           | about
        
           | kenjackson wrote:
           | You can use them for whatever you like, or not use them.
           | Everyone has a different bar for when technology is useful.
           | My dad doesn't think EVs are useful due to the long charge
           | times, but there are others who find it fully acceptable.
        
           | dumbfounder wrote:
           | Why not just verify the output? It's faster than generating
           | the entire thing yourself. Why do you need perfection in a
           | productivity tool?
        
             | toasteros wrote:
             | At that point why not just... I dunno, do the research
             | yourself?
        
               | tuckerman wrote:
               | Perhaps because the time to proofread/correct is less
               | than to do it from scratch? That would still make it a
               | valuable tool
        
               | Yoric wrote:
               | But is it?
        
               | toasteros wrote:
               | How?
               | 
               | It's given you some information and now you have to seek
               | out a source to verify that it's correct.
               | 
               | Finding information is hard work. It's why librarian is a
               | valuable skilled profession. What you've done by
               | suggesting that I should "verify" or "proofread" what a
               | glorified, water-wasting Markov chain has given me now
               | entails me looking up that information to verify that
               | it's correct. That's...not quite doubling the work
               | involved but it's adding an unnecessary step.
               | 
               | I could have searched for the source in the first
               | instance. I could have gone to the library and asked for
               | help.
               | 
               | We spent time coming up with a question ("prompt
               | engineering"! hah!), we used up a bunch of electricity
               | for an answer to be generated and now you...want me to
               | search up that answer to find the source? Why did we do
               | the first step?
               | 
               | People got undergraduate degrees - hell, even PhDs -
               | before generative AI.
               | 
               | Look up the tweet from someone who said "Sometimes when
               | coming up with a good prompt for ChatGPT, I sometimes
               | come up with the answer myself without needing to
               | submit".
        
           | roflyear wrote:
           | > We'd never hire someone who just makes stuff up
           | 
           | We do all the time - of course we do, all the time.
        
           | nomel wrote:
           | LLM are _" great"_ in some use cases, "ok" in others, and
           | "laughable" in more.
           | 
           | Some people might find $500 worth of value, in their specific
           | use case, in those "great" and "ok" categories, where they
           | get more value than "lies" out of it.
           | 
           | A few verifiable lies, vs hours of time, could be worth it
           | for some people, with use cases outside of your perspective.
        
           | rybosome wrote:
           | This doesn't make LLMs worthless, you just need to structure
           | your processes around fallibility. Much like a well designed
           | release pipeline is built with the expectation that devs will
           | write bugs that shouldn't ship.
        
         | vessenes wrote:
         | Interesting!
         | 
         | I wonder if it's carried over too much of that 'helpful' DNA
         | from 4o's RLHF. In that case, maybe asking for 500 words was
         | the difficult part -- it just didn't have enough to say based
         | on one SO post and one article, but the overall directives
         | assume there is, and so the model is put into a place where it
         | must publish..
         | 
         | Put another way, it seems this model faithfully replicates the
         | incentives most academics have -- publish a positive result, or
         | get dinged. :)
         | 
         | Did it pick up your HN comments? Kadua claims that's more than
         | enough to roast me, ... and it's not wrong. It seems like
         | there's enough detail about you (or me) there to do a better
         | job summarizing.
        
           | timabdulla wrote:
           | I didn't actually give it a goal of writing any particular
           | length, but I do think that perhaps given my not-so-large
           | online footprint, it may have felt "pressured" to generate
           | content that simply isn't there.
           | 
           | It didn't pick up my HN comments, probably because my first
           | and last name are not in my profile, though obviously that is
           | my handle in a smooshed-together form.
        
         | brushfoot wrote:
         | I disagree that this is a useful springboard. And I say that as
         | an AI optimist.
         | 
         | A report full of factual errors that a careful intern wouldn't
         | make is worse than useless (yes, yes, I've mentored interns).
         | 
         | If the hard part is the language, then do the research
         | yourself, write an outline, and have the LLM turn it into
         | complete sentences. That would at least be faster.
         | 
         | Here's the thing, though: If you do that, you're effectively
         | proving that prose style is the low-value part of the work, and
         | may be unnecessary. Which, as much as it pains to me say as a
         | former English major, is largely true.
        
         | machiaweliczny wrote:
         | This is very bearish for current AI. Seems like 99% reliability
         | is still too small with compounding errors. But I wonder of
         | this is inherently specific to longer context or if this just
         | depends on how it's trained. In theory longer context => more
         | errors
         | 
         | Although I think people are the same, too big problem and you
         | are getting lost unless taking it in bites, so seems like
         | OpenAI implementation is just bad because o3 hallucination
         | benchmark shouldn't lead to such poor performance
        
         | Bjorkbat wrote:
         | So, I still think this is a cool tool for search reasons, but
         | otherwise the tendency to hallucinate makes it questionable as
         | a researcher.
         | 
         | Hypothetically speaking, if the time you saved is now spent
         | verifying the statements of your AI researcher, then did you
         | really save any time at all?
         | 
         | If the answers aren't important enough to verify, then was it
         | ever even important enough to actually research to begin with?
        
       | smusamashah wrote:
       | can sometimes hallucinate facts in responses or make incorrect
       | inferences, though at a notably lower rate than existing ChatGPT
       | models, according to internal evaluations. It may struggle with
       | distinguishing authoritative information from rumors, and
       | currently shows weakness in confidence calibration, often failing
       | to convey uncertainty accurately
       | 
       | Taken from the limitations section.
       | 
       | These tools are just good at creating pollution. I don't see the
       | point of delegating a (not just) research where 1% blatant
       | mistakes are acceptable. These need much better grounding before
       | handing out to masses.
       | 
       | I can not take any output by these tools (google summaries,
       | comment summaries by amazon, youtube summaries etc etc) while
       | knowing for a fact some of that is a total lie. I can not tell
       | which part is a lie. e.g. If LLM says that in any given text the
       | sentiment is divided, it could be just one person with an
       | opposing view.
       | 
       | If same task was given to a person, I could reason with that
       | person on _any_ conclusion. These tools will reason on their
       | hallucinations.
        
       | joanfihu wrote:
       | There is no way I'll read all that text from the demos...
       | 
       | AskPandi has a similar feature called "Super Search" that
       | essentially checks more sources and self validates it's own
       | answers.
       | 
       | iT's AgEnTic.
       | 
       | The answers are easier to digest, if you search for products,
       | you'll get a list of products with images, prices and retailers.
        
       | monkeydust wrote:
       | What a decent setup to replicate via open model and agent
       | framework? One thing I have struggled with is getting
       | comprehensive web searches using an agentic framework.
        
       | regularjack wrote:
       | Of course, they had to weasel the word "deep" in there.
        
       | sharpshadow wrote:
       | Are they launching a new feature after some other AI got the
       | attention to get the attention back?
        
       | titzer wrote:
       | It's great that none of these AI models are being foisted on us
       | by advertising companies.
        
       | axpy906 wrote:
       | Don't most researchers have a local setup plugged into Olama so
       | that they do NOT share their search information?
        
       | sivm wrote:
       | I used it once to research language learning and had my pro mode
       | taken away pending review for abuse.
        
       | Alifatisk wrote:
       | When I saw new to llms, I used Bing ai in a fun way. So when I
       | was writing my report, it was sometimes hard to find discussions
       | or material about a certain topic.
       | 
       | What I did was to ask Bing ai about that topic and it returned
       | information aswell as sources to where it found those, so I
       | picked up all those links and researched them myself.
       | 
       | Bing ai was a great resource for finding relevant links, this was
       | until I found out about perplexity, my life haven't been the same
       | since.
        
       | throwaway123lol wrote:
       | This is so lame. This feels like another desperate attempt to
       | stay relevant cobbled together after the DeepSeek announcement
       | last week. What was the other attempt they made? Skip a version
       | number to seem like more progress was made (o1->o3)? From what I
       | can tell "o3" is just the same as o1 with an extra reasoning-
       | effort parameter.
       | 
       | Oh and "Deep research" is available to people on the $200 per
       | month plan? Lol - cool. I've been using DeepSeek a lot more
       | recently and it's so incredibly good even with all the scaling
       | issues.
        
       | z7 wrote:
       | Business and technical analysis of DeepSeek's entire R&D history
       | with extrapolations:
       | 
       | https://chatgpt.com/share/67a0d59b-d020-8001-bb88-dc9869d52b...
        
       | DoctorOetker wrote:
       | Would formalizing Wiles' proof of Fermat's Last Theorem be
       | considered deep research? Is it able to formalize it in say
       | metamath's set.mm?
       | 
       | Or is the position of OpenAI that Wiles' proof is incomplete?
        
       | bilater wrote:
       | Not quite the agent they are building but I have an open source
       | alternative that lets you use a variety of models, based on links
       | of your choice to generate reports:
       | https://github.com/btahir/open-deep-research
        
       | TheGradfather wrote:
       | The OpenAI Deep Research graph showing tool calls vs pass rate
       | reveals something fascinating about how these models handle
       | increasing amounts of information. The relationship follows a
       | logistic curve that plateaus around 16% pass rate, even as we
       | allow more tool calls.
       | 
       | This plateau behavior reflects something deeper about our current
       | approach to AI. We've built transformer architectures partly
       | inspired by simplified observations of human cognition -
       | particularly how our brains use attention mechanisms to filter
       | and process information. And like human attention, these models
       | have inherent constraints: each attention layer normalizes scores
       | to sum to 1, creating a fixed "attention budget" that must be
       | distributed across all inputs.
       | 
       | A recent paper (https://arxiv.org/abs/2501.19399) explores this
       | limitation, showing how standard attention becomes increasingly
       | diffuse with longer contexts. Their proposed "Scalable-Softmax"
       | helps maintain focused attention at longer ranges, but still
       | shows diminishing returns - pushing the ceiling higher rather
       | than eliminating it.
       | 
       | But here's the deeper question: As we push toward AGI and
       | potentially superintelligent systems, should we remain bound by
       | architectures modeled on our current understanding of human
       | cognition? The human brain's limited attention mechanism evolved
       | under specific constraints and for specific purposes. While it's
       | remarkably effective for human-level intelligence, it might be
       | fundamentally limiting for artificial systems that could
       | theoretically process information in radically different ways.
       | 
       | Looking at the Deep Research results through this lens, the
       | plateau might not just be a technical limitation to overcome, but
       | a sign that we need to fundamentally rethink how artificial
       | systems could process and integrate information. Instead of
       | trying to stretch the capabilities of attention-based
       | architectures, perhaps we need to explore entirely different
       | paradigms of information processing that aren't constrained by
       | biological analogues.
       | 
       | This isn't to dismiss the remarkable achievements of transformer
       | architectures, but rather to suggest that the path to AGI might
       | require breaking free from some of our biologically-inspired
       | assumptions. What would an architecture that processes
       | information in ways fundamentally different from human cognition
       | look like? How might it integrate and reason about information
       | without the constraints of normalized attention?
       | 
       | Would love to hear thoughts from others working on these
       | problems, particularly around novel approaches that move beyond
       | our current biological inspirations.
        
       | enknamel wrote:
       | So they did a RAG on the whole internet? Basically Google search
       | results summary but better?
        
       ___________________________________________________________________
       (page generated 2025-02-03 23:02 UTC)