[HN Gopher] I genuinely don't understand why some people are sti...
       ___________________________________________________________________
        
       I genuinely don't understand why some people are still bullish
       about LLMs
        
       Author : ksec
       Score  : 93 points
       Date   : 2025-03-27 21:22 UTC (1 hours ago)
        
 (HTM) web link (twitter.com)
 (TXT) w3m dump (twitter.com)
        
       | encypherai wrote:
       | We've had the opposite experience, especially with o3-mini using
       | Deep Research for market research & topic deep-dive tasks. The
       | sources that are pulled have never been 404 for us, and typically
       | have been highly relevant to the search prompt. It's been a huge
       | time-saver. We are just scratching the surface of how good these
       | LLMs will become at research tasks.
        
       | retrac wrote:
       | You're using them wrong. Everyone is though I can't fault you
       | specifically. Chatbot is like the worst possible application of
       | these technologies.
       | 
       | Of late, deaf tech forums are taken over by language model
       | debates over which works best for speech transcription.
       | (Multimodal language models are the the state of the art in
       | machine transcription. Everyone seems to forget that when
       | complaining they can't cite sources for scientific papers yet.)
       | The debates are sort of to the point that it's become annoying
       | how it has taken over so much space just like it has here on HN.
       | 
       | But then I remember, oh yeah, there was no such thing as live
       | machine transcription ten years ago. And now there is. And it's
       | going to continue to get better. It's already good enough to be
       | very useful in many situations. I have elsewhere complained about
       | the faults of AI models for machine transcription - in particular
       | when they make mistakes they tend to hallucinate something that
       | is superficially grammatical and coherent instead - but for a
       | single phrase in an audio transcription sporadically that's
       | sometimes tolerable. In many cases you still want a human
       | transcriber but the cost of that means that the amount of
       | transcription needed can never be satisfied.
       | 
       | It's a revolutionary technology. I think in a few years I'm going
       | have glasses that continuously narrate the sounds around me and
       | transcribe speech and it's going to be so good I can probably
       | "pass" as a hearing person in some contexts. It's hard not to get
       | a bit giddy and carried away sometimes.
        
         | Shank wrote:
         | > You're using them wrong. Everyone is though I can't fault you
         | specifically.
         | 
         | If everyone is using them wrong, I would argue that says
         | something more about them than the users. Chat-based interfaces
         | are _the_ thing that kicked LLMs into the mainstream
         | consciousness and started the cycle /trajectory we're on now.
         | If this is the wrong use case, everything the author said is
         | still true.
         | 
         | There are still applications made better by LLMs, but they are
         | a far cry from AGI/ASI in terms of being all-knowing problem
         | solvers that don't make mistakes. Language tasks like
         | transcription and translation are valuable, but by no stretch
         | do they account for the billions of dollars of spend on these
         | platforms, I would argue.
        
           | minimaxir wrote:
           | LLM providers actually have an incentive _not_ to write
           | literature on how to use LLM optimally, as that causes
           | friction which means less engagement /money spent on the
           | provider. There's also the typical tin-foil hat explanation
           | of "it's bad so you'll keep retrying it to get the LLM to
           | work which means more money for us."
        
         | sidewndr46 wrote:
         | If the goal is to layoff all the customer support and trap the
         | customer in a tarpit with no exit, LLMs are likely the best
         | choice.
        
           | Hikikomori wrote:
           | US can have fun with that. In EU well likely get laws that
           | force companies to let us talk to a human if it gets bad
           | enough.
        
         | whazor wrote:
         | I think all the technology is already in place. There are
         | already smart glasses with tiny text displays. Also smartphones
         | have more than enough processing capacity to handle live speech
         | transcription.
        
         | azemetre wrote:
         | What is the best open source live machine transcription tools
         | would you say? Know of any guides that make it easy to setup
         | locally if so?
        
       | MostlyStable wrote:
       | My experience (almost exclusively Claude), has just been so
       | different that I don't know what to say. Some of the examples are
       | the kinds of things I explicitly wouldn't expect LLMs to be
       | particularly good at so I wouldn't use them for, and others, she
       | says that it just doesn't work for her, and that experience is
       | just so different than mine that I don't know how to respond.
       | 
       | I think that there are two kinds of people who use AI: people who
       | are looking for the ways in which AIs fail (of which there are
       | still many) and people who are looking for the ways in which AIs
       | succeed (of which there are also many).
       | 
       | A lot of what I do is relatively simple one off scripting. Code
       | that doesn't need to deal with edge cases, won't be widely
       | deployed, and whose outputs are very quickly and easily
       | verifiable.
       | 
       | LLMs are almost perfect for this. It's generally faster than me
       | looking up syntax/documentation, when it's wrong it's easy to
       | tell and correct.
       | 
       | Look for the ways that AI works, and it can be a powerful tool.
       | Try and figure out where it still fails, and you will see nothing
       | but hype and hot air. Not every use case is like this, but there
       | are many.
       | 
       | -edit- Also, when she says "none of my students has ever invented
       | references that just don't exist"...all I can say is "press X to
       | doubt"
        
         | CrossVR wrote:
         | The point is that given the current valuations, being good at a
         | bunch of narrow use cases is just not good enough. It needs to
         | be able to replace humans in every role where the primary
         | output is text or speech to meet expectations.
        
           | MostlyStable wrote:
           | I don't think that "replacing humans in every role" is the
           | line for "being bullish on AI models". I think they could
           | stop development exactly where they are, and they would still
           | make pretty dramatic improvements to productivity in a lot of
           | places. For me at least, their value already exceeds the
           | $20/month I'm paying, and I'm pretty sure that way more than
           | covers inference costs.
        
             | TeMPOraL wrote:
             | > _I think they could stop development exactly where they
             | are, and they would still make pretty dramatic improvements
             | to productivity in a lot of places._
             | 
             | Yup. Not to mention, we don't even have time to figure out
             | how to effectively work with one generation of models,
             | before the next generation of models get released and rises
             | the bar. If development stopped right now, I'd still expect
             | LLMs to get better for _years_ , as people slowly figure
             | out how to use them well.
        
         | bluefirebrand wrote:
         | > Look for the ways that AI works, and it can be a powerful
         | tool. Try and figure out where it still fails, and you will see
         | nothing but hype and hot air. Not every use case is like this,
         | but there are many.
         | 
         | The problem is that I feel I am constantly being bombarded by
         | people bullish on AI saying "look how great this is" but when I
         | try to do the exact same things they are doing, it doesn't work
         | very well for me
         | 
         | Of course I am skeptical of positive claims as a result.
        
           | mattigames wrote:
           | Exactly, thanks to all the money involved in such hype the
           | incentives will always skew towards over spamming naive
           | optimism about it's features.
        
           | MostlyStable wrote:
           | I don't know what you are doing or why it's failed. Maybe my
           | primary use cases really are in the top whatever percentile
           | for AI usefulness, but it doesn't feel like it. All I know is
           | that frontier models have already been good enough for more
           | than a year to increase my productivity by a fair bit.
        
             | nonchalantsui wrote:
             | Your use case is in fact in the top whatever percentile for
             | AI usefulness. Short simple scripting that won't have to be
             | relied on due to never being widely deployed. No large
             | codebase it has to comb through, no need for thorough
             | maintenance and update management, no need for efficient
             | (and potentially rare) solutions.
             | 
             | The only use case that would beat yours is the type of
             | office worker that cannot write professional sounding
             | emails but has to send them out regularly manually.
        
               | MostlyStable wrote:
               | I fully believe it's far better at the kind of
               | coding/scripting that I do than the kind that real SWEs
               | do. If for no other reason than the coding itself that I
               | do is far far simpler and easier, so of course it's going
               | to do better at it. However, I don't really believe that
               | coding is the only use case. I think that there are a
               | whole universe of other use cases that probably also get
               | a lot of value from LLMs.
               | 
               | I think that HN has a lot of people who are working on
               | large software projects that are incredibly complex and
               | have a huge numbers of interdependencies etc., and LLMs
               | aren't quite to the point that they can very usefully
               | contribute to that except around the edges.
               | 
               | But I don't think that generalizing from that failure is
               | very useful either. Most things humans do aren't that
               | hard. There is a reason that SWE is one of the best paid
               | jobs in the country.
        
           | turtletontine wrote:
           | I literally had a developer of an open source package I'm
           | working with tell me "yeah that's a known problem, I gave up
           | on trying to fix it. You should just ask ChatGPT to fix it, I
           | bet it will immediately know the answer."
           | 
           | Annoying response of course. But I'd never used an LLM to
           | debug before, so I figured I'd give it a try.
           | 
           | First: it regurgitated a bunch of documentation and basic
           | debugging tips, which might have actually been helpful if I
           | had just encountered this problem and had put no thought into
           | debugging it yet. In reality, I had already spent hours on
           | the problem. So not helpful
           | 
           | Second: I provided some further info on environment variables
           | I thought might be the problem. It latched on to that. "Yes
           | that's your problem! These environment variables are (causing
           | the problem) because (reasons that don't make sense). Delete
           | them and that should fix things." I deleted them. It changed
           | nothing.
           | 
           | Third: It hallucinated a magic numpy function that would
           | solve my problem. I informed it this function did not exist,
           | and it wrote me a flowery apology.
           | 
           | Clearly AI coding works great for some people, but this was
           | purely an infuriating distraction. Not only did it not solve
           | my problem, it wasted my time and energy, and threw tons of
           | useless and irrelevant information at me. Bad experience.
        
             | ajkdhcb2 wrote:
             | My experiences have all been like this too. I am puzzled by
             | how some people say it works for them
        
               | simonw wrote:
               | I wrote this article precisely for people who are having
               | trouble getting good results out of LLMs for coding:
               | https://simonwillison.net/2025/Mar/11/using-llms-for-
               | code/
        
             | sdenton4 wrote:
             | This morning I was using an LLM to develop some SQL queries
             | against a database it had never seen before. I gave it a
             | starting point, and outlined what I wanted to do. It
             | proposed a solution, which was a bit wrong, mostly because
             | I hadn't given it the full schema to work with. Small
             | nudges and corrections, and we had something that worked.
             | From there, I iterated and added more features to the
             | outputs.
             | 
             | At many points, the code would have an error; to deal with
             | this, I just supply the error message, as-is to the LLM,
             | and it proposes a fix. Sometimes the fix works, and
             | sometimes I have to intervene to push the fix in the right
             | direction. It's OK - the whole process took a couple hours,
             | and probably would have been a whole day if I were doing it
             | on my own, since I usually only need to remember anything
             | about SQL syntax once every year or three.
             | 
             | A key part of the workflow, imo, was that we were working
             | in the medium of the actual code. If the code is broken, we
             | get an error, and can iterate. Asking for opinions doesn't
             | really help...
        
               | simonw wrote:
               | I often wonder if people who report that LLMs are useless
               | for code haven't cracked the fact that you need to to
               | have a conversation with it - expecting a perfect result
               | after your first prompt is setting it up for failure, the
               | real test is if you can get to a working solution after
               | iterating with it for a few rounds.
        
               | bluefirebrand wrote:
               | Why would I do this, when I can just write it from
               | scratch in less time than it takes you to have this
               | conversation with the LLM?
        
               | QuantumGood wrote:
               | It really requires self-discipline to ignore the
               | enthusiasm of the LLM as a signal for whether you are
               | moving in the direction of a solution. I blame myself for
               | lazy prompting, but have a hard time not just jumping in
               | with a quick project, hoping the LLM can get somewhere
               | with it, and not attempt things that are impossible, etc.
        
               | turtletontine wrote:
               | That makes sense, and from what I've heard this sort of
               | simple quick prototyping is where LLM coding works well.
               | The problem with my case was I'm working with multiple
               | large code bases, and couldn't pinpoint the problem to a
               | specific line, or even file. So I wasn't gonna just copy
               | multiple git repos into the chat
               | 
               | (The details: I was working with running a Bayesian
               | sampler across multiple compute nodes with MPI. There
               | seemed to be a pathological interaction between the code
               | and MPI where things looked like they were working, but
               | never actually progressed.)
        
               | SoftTalker wrote:
               | I wonder if it breaks like this: people who don't know
               | how to code find LLMs very helpful and don't realize
               | where they are wrong. People who do know immediately see
               | all the things they get wrong and they just give up and
               | say "I'll do it myself".
        
               | bluefirebrand wrote:
               | > OK - the whole process took a couple hours, and
               | probably would have been a whole day if I were doing it
               | on my own, since I usually only need to remember anything
               | about SQL syntax once every year or three
               | 
               | If you have any reasonable understanding of SQL, I
               | guarantee you could brush up on it and write it yourself
               | in less than a couple of hours unless you're trying to do
               | something _very_ complex
               | 
               | SQL is absolutely trivial to write by hand
        
             | dale_glass wrote:
             | On the other hand, when it works it's darn near magic.
             | 
             | I spent like a week trying to figure out why a livecd image
             | I was working on wasn't initializing devices correctly.
             | Read the docs, read source code, tried strace, looked at
             | the logs, found forums of people with the same problem but
             | no solution, you know the drill. In desperation I asked
             | ChatGPT. ChatGPT said "Use udevadm trigger". I did. Things
             | started working.
             | 
             | For some problems it's just very hard to express them in a
             | googleable form, especially if you're doing something weird
             | almost nobody else does.
        
               | bluefirebrand wrote:
               | Honestly this says more about how bad Google has become
               | than about how good GPT is
        
         | sidewndr46 wrote:
         | every time someone brings up "Code that doesn't need to deal
         | with edge cases" I like to point at that such code is not
         | likely to be used for anything that matters
        
           | Panzer04 wrote:
           | Is such code hard to write in the first place?
           | 
           | Automating the easy 80% _sounds_ useful, but in practice I 'm
           | not convinced that's all that helpful. Reading and putting
           | together code you didn't write is hard enough to begin with.
        
             | Ferret7446 wrote:
             | It's not hard, but it's time consuming.
        
           | brulard wrote:
           | Oh, but it is. I can have code that does something nice to
           | have, needs not to be 100% correct etc. For example, I want a
           | background for my playful webpage. Maybe a WebGL shader. It
           | might not be exactly what I asked for, but I can have it in
           | few minutes up and running. Or some non-critical internal
           | tools - like scraper for lunch menus from restaurants around
           | office. Or simple parking spot sharing app. Or any kind of
           | prototypes which in some companies are being created all the
           | time. There are so many use cases that are forgiving
           | regarding correctness and are much more sensitive to
           | development effort.
        
             | nonchalantsui wrote:
             | There is a cost burden to not being 100% correct when it
             | comes to programming. You simply have chosen to ignore that
             | burden, but it still exists for others. Whether it's for
             | example a percent of your users now getting stalled pages
             | due to the webgl shader, or your lunch scraper ddosing
             | local restaurants. They aren't actually forgiving regarding
             | correctness.
             | 
             | Which is fine for actual testing you're doing internally,
             | since that cost burden is then remedied by you fixing those
             | issues. However, no feature is as free as you're making it
             | sound, not even the "nice to have" additions that seem so
             | insignificant.
        
           | SoftTalker wrote:
           | I write code like that all the time. It's used for very
           | specific use cases, only by myself or something I've also
           | written. It's not exposed to random end users or inputs.
        
         | Sohcahtoa82 wrote:
         | > LLMs are almost perfect for this. It's generally faster than
         | me looking up syntax/documentation, when it's wrong it's easy
         | to tell and correct.
         | 
         | Exactly this.
         | 
         | I once had a function that would generate several .csv reports.
         | I wanted these reports to then be uploaded to
         | s3://my_bucket/reports/{timestamp}/ _.csv
         | 
         | I asked ChatGPT "Write a function that moves all .csv files in
         | the current directory to and old_reports directory, calls a
         | create_reports function, then uploads all the csv files in the
         | current directory to s3://my_bucket/reports/{timestamp}/_.csv
         | with the timestamp in YYYY-MM-DD format""
         | 
         | And it created the code perfectly. I knew what the correct code
         | would look like, I just couldn't be fucked to look up the exact
         | calls to boto3, whether moving files was os.move or os.rename
         | or something from shutil, and the exact way to format a
         | datetime object.
         | 
         | It created the code far faster that I would have.
         | 
         | Like, I certainly wouldn't use it to write a whole app, or even
         | a whole class, but individual blocks like this, it's great.
        
           | simonw wrote:
           | I've had so many cases exactly like your example here. If you
           | build up an intuition that knows that e.g. Claude 3.7 Sonnet
           | can write code that uses boto3, and boto3 hasn't had any
           | breaking changes that would affect S3 usage in the past ~24
           | months, you can jump straight into a prompt for this kind of
           | task.
           | 
           | It doesn't just save me a ton of time, it results in me
           | building automations that I normally wouldn't have taken on
           | at all because the time spent fiddling with os.move/boto3/etc
           | wouldn't have been worthwhile compared to other things on my
           | plate.
        
         | runjake wrote:
         | More often than not, when I inquire deeper, I often find their
         | prompting isn't very good at all.
         | 
         | "Garbage in, garbage out" as the law says.
         | 
         | Of course, it took a lot of trial and error for me to get to my
         | current level of effectiveness with LLMs. It's probably our
         | responsibility to teach these who are willing.
        
           | cool_dude85 wrote:
           | It seems hard to be bullish on LLMs as a generally useful
           | tool if the solution to problems people have is "use trial
           | and error to improve how you write your prompts, no, it's not
           | obvious how to do so, yes, it depends heavily on the exact
           | model you use."
        
             | simonw wrote:
             | You could say that about any power tool.
             | 
             | A Mitre Saw is an amazing thing to have in a woodshop, but
             | if you don't learn how to use it you're probably going to
             | cut off a finger.
             | 
             | The problem is that LLMs are power tools that are sold as
             | being so easy to use that you don't need to invest any
             | effort in learning them at all. That's extremely
             | misleading.
        
               | zeitgeistcowboy wrote:
               | So is using an LLM to write SQL for you like using a
               | mitre saw instead of a table saw? I guess the crux is
               | that you still need to do work either way.
        
         | heraldgeezer wrote:
         | >A lot of what I do is relatively simple one off scripting.
         | Code that doesn't need to deal with edge cases, won't be widely
         | deployed, and whose outputs are very quickly and easily
         | verifiable.
         | 
         | Yes somewhat. Its good for powershell/bash/cmd scripts and
         | configs but early it would make stuff up
        
         | bbor wrote:
         | Look for the ways that AI works, and it can be a powerful tool.
         | Try and figure out where it still fails, and you will see
         | nothing but hype and hot air.
         | 
         | Perfectly put, IMO.
         | 
         | I know arguments from authority aren't primary, but I think
         | this point highlights some important context: Dr. Hossenfelder
         | has gained international renown by publishing clickbait-y
         | YouTube videos that ostensibly debunk scientific and
         | technological advances of all kinds. She's clearly educated and
         | thoughtful (not to mention otherwise gainfully employed), but
         | her whole public persona kinda relies on assuming the
         | exclusively-critical standpoint you mention.
         | 
         | I doubt she necessarily feels _indebted_ to her large audience
         | expecting this take (it 's not new...), but that certainly does
         | seem like a hard cognitive habit to break.
        
         | cjf101 wrote:
         | It's a weird circle with these things. If you _can't_ do the
         | task you are using the LLM for, you probably shouldn't.
         | 
         | But if you can do the task well enough to at least recognize
         | likely-to-be-correct output, then you can get a lot done in
         | less time than you would do it without their assistance.
         | 
         | Is that worth the second order effects we're seeing? I'm not
         | convinced, but it's definitely changed the way we do work.
        
         | asdev wrote:
         | all fun and games until your AI generated script deletes the
         | production database. I think that's the point, fault tolerance
         | in academic and financial settings is too high for LLMs to be
         | useful
        
       | latemedium wrote:
       | My experience is starkly different. Today I used LLMs to:
       | 
       | 1. Write python code for a new type of loss function I was
       | considering
       | 
       | 2. Perform lots of annoying CSV munging ("split this CSV into 4
       | equal parts", "convert paths in this column into absolute paths",
       | "combine these and then split into 4 distinct subsets based on
       | this field.." - they're great for that)
       | 
       | 3. Expedite some basic shell operations like "generate softlinks
       | for 100 randomly selected files in this directory"
       | 
       | 4. Generate some summary plots of the data in the files I was
       | working with
       | 
       | 5. Not to mention extensive use in Cursor & GH Copilot
       | 
       | The tool (Claude 3.7 mostly, integrated with my shell so it can
       | execute shell commands and run python locally) worked great in
       | all cases. Yes I could've done most of it myself, but I
       | personally hate CSV munging and bulk file manipulations and its
       | super nice to delegate that stuff to an LLM agent
       | 
       | edit: formatting
        
         | mnky9800n wrote:
         | How did you integrate Claude into your shell
        
           | airstrike wrote:
           | Claude Code is available directly from Anthropic, but you
           | have to request an invite as it's in "Research Preview"
           | 
           | There are third party tools that do the same, though
        
           | latemedium wrote:
           | I hacked something together a while back - a hotkey toggles
           | between standard terminal mode and LLM mode. LLM mode
           | interacts with Claude, and has functions / tool calls to run
           | shell commands, python code, web search, clipboard, and a few
           | other things. For routine data science tasks it's been super
           | useful. Claude 3.7 was a big step forward because it will
           | often examine files before it begins manipulating them and
           | double-checks that things were done correctly afterwards
           | (without prompting!). For me this works a lot better than
           | other shell-integration solutions like Warp
        
           | simonw wrote:
           | I wrote my own tool for that a while back as an LLM plugin,
           | so I can do this:                   llm cmd extract first
           | frame of movie.mp4 as a jpeg using ffmpeg
           | 
           | I use that all the time, it works really well (defaulting to
           | GPT-4o-mini because it's so cheap, but it works with Claude
           | too): https://simonwillison.net/2024/Mar/26/llm-cmd/
        
         | kwertyoowiyop wrote:
         | These seem like fine use cases: trivial boilerplate stuff you'd
         | otherwise have to search for and then munge to fit your exact
         | need. An LLM can often do both steps for you. If it doesn't
         | work, you'll know immediately and you can probably figure out
         | whether it's a quick fix or if the LLM is completely off-base.
        
         | zeroonetwothree wrote:
         | That's fair but it's totally different use cases than the
         | linked post discusses.
        
       | jongjong wrote:
       | People who don't work in tech have no idea how hard it is to do
       | certain things at scale. Skilled tech people are severely
       | underappreciated.
       | 
       | From a sub-tweet:
       | 
       | >> no LLM should ever output a url that gives a 404 error. How
       | hard can it be?
       | 
       | As a developer, I'm just imagining a server having to call up all
       | the URLs to check that they still exist (and the extra
       | costs/latency incurred there)... And if any URLs are missing,
       | getting the AI to re-generate a different variant of the
       | response, until you find one which does not contain the missing
       | links.
       | 
       | And no, you can't do it from the client side either... It would
       | just be confusing if you removed invalid URLs from the middle of
       | the AI's sentence without re-generating the sentence.
       | 
       | You almost need to get the LLM to engineer/pre-process its own
       | prompts in a way which guesses what the user is thinking in order
       | to produce great responses...
       | 
       | Worse than that though... A fundamental problem of 'prompt
       | engineering' is that people (especially non-tech people) often
       | don't actually fully understand what they're asking.
       | Contradictions in requirements are extremely common. When
       | building software especially, people often have a vague idea of
       | what they want... They strongly believe that they have a
       | perfectly clear idea but once you scope out the feature in
       | detail, mapping out complex UX interactions, they start to see
       | all these necessary tradeoffs and limitations rise to the surface
       | and suddenly they realize that they were asking for something
       | they don't want.
       | 
       | It's hard to understand your own needs precisely; even harder to
       | communicate them.
        
         | readthenotes1 wrote:
         | "How hard can it be?"
         | 
         | If I recall correctly, that is one of Dilbert's management
         | axioms: if I don't understand it it cannot be difficult
        
       | airstrike wrote:
       | I wrote an AI assistant which generates working spreadsheets with
       | formulas and working presentations with neatly laid out elements
       | and styles. It's a huge productivity gain relative to starting
       | from a blank page.
       | 
       | I think LLMs work best when they are used as a "creative" tool.
       | They're good for the brainstorming part of a task, not for the
       | finishing touches.
       | 
       | They are too unreliable to be put in front of your users. People
       | don't want to talk to unpredictable chatbots. Yes, they can be
       | useful in customer service chats because you can put them on
       | rails and map natural language to predetermined actions. But
       | generally speaking I think LLMs are most effective when used _by_
       | someone who's piloting them instead of wrapped in a service
       | offered _to_ someone.
       | 
       | I do think we've squeezed 90%+ of what we could from current
       | models. Throwing _more_ dollars of compute at training or
       | inference won 't make much difference. The next "GPT moment" will
       | come from some sufficiently novel approach.
        
       | joegibbs wrote:
       | Because it's not a scientific research tool, it's a most likely
       | next text generator. It doesn't keep a database of ingested
       | information with source URLs. There are plenty of scientific
       | research tools but something that just outputs text based on your
       | input is no good for it.
       | 
       | I'm sure that in the future there will be a really good search
       | tool that utilises an LLM but for now a plain model just isn't
       | designed for that. There are a ton of other uses for them, so I
       | don't think that we should discount them entirely based on their
       | ability to output citations.
        
       | GaggiX wrote:
       | I use them everyday and they work greatly, I even made a command
       | (using Claude, actually Claude made everything in that script)
       | that calls Gemini from the terminal so that I can ask for
       | question related to the shell directly there, just doing a: ai
       | "how can I convert a webp to a png", the system prompt asks to be
       | brief, using markdown (it does display nicely), that most
       | question are related to Linux and it provides information about
       | my OS (uname -a), the last code block is also copied in the
       | clipboard, super useful, I imagine there are plenty online of
       | similar utilities.
        
       | saaaaaam wrote:
       | I've used Claude today to:
       | 
       | Write code to pull down a significant amount of public data using
       | an open API. (That took about 30 seconds - I just gave it the
       | swagger file and said "here's what I want")
       | 
       | Get the data (an hour or so), clean the data (barely any time,
       | gave it some samples, it wrote the code), used the cleaned data
       | to query another API, combined the data sources, pulled down a
       | bunch of PDFs relating to the data, had the AI write code to use
       | tesseract to extract data from the PDFs, and used that to build a
       | dashboard. That's a mini product for my users.
       | 
       | I also had a play with Mistral's OCR and have tested a few things
       | using that against the data. When I was out walking my dogs I
       | thought about that more, and have come up with a nice workflow
       | for a problem I had, which I'll test in more detail next week.
       | 
       | That was all whole doing an entirely different series of tasks,
       | on calls, in meetings. I literally checked the progress a few
       | times and wrote a new prompt or copy/pasted some stuff in from
       | dev tools.
       | 
       | For the calls I was on, I took the recording of those calls,
       | passed them into my local instance whisper, fed the transcript
       | into Claude with a prompt I use to extract action points, pasted
       | those into a google doc, circulated them.
       | 
       | One of the calls was an interview with an expert. The transcript
       | + another prompt has given me the basis for an article (bulleted
       | narrative + key quotes) - I will refine that tomorrow, and write
       | the article, using a detailed prompt based on my own writing
       | style and tone.
       | 
       | I needed to gather data for a project I'm involved in, so had
       | Claude write a handful of scrapers for me (HTML source > here is
       | what I need).
       | 
       | I downloaded two podcasts I need to listen to - but only need to
       | listen to five minutes of each - and fed them into whisper then
       | found the exact bits I needed and read the extracts rather than
       | listening to tedious podcast waffle.
       | 
       | I turned an article I'd written into an audio file using
       | elevenlabs, as a test for something a client asked me about
       | earlier this week.
       | 
       | I achieved about three times as much today as I would have done a
       | year ago. And finished work at 3pm.
       | 
       | So yeah, I don't understand why people are so bullish about LLMs.
       | Who knows?
        
         | kilolima wrote:
         | Yuck. Do your users know that they are reading recycled LLM
         | content? Is this long winded post generated by an LLM?
        
         | Paratoner wrote:
         | Did you also do that while mewing and listening to an AI
         | abridged audiobook version of the laws of power in chinese?
         | Don't forget your morning ice face dunks.
        
       | doctoboggan wrote:
       | Why do people who don't like using LLMs keep insisting they are
       | useless for the rest of us? If you don't like to use them, then
       | simply don't use them.
       | 
       | I use them almost daily in my job and get tremendous use out of
       | them. I guess you could accuse me of lying, but what do I stand
       | to gain from that?
       | 
       | I've also seem people claim that only people who don't know how
       | to code or people doing super simple done a million times apps
       | can get value out of LLMs. I don't believe that applies to my
       | situation, but even if it did, so what? I do real work for a real
       | company delivering real value, and the LLM delivers value to me.
       | It's really as simple as that.
        
         | biker142541 wrote:
         | I don't think it's a matter or liking or not. The use cases
         | just differ considerably, and tools and not as useful or
         | applicable across those. THe OP's use case is probably one of
         | the worst possible for LLMs right now, imo...
        
       | harrall wrote:
       | I am neither bullish or bearish. LLM is a tool.
       | 
       | It's a hammer -- sometimes it works well. It summarizes the user
       | reviews on a site... cool, not perfect, but useful.
       | 
       | And like every tool, it is useless for 90% of life's situations.
       | 
       | And I know when it's useful because I've already tried a hammer
       | on 1000 things and have figured out what I should be using a
       | hammer on.
        
         | rufus_foreman wrote:
         | >> I am neither bullish or bearish. LLM is a tool...It's a
         | hammer
         | 
         | If someone says, "This new type of hammer will increase
         | productivity in the construction industry by 25%", it's
         | something else in addition to being a tool. It's either a lie,
         | or it's an incredible advance in technology.
        
       | belter wrote:
       | "Yes, I have tried Gemini, and actually it was even worse in that
       | it frequently refuses to even search for a source and instead
       | gives me instructions for how to do it myself. Stopped using it
       | for that reason."
       | 
       | Thank you Sabine. Every time I have mentioned Gemini is the
       | worst, and not even worth of consideration, I have been bombarded
       | with downvotes, and told I am using it wrong.
        
       | throwawa14223 wrote:
       | My experience mirrors hers. Asking questions is worthless because
       | the answers are either 404 links, telling me how to use a search
       | engine, or just flat out wrong and the code generated compiles
       | maybe one time out of ten and when it does the implementation is
       | usually poor.
       | 
       | When I evaluate against areas I possess professional expertise I
       | become convinced LLMs produce the Gell Mann amnesia effect for
       | any area I don't know.
        
       | islewis wrote:
       | > I genuinely don't understand why some people are still bullish
       | about LLMs.
       | 
       | I don't believe OP's thesis is properly backed by the rest of his
       | tweet, which seems to boil down to "LLM's can't properly cite
       | links".
       | 
       | If LLM's performing poorly on an arbitrary small-scoped test case
       | makes you bearish on the whole field, I don't think that falls on
       | the LLM's.
        
       | crazygringo wrote:
       | If there's one common thread across LLM criticisms, it's that
       | they're not perfect.
       | 
       | These critics don't seem to have learned the lesson that _the
       | perfect is the enemy of the good_.
       | 
       | I use ChatGPT all the time for academic research. Does it
       | fabricate references? Absolutely, maybe about a third of the
       | time. But has it pointed me to important research papers I might
       | never have found otherwise? _Absolutely_.
       | 
       | The rate of inaccuracies and falsehoods doesn't matter. What
       | matters is, is it saving you time and increasing your
       | productivity. Verifying the accuracy of its statements is _easy_.
       | While finding the knowledge it spits out in the first place is
       | _hard_. The net balance is a _huge positive_.
       | 
       | People are bullish on LLM's because they can save you days' worth
       | of work, like every day. My research productivity has gone _way_
       | up with ChatGPT -- asking it to explain ideas, related concepts,
       | relevant papers, and so forth. It 's amazing.
        
       | simonw wrote:
       | The most interesting thing about this post is how it reinforces
       | how terrible the usability of LLMs still is today:
       | 
       | "I ask them to give me a source for an alleged quote, I click on
       | the link, it returns a 404 error. I Google for the alleged quote,
       | it doesn't exist. They reference a scientific publication, I look
       | it up, it doesn't exist."
       | 
       | To experienced LLM users that's not surprising at all - providing
       | citations, sources for quotes, useful URLs are all things that
       | they are demonstrably terrible at.
       | 
       | But it's a computer! Telling people "this advanced computer
       | system cannot reliably look up facts" goes against everything
       | computers have been good at for the last 40+ years.
        
       | puppycodes wrote:
       | Here's one simple reason:
       | 
       | I have a very specific esoteric question like: "What material is
       | both electrically conductive and good at blocking sound?" I could
       | type this into google and sift through the titles and short
       | descriptions of websites and eventually maybe find an answer, or
       | I can put the question to the LLM and instantly get an answer
       | that I can then research further to confirm.
       | 
       | This is significantly faster, more informative, more efficient,
       | and a rewarding experience.
       | 
       | As others have said, its a tool. A tool is as good as how you use
       | it. If you expect to build a house by yelling at your tools I
       | wouldn't be bullish either.
        
         | krapp wrote:
         | I mean, the first link I got when I pasted that in is probably
         | the Stack Exchange thread you would use to research further,
         | along with other sources, which do seem relevant to the query.
         | 
         | I don't see how an LLM is significantly faster or more
         | informative, since you still have to do the legwork to validate
         | the answer. I guess if you're google-phobic (which a lot of
         | people seem to be, especially on HN) then I can see how it's
         | more rewarding to put it off until later in the process.
        
           | puppycodes wrote:
           | Its just an example and you can "well actually" until the
           | cows come home, but your missing the point. I'm sure there
           | are things you have found hard to google. Also you are most
           | likely (as your commenting on hacker news) _not_ a good
           | representation of the majority of the world.
           | 
           | Also idk if you have noticed, but google is clearly using LLM
           | technology in conjunction with its search results, so the
           | assumption they are just using not LLMs to inform or modify
           | its result set I think is naive or at very best just a
           | personal opinion.
        
           | puppycodes wrote:
           | you can boil it down to this, whats easier for most people?
           | looking through websites and search engine results for
           | answers or speaking in plain language? The answer is pretty
           | obvious.
           | 
           | The validity of the answers is not 1:1 with its potential
           | profitability.
           | 
           | Like James Baldwin said "people love answers, but hate
           | questions."
           | 
           | getting an answer faster is exponentially better than getting
           | the more precise, more right, more nuanced answer for most
           | people every time. Doing the due dilligence is smart but its
           | also after the fact.
        
       | xiphias2 wrote:
       | ,,By my personal estimate currently GPT 4o DeepResearch is the
       | best one. ''
       | 
       | If the o3 based 3 month old strongest model is the best one, it's
       | a proof that there were quite significant improvements in the
       | last 2 years.
       | 
       | I can't name any other technology that improved as much in 2
       | years.
       | 
       | O1 and o1 pro helped me with filing tax returns and answered me
       | questions that (probably quite bad) tax accountants (and less
       | smart models) weren't able to (of course I read the referenced
       | laws, I don't trust the output either).
        
       | meowface wrote:
       | AI coding overall still seems to be underrated by the average
       | developer.
       | 
       | They try to add a new feature or change some behavior in a large
       | existing codebase and it does something dumb and they write it
       | off as a waste of time for that use case. And that's
       | understandable. But if they had tweaked the prompt just a bit it
       | actually might've done it flawlessly.
       | 
       | It requires patience and learning the best way to guide it and
       | iterate with it when it does something silly.
       | 
       | Although you undoubtedly will lose some time re-attempting
       | prompts and fixing mistakes and poor design choices, on net I
       | believe the frontier models can currently make development much
       | more productive in almost any codebase.
        
       | jwrallie wrote:
       | The author mentioned Gemini sometimes refusing to do something.
       | 
       | I've recently been using Gemini (mostly 2.0 flash) a lot and I've
       | noticed it sometimes will challenge me to try doing something by
       | myself. Maybe it's something in my system prompt or the way I
       | worded the request itself. I am a long time user of 4o so it felt
       | annoying at first.
       | 
       | Since my purpose was to learn how to do something, being open
       | minded I tried to comply with the request and I can say that...
       | it's being a really great experience in terms of retention of
       | knowledge. Even if I'm making mistakes Gemini will point them out
       | and explain it nicely.
        
       | BSOhealth wrote:
       | Like others here, I use it to code (no longer a professional
       | engineer, but keep side projects).
       | 
       | As soon as LLMs were introduced into the IDE it began to feeling
       | like LLM autocomplete was almost reading my mind. With some
       | context built up over a few hundred lines of initial
       | architecture, autocomplete now sees around the same corners I am.
       | It's more than just "solve this contrived puzzle" or "write
       | snake". It combines the subject matter use case (informed by
       | variable and type naming) underlying the architecture and
       | sometimes produces really breathtaking and productive results.
       | Like I said, it took some time but when it happened, it was
       | pretty shocking.
        
       | kriro wrote:
       | She's a scientist. In that area LLMs are quite useful in my
       | opinion and part of my daily workflow. Quick scripts that use
       | APIs to get data, cleaning the data and converting it. Quickly
       | working with polar data frames. Dumb daily stuff like "take this
       | data from my CSV file and turn it into a Latex table"...but most
       | importantly freeing up time from tedious administrative tasks
       | (let's not go into detail here).
       | 
       | Also great for brainstorming and quick drafting grant proposals.
       | Anything prototyping and quickly glued together I'll go for LLMs
       | (or LLM agents). They are no substitute for your own brain
       | though.
       | 
       | I'm also curious about the hallucinated sources. I've recently
       | read some papers on using LLM-agents to conduct structured
       | literature reviews and they do it quite well and fairly
       | reproducible. I'm quite willing to build some LLM-agents to
       | reproduce my literature review process in the near future since
       | it's fairly algorithmic. Check for surveys and reviews on the
       | topic, scan for interesting papers within, check sources of
       | sources, go through A tier conference proceedings for the last X
       | years and find relevant papers. Rinse, repeat.
       | 
       | I'm mostly bullish because of LLM-agents, not because of using
       | stock models with the default chat interface.
        
       | marcuschong wrote:
       | The key for LLM productivity, it seems to me, is grounding. Let
       | me give you my last example, from something I've been working on.
       | 
       | I just updated my company commercial PPT. ChatGPT helped me with:
       | - Deep Research great examples and references of such
       | presentations. - Restructure my argument and slides according to
       | some articles I found on the previous step, and thought were
       | pretty good. - Come up with copy for each slide. - Iterate new
       | ideas as I was progressing.
       | 
       | Now, without proper context and grounding, LLMs wouldn't be so
       | helpful at this task, because they don't know my company,
       | clients, product and strategy, and would be generic at best. The
       | key: I provided it with my support portal documentation and a
       | brain dump I recorded to text on ChatGPT with key strategic
       | information about my company. Those are two bits of info I keep
       | always around, so ChatGPT can help me with many tasks in the
       | company.
       | 
       | From that grounding to the final PPT, it's pretty much a trivial
       | and boring _transformation_ task that would have cost me many,
       | many hours to do.
        
       ___________________________________________________________________
       (page generated 2025-03-27 23:01 UTC)