[HN Gopher] I genuinely don't understand why some people are sti...
___________________________________________________________________
I genuinely don't understand why some people are still bullish
about LLMs
Author : ksec
Score : 93 points
Date : 2025-03-27 21:22 UTC (1 hours ago)
(HTM) web link (twitter.com)
(TXT) w3m dump (twitter.com)
| encypherai wrote:
| We've had the opposite experience, especially with o3-mini using
| Deep Research for market research & topic deep-dive tasks. The
| sources that are pulled have never been 404 for us, and typically
| have been highly relevant to the search prompt. It's been a huge
| time-saver. We are just scratching the surface of how good these
| LLMs will become at research tasks.
| retrac wrote:
| You're using them wrong. Everyone is though I can't fault you
| specifically. Chatbot is like the worst possible application of
| these technologies.
|
| Of late, deaf tech forums are taken over by language model
| debates over which works best for speech transcription.
| (Multimodal language models are the the state of the art in
| machine transcription. Everyone seems to forget that when
| complaining they can't cite sources for scientific papers yet.)
| The debates are sort of to the point that it's become annoying
| how it has taken over so much space just like it has here on HN.
|
| But then I remember, oh yeah, there was no such thing as live
| machine transcription ten years ago. And now there is. And it's
| going to continue to get better. It's already good enough to be
| very useful in many situations. I have elsewhere complained about
| the faults of AI models for machine transcription - in particular
| when they make mistakes they tend to hallucinate something that
| is superficially grammatical and coherent instead - but for a
| single phrase in an audio transcription sporadically that's
| sometimes tolerable. In many cases you still want a human
| transcriber but the cost of that means that the amount of
| transcription needed can never be satisfied.
|
| It's a revolutionary technology. I think in a few years I'm going
| have glasses that continuously narrate the sounds around me and
| transcribe speech and it's going to be so good I can probably
| "pass" as a hearing person in some contexts. It's hard not to get
| a bit giddy and carried away sometimes.
| Shank wrote:
| > You're using them wrong. Everyone is though I can't fault you
| specifically.
|
| If everyone is using them wrong, I would argue that says
| something more about them than the users. Chat-based interfaces
| are _the_ thing that kicked LLMs into the mainstream
| consciousness and started the cycle /trajectory we're on now.
| If this is the wrong use case, everything the author said is
| still true.
|
| There are still applications made better by LLMs, but they are
| a far cry from AGI/ASI in terms of being all-knowing problem
| solvers that don't make mistakes. Language tasks like
| transcription and translation are valuable, but by no stretch
| do they account for the billions of dollars of spend on these
| platforms, I would argue.
| minimaxir wrote:
| LLM providers actually have an incentive _not_ to write
| literature on how to use LLM optimally, as that causes
| friction which means less engagement /money spent on the
| provider. There's also the typical tin-foil hat explanation
| of "it's bad so you'll keep retrying it to get the LLM to
| work which means more money for us."
| sidewndr46 wrote:
| If the goal is to layoff all the customer support and trap the
| customer in a tarpit with no exit, LLMs are likely the best
| choice.
| Hikikomori wrote:
| US can have fun with that. In EU well likely get laws that
| force companies to let us talk to a human if it gets bad
| enough.
| whazor wrote:
| I think all the technology is already in place. There are
| already smart glasses with tiny text displays. Also smartphones
| have more than enough processing capacity to handle live speech
| transcription.
| azemetre wrote:
| What is the best open source live machine transcription tools
| would you say? Know of any guides that make it easy to setup
| locally if so?
| MostlyStable wrote:
| My experience (almost exclusively Claude), has just been so
| different that I don't know what to say. Some of the examples are
| the kinds of things I explicitly wouldn't expect LLMs to be
| particularly good at so I wouldn't use them for, and others, she
| says that it just doesn't work for her, and that experience is
| just so different than mine that I don't know how to respond.
|
| I think that there are two kinds of people who use AI: people who
| are looking for the ways in which AIs fail (of which there are
| still many) and people who are looking for the ways in which AIs
| succeed (of which there are also many).
|
| A lot of what I do is relatively simple one off scripting. Code
| that doesn't need to deal with edge cases, won't be widely
| deployed, and whose outputs are very quickly and easily
| verifiable.
|
| LLMs are almost perfect for this. It's generally faster than me
| looking up syntax/documentation, when it's wrong it's easy to
| tell and correct.
|
| Look for the ways that AI works, and it can be a powerful tool.
| Try and figure out where it still fails, and you will see nothing
| but hype and hot air. Not every use case is like this, but there
| are many.
|
| -edit- Also, when she says "none of my students has ever invented
| references that just don't exist"...all I can say is "press X to
| doubt"
| CrossVR wrote:
| The point is that given the current valuations, being good at a
| bunch of narrow use cases is just not good enough. It needs to
| be able to replace humans in every role where the primary
| output is text or speech to meet expectations.
| MostlyStable wrote:
| I don't think that "replacing humans in every role" is the
| line for "being bullish on AI models". I think they could
| stop development exactly where they are, and they would still
| make pretty dramatic improvements to productivity in a lot of
| places. For me at least, their value already exceeds the
| $20/month I'm paying, and I'm pretty sure that way more than
| covers inference costs.
| TeMPOraL wrote:
| > _I think they could stop development exactly where they
| are, and they would still make pretty dramatic improvements
| to productivity in a lot of places._
|
| Yup. Not to mention, we don't even have time to figure out
| how to effectively work with one generation of models,
| before the next generation of models get released and rises
| the bar. If development stopped right now, I'd still expect
| LLMs to get better for _years_ , as people slowly figure
| out how to use them well.
| bluefirebrand wrote:
| > Look for the ways that AI works, and it can be a powerful
| tool. Try and figure out where it still fails, and you will see
| nothing but hype and hot air. Not every use case is like this,
| but there are many.
|
| The problem is that I feel I am constantly being bombarded by
| people bullish on AI saying "look how great this is" but when I
| try to do the exact same things they are doing, it doesn't work
| very well for me
|
| Of course I am skeptical of positive claims as a result.
| mattigames wrote:
| Exactly, thanks to all the money involved in such hype the
| incentives will always skew towards over spamming naive
| optimism about it's features.
| MostlyStable wrote:
| I don't know what you are doing or why it's failed. Maybe my
| primary use cases really are in the top whatever percentile
| for AI usefulness, but it doesn't feel like it. All I know is
| that frontier models have already been good enough for more
| than a year to increase my productivity by a fair bit.
| nonchalantsui wrote:
| Your use case is in fact in the top whatever percentile for
| AI usefulness. Short simple scripting that won't have to be
| relied on due to never being widely deployed. No large
| codebase it has to comb through, no need for thorough
| maintenance and update management, no need for efficient
| (and potentially rare) solutions.
|
| The only use case that would beat yours is the type of
| office worker that cannot write professional sounding
| emails but has to send them out regularly manually.
| MostlyStable wrote:
| I fully believe it's far better at the kind of
| coding/scripting that I do than the kind that real SWEs
| do. If for no other reason than the coding itself that I
| do is far far simpler and easier, so of course it's going
| to do better at it. However, I don't really believe that
| coding is the only use case. I think that there are a
| whole universe of other use cases that probably also get
| a lot of value from LLMs.
|
| I think that HN has a lot of people who are working on
| large software projects that are incredibly complex and
| have a huge numbers of interdependencies etc., and LLMs
| aren't quite to the point that they can very usefully
| contribute to that except around the edges.
|
| But I don't think that generalizing from that failure is
| very useful either. Most things humans do aren't that
| hard. There is a reason that SWE is one of the best paid
| jobs in the country.
| turtletontine wrote:
| I literally had a developer of an open source package I'm
| working with tell me "yeah that's a known problem, I gave up
| on trying to fix it. You should just ask ChatGPT to fix it, I
| bet it will immediately know the answer."
|
| Annoying response of course. But I'd never used an LLM to
| debug before, so I figured I'd give it a try.
|
| First: it regurgitated a bunch of documentation and basic
| debugging tips, which might have actually been helpful if I
| had just encountered this problem and had put no thought into
| debugging it yet. In reality, I had already spent hours on
| the problem. So not helpful
|
| Second: I provided some further info on environment variables
| I thought might be the problem. It latched on to that. "Yes
| that's your problem! These environment variables are (causing
| the problem) because (reasons that don't make sense). Delete
| them and that should fix things." I deleted them. It changed
| nothing.
|
| Third: It hallucinated a magic numpy function that would
| solve my problem. I informed it this function did not exist,
| and it wrote me a flowery apology.
|
| Clearly AI coding works great for some people, but this was
| purely an infuriating distraction. Not only did it not solve
| my problem, it wasted my time and energy, and threw tons of
| useless and irrelevant information at me. Bad experience.
| ajkdhcb2 wrote:
| My experiences have all been like this too. I am puzzled by
| how some people say it works for them
| simonw wrote:
| I wrote this article precisely for people who are having
| trouble getting good results out of LLMs for coding:
| https://simonwillison.net/2025/Mar/11/using-llms-for-
| code/
| sdenton4 wrote:
| This morning I was using an LLM to develop some SQL queries
| against a database it had never seen before. I gave it a
| starting point, and outlined what I wanted to do. It
| proposed a solution, which was a bit wrong, mostly because
| I hadn't given it the full schema to work with. Small
| nudges and corrections, and we had something that worked.
| From there, I iterated and added more features to the
| outputs.
|
| At many points, the code would have an error; to deal with
| this, I just supply the error message, as-is to the LLM,
| and it proposes a fix. Sometimes the fix works, and
| sometimes I have to intervene to push the fix in the right
| direction. It's OK - the whole process took a couple hours,
| and probably would have been a whole day if I were doing it
| on my own, since I usually only need to remember anything
| about SQL syntax once every year or three.
|
| A key part of the workflow, imo, was that we were working
| in the medium of the actual code. If the code is broken, we
| get an error, and can iterate. Asking for opinions doesn't
| really help...
| simonw wrote:
| I often wonder if people who report that LLMs are useless
| for code haven't cracked the fact that you need to to
| have a conversation with it - expecting a perfect result
| after your first prompt is setting it up for failure, the
| real test is if you can get to a working solution after
| iterating with it for a few rounds.
| bluefirebrand wrote:
| Why would I do this, when I can just write it from
| scratch in less time than it takes you to have this
| conversation with the LLM?
| QuantumGood wrote:
| It really requires self-discipline to ignore the
| enthusiasm of the LLM as a signal for whether you are
| moving in the direction of a solution. I blame myself for
| lazy prompting, but have a hard time not just jumping in
| with a quick project, hoping the LLM can get somewhere
| with it, and not attempt things that are impossible, etc.
| turtletontine wrote:
| That makes sense, and from what I've heard this sort of
| simple quick prototyping is where LLM coding works well.
| The problem with my case was I'm working with multiple
| large code bases, and couldn't pinpoint the problem to a
| specific line, or even file. So I wasn't gonna just copy
| multiple git repos into the chat
|
| (The details: I was working with running a Bayesian
| sampler across multiple compute nodes with MPI. There
| seemed to be a pathological interaction between the code
| and MPI where things looked like they were working, but
| never actually progressed.)
| SoftTalker wrote:
| I wonder if it breaks like this: people who don't know
| how to code find LLMs very helpful and don't realize
| where they are wrong. People who do know immediately see
| all the things they get wrong and they just give up and
| say "I'll do it myself".
| bluefirebrand wrote:
| > OK - the whole process took a couple hours, and
| probably would have been a whole day if I were doing it
| on my own, since I usually only need to remember anything
| about SQL syntax once every year or three
|
| If you have any reasonable understanding of SQL, I
| guarantee you could brush up on it and write it yourself
| in less than a couple of hours unless you're trying to do
| something _very_ complex
|
| SQL is absolutely trivial to write by hand
| dale_glass wrote:
| On the other hand, when it works it's darn near magic.
|
| I spent like a week trying to figure out why a livecd image
| I was working on wasn't initializing devices correctly.
| Read the docs, read source code, tried strace, looked at
| the logs, found forums of people with the same problem but
| no solution, you know the drill. In desperation I asked
| ChatGPT. ChatGPT said "Use udevadm trigger". I did. Things
| started working.
|
| For some problems it's just very hard to express them in a
| googleable form, especially if you're doing something weird
| almost nobody else does.
| bluefirebrand wrote:
| Honestly this says more about how bad Google has become
| than about how good GPT is
| sidewndr46 wrote:
| every time someone brings up "Code that doesn't need to deal
| with edge cases" I like to point at that such code is not
| likely to be used for anything that matters
| Panzer04 wrote:
| Is such code hard to write in the first place?
|
| Automating the easy 80% _sounds_ useful, but in practice I 'm
| not convinced that's all that helpful. Reading and putting
| together code you didn't write is hard enough to begin with.
| Ferret7446 wrote:
| It's not hard, but it's time consuming.
| brulard wrote:
| Oh, but it is. I can have code that does something nice to
| have, needs not to be 100% correct etc. For example, I want a
| background for my playful webpage. Maybe a WebGL shader. It
| might not be exactly what I asked for, but I can have it in
| few minutes up and running. Or some non-critical internal
| tools - like scraper for lunch menus from restaurants around
| office. Or simple parking spot sharing app. Or any kind of
| prototypes which in some companies are being created all the
| time. There are so many use cases that are forgiving
| regarding correctness and are much more sensitive to
| development effort.
| nonchalantsui wrote:
| There is a cost burden to not being 100% correct when it
| comes to programming. You simply have chosen to ignore that
| burden, but it still exists for others. Whether it's for
| example a percent of your users now getting stalled pages
| due to the webgl shader, or your lunch scraper ddosing
| local restaurants. They aren't actually forgiving regarding
| correctness.
|
| Which is fine for actual testing you're doing internally,
| since that cost burden is then remedied by you fixing those
| issues. However, no feature is as free as you're making it
| sound, not even the "nice to have" additions that seem so
| insignificant.
| SoftTalker wrote:
| I write code like that all the time. It's used for very
| specific use cases, only by myself or something I've also
| written. It's not exposed to random end users or inputs.
| Sohcahtoa82 wrote:
| > LLMs are almost perfect for this. It's generally faster than
| me looking up syntax/documentation, when it's wrong it's easy
| to tell and correct.
|
| Exactly this.
|
| I once had a function that would generate several .csv reports.
| I wanted these reports to then be uploaded to
| s3://my_bucket/reports/{timestamp}/ _.csv
|
| I asked ChatGPT "Write a function that moves all .csv files in
| the current directory to and old_reports directory, calls a
| create_reports function, then uploads all the csv files in the
| current directory to s3://my_bucket/reports/{timestamp}/_.csv
| with the timestamp in YYYY-MM-DD format""
|
| And it created the code perfectly. I knew what the correct code
| would look like, I just couldn't be fucked to look up the exact
| calls to boto3, whether moving files was os.move or os.rename
| or something from shutil, and the exact way to format a
| datetime object.
|
| It created the code far faster that I would have.
|
| Like, I certainly wouldn't use it to write a whole app, or even
| a whole class, but individual blocks like this, it's great.
| simonw wrote:
| I've had so many cases exactly like your example here. If you
| build up an intuition that knows that e.g. Claude 3.7 Sonnet
| can write code that uses boto3, and boto3 hasn't had any
| breaking changes that would affect S3 usage in the past ~24
| months, you can jump straight into a prompt for this kind of
| task.
|
| It doesn't just save me a ton of time, it results in me
| building automations that I normally wouldn't have taken on
| at all because the time spent fiddling with os.move/boto3/etc
| wouldn't have been worthwhile compared to other things on my
| plate.
| runjake wrote:
| More often than not, when I inquire deeper, I often find their
| prompting isn't very good at all.
|
| "Garbage in, garbage out" as the law says.
|
| Of course, it took a lot of trial and error for me to get to my
| current level of effectiveness with LLMs. It's probably our
| responsibility to teach these who are willing.
| cool_dude85 wrote:
| It seems hard to be bullish on LLMs as a generally useful
| tool if the solution to problems people have is "use trial
| and error to improve how you write your prompts, no, it's not
| obvious how to do so, yes, it depends heavily on the exact
| model you use."
| simonw wrote:
| You could say that about any power tool.
|
| A Mitre Saw is an amazing thing to have in a woodshop, but
| if you don't learn how to use it you're probably going to
| cut off a finger.
|
| The problem is that LLMs are power tools that are sold as
| being so easy to use that you don't need to invest any
| effort in learning them at all. That's extremely
| misleading.
| zeitgeistcowboy wrote:
| So is using an LLM to write SQL for you like using a
| mitre saw instead of a table saw? I guess the crux is
| that you still need to do work either way.
| heraldgeezer wrote:
| >A lot of what I do is relatively simple one off scripting.
| Code that doesn't need to deal with edge cases, won't be widely
| deployed, and whose outputs are very quickly and easily
| verifiable.
|
| Yes somewhat. Its good for powershell/bash/cmd scripts and
| configs but early it would make stuff up
| bbor wrote:
| Look for the ways that AI works, and it can be a powerful tool.
| Try and figure out where it still fails, and you will see
| nothing but hype and hot air.
|
| Perfectly put, IMO.
|
| I know arguments from authority aren't primary, but I think
| this point highlights some important context: Dr. Hossenfelder
| has gained international renown by publishing clickbait-y
| YouTube videos that ostensibly debunk scientific and
| technological advances of all kinds. She's clearly educated and
| thoughtful (not to mention otherwise gainfully employed), but
| her whole public persona kinda relies on assuming the
| exclusively-critical standpoint you mention.
|
| I doubt she necessarily feels _indebted_ to her large audience
| expecting this take (it 's not new...), but that certainly does
| seem like a hard cognitive habit to break.
| cjf101 wrote:
| It's a weird circle with these things. If you _can't_ do the
| task you are using the LLM for, you probably shouldn't.
|
| But if you can do the task well enough to at least recognize
| likely-to-be-correct output, then you can get a lot done in
| less time than you would do it without their assistance.
|
| Is that worth the second order effects we're seeing? I'm not
| convinced, but it's definitely changed the way we do work.
| asdev wrote:
| all fun and games until your AI generated script deletes the
| production database. I think that's the point, fault tolerance
| in academic and financial settings is too high for LLMs to be
| useful
| latemedium wrote:
| My experience is starkly different. Today I used LLMs to:
|
| 1. Write python code for a new type of loss function I was
| considering
|
| 2. Perform lots of annoying CSV munging ("split this CSV into 4
| equal parts", "convert paths in this column into absolute paths",
| "combine these and then split into 4 distinct subsets based on
| this field.." - they're great for that)
|
| 3. Expedite some basic shell operations like "generate softlinks
| for 100 randomly selected files in this directory"
|
| 4. Generate some summary plots of the data in the files I was
| working with
|
| 5. Not to mention extensive use in Cursor & GH Copilot
|
| The tool (Claude 3.7 mostly, integrated with my shell so it can
| execute shell commands and run python locally) worked great in
| all cases. Yes I could've done most of it myself, but I
| personally hate CSV munging and bulk file manipulations and its
| super nice to delegate that stuff to an LLM agent
|
| edit: formatting
| mnky9800n wrote:
| How did you integrate Claude into your shell
| airstrike wrote:
| Claude Code is available directly from Anthropic, but you
| have to request an invite as it's in "Research Preview"
|
| There are third party tools that do the same, though
| latemedium wrote:
| I hacked something together a while back - a hotkey toggles
| between standard terminal mode and LLM mode. LLM mode
| interacts with Claude, and has functions / tool calls to run
| shell commands, python code, web search, clipboard, and a few
| other things. For routine data science tasks it's been super
| useful. Claude 3.7 was a big step forward because it will
| often examine files before it begins manipulating them and
| double-checks that things were done correctly afterwards
| (without prompting!). For me this works a lot better than
| other shell-integration solutions like Warp
| simonw wrote:
| I wrote my own tool for that a while back as an LLM plugin,
| so I can do this: llm cmd extract first
| frame of movie.mp4 as a jpeg using ffmpeg
|
| I use that all the time, it works really well (defaulting to
| GPT-4o-mini because it's so cheap, but it works with Claude
| too): https://simonwillison.net/2024/Mar/26/llm-cmd/
| kwertyoowiyop wrote:
| These seem like fine use cases: trivial boilerplate stuff you'd
| otherwise have to search for and then munge to fit your exact
| need. An LLM can often do both steps for you. If it doesn't
| work, you'll know immediately and you can probably figure out
| whether it's a quick fix or if the LLM is completely off-base.
| zeroonetwothree wrote:
| That's fair but it's totally different use cases than the
| linked post discusses.
| jongjong wrote:
| People who don't work in tech have no idea how hard it is to do
| certain things at scale. Skilled tech people are severely
| underappreciated.
|
| From a sub-tweet:
|
| >> no LLM should ever output a url that gives a 404 error. How
| hard can it be?
|
| As a developer, I'm just imagining a server having to call up all
| the URLs to check that they still exist (and the extra
| costs/latency incurred there)... And if any URLs are missing,
| getting the AI to re-generate a different variant of the
| response, until you find one which does not contain the missing
| links.
|
| And no, you can't do it from the client side either... It would
| just be confusing if you removed invalid URLs from the middle of
| the AI's sentence without re-generating the sentence.
|
| You almost need to get the LLM to engineer/pre-process its own
| prompts in a way which guesses what the user is thinking in order
| to produce great responses...
|
| Worse than that though... A fundamental problem of 'prompt
| engineering' is that people (especially non-tech people) often
| don't actually fully understand what they're asking.
| Contradictions in requirements are extremely common. When
| building software especially, people often have a vague idea of
| what they want... They strongly believe that they have a
| perfectly clear idea but once you scope out the feature in
| detail, mapping out complex UX interactions, they start to see
| all these necessary tradeoffs and limitations rise to the surface
| and suddenly they realize that they were asking for something
| they don't want.
|
| It's hard to understand your own needs precisely; even harder to
| communicate them.
| readthenotes1 wrote:
| "How hard can it be?"
|
| If I recall correctly, that is one of Dilbert's management
| axioms: if I don't understand it it cannot be difficult
| airstrike wrote:
| I wrote an AI assistant which generates working spreadsheets with
| formulas and working presentations with neatly laid out elements
| and styles. It's a huge productivity gain relative to starting
| from a blank page.
|
| I think LLMs work best when they are used as a "creative" tool.
| They're good for the brainstorming part of a task, not for the
| finishing touches.
|
| They are too unreliable to be put in front of your users. People
| don't want to talk to unpredictable chatbots. Yes, they can be
| useful in customer service chats because you can put them on
| rails and map natural language to predetermined actions. But
| generally speaking I think LLMs are most effective when used _by_
| someone who's piloting them instead of wrapped in a service
| offered _to_ someone.
|
| I do think we've squeezed 90%+ of what we could from current
| models. Throwing _more_ dollars of compute at training or
| inference won 't make much difference. The next "GPT moment" will
| come from some sufficiently novel approach.
| joegibbs wrote:
| Because it's not a scientific research tool, it's a most likely
| next text generator. It doesn't keep a database of ingested
| information with source URLs. There are plenty of scientific
| research tools but something that just outputs text based on your
| input is no good for it.
|
| I'm sure that in the future there will be a really good search
| tool that utilises an LLM but for now a plain model just isn't
| designed for that. There are a ton of other uses for them, so I
| don't think that we should discount them entirely based on their
| ability to output citations.
| GaggiX wrote:
| I use them everyday and they work greatly, I even made a command
| (using Claude, actually Claude made everything in that script)
| that calls Gemini from the terminal so that I can ask for
| question related to the shell directly there, just doing a: ai
| "how can I convert a webp to a png", the system prompt asks to be
| brief, using markdown (it does display nicely), that most
| question are related to Linux and it provides information about
| my OS (uname -a), the last code block is also copied in the
| clipboard, super useful, I imagine there are plenty online of
| similar utilities.
| saaaaaam wrote:
| I've used Claude today to:
|
| Write code to pull down a significant amount of public data using
| an open API. (That took about 30 seconds - I just gave it the
| swagger file and said "here's what I want")
|
| Get the data (an hour or so), clean the data (barely any time,
| gave it some samples, it wrote the code), used the cleaned data
| to query another API, combined the data sources, pulled down a
| bunch of PDFs relating to the data, had the AI write code to use
| tesseract to extract data from the PDFs, and used that to build a
| dashboard. That's a mini product for my users.
|
| I also had a play with Mistral's OCR and have tested a few things
| using that against the data. When I was out walking my dogs I
| thought about that more, and have come up with a nice workflow
| for a problem I had, which I'll test in more detail next week.
|
| That was all whole doing an entirely different series of tasks,
| on calls, in meetings. I literally checked the progress a few
| times and wrote a new prompt or copy/pasted some stuff in from
| dev tools.
|
| For the calls I was on, I took the recording of those calls,
| passed them into my local instance whisper, fed the transcript
| into Claude with a prompt I use to extract action points, pasted
| those into a google doc, circulated them.
|
| One of the calls was an interview with an expert. The transcript
| + another prompt has given me the basis for an article (bulleted
| narrative + key quotes) - I will refine that tomorrow, and write
| the article, using a detailed prompt based on my own writing
| style and tone.
|
| I needed to gather data for a project I'm involved in, so had
| Claude write a handful of scrapers for me (HTML source > here is
| what I need).
|
| I downloaded two podcasts I need to listen to - but only need to
| listen to five minutes of each - and fed them into whisper then
| found the exact bits I needed and read the extracts rather than
| listening to tedious podcast waffle.
|
| I turned an article I'd written into an audio file using
| elevenlabs, as a test for something a client asked me about
| earlier this week.
|
| I achieved about three times as much today as I would have done a
| year ago. And finished work at 3pm.
|
| So yeah, I don't understand why people are so bullish about LLMs.
| Who knows?
| kilolima wrote:
| Yuck. Do your users know that they are reading recycled LLM
| content? Is this long winded post generated by an LLM?
| Paratoner wrote:
| Did you also do that while mewing and listening to an AI
| abridged audiobook version of the laws of power in chinese?
| Don't forget your morning ice face dunks.
| doctoboggan wrote:
| Why do people who don't like using LLMs keep insisting they are
| useless for the rest of us? If you don't like to use them, then
| simply don't use them.
|
| I use them almost daily in my job and get tremendous use out of
| them. I guess you could accuse me of lying, but what do I stand
| to gain from that?
|
| I've also seem people claim that only people who don't know how
| to code or people doing super simple done a million times apps
| can get value out of LLMs. I don't believe that applies to my
| situation, but even if it did, so what? I do real work for a real
| company delivering real value, and the LLM delivers value to me.
| It's really as simple as that.
| biker142541 wrote:
| I don't think it's a matter or liking or not. The use cases
| just differ considerably, and tools and not as useful or
| applicable across those. THe OP's use case is probably one of
| the worst possible for LLMs right now, imo...
| harrall wrote:
| I am neither bullish or bearish. LLM is a tool.
|
| It's a hammer -- sometimes it works well. It summarizes the user
| reviews on a site... cool, not perfect, but useful.
|
| And like every tool, it is useless for 90% of life's situations.
|
| And I know when it's useful because I've already tried a hammer
| on 1000 things and have figured out what I should be using a
| hammer on.
| rufus_foreman wrote:
| >> I am neither bullish or bearish. LLM is a tool...It's a
| hammer
|
| If someone says, "This new type of hammer will increase
| productivity in the construction industry by 25%", it's
| something else in addition to being a tool. It's either a lie,
| or it's an incredible advance in technology.
| belter wrote:
| "Yes, I have tried Gemini, and actually it was even worse in that
| it frequently refuses to even search for a source and instead
| gives me instructions for how to do it myself. Stopped using it
| for that reason."
|
| Thank you Sabine. Every time I have mentioned Gemini is the
| worst, and not even worth of consideration, I have been bombarded
| with downvotes, and told I am using it wrong.
| throwawa14223 wrote:
| My experience mirrors hers. Asking questions is worthless because
| the answers are either 404 links, telling me how to use a search
| engine, or just flat out wrong and the code generated compiles
| maybe one time out of ten and when it does the implementation is
| usually poor.
|
| When I evaluate against areas I possess professional expertise I
| become convinced LLMs produce the Gell Mann amnesia effect for
| any area I don't know.
| islewis wrote:
| > I genuinely don't understand why some people are still bullish
| about LLMs.
|
| I don't believe OP's thesis is properly backed by the rest of his
| tweet, which seems to boil down to "LLM's can't properly cite
| links".
|
| If LLM's performing poorly on an arbitrary small-scoped test case
| makes you bearish on the whole field, I don't think that falls on
| the LLM's.
| crazygringo wrote:
| If there's one common thread across LLM criticisms, it's that
| they're not perfect.
|
| These critics don't seem to have learned the lesson that _the
| perfect is the enemy of the good_.
|
| I use ChatGPT all the time for academic research. Does it
| fabricate references? Absolutely, maybe about a third of the
| time. But has it pointed me to important research papers I might
| never have found otherwise? _Absolutely_.
|
| The rate of inaccuracies and falsehoods doesn't matter. What
| matters is, is it saving you time and increasing your
| productivity. Verifying the accuracy of its statements is _easy_.
| While finding the knowledge it spits out in the first place is
| _hard_. The net balance is a _huge positive_.
|
| People are bullish on LLM's because they can save you days' worth
| of work, like every day. My research productivity has gone _way_
| up with ChatGPT -- asking it to explain ideas, related concepts,
| relevant papers, and so forth. It 's amazing.
| simonw wrote:
| The most interesting thing about this post is how it reinforces
| how terrible the usability of LLMs still is today:
|
| "I ask them to give me a source for an alleged quote, I click on
| the link, it returns a 404 error. I Google for the alleged quote,
| it doesn't exist. They reference a scientific publication, I look
| it up, it doesn't exist."
|
| To experienced LLM users that's not surprising at all - providing
| citations, sources for quotes, useful URLs are all things that
| they are demonstrably terrible at.
|
| But it's a computer! Telling people "this advanced computer
| system cannot reliably look up facts" goes against everything
| computers have been good at for the last 40+ years.
| puppycodes wrote:
| Here's one simple reason:
|
| I have a very specific esoteric question like: "What material is
| both electrically conductive and good at blocking sound?" I could
| type this into google and sift through the titles and short
| descriptions of websites and eventually maybe find an answer, or
| I can put the question to the LLM and instantly get an answer
| that I can then research further to confirm.
|
| This is significantly faster, more informative, more efficient,
| and a rewarding experience.
|
| As others have said, its a tool. A tool is as good as how you use
| it. If you expect to build a house by yelling at your tools I
| wouldn't be bullish either.
| krapp wrote:
| I mean, the first link I got when I pasted that in is probably
| the Stack Exchange thread you would use to research further,
| along with other sources, which do seem relevant to the query.
|
| I don't see how an LLM is significantly faster or more
| informative, since you still have to do the legwork to validate
| the answer. I guess if you're google-phobic (which a lot of
| people seem to be, especially on HN) then I can see how it's
| more rewarding to put it off until later in the process.
| puppycodes wrote:
| Its just an example and you can "well actually" until the
| cows come home, but your missing the point. I'm sure there
| are things you have found hard to google. Also you are most
| likely (as your commenting on hacker news) _not_ a good
| representation of the majority of the world.
|
| Also idk if you have noticed, but google is clearly using LLM
| technology in conjunction with its search results, so the
| assumption they are just using not LLMs to inform or modify
| its result set I think is naive or at very best just a
| personal opinion.
| puppycodes wrote:
| you can boil it down to this, whats easier for most people?
| looking through websites and search engine results for
| answers or speaking in plain language? The answer is pretty
| obvious.
|
| The validity of the answers is not 1:1 with its potential
| profitability.
|
| Like James Baldwin said "people love answers, but hate
| questions."
|
| getting an answer faster is exponentially better than getting
| the more precise, more right, more nuanced answer for most
| people every time. Doing the due dilligence is smart but its
| also after the fact.
| xiphias2 wrote:
| ,,By my personal estimate currently GPT 4o DeepResearch is the
| best one. ''
|
| If the o3 based 3 month old strongest model is the best one, it's
| a proof that there were quite significant improvements in the
| last 2 years.
|
| I can't name any other technology that improved as much in 2
| years.
|
| O1 and o1 pro helped me with filing tax returns and answered me
| questions that (probably quite bad) tax accountants (and less
| smart models) weren't able to (of course I read the referenced
| laws, I don't trust the output either).
| meowface wrote:
| AI coding overall still seems to be underrated by the average
| developer.
|
| They try to add a new feature or change some behavior in a large
| existing codebase and it does something dumb and they write it
| off as a waste of time for that use case. And that's
| understandable. But if they had tweaked the prompt just a bit it
| actually might've done it flawlessly.
|
| It requires patience and learning the best way to guide it and
| iterate with it when it does something silly.
|
| Although you undoubtedly will lose some time re-attempting
| prompts and fixing mistakes and poor design choices, on net I
| believe the frontier models can currently make development much
| more productive in almost any codebase.
| jwrallie wrote:
| The author mentioned Gemini sometimes refusing to do something.
|
| I've recently been using Gemini (mostly 2.0 flash) a lot and I've
| noticed it sometimes will challenge me to try doing something by
| myself. Maybe it's something in my system prompt or the way I
| worded the request itself. I am a long time user of 4o so it felt
| annoying at first.
|
| Since my purpose was to learn how to do something, being open
| minded I tried to comply with the request and I can say that...
| it's being a really great experience in terms of retention of
| knowledge. Even if I'm making mistakes Gemini will point them out
| and explain it nicely.
| BSOhealth wrote:
| Like others here, I use it to code (no longer a professional
| engineer, but keep side projects).
|
| As soon as LLMs were introduced into the IDE it began to feeling
| like LLM autocomplete was almost reading my mind. With some
| context built up over a few hundred lines of initial
| architecture, autocomplete now sees around the same corners I am.
| It's more than just "solve this contrived puzzle" or "write
| snake". It combines the subject matter use case (informed by
| variable and type naming) underlying the architecture and
| sometimes produces really breathtaking and productive results.
| Like I said, it took some time but when it happened, it was
| pretty shocking.
| kriro wrote:
| She's a scientist. In that area LLMs are quite useful in my
| opinion and part of my daily workflow. Quick scripts that use
| APIs to get data, cleaning the data and converting it. Quickly
| working with polar data frames. Dumb daily stuff like "take this
| data from my CSV file and turn it into a Latex table"...but most
| importantly freeing up time from tedious administrative tasks
| (let's not go into detail here).
|
| Also great for brainstorming and quick drafting grant proposals.
| Anything prototyping and quickly glued together I'll go for LLMs
| (or LLM agents). They are no substitute for your own brain
| though.
|
| I'm also curious about the hallucinated sources. I've recently
| read some papers on using LLM-agents to conduct structured
| literature reviews and they do it quite well and fairly
| reproducible. I'm quite willing to build some LLM-agents to
| reproduce my literature review process in the near future since
| it's fairly algorithmic. Check for surveys and reviews on the
| topic, scan for interesting papers within, check sources of
| sources, go through A tier conference proceedings for the last X
| years and find relevant papers. Rinse, repeat.
|
| I'm mostly bullish because of LLM-agents, not because of using
| stock models with the default chat interface.
| marcuschong wrote:
| The key for LLM productivity, it seems to me, is grounding. Let
| me give you my last example, from something I've been working on.
|
| I just updated my company commercial PPT. ChatGPT helped me with:
| - Deep Research great examples and references of such
| presentations. - Restructure my argument and slides according to
| some articles I found on the previous step, and thought were
| pretty good. - Come up with copy for each slide. - Iterate new
| ideas as I was progressing.
|
| Now, without proper context and grounding, LLMs wouldn't be so
| helpful at this task, because they don't know my company,
| clients, product and strategy, and would be generic at best. The
| key: I provided it with my support portal documentation and a
| brain dump I recorded to text on ChatGPT with key strategic
| information about my company. Those are two bits of info I keep
| always around, so ChatGPT can help me with many tasks in the
| company.
|
| From that grounding to the final PPT, it's pretty much a trivial
| and boring _transformation_ task that would have cost me many,
| many hours to do.
___________________________________________________________________
(page generated 2025-03-27 23:01 UTC)