[HN Gopher] Agents Are Not Enough
___________________________________________________________________
Agents Are Not Enough
Author : awaxman11
Score : 131 points
Date : 2025-01-06 15:00 UTC (3 days ago)
(HTM) web link (www.arxiv.org)
(TXT) w3m dump (www.arxiv.org)
| bob1029 wrote:
| I think the goldilocks path is to make the user the agent and use
| the LLM simply as their UI/UX for working with the system. Human
| (domain expert) in the loop gives you a reasonable chance of
| recovering from hallucinations before they spiral entirely out of
| control.
|
| "LLM as UI" seems to be something hanging pretty low on the tree
| of opportunity. Why spent months struggling with complex admin
| dashboard layouts and web frameworks when you could wire the
| underlying CRUD methods directly into LLM prompt callbacks? You
| could hypothetically make the LLM the _exclusive_ interface for
| managing your next SaaS product. There are ways to make this just
| as robust and secure as an old school form punching application.
| GiorgioG wrote:
| re: LLM as UI: Given that I don't trust LLMs to be
| deterministic, I wouldn't trust them to make the correct API
| call every time I tell it to do X.
| kgeist wrote:
| I think most users have a fixed set of workflows which
| usually don't change from day to day, so why not just use
| LLMs as a macro builder with a natural language interface
| (and which doesn't require you to know the product's UI well
| beforehand):
|
| - you ask LLM to build a workflow for your problem
|
| - the LLM builds the workflow (macro) using predefined
| commands
|
| - you review the workflow (can be an intuitive list of
| commands, understandable by non-specialist) - to weed out
| hallucinations and misunderstanding
|
| - you save the workflow and can use it without any LLM
| agents, just clicking a button - pretty determenistic and
| reliable
|
| Advantages:
|
| - reliable, deterministic
|
| - you don't need to learn a product's UI, you just formulate
| your problem using natural language
| dingnuts wrote:
| >- you review the workflow (can be an intuitive list of
| commands, understandable by non-specialist)
|
| so you define a DSL that the LLM outputs, and that's the
| real UI
|
| >- you don't need to learn a product's UI, you just
| formulate your problem using natural language
|
| yes, you do. You have to learn the DSL you just manifested
| so that you can check it for errors. Once you have the
| ability to review the LLM's output, you will also have the
| ability to just write the DSL to get the desired behavior,
| at which point that will be faster unless it's a
| significant amount of typing, and even then, you will still
| need to review the code generated by the LLM, which means
| you have to learn and understand the DSL. I would much
| rather learn a GUI than a DSL.
|
| You haven't removed the UI, nor have you made the LLM the
| UI, in this example. The DSL ("intuitive list of commands..
| I guess it'll look like the Robot Framework right? that's
| what human-readable DSLs tend to look like in practice) is
| the actual UI.
|
| This is vastly more complicated than having a GUI to
| perform an action.
| kgeist wrote:
| I never said the user must be exposed to a DSL, I think
| you're overcomplicating it for the sake of
| overcomplicating. DSL can be used under the hood by the
| execution engine, but the user can be exposed to a
| simpler variant of it, either by clever hardcoded
| postprocessing of known commands when rendering the final
| result for human review, or maybe use the LLM itself to
| summarize the planned actions (although it can
| hallucinate while summarizing, but the chance is
| miniscule, especially if a user can test a saved
| workflow). My point was mostly about two things:
|
| 1) "it's unpredictable each time" - it won't be, if a
| workflow is saved and tested, because when it's run, no
| LLM is involved anymore in decision making
|
| 2) I did remove the UI, because I don't need to learn the
| UI, I just formulate my problem and the LLM constructs a
| possible workflow which solves my problem out of
| predefined commands known to the system.
|
| Sure this is most useful for more complex apps. In our
| homegrown CRM/ERP, users have lots of different workflows
| depending on their department, and they often experiment
| with workflows, and today they either have to click
| through everything manually (wasting time) or ask devs to
| implement the needed workflow for them (wasting time). If
| your app has 3 commands on 1 page then sure, it's easier
| to do it using GUI.
|
| Also IMHO it can be used alongside with GUI, it doesn't
| need to replace it, I think it's great for
| discoverability/onboarding and automation, but if you
| want to click through everything manually, why not.
| svieira wrote:
| The bit you are missing is that "known to the system" is
| not enough, as the consumer I need to _verify the logic_,
| which means that at some level, I do have to read the DSL
| (just as I have to read the Java, not, in general, the
| actual assembly emitted by the JIT). Which means that
| _the DSL is actually the product here_ (though the LLM
| may make it easier to _learn_ that DSL and in some cases
| to write something in it).
| kgeist wrote:
| 1) You don't need to read the DSL in the raw form if you
| use a language model to convert it to a few paragraphs in
| natural language.
|
| 2) You can test the created workflow on a bunch of test
| data to verify it works as intended. After a workflow is
| created, it's deterministic (since we don't use LLMs
| anymore for decision making), so it will always work the
| same.
|
| Sure we can expose DSL to power users as an option, but
| is reading the raw DSL really required for the majority
| of cases?
| svieira wrote:
| 1. Now you have two problems (did the writer translate
| what I said correctly and did the summarizer translate
| what the writer wrote correctly).
|
| 2. This is absolutely true and it does help somewhat.
| However, writing the test cases is now your bottleneck
| (and you're writing them as a substitute for being able
| to read a reliable high-level summary of what the
| workflow actually is).
| nyrikki wrote:
| So visual programming x.0?
|
| I am pretty sure PLCs with ladder logic are about the
| limits of the traditional visual/macro model?
|
| Word-sense disambiguation is going to be problematic with
| the 'don't need to learn' part above.
|
| Consider this sentence:
|
| 'I never said she stole my money'
|
| Now read that sentence multiple times, puting emphasis on
| each word, one at a time and notice how the symantic
| meaning changes.
|
| LLMs are great at NLP, but we still don't have solutions to
| those NLU problems that I am aware of.
|
| I think to keep maximum generality without severely
| restricted use cases that a common DSL would need to be
| developed.
|
| There will have to be tradeoffs made, specific to
| particular use cases, even if it is better than Alexa.
|
| But I am thinking about Rice's theorm and what happens when
| you lose PEM.
|
| Maybe I just am too embedded in an area where these
| problems are a large part of the difficulty for macro style
| logic to provide much use.
| bob1029 wrote:
| > you review the workflow (can be an intuitive list of
| commands, understandable by non-specialist) - to weed out
| hallucinations and misunderstanding
|
| This is the idea that is most valuable from my perspective
| of having tried to extract accurate requirements from the
| customer. Getting them to learn your product UI and
| capabilities is an uphill battle if you are in one of the
| cursed boring domains (banking, insurance, healthcare,
| etc.).
|
| Even if the customer doesn't get the LLM-defined path to
| provide their desired final result, you still have their
| entire conversation history available to review. This seems
| more likely to succeed in practice than hoping the customer
| provides accurate requirements up-front in some
| unconstrained email context.
| shekhargulati wrote:
| This is the same approach we took when we added LLM
| capability to a low code tool Appian. LLM helped us
| generate the Appian workflow configuration file, user
| reviews it and make changes if required, and then finally
| publishes it.
| deadbabe wrote:
| They are deterministic at 0 temperature
| BalinKing wrote:
| (Disclaimer: I know literally nothing about LLMs.) Wouldn't
| there still be issues of sensitivity, though? Like,
| wouldn't you still have to ensure that the wording of your
| commands stays _exactly_ the same every time? And with
| models that take less discrete data (e.g. ChatGPT 's new
| "advanced voice model" that works on audio directly), this
| seems even harder.
| BalinKing wrote:
| s/advanced voice model/advanced voice mode/ (too late for
| me to edit my original comment)
| lokhura wrote:
| At zero temp there is still non-determism due to sampling
| and the fact that floating point addition is not
| commutative so you will get varying results due to
| parallelism.
| wkat4242 wrote:
| They are pretty deterministic then but they are also pretty
| useless at 0 temperature.
| ukuina wrote:
| Not for the leading LLMs from OpenAI and Anthropic.
| hitchstory wrote:
| I dont either, but this can be mitigated by adding guard
| rails (strictly validating input), double checking actions
| with the user and using it for tasks where a mistake isnt
| world ending.
|
| Even then mistakes can slip through, but it could still be
| more reliable than a visual UI.
|
| There are lots of horrible web UIs i would LOVE to replace
| with a conversational LLM agent. No #1 is jira and so is no
| #2 and #3.
| barrkel wrote:
| It's quite tedious to have to write (or even say) full
| sentences to express intent. Imagine driving a car with a voice
| interface, including accelerator, brake, indicators and so on.
| Controls are less verbose and dashboards are more information
| rich than linear text.
|
| It's difficult to be precise. Often it's easier to gauge things
| by looking at them while giving motor feedback (e.g. turning a
| dial, pushing a slider) than to say "a little more X" or "a bit
| less Y".
|
| Language is poorly suited to expressing things in continuous
| domains, especially when you don't have relevant numbers that
| you can pick out of your head - size, weight, color etc.
| Quality-price ratio is a particularly tough one - a hard
| numeric quantity traded off against something subjective.
|
| Most people can't specify up front what they want. They don't
| know what they want until they know what's possible, what other
| people have done, started to realize what getting what they
| want will entail, and then changed what they want. It's why we
| have iterative development instead of waterfall.
|
| LLMs are a good start and a tool we can integrate into systems.
| They're a long, long way short of what we need.
| klabb3 wrote:
| > "LLM as UI" seems to be something hanging pretty low on the
| tree of opportunity.
|
| Yes if you want to annoy your users and deliberately put
| roadblocks to make progress on a task. Exhibit A: customer
| support. They put the LLM in between to waste your time. It's
| not even a secret.
|
| > Why spent months struggling with complex admin dashboard
| layouts
|
| You can throw something together, and even auto generate forms
| based on an API spec. People don't do this too often because
| the UX is insufficient even for many internal/domain expert
| support applications. But you could and it would be
| deterministic, unlike an LLM. If the API surface is simple, you
| can make it manually with html & css quickly.
|
| Overuse of web frameworks has completely different causes than
| "I need a functional thing" and thus it cannot be solved with a
| different layer of tech like LLMs, NFTs or big data.
| wkat4242 wrote:
| > Yes if you want to annoy your users and deliberately put
| roadblocks to make progress on a task. Exhibit A: customer
| support. They put the LLM in between to waste your time. It's
| not even a secret.
|
| No this is because they use the LLM not only as human
| interface but also as a reasoning engine for troubleshooting.
| And give it way less capability than a human agent to boot.
| So all it can really do is serve FAQs and route to real
| support.
|
| In this case the fault is not with the LLM but with the
| people that put it there.
| diggan wrote:
| > I think the goldilocks path is to make the user the agent and
| use the LLM simply as their UI/UX for working with the system
|
| That's a funny definition to me, because doing so would mean
| the LLM is the agent, if you use the classic definition for
| "user-agent" (as in what browsers are). You're basically
| inverting that meaning :)
| pwillia7 wrote:
| I had the same epiphany about LLM as UI trying to build a front
| end for a image enhancer workflow I built with Stable
| Diffusion. I just about fully built out a Chrome extension and
| then realized I should just build a 'tool' that llama can
| interact with and use open webui as the front end.
|
| quick demo: https://youtu.be/2zvbvoRCmrE
| simonw wrote:
| This paper does at least lead with its version of what "agents"
| means (I get very frustrated when people talk about agents
| without clarifying which of the many potential definitions they
| are using):
|
| > An agent, in the context of AI, is an autonomous entity or
| program that takes preferences, instructions, or other forms of
| inputs from a user to accomplish specific tasks on their behalf.
| Agents can range from simple systems, such as thermostats that
| adjust ambient temperature based on sensor readings, to complex
| systems, such as autonomous vehicles navigating through traffic.
|
| This appears to be the broadest possible definition, encompassing
| thermostats all the way through to Waymos.
| adpirz wrote:
| You posted on X a while back asking for a crowdsourced
| definition of what an "agent" was and I regularly cite that
| thread as an example of the fact that this word is so blurry
| right now.
| simonw wrote:
| I really need to write that up in one place - closest I've
| got is this section from my 2024 review
| https://simonwillison.net/2024/Dec/31/llms-
| in-2024/#-agents-...
| jvans wrote:
| this is a great write up, thank you
| swyx wrote:
| didnt your summary https://gist.github.com/simonw/beaa5f901
| 33b30724c5cc1c4008d0... pretty much cover it?
| adpirz wrote:
| Whoa missed this! Love it.
| adpirz wrote:
| This write up was also fantastic and has made the rounds at
| our org!
| sanjin wrote:
| I took a crack at it here that tries to bridge the gap from
| "autonomous" which is just software to that Agentic
| autonomy - https://www.aiimpactfrontier.com/p/framework-
| for-ai-agents
| DebtDeflation wrote:
| >"The two main categories I see are people who think AI
| agents are obviously things that go and act on your behalf
| --the travel agent model--and people who think in terms of
| LLMs that have been given access to tools which they can
| run in a loop as part of solving a problem."
|
| This is exactly the problem and these two categories nicely
| sum up the source of the confusion.
|
| I consider myself in the former camp. The AI needs to
| determine my intent (book a flight) which is a
| classification problem, extract out the relevant
| information (travel date, return date, origin city,
| destination city, preferred airline) which is a Named
| Entity Recognition problem, and then call the appropriate
| API and pass this information as the parameters (tool
| usage). I'm asking the agent to perform an action on my
| behalf, and then it's taking my natural language and going
| from there. The overall workflow is deterministic, but
| there are elements within it that require some
| probabilistic reasoning.
|
| Unfortunately, the second camp seems to be winning the day.
| Creating unrealistic expectations of what can be
| accomplished by current day LLMs running in a loop while
| simultaneously providing toy examples of it.
| mindcrime wrote:
| It's been blurry for a long time, FWIW. I have books on
| "Agents" dating back to the late 90's or early 2000's in
| which the "Intro" chapter usually has a section that tries to
| define what an "agent" is, and laments that there is no
| universally accepted definition.
|
| To illustrate: here's a paper from 1996 that tries to lay out
| a taxonomy of the different kinds of agents and provide some
| definitions:
|
| https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&d.
| ..
|
| And another from the same time-frame, which makes a similar
| effort:
|
| https://www.researchgate.net/profile/Stan-
| Franklin/publicati...
| williamcotton wrote:
| So basically just the concept of feedback in a cybernetic
| system.
|
| https://en.wikipedia.org/wiki/Cybernetics
| bob1029 wrote:
| > The field is named after an example of circular causal
| feedback--that of steering a ship (the ancient Greek
| kubernetes (kybernetes)...
|
| Now that name makes a lot more sense to me.
| openrisk wrote:
| Which is also the root of the word 'government', so a
| government agent is doubly cybernetic in a sense
| 8338550bff96 wrote:
| Then is "agents" just non-spooky coded language for "cyborgs"
| throw5959 wrote:
| I studied cybernetics. Our teachers called us "cybernets".
| htrp wrote:
| agents are the 2020s version of data science in the 2010s
| Kerbonut wrote:
| Do you mean that agents are being hyped in the same way data
| science was in the 2010s, or that they'll have a similar
| impact over time? Would love to hear more of your thoughts.
| snapcaster wrote:
| I think he meant it's a similarly blurry term
| simonw wrote:
| Yeah, what does "data science" mean, exactly?
| throw5959 wrote:
| Using the scientific method to handle data.
| curious_cat_163 wrote:
| Yes, and the definition works reasonably well for the core
| arguments they are making in Section 5.
|
| I suspect they'll follow up with a full paper with more details
| (and artifacts) of their proposed approach.
| behnamoh wrote:
| People have been talking about agents for at least 2 years.
| Remember when AgentGPT came out? How's that going so far?
| Agents are just LLMs with structured output, which often
| happens to be a JSON with info about a function arguments to be
| called.
| mindcrime wrote:
| > People have been talking about agents for at least 2 years.
|
| WAY longer than that. What's come to the forefront
| specifically in the last year or two is very specific subset
| of the overall agent landscape. What I like to call "LLM
| Agents". But "Agents" at large date back to at least the
| 1980's if not before. For some of the history of all of this,
| see this page and some of the listed citations:
|
| https://en.wikipedia.org/wiki/Software_agent
|
| > Agents are just LLMs with structured output
|
| That's only true for the "LLM Agent" version. There are
| Agents that have nothing to do with LLM's at all.
| simonw wrote:
| Right - the term "user-agent" shows up in the HTTP/1.0 spec
| from 1996: https://datatracker.ietf.org/doc/html/rfc1945
| and there's plenty of history of debates about the meaning
| of the term before then.
|
| In 1994 people were already complaining that the term that
| had no universal agreed definition:
| https://simonwillison.net/2024/Oct/12/michael-wooldridge/
| mindcrime wrote:
| Yes. I am fond of saying "If you're talking about agents
| and think the term is something new, go back and read
| everything Michael Wooldridge ever wrote before talking
| any further". :-)
| cratermoon wrote:
| We're at the phase of the hype cycle where "agent" means
| whatever the marketing materials want it to mean.
| baxtr wrote:
| To me "Agents" sound like computer programs that interact
| through APIs?
| bsenftner wrote:
| Oh come on! You and I know very well an AI Agent is anything
| marketing says they are, and that is the absolute final truth.
| TaurenHunter wrote:
| "More Agents is all you need" https://arxiv.org/abs/2402.05120
|
| I could not find a "Agents considered harmful" related to AI, but
| there is this one: "AgentHarm: A benchmark for measuring
| harmfulness of LLM agents" https://arxiv.org/pdf/2410.09024
|
| This "Agents considered harmful" is not AI-related:
| https://www.scribd.com/document/361564026/Math-works-09
| kridsdale1 wrote:
| Morpheus taught me they are quite harmful.
| ksplicer wrote:
| When reading anthropics blog on agents I basically took away
| that their advice is you shouldn't use them to solve most
| problems.
|
| https://www.anthropic.com/research/building-effective-agents
|
| "For many applications, however, optimizing single LLM calls
| with retrieval and in-context examples is usually enough."
| dist-epoch wrote:
| Real agents have never been tried
| zombiwoof wrote:
| Agent is a funding and marketing term imho
|
| Soon it will be AI Microservices
| bad_haircut72 wrote:
| Who wants to invest in my startup, its a Microagent service
| architectures orchestration platform. All you do is define the
| inputs, write the agents algorithms, apply agency by inputting
| a decision tree (ifs and conditionals) and then a function to
| format output! And the best part? You do all of it in YAML!
|
| /sarcasm, hopefully obviously
| mindcrime wrote:
| I was thinking "shut up and take my money" until you brought
| YAML into it. Hard pass. ;p
| ramesh31 wrote:
| >Agent is a funding and marketing term imho
|
| So was "mobile" 15 years ago. Companies are deploying hundreds
| of billions in capital for this. It's not going anywhere, and
| you'd be best off upskilling now instead of dismissing things.
| nowittyusername wrote:
| With time, they will get a lot better. IMO, the biggest hurdles
| the agents currently lack is good implementation of function
| calling capabilities. LLM's should be used as reasoning engines
| and everything else should be offloaded to tool use. This will
| drastically reduce hallucinations and errors in math and all the
| other areas.
| lionkor wrote:
| Do they reason, though?
| ripped_britches wrote:
| I can imagine really powerful agents this year or next in theory.
| Agents meaning (not a thermostat) a system that can go complete
| some async tasks on your behalf. But in practice I don't have any
| idea how we will solve for prompt injection attacks. Hopefully
| someone cracks it.
| cratermoon wrote:
| "AI will soon be able too..."
| Jerrrry wrote:
| >solve for prompt injection attacks
|
| It is essentially the same Code as Data problem as always.
| georgestrakhov wrote:
| IMHO, the word agent is quickly becoming meaningless. The amount
| of agency that sits with the program vs. the user is something
| that changes gradually.
|
| So we should think about these things in terms of how much agency
| are we willing to give away in each case and for what gain[1].
|
| Then the ecosystem question that the paper is trying to solve
| will actually solve itself, because it is already the case today
| that in many processes agency has been outsourced almost fully
| and in others - not at all. I posit that this will continue, just
| expect a big change of ratios and types of actions.
|
| [1] https://essays.georgestrakhov.com/artificial-agency-ladder/
| w10-1 wrote:
| > IMHO, the word agent is quickly becoming meaningless. The
| amount of agency that sits with the program vs. the user is
| something that changes gradually
|
| Yes, the term is becoming ambiguous, but that's because it's
| abstracting out the part of AI that is most important and
| activating: the ability to work both independently and per
| intention/need.
|
| Per the paper: "Key characteristics of agents include autonomy,
| programmability, reactivity, and proactiveness.[...] high
| degree of autonomy, making decisions and taking actions
| independently of human intervention."
|
| Yes, "the ecosystem will evolve," but to understand and
| anticipate the evolution, one needs a notion of fitness, which
| is based on agency.
|
| > So we should think about these things in terms of how much
| agency are we willing to give away in each case
|
| It's unclear there can be any "we" deciding. For resource-
| limited development, the ecosystem will evolve regardless of
| our preferences or ethics according to economic advantage and
| capture of value. (Manufacturing went to China against the
| wishes of most everyone involved.)
|
| More generally, the value is AI is not just replacing work.
| It's giving more agency to one person, avoiding the cost and
| messiness of delegation and coordination. It's gaining the same
| advantages seen where smaller team can be much more effective
| than a larger one.
|
| Right now people are conflating these autonomy/delegation
| features with the extension features of AI agents (permitting
| them to interact with databases or web browsers). The extension
| vendors will continue to claim agency because it's much more
| alluring, but the distinction will likely become clear in a
| year or so.
| HarHarVeryFunny wrote:
| An agent, or something that has agency, is just something that
| takes some action, which could be anything from a thermostat
| regulating the temperature all the way up to an autonomous
| entity such as an animal going about it's business.
|
| Hugging Face have their own definitions of a few different
| types of agent/agentic system here:
|
| https://huggingface.co/docs/smolagents/en/conceptual_guides/...
|
| As related to LLMs, it seems most people are using "agent" to
| refer to systems that use LLMs to achieve some goal - maybe a
| fairly narrow business objective/function that can be
| accomplished by using one or more LLMs as a tool to accomplish
| various parts of the task.
| cratermoon wrote:
| Here's a link to arxiv page for the paper, in case you want to
| look over the abstract and citation metadata before downloading
| the PDF.
|
| https://arxiv.org/abs/2412.16241
| jokethrowaway wrote:
| I don't get the hype about Agents.
|
| It's just calling a LLM n-times with slightly different prompts
|
| Sure, you get the ability to correct previous mistakes, it's
| basically a custom chain of thought - but errors compound and the
| results coming from agents have a pretty low success rate.
|
| Bruteforcing your way out of problems can work sometimes (as
| evinced by the latest o3 benchmarks) but it's expensive and
| rarely viable for production use.
| grahamj wrote:
| > It's just calling a LLM n-times with slightly different
| prompts
|
| It can be, but ideally each agent's model, prompts and tools
| are tailored to a particular knowledge domain. That way tasks
| can be broken down into subtasks which are classified and
| passed to the agents best suited to them.
|
| Agree RE it being bruteforce and expensive but it does look
| like it can improve some aspects of LLM use.
| mindcrime wrote:
| > It's just calling a LLM n-times with slightly different
| prompts
|
| That's one way of building something you could call an "agent".
| It's far from the only way. It's certainly possible to build
| agents where the LLM plays a very small role, or even one that
| uses no LLM at all.
| coro_1 wrote:
| The paper covers technical details and the logistics of AI Agents
| to come. But how are humans going to react to mass AI Agents
| replacing other human emotion and connection? Bias is central in
| tech-culture to only agents, but this could become an issue.
| danielmarkbruce wrote:
| Why post this paper? It says nothing, it's a waste of people's
| time to read.
| duxup wrote:
| Even just the definition of an Agent (maybe imperfect) made it
| worthwhile for me.
| danielmarkbruce wrote:
| I'm not sure it's even good though... the input doesn't need
| to come from a user. I have an "agent" which listens for an
| event in financial markets and then goes and does some stuff.
|
| In practice the current usage of "agent" is just: a program
| which does a task and uses an LLM somewhere to help make a
| decision as to what to do and maybe uses an LLM to help do
| it.
| j45 wrote:
| Math that can't be too warm and too accurate to work may have
| challenges being too accurate and reliable with repeating
| processes.
| pwillia7 wrote:
| How would the SIMS that contain the user prefs and whatnot not
| have the same issues described in the paper as the agents
| themselves?
| tonetegeatinst wrote:
| Somewhat related but here's my take on super intelligence or AGI.
| I have worked with CNN,GNN and other old school AI methods, but
| don't have the resources to build a real SOT LLM, but I do use
| and tinker with LLM's occasionally.
|
| If AGI or SI(super intelligence)/is possible, and that is an
| if...I don't think LLM's are going to be this silver bullet
| solution Just as we have in the real world of people who are
| dedicated to a single task in their field like a lawyer or
| construction workers or doctors and brain surgeons, I see the
| current best path forward as being a "mixture of experts". We
| know LLM's are pretty good for what iv seen some refer to as NLP
| problems, where the model input is the tokenized string input.
| However I would argue an LLM will never built a trained model
| like stockfish or deepseek. Certain model types seem to be suited
| to certain issues/types of problems or inputs. True AGI or SI
| would stop trying to be a grand master of everything but rather
| know what best method/model should be applied to a given problem.
| We still do not know if it is possible to combine the knowledge
| of different types of neural networks like LLMs, convolutional
| neural networks, and deep learning...and while its certainly
| worth exploring, it is foolish to throw all hope on a single
| solution approach. I think the first step would be to create a
| new type of model where given a problem of any type. It knows the
| best method to solve it. And it doesn't rely on itself but rather
| the mixture of agents or experts. And they don't even have to be
| LLMs. They could be anything.
|
| Where this really would explode is, if the AI was able to
| identify a problem that it can't solve and invent or come up with
| a new approach, multiple approaches, because we don't have to be
| the ones who develop every expert.
| wkat4242 wrote:
| Totally agree. An LLM won't be an AGI.
|
| It could be part of an AGI, specifically the human interface
| part. That's what an LLM is good at. The rest (knowledge
| oracle, reasoning etc) are just things that kinda work as a
| side-effect. Other types of AI models are going to be better at
| that.
|
| It's just that since the masses found that they can talk to an
| AI like a human they think that it's got human capabilities
| too. But it's more like fake it till you make it :) An LLM is a
| professional bullshitter.
| lugu wrote:
| I am not sure what you mean by LLM when you say they are
| professional bullshitter. While it was certainly true for
| model based on transformers just doing inference, recent
| models have progressed significantly.
| Terr_ wrote:
| > I am not sure what you mean by LLM when you say they are
| professional bullshitter.
|
| Not parent-poster, but an LLM is a tool for extending a
| document by choosing whatever statistically-seems-right
| based on other documents, and it does so with no
| consideration of worldly facts and no modeling of logical
| prepositions or contradictions. (Which also relates to math
| problems.) If it has been fed on documents with logic
| puzzles and prior tests, it may give plausible answers, but
| tweaking the test to avoid the pattern-marching can still
| reveal that it was a sham.
|
| The word "bullshit" is appropriate because human
| bullshitter is someone who picks whatever "seems right"
| with no particular relation to facts or logical
| consistency. It just doesn't matter to them. Meanwhile, a
| "liar" can actually have a harder job, since they must
| track what is/isn't true and craft a story that is as
| internally-consistent as possible.
|
| Adding more parts around and LLM won't change that: Even if
| you add some external sensors, a calculator, a SAT solver,
| etc. to create a document with facts _in_ it, once you ask
| the LLM to make the document bigger, it 's going to be
| bullshitting the additions.
| Terr_ wrote:
| > It's just that since the masses found that they can talk to
| an AI like a human
|
| In a way it's worse: Even the "talking to" part is an
| illusion, and unfortunately a lot of technical people have
| trouble remembering it too.
|
| In truth, the LLM is an idiot-savant which dreams up
| "fitting" additions to a given document. Some humans have
| prepared a document which is in the form of a a theater-play
| or a turn-based chat transcript, with a pre-written character
| that is often described as a helpful robot. Then the humans
| launch some code that "acts out" any text that looks like it
| came from that fictional character, and inserts whatever the
| real-human-user types as dialogue for the document's human-
| character.
|
| There's zero reason to believe that the LLM is "recognizing
| itself" in the story, or that is is choosing to self-insert
| itself into one of the characters. It's not having a
| conversation. It's not interacting with the world. It's just
| coded to Make Document Bigger Somehow.
|
| > they think that it's got human capabilities too
|
| Yeah, we easily confuse the character with the author. If I
| write an obviously-dumb algorithm which slaps together a
| story, it's still a dumb algorithm no matter how smart the
| robot in the story is.
| pton_xd wrote:
| > However I would argue an LLM will never built a trained model
| like stockfish or deepseek.
|
| It doesn't have to, the LLM just needs access to a computer.
| Then it can write the code for Stockfish and execute it. Or
| just download it, the same way you or I would.
|
| > True AGI or SI would stop trying to be a grand master of
| everything but rather know what best method/model should be
| applied to a given problem.
|
| Yep, but I don't see how that relates to LLMs not reaching AGI.
| They can already write basic Python scripts to answer
| questions, they just need (vastly) more advanced scripting
| capabilities.
| nuancebydefault wrote:
| Indeed! That's what I have been thinking for a while but I
| never had the occasion and or breath to write it down, and you
| explained it concisely. Finally some 'confirmation' 'bias'...
| lukeplato wrote:
| I don't see why a mixture of experts couldn't be distilled into
| a single model and unified latent space
| phaedrus wrote:
| But the G in AGI stands for General. I think the hope is that
| there is some as-yet-undiscovered algorithm for general
| intelligence. While I agree that deferring to a subsystem that
| is an expert in that type of problem is the best way to handle
| problems, I would hope that it is possible that that central
| coordinator not just be able to delegate but design new
| subsystems as needed. Otherwise what happens when you run out
| of types of expert problem solvers to use (and still haven't
| solved the problem well)?
|
| One might argue maybe a mixture of experts is just the best
| that can be done - and that it's unlikely the AGI be able to
| design new experts itself. However where do the limited
| existing expert problem solvers come from? Well - we invented
| them. Human intelligences. So to argue that an AGI could NOT
| come up with its own novel expert problem solvers implies there
| is something ineffable about human general intelligence that
| can't be replicated by machine intelligence (which I don't
| agree with).
| ocean_moist wrote:
| Maybe I just don't understand the article but I really have 0
| clue how they go about making their conclusions and really don't
| understand what they are saying.
|
| I think the 5 issues they provide under "Cognitive Architectures"
| are severely underspecified to the point where they really don't
| _mean_ anything. Because the issues are so underspeficifed I
| don't know how their proposed solution solves their proposed
| problems. If I understand it correctly, they just want agents
| (Assistants/Agents) with user profiles (Sims) on an app store?
| I'm pretty sure this already exists on the ChatGPT store.
| (sims==memories/user profiles, agents==tools/plugins,
| assistants==chat interface)
|
| This whole thing is so broad and full of academic (pejorative)
| platitudes that it's practically meaningless to me. And of course
| although completely unrelated they through a reference into
| symbolic systems. Academic theater.
| spiderfarmer wrote:
| This is publishing for the sake of publishing.
| antisthenes wrote:
| It's a 4-page paper trying to give a summary of 40+ years of
| research on AI.
|
| Of course it's going to be vague and presumptuous. It's more of
| a high-level executive summary for tech-adjacent folks than an
| actual research paper.
| DebtDeflation wrote:
| This whole idea of prompting an LLM and piping the output as the
| input (prompt) of another LLM and asking it to do something with
| it (like critique/edit it) and then piping the output of that LLM
| back to the first LLM along with instructions to keep repeating
| the process until some stop criteria is met seems to me to just
| be a money-making scheme to drive up token consumption.
| beezle wrote:
| For those who _dont_ want to down load the PDF directly and
| prefer to start with the abstract:
| https://arxiv.org/abs/2412.16241
| syntex wrote:
| Why does this have so many upvotes? Is this the current state of
| research nowadays?
___________________________________________________________________
(page generated 2025-01-09 23:00 UTC)