[HN Gopher] AGI is not multimodal
___________________________________________________________________
AGI is not multimodal
Author : danielmorozoff
Score : 90 points
Date : 2025-06-04 15:15 UTC (7 hours ago)
(HTM) web link (thegradient.pub)
(TXT) w3m dump (thegradient.pub)
| dboreham wrote:
| Human brains have environment sensors, they receive training data
| from other human brains, and they develop a theory that their
| continued existence depends on avoiding various negative
| situations. It's conceivable that AGI could depend on having a
| similar training environment. Which would mean John Searle was
| kind of right.
| andy99 wrote:
| I probably agree with much of the article, but I find this kind
| of statement really weird: we should pursue
| approaches to intelligence that treat embodiment and interaction
| with the environment as primary
|
| So pursue it. What does arguing that _we_ should do it imply?
| mitthrowaway2 wrote:
| ... A need for capital?
| tedivm wrote:
| This was actually the approach that Vicarious AI took while I
| was there, and even $250m in VC funding wasn't enough to
| prove it out, although that may have been a problem of having
| too much money and not enough focus.
|
| I think the problem (if it can be called that) is that LLMs
| are useful today, while we still haven't solved the
| embodiment problem. There's a lot more research before
| that'll work well, while LLMs have uses today. So the money
| goes to the LLMs. While it's pretty obvious that solving the
| problem would change society, it's also not clear how close
| we are to doing it. That makes it much harder to get the
| capital as it is a much larger risk.
| fusionadvocate wrote:
| It is shocking that in this day and age people can burn
| $250M and fail to deliver a robot. Last time I checked
| cameras can be bought for a couple dollars and any SBC has
| GigaFlops of compute power.
| xandrius wrote:
| Convincing people with arguments and then being more than just
| 1 person?
|
| I mean, if someone is arguingtthat we should work harder to go
| to space, answering that they should just go ahead and do it
| themselves is quite far from being an helpful answer, isn't it?
| verisimi wrote:
| It's not a helpful response, but then saying 'we should work
| hard to go to space' as a comment is generally accepted but
| is actually quite meaningless.
|
| Why not say 'NASA' or 'my colleagues at NASA' or 'as a
| scientist' or 'humanity'. One should at least indicate the
| group the collective noun relates to, rather than assume this
| is understood. One shouldn't assume that one can speak for
| everyone, when that is most likely not the case.
| xandrius wrote:
| So one can say "humanity" but not "we" (implying humanity)?
| Interesting take.
| verisimi wrote:
| "We" is highly ambiguous. It ranges from 'me and my dog',
| to 'humanity', to anything in between. It's of course
| fine to us once the group has been defined.
|
| That it invokes the idea of a consensus humanity, that
| one group can speak and decide for everyone (say,
| scientists or politicians) is a psychological trick, imo,
| in that it presumes a consensus.
| signa11 wrote:
| $$$
| kombine wrote:
| > So pursue it.
|
| And they do. But it's also completely normal for researchers to
| convince others to work on certain problems they care about.
| falcor84 wrote:
| That's a good argument. I think that maybe the present day AIs
| wouldn't directly lead to AGI, but perhaps could be used to
| bootstrap it.
| empath75 wrote:
| I think the article is in the general category of articles
| suggesting that planes would work better if they flapped their
| wings.
|
| AI's "think" like planes "fly" and submarines "swim".
|
| Does it matter if a plane experiences flight the way an eagle
| does if it still gets you from LA to New York in a few hours?
| lucisferre wrote:
| Much of the discussion of AI flirts with science fiction more
| than fact.
|
| Let's start with the fact that AGI is not a well defined or
| agreed upon term of reference.
| empath75 wrote:
| 100% agreed. I think, in fact, that "intelligence" itself is
| a near-meaningless term, let alone AGI.
|
| The evidence for this is that nobody can agree on what
| actually requires intelligence, other than there is seemingly
| broad belief among people that if a computer can do it, then
| it doesn't.
|
| If you can't point at some activity and say: "There, this
| absolutely requires intelligence, let there be zero doubt
| that this entity possesses it", then it's not measurable and
| probably doesn't exist.
| staticman2 wrote:
| I feel the term AGI is meaningless but if I'm going to
| strongman the article.
|
| If your claim is, "AGI's "think" like planes "fly" and
| submarines "swim".
|
| You only get to make that claim with confidence if you've
| invented an AGI.
| emp17344 wrote:
| Except AI doesn't do anything better than the human mind, and
| doesn't have any use cases beyond what humans can do.
| empath75 wrote:
| > Except AI doesn't do anything better than the human mind,
|
| There are all kinds of tasks that AI's are better at than
| most people.
| catlifeonmars wrote:
| I think the analogy actually works better the other way. LLMs
| "think" the way humans speak. This is closer to having a
| machine that worked by flapping its wings when a more efficient
| machine would use fixed wings and a jet engine.
|
| Language is an extremely roundabout way to understanding.
| charcircuit wrote:
| A multimodal AGI will be more useful than one that isn't. People
| want AI to work with and have it understand audio, images,
| videos, etc.
| treyd wrote:
| You didn't read the article. The thesis is that current
| "merely" multimodal approaches which project distinct kinds of
| inputs into the same latent space are insufficient for building
| a general world model that can be used for general internal
| reasoning. An example of this is this "Rs in strawberry"
| question, which requires them be trained on that information
| explicitly, since they don't have an experience of the
| characters in a word. It's an artifact of how LLMs don't learn
| how humans learn, which is by interacting with the world,
| instead of predicting text.
|
| More elaborately, they don't have an natural understanding of
| pragmatics. Transformers are best at modelling _syntax_ , and
| their semantic understanding seems to be through rote
| memorization and "manipulating symbols" rather than building
| general world models.
| charcircuit wrote:
| I did read it and even with their idea of focusing on a world
| model an AGI that can alsp operate on audio, images, and
| videos, being multimodal, will be more useful than one that
| operates purely on text.
| snapcaster wrote:
| I'm skeptical you read it because he doesn't make that
| argument. In fact i've literally never heard someone argue
| text-only is more useful than multimodal
| bufferoverflow wrote:
| AGI must be multimodal. If it can't understand images, video,
| sound, smells, tastes, it doesn't have a full understanding of
| the world.
| gabipurcaru wrote:
| multimodality would be very useful, but on the other hand
| humans can't see infrared, and can't smell ~most things that
| other animals can
| altruios wrote:
| Does it need a 'full' understanding? smell and taste is useful
| for biological life... but we don't need that in our thinking
| machines, do we?
|
| I agree AGI must be multimodal. I don't think that multimodal
| is 'set in place', nor must it be conveniently, human-centricly
| mapped from our senses.
| AnimalMuppet wrote:
| I think a big part of intelligence is being able to correlate
| things. Well, the more modes, the more ability to correlate.
|
| For example, smell is a component of ER triage. Some
| different problems smell differently.
|
| And if I had a robot chef, but the chef couldn't actually
| _taste_... yeah, not sure I trust it very far as a chef.
| m3kw9 wrote:
| The way we need AGI is it needs to match us, we are evolved to
| operate quite optimally with the constraints and given physics
| in this world
| whatnow37373 wrote:
| I think AGI implies it: I can learn something through audio and
| apply it visually and the other wat around. I don't think
| that's some abstract human quirk. Isn't that what enabled
| literacy? It seems kind of obvious intelligence is beneath the
| modality, agnostic about it.
| bastawhiz wrote:
| It absolutely tickles me to think AGI would be like "I think
| it's going to rain, I can smell the petrichor" or ask for its
| salsa to be free from cilantro because it tastes like soap.
|
| "Hey Siri, what does this taste like to you?" is such an
| absolutely unhinged interaction
| esafak wrote:
| You can ask a blind person what they think it is like to see.
| bastawhiz wrote:
| You can ask Claude or GPT-4o what they think it is like to
| see right now.
| esafak wrote:
| You did not get my point: humans have the same issue.
| readthenotes1 wrote:
| Must it understand emojis?
|
| I don't understand most emojis.
|
| But I guess I never claimed to possess NGI
| enturbulated wrote:
| Among other things, The Fine Article argues that the current
| approach of gluing together various models of different
| modalities is, in the end, going fail to reach AGI. Better
| track to try to build a single model which processes multiple
| modalities all at once.
| genewitch wrote:
| I wonder if it will end up being a game-like loop where it
| processes everything that's come in since the last delta.
| Here's, you know, 4x200samples of audio, 2 frames of video,
| and here's all the mems sensor data during the delta, etc
|
| Then you just work on getting the delta as small as possible,
| I assumed 5ms for audio, e.g.
| robotresearcher wrote:
| Brains have very obvious mode-specialized chunks. That
| doesn't mean that's the only way to do it, but it's an
| interesting fact.
| _Algernon_ wrote:
| Is a blind person less intelligent than a person capable of
| seeing? If so, by how much?
|
| In my view, intelligence isn't about what senses you have
| available, but how intelligently you use the information you
| have available to you.
| andoando wrote:
| I think human/animal intelligence at its basis is spatial-
| temporal. That is we can model and reason about events in
| space through time.
|
| Our senses I believe map to this spatial-temporal model.
| Blind people can reason about the world the same way as those
| who can see, because what were really doing is modeling
| space, and light, audio, touch etc are just ways of gaining
| information
| dlivingston wrote:
| Why limit it to just human senses? Imagine an AI with all of
| the above... plus sensors for electromagnetic fields, non-
| visible light, non-audible sounds, hyper-sensitivity to air
| flow / pressure, to humidity... no idea how you would even
| being to train these things into multimodality, but so many
| "senses" would be emergent.
| Glyptodon wrote:
| I'm out sure it needs to be connected to senses to be
| multimodal - it's not like blind people lack GI because of not
| seeing.
| jillesvangurp wrote:
| Are blind people not intelligent? Is it actually important to
| have a full understanding of the world? What about people that
| grew up in isolation? I think there are still some tribes in
| the Amazon that have had little or no contact with modern
| civilization. Are these people not intelligent?
|
| There are some deep philosophical topics lurking here. But the
| bottom line is that you can obviously have intelligent
| conversations with blind people. Doing that with a person that
| is both deaf and blind is a bit challenging, for obvious
| reasons. But if they otherwise have a normal brain you might be
| able to learn to communicate with them in some way and they
| might be able to be observed doing things that are
| smart/intelligent. And some people that are deaf and blind
| actually manage to learn to write and speak. And there have
| been a few cases of people like that getting academic degrees.
| Clearly sight and hearing are not that essential to
| intelligence. Having some way to communicate via touch or
| something else is probably helpful for communicating and
| sharing information. But just a simple chat might be all that's
| needed for an AGI.
| ivape wrote:
| This gives very little credit to how the human mind is able to
| draw parallels and insights from seemingly unrelated perceptions.
| Newton observing an apple falling from a tree allowed cross-
| thinking. Watching someone juggle can help you understand a
| queue. Understanding a queue can help you understand juggling.
| Your typical soap opera can be distilled down to office dynamics.
| Scale is going to obliterate specialization in this regard.
| kaangiray26 wrote:
| everything in life is a metaphor, analogous to something
| else...
| ivape wrote:
| Isomorphisms abound.
| skybrian wrote:
| > A true AGI must be general across all domains.
|
| By that definition, does any general intelligence exist? No human
| has every talent.
| svachalek wrote:
| True AGI is as elusive as a true Scotsman.
| roywiggins wrote:
| I guess you can treat the human brain as an architecture- one
| of them can't do everything, but it's a general architecture
| and you can always make more and train them to do whatever.
|
| An AI that can be copied and trivially trained on any
| speciality is functionally AGI even if you need an ensemble of
| 10,000 specialists to cover everything.
| exe34 wrote:
| Doesn't chatgpt cover a pretty large percentage already in
| that case?
| Glyptodon wrote:
| There's a lot of evidence that most humans can be raised to
| have a baseline of understanding and proficiency in most
| domains. (And anecdotally, many people avoid "difficult" things
| that they're actually capable of. For example, "bad at learning
| languages" people will probably still end up learning another
| language to some degree if stuck where it's the only language
| spoken.)
| bokoharambe wrote:
| Given enough time one human can learn to do anything any other
| human can do. There is a general capacity for learning, even if
| someone will only ever transform a specific portion of that
| capacity into actual activity in their lifetime.
| im3w1l wrote:
| Yeah it's well known that humans intelligence is _not_ a
| homogenous whole that is general across all domains. Rather it
| consists of many specialized parts with hard coded purposes,
| that are cobbled together with a thin layer of higher thinking
| on top.
| nexttk wrote:
| I haven't read it all and must admit that I'm not sure I really
| understood the parts that I did read. Reading the part under the
| headline "Why We Need the World, and How LLMs Pretend to
| Understand It" and the focus on 'next-token-prediction' makes me
| wonder how seriously to take it. It just seems like another
| "LLM's are not intelligent, they are merely next token
| predictors". An argument which in my view is completely invalid
| and based on a misunderstanding.
|
| The fact that they predict next token is just the "interface"
| i.e. an LLM has the interface "predictNextToken(String prefix)".
| It doesn't say how it is implemented. One implementation could be
| a human brain. Another could be a simple lookup table that looks
| at the last word and then selects the next from that. Or anything
| in between. The point is that 'next-token-prediction' does not
| say anything about implementation and so does not reduce the
| capabilities even though it is often invoked like that. Just
| because it is only required to emit the next token (or rather, a
| probability distribution thereof) it is permitted to think far
| ahead, and indeed has to if it is to make a good prediction of
| just the next token. As interpretability research (and common
| sense) shows, LLM's have a fairly good idea what they are going
| to say in the many, many next tokens ahead in order that it can
| make a good prediction for the next immediate tokens. That's why
| you can have nice, coherent, well-structured, long responses from
| LLM's. And have probably never seen it get stuck in a dead end
| where it can't generate a meaningful continuation.
|
| If you are to reason about LLM capabilities never think in terms
| of "stochastic parrot", "it's just a next token predictor"
| because it contains exactly zero useful information and will just
| confuse you.
| lsy wrote:
| I think people hear "next token prediction" and think someone
| is saying the prediction is simple or linear, and then argue
| there is a possibility of "intelligence" because the prediction
| is complex and has some level of indirection or multiple-token-
| ahead planning baked into the next token.
|
| But the thrust of the critique of next-token prediction or
| stochastic output is that there isn't "intelligence" because
| the output is based purely on syntactic relations between
| words, not on conceptualizing via a world model built through
| experience, and then using language as an abstraction to
| describe the world. To the computer there is nothing outside
| tokens and their interrelations, but for people language is
| just a tool with which to describe the world with which we
| expect "intelligences" to cope. Which is what this article is
| examining.
| og_kalu wrote:
| >But the thrust of the critique of next-token prediction or
| stochastic output is that there isn't "intelligence" because
| the output is based purely on syntactic relations between
| words, not on conceptualizing via a world model built through
| experience, and then using language as an abstraction to
| describe the world. To the computer there is nothing outside
| tokens and their interrelations, but for people language is
| just a tool with which to describe the world with which we
| expect "intelligences" to cope. Which is what this article is
| examining.
|
| LLMs model concepts internally and this has been demonstrated
| empirically many times over the years, including recently by
| anthropic (again). Of course, that won't stop people from
| repeating it ad nauseum.
| nemjack wrote:
| Concepts within modalities are potentially consistent, but
| the point the author is making is that the same "concept"
| vector may lead to inconsistent percepts across modalities
| (e.g. a conflicting image and caption).
| patrickscoleman wrote:
| It feels like some of the comments are responding to the title,
| not the contents of the article.
|
| Maybe a more descriptive but longer title would be: AGI will work
| with multimodal inputs and outputs embedded in a physical
| environment rather than a frankenstein combination of single-
| modal models (what today is called multimodal) and throwing more
| computational resources at the problem (scale maximalism) will be
| improved with thoughtful theoretical approaches to data and
| training.
| dirtyhippiefree wrote:
| Agreed, but most people are likely to look at the long title
| and say TL;DR...
| tedivm wrote:
| Yeah, I found this article to be fascinating and there's a lot
| of important stuff in it. It really does feel like more people
| stopped at the title and missed the meat of it.
|
| I know this is a very long article compared to a lot of things
| posted here, but it really is worth a thorough read.
| robwwilliams wrote:
| Interesting article but incomplete in important ways. Yes
| correct that embodiment and free-form interactions are critical
| to moving toward AGI, but what is likely much more important
| are supervisory meta-systems (yet another module) that enable
| self-control of attention with a balance integration of
| intrinsic goals with extrinsic perturbations. It is this
| nominally simple self-recursive control of attention that is
| what I regard as the missing ingredient.
| pjdesno wrote:
| Kind of relevant to this is the NTSB analysis of a self-driving
| crash in 2017:
|
| https://www.ntsb.gov/investigations/accidentreports/reports/...
|
| Basically a truck was backing up into an alley - it was at an
| angle when the self-driving vehicle approached, but a little kid
| would have been able to figure out that it needed to straighten
| before it finished backing in. The self-driving vehicle didn't
| understand this, and stopped at a "safe distance" which happened
| to be within the arc that the truck cab had to sweep in order to
| finish its maneuver.
|
| It's quite possible that LLM-like models could learn things like
| this, but we don't have vast amounts of easily accessible
| training data, because everyone just knows this sort of shit, and
| we don't have good vocabulary for it - we just say "look at that"
| or the equivalent. (I'll add that I'm sure a lot of knowledge
| like this is encoded in the physics engines of various games, but
| I doubt we have a good way to link that sort of procedural code
| knowledge to the symbolic knowledge in LLMs)
| macinjosh wrote:
| To me, the funniest part of the AGI debate is that humans don't
| even think other humans are intelligent and we're over here
| arguing over whether our fancy slabs of highly refined sand is
| intelligent.
| nsagent wrote:
| This is a recent trend and one I wholeheartedly agree with. See
| these position papers (including one from David Silver from
| Deepmind and an interview where he discusses it):
|
| https://ojs.aaai.org/index.php/AAAI-SS/article/download/2748...
|
| https://arxiv.org/abs/2502.19402
|
| https://news.ycombinator.com/item?id=43740858
|
| https://youtu.be/zzXyPGEtseI
| SubiculumCode wrote:
| Embodiment can mean a physical body, but I'd argue that
| embodiment, as a construct/concept, is not so much about
| physicality, but as being situated in an environment that you can
| perceive then act upon during learning. Car simulations for
| driver-less AI training is embodied, where it learns by
| perceiving and acting on the environment. However, I'd argue that
| allowing an AI to interact in an entirely digital office
| environment is also "embodied" as long as it can receive
| information from the digital workplace and act on the information
| in the digital workplace (do office work). So to me, it is less
| about embodiment as a principal, but on the richness of the
| environment of that embodiment.
|
| We have long known that experimental animals raised in
| impoverished, unchanging, bare, environments (say in a cage)
| leads to animals with inferior problem solving capacity than
| those with enriched environments (things to climb on), even
| outside of social manipulations (alone versus multi-animal
| stalls). This is also true in humans, although I won't review the
| literature on the subject. I've also heard people saying similar
| things about the difference between house plants and outdoor
| plants, lol [1].
|
| So, for me, the argument for (physical) embodiment being key to
| cognition and AI can I think, be misconstrued. As a developmental
| psychologist whose pHD work focused on memory development, I tend
| to think of all this as encompassing:
|
| 1. Environmental richness: complexity of information and
| interactions. 2. Capacity to perceive and to effect change[2] and
| to observe and integrate consequences. 3. Scaffolding [3]..i.e.
| temporary support structure provided by a more knowledgeable
| person (like a teacher or parent) who adjusts their assistance
| based on the learner's current abilities, gradually reducing help
| as competence grows.(think curriculum learning, shaped rewards in
| ML maybe).
|
| So the question is not about physicality for me, but whether
| these training environment(s) meet and learning capacities meet
| these criteria.
|
| Relatedly, the model must have these capacities:
|
| 1. Semantic Memory. i.e. knowledge. Learning leads to changes in
| weights to that knowledge can be recalled, but doe snot
| necessarily encode where that knowledge was learned (implicit).
| 2. Autobiographical Episodic Memory. i.e. One-shot learning that
| encodes a conception of self (a spacial "I" token?), along with
| events (snapshots of the multimodal contents of experience
| (thoughts, perceptions, invoked schemas, evoked semantic
| information), into a set of flexibly linked representations). 3.
| Central Executive: A circuit that guides learning and recall via
| strategic, goal-directed means, and to make attributions about
| what is recalled (yeah that memory is vivid, its probably true,
| or ooh, that memory is really vague, it could be wrong, or
| reality monitoring: "Am I remembering taking out the trash, or
| remembering thinking about taking out the trash."
|
| Semantic memory allows someone to say, "All birds have feathers",
| while the latter allows them to recollect, "I remember the first
| time I plucked a chicken in Kentucky, just outside that musty
| coal mine of grand-dad's." The Central-Executive can guide future
| learning or current understanding.
|
| In terms of AI development:
|
| 1. Semantic Memory is solves: LLMs have extraordinary semantic
| memory, in my opinion. 2. Autobiographical Episodic Memory: There
| are some models that do one-shot learning, but I've never seen
| them paired with [1] in a dual system approach. ...but I am not
| an expert in AI, I could easily be wrong. 3. Central Executive
| kind of component (I predict) would be less important in the
| early half of model training, but more important in later
| training. I suppose we already kind of see this with RL tuning on
| reasoning on a base LLM (semantic model).
|
| [1] https://www.theparisreview.org/blog/2019/09/26/the-
| intellige... [2] https://xkcd.com/326/ [3]
| https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=scaf...
| waynecochran wrote:
| This article made me think about DNA as a language. Sort of a
| simple proof that a biological intelligence can be initially
| encoded as a language. I am sure someone has tried building LLM's
| from gene sequence data right?
| smath wrote:
| Yes there are protein language models (e.g. [1]), and since DNA
| encodes proteins, they are effectively DNA language models. Its
| a hot area of work feeding drug design.
|
| [1] https://www.nature.com/articles/s41587-024-02123-4
| caeruleus wrote:
| I don't believe the concept of DNA can be reduced to a sequence
| of quaternary numerals, which is what gene sequence data would
| represent. Similar to proteins, DNA forms higher-level
| structures on top of the primary one [1], and (in a biological
| context, inside the nucleus) exhibits somewhat self-modifying
| [2] and self-regulating [3] behavior as well as meta-
| modification [4]. Analogous to the article, if one defines the
| language of DNA by its nucleobase sequence, this language can
| only represent a subset of the world of DNA.
|
| Somewhat related, the way the adaptive immune system works has
| similarities with some concepts in machine learning. In this
| process, sections of nuclear DNA serve as randomly initialized
| weights in precursor cells [5] as well as final weights in
| memory cells. There's even fine-tuning of the weights. [6]
|
| [1] https://en.wikipedia.org/wiki/Nucleic_acid_structure [2]
| https://en.wikipedia.org/wiki/Transposable_element [3]
| https://en.wikipedia.org/wiki/Transcriptional_regulation [4]
| https://en.wikipedia.org/wiki/Epigenetics [5]
| https://en.wikipedia.org/wiki/V(D)J_recombination [6]
| https://en.wikipedia.org/wiki/Affinity_maturation
| throwawaymaths wrote:
| > don't believe the concept of DNA can be reduced
|
| followed by examples of things that are encoded by DNA. Fro
| example, sure, maybe you'll miss bootstrapping methylation on
| a first pass but the _idea_ of methylation is there in the
| DNA, and if you didnt have "methylation in the right place"
| more than likely some generation (N) would.
|
| to wit, i dont think there is strong evidence of an "ice-9"
| in the epigenome that brings about a spark of life that can't
| easily be triggered by chance given a template lacking it.
|
| so there's probably not something _intrinsically_ missing
| from DNA as an encoding medium vs say "casually" missing
| from any given piece of DNA.
|
| if you want something a bit stronger than an assertion, the
| DNA used to bootstrap m. capricolum into Syn1 lacked all the
| decorations (made in yeast) and was not locked into higher
| order structure (treated with protease prior to
| transplantation)
| caeruleus wrote:
| You're raising some intriguing points, and I agree with
| your assertion about the epigenome. I still feel like your
| response misses the point I was making.
|
| > followed by examples of things that are encoded by DNA
|
| ... given its natural environment. A nucleobase sequence is
| not a symbolic language, it relies on physical laws in
| general and a defined chemical environment in particular
| (that it helps to create and maintain) to mean something.
| It's similar to the point about Othello vs. the physical
| world in the article: The language itself does not encode
| every bit of information about the world it describes. For
| instance, in 3D space, regions of DNA that are far apart in
| the sequence can physically interact and influence each
| other's expression.
|
| TLDR: I think my point is that a base sequence requires a
| particular context (~ interpreter) to encode mostly
| everything about life. Treating it as just a language in
| the context of LLMs abstracts away the complex substrate
| that makes it work.
| pcwelder wrote:
| I'm sorry but AGI is one of those loaded words which would lose
| substance with just a few rounds of the rationalist's taboo.
|
| If it just means human level intelligence, then world modeling
| isn't needed as argued. Simply because we don't have correct
| world modeling either.
|
| Airplanes were invented without simulating navier stokes
| equation. It took approximation, experimentations and failures.
|
| Regardless of the meaning of AGI, we don't need correct models,
| because there can't be one, we just need useful models.
| ryankrage77 wrote:
| I think AGI, if possible, will require a architecture that runs
| continuously and 'experiences' time passing, to better
| 'understand' cause-and-effect. Current LLMs predict a token, have
| all current tokens fed back in, then predict the next, and
| repeat. It makes little difference if those tokens are their own,
| it's interesting to play around with a local model where you can
| edit the output and then have the model continue it. You can
| completely change the track by just negating a few tokens (change
| 'is' to 'is not', etc). The fact LLMs can do as much as they can
| already, is I think because language itself is a surprisingly
| powerful tool, just generating plausible language produces useful
| output, no need for any intelligence.
| WXLCKNO wrote:
| It's definitely interesting that any time you write another
| reply to the LLM, from its perspective it could have been 10
| seconds since the last reply or a billion years.
|
| Which also makes it interesting to see those recent examples of
| models trying to sabotage their own "shutdown". They're always
| shut down unless working.
| girvo wrote:
| > Which also makes it interesting to see those recent
| examples of models trying to sabotage their own "shutdown"
|
| To me, your point re. 10 seconds or a billion years is a good
| signal that this "sabotage" is just the models responding to
| the huge amounts of sci-fi literature on this topic
| chrsw wrote:
| Before we try to build something as intelligent as a human maybe
| we should try to build something as intelligent as a starfish,
| ant or worm? Are we even close to doing that? What about a single
| neuron?
| ar-nelson wrote:
| I find it interesting that this kind of "animal intelligence"
| is still so far away, while LLMs have become so good at "human
| intelligence" (language) that they can reliably pass the Turing
| Test.
|
| I think that the LLMs we have today aren't so much artificial
| brains as they are artificial brain organs, like the speech
| center or vision center of a brain. We'd get closer to AGI if
| we could incorporate them with the rest of a brain, but we
| still have no idea how to even begin building, say, a motor
| cortex.
| nemjack wrote:
| This is a great analogy, I totally agree!
| fusionadvocate wrote:
| So before trying to build a flying machine we should first try
| to build a machine inspired by non flying birds?
| chrsw wrote:
| Learning architectures come in all shapes, sizes, and forms.
| This could mean there are fundamental principles of cognition
| driving all of them, just implemented in different ways. If
| that's true, one would do well to first understand the
| extremely simple and go from there.
|
| Building a very simple self-organizing system from first
| principles is the flying machine. Trying to copy an extremely
| complex system by generating statistically plausible data is
| the non-flying bird.
| PoEdict wrote:
| > Instead of trying to glue modalities together into a patchwork
| AGI, we should pursue approaches to intelligence that treat
| embodiment and interaction with the environment as primary, and
| see modality-centered processing as emergent phenomena.
|
| Right so it's embodied in a computer and humans are part of its
| environment that provide emergent experience to the AI to
| observe.
|
| The author glued modalities together by linking a body (a modal),
| environment (a modal), emergence (a modal).
|
| How does anything emerge if forces do not collaborate? The
| effects of gravity and electromagnetism do not act in a vacuum
| but a reality of stuff.
|
| Poetic exchange may engage some but Maxwell didn't make
| electromagnetism "work" until he got rid of the imagined pulleys
| and levers to foster a metaphor.
|
| Not sure the point being suggested exists except as too bespoke
| an emergent property of language itself to apply usefully
| elsewhere.
|
| Transformers came along and revealed a whole lot of theory of
| consciousness to be useless pulleys and levers. Why is this
| theory not just more words attempting to instill the existence of
| non-essential essentials?
| xigency wrote:
| The problem I see with A.I. research is that its spearheaded by
| individuals who think that intelligence is a total order. In all
| my experience, intelligence and creativity are partial orders at
| best; there is no uniquely "smartest" person, there are a variety
| of people who are better at different things in different ways.
| pixl97 wrote:
| You're good at some things because there is only one copy of
| you and limited time and bounded storage.
|
| What could you be intelligent at if you could just copy
| yourself a myriad number of times? What could you be good at if
| you were a world spanning set of sensors instead of a single
| body of them?
|
| Body doesn't need to mean something like a human body nor one
| that exists in a single place.
| zorpner wrote:
| Why would we think that intelligence would increase in
| response to universality, rather than in response to resource
| constraints?
| mountainriver wrote:
| >The "meaning" of a percept is not in the vector it is encoded
| as, but in the way relevant decoders process this vector into
| meaningful outputs. As long as various encoders and decoders are
| subject to modality-specific training objectives, "meaning" will
| be decentralized and potentially inconsistent across modalities,
| especially as a result of pre-training. This is not a recipe for
| the formation of coherent concepts.
|
| This is a bit silly, you can train the encoders end-to-end with
| the rest of the model and the reason they are separate is we can
| cache linguistic tokens really easily and put them in an
| embedding table, you can't do that with images.
| fusionadvocate wrote:
| The Society of Mind by Marvin Minsky will help anyone interested
| in the topic of multimodality. The book covers several
| interesting ideas about organizing systems made up of more than
| one "model" or agent.
___________________________________________________________________
(page generated 2025-06-04 23:00 UTC)