[HN Gopher] AGI is not multimodal
       ___________________________________________________________________
        
       AGI is not multimodal
        
       Author : danielmorozoff
       Score  : 90 points
       Date   : 2025-06-04 15:15 UTC (7 hours ago)
        
 (HTM) web link (thegradient.pub)
 (TXT) w3m dump (thegradient.pub)
        
       | dboreham wrote:
       | Human brains have environment sensors, they receive training data
       | from other human brains, and they develop a theory that their
       | continued existence depends on avoiding various negative
       | situations. It's conceivable that AGI could depend on having a
       | similar training environment. Which would mean John Searle was
       | kind of right.
        
       | andy99 wrote:
       | I probably agree with much of the article, but I find this kind
       | of statement really weird:                 we should pursue
       | approaches to intelligence that treat embodiment and interaction
       | with the environment as primary
       | 
       | So pursue it. What does arguing that _we_ should do it imply?
        
         | mitthrowaway2 wrote:
         | ... A need for capital?
        
           | tedivm wrote:
           | This was actually the approach that Vicarious AI took while I
           | was there, and even $250m in VC funding wasn't enough to
           | prove it out, although that may have been a problem of having
           | too much money and not enough focus.
           | 
           | I think the problem (if it can be called that) is that LLMs
           | are useful today, while we still haven't solved the
           | embodiment problem. There's a lot more research before
           | that'll work well, while LLMs have uses today. So the money
           | goes to the LLMs. While it's pretty obvious that solving the
           | problem would change society, it's also not clear how close
           | we are to doing it. That makes it much harder to get the
           | capital as it is a much larger risk.
        
             | fusionadvocate wrote:
             | It is shocking that in this day and age people can burn
             | $250M and fail to deliver a robot. Last time I checked
             | cameras can be bought for a couple dollars and any SBC has
             | GigaFlops of compute power.
        
         | xandrius wrote:
         | Convincing people with arguments and then being more than just
         | 1 person?
         | 
         | I mean, if someone is arguingtthat we should work harder to go
         | to space, answering that they should just go ahead and do it
         | themselves is quite far from being an helpful answer, isn't it?
        
           | verisimi wrote:
           | It's not a helpful response, but then saying 'we should work
           | hard to go to space' as a comment is generally accepted but
           | is actually quite meaningless.
           | 
           | Why not say 'NASA' or 'my colleagues at NASA' or 'as a
           | scientist' or 'humanity'. One should at least indicate the
           | group the collective noun relates to, rather than assume this
           | is understood. One shouldn't assume that one can speak for
           | everyone, when that is most likely not the case.
        
             | xandrius wrote:
             | So one can say "humanity" but not "we" (implying humanity)?
             | Interesting take.
        
               | verisimi wrote:
               | "We" is highly ambiguous. It ranges from 'me and my dog',
               | to 'humanity', to anything in between. It's of course
               | fine to us once the group has been defined.
               | 
               | That it invokes the idea of a consensus humanity, that
               | one group can speak and decide for everyone (say,
               | scientists or politicians) is a psychological trick, imo,
               | in that it presumes a consensus.
        
         | signa11 wrote:
         | $$$
        
         | kombine wrote:
         | > So pursue it.
         | 
         | And they do. But it's also completely normal for researchers to
         | convince others to work on certain problems they care about.
        
       | falcor84 wrote:
       | That's a good argument. I think that maybe the present day AIs
       | wouldn't directly lead to AGI, but perhaps could be used to
       | bootstrap it.
        
       | empath75 wrote:
       | I think the article is in the general category of articles
       | suggesting that planes would work better if they flapped their
       | wings.
       | 
       | AI's "think" like planes "fly" and submarines "swim".
       | 
       | Does it matter if a plane experiences flight the way an eagle
       | does if it still gets you from LA to New York in a few hours?
        
         | lucisferre wrote:
         | Much of the discussion of AI flirts with science fiction more
         | than fact.
         | 
         | Let's start with the fact that AGI is not a well defined or
         | agreed upon term of reference.
        
           | empath75 wrote:
           | 100% agreed. I think, in fact, that "intelligence" itself is
           | a near-meaningless term, let alone AGI.
           | 
           | The evidence for this is that nobody can agree on what
           | actually requires intelligence, other than there is seemingly
           | broad belief among people that if a computer can do it, then
           | it doesn't.
           | 
           | If you can't point at some activity and say: "There, this
           | absolutely requires intelligence, let there be zero doubt
           | that this entity possesses it", then it's not measurable and
           | probably doesn't exist.
        
         | staticman2 wrote:
         | I feel the term AGI is meaningless but if I'm going to
         | strongman the article.
         | 
         | If your claim is, "AGI's "think" like planes "fly" and
         | submarines "swim".
         | 
         | You only get to make that claim with confidence if you've
         | invented an AGI.
        
         | emp17344 wrote:
         | Except AI doesn't do anything better than the human mind, and
         | doesn't have any use cases beyond what humans can do.
        
           | empath75 wrote:
           | > Except AI doesn't do anything better than the human mind,
           | 
           | There are all kinds of tasks that AI's are better at than
           | most people.
        
         | catlifeonmars wrote:
         | I think the analogy actually works better the other way. LLMs
         | "think" the way humans speak. This is closer to having a
         | machine that worked by flapping its wings when a more efficient
         | machine would use fixed wings and a jet engine.
         | 
         | Language is an extremely roundabout way to understanding.
        
       | charcircuit wrote:
       | A multimodal AGI will be more useful than one that isn't. People
       | want AI to work with and have it understand audio, images,
       | videos, etc.
        
         | treyd wrote:
         | You didn't read the article. The thesis is that current
         | "merely" multimodal approaches which project distinct kinds of
         | inputs into the same latent space are insufficient for building
         | a general world model that can be used for general internal
         | reasoning. An example of this is this "Rs in strawberry"
         | question, which requires them be trained on that information
         | explicitly, since they don't have an experience of the
         | characters in a word. It's an artifact of how LLMs don't learn
         | how humans learn, which is by interacting with the world,
         | instead of predicting text.
         | 
         | More elaborately, they don't have an natural understanding of
         | pragmatics. Transformers are best at modelling _syntax_ , and
         | their semantic understanding seems to be through rote
         | memorization and "manipulating symbols" rather than building
         | general world models.
        
           | charcircuit wrote:
           | I did read it and even with their idea of focusing on a world
           | model an AGI that can alsp operate on audio, images, and
           | videos, being multimodal, will be more useful than one that
           | operates purely on text.
        
             | snapcaster wrote:
             | I'm skeptical you read it because he doesn't make that
             | argument. In fact i've literally never heard someone argue
             | text-only is more useful than multimodal
        
       | bufferoverflow wrote:
       | AGI must be multimodal. If it can't understand images, video,
       | sound, smells, tastes, it doesn't have a full understanding of
       | the world.
        
         | gabipurcaru wrote:
         | multimodality would be very useful, but on the other hand
         | humans can't see infrared, and can't smell ~most things that
         | other animals can
        
         | altruios wrote:
         | Does it need a 'full' understanding? smell and taste is useful
         | for biological life... but we don't need that in our thinking
         | machines, do we?
         | 
         | I agree AGI must be multimodal. I don't think that multimodal
         | is 'set in place', nor must it be conveniently, human-centricly
         | mapped from our senses.
        
           | AnimalMuppet wrote:
           | I think a big part of intelligence is being able to correlate
           | things. Well, the more modes, the more ability to correlate.
           | 
           | For example, smell is a component of ER triage. Some
           | different problems smell differently.
           | 
           | And if I had a robot chef, but the chef couldn't actually
           | _taste_... yeah, not sure I trust it very far as a chef.
        
         | m3kw9 wrote:
         | The way we need AGI is it needs to match us, we are evolved to
         | operate quite optimally with the constraints and given physics
         | in this world
        
         | whatnow37373 wrote:
         | I think AGI implies it: I can learn something through audio and
         | apply it visually and the other wat around. I don't think
         | that's some abstract human quirk. Isn't that what enabled
         | literacy? It seems kind of obvious intelligence is beneath the
         | modality, agnostic about it.
        
         | bastawhiz wrote:
         | It absolutely tickles me to think AGI would be like "I think
         | it's going to rain, I can smell the petrichor" or ask for its
         | salsa to be free from cilantro because it tastes like soap.
         | 
         | "Hey Siri, what does this taste like to you?" is such an
         | absolutely unhinged interaction
        
           | esafak wrote:
           | You can ask a blind person what they think it is like to see.
        
             | bastawhiz wrote:
             | You can ask Claude or GPT-4o what they think it is like to
             | see right now.
        
               | esafak wrote:
               | You did not get my point: humans have the same issue.
        
         | readthenotes1 wrote:
         | Must it understand emojis?
         | 
         | I don't understand most emojis.
         | 
         | But I guess I never claimed to possess NGI
        
         | enturbulated wrote:
         | Among other things, The Fine Article argues that the current
         | approach of gluing together various models of different
         | modalities is, in the end, going fail to reach AGI. Better
         | track to try to build a single model which processes multiple
         | modalities all at once.
        
           | genewitch wrote:
           | I wonder if it will end up being a game-like loop where it
           | processes everything that's come in since the last delta.
           | Here's, you know, 4x200samples of audio, 2 frames of video,
           | and here's all the mems sensor data during the delta, etc
           | 
           | Then you just work on getting the delta as small as possible,
           | I assumed 5ms for audio, e.g.
        
           | robotresearcher wrote:
           | Brains have very obvious mode-specialized chunks. That
           | doesn't mean that's the only way to do it, but it's an
           | interesting fact.
        
         | _Algernon_ wrote:
         | Is a blind person less intelligent than a person capable of
         | seeing? If so, by how much?
         | 
         | In my view, intelligence isn't about what senses you have
         | available, but how intelligently you use the information you
         | have available to you.
        
           | andoando wrote:
           | I think human/animal intelligence at its basis is spatial-
           | temporal. That is we can model and reason about events in
           | space through time.
           | 
           | Our senses I believe map to this spatial-temporal model.
           | Blind people can reason about the world the same way as those
           | who can see, because what were really doing is modeling
           | space, and light, audio, touch etc are just ways of gaining
           | information
        
         | dlivingston wrote:
         | Why limit it to just human senses? Imagine an AI with all of
         | the above... plus sensors for electromagnetic fields, non-
         | visible light, non-audible sounds, hyper-sensitivity to air
         | flow / pressure, to humidity... no idea how you would even
         | being to train these things into multimodality, but so many
         | "senses" would be emergent.
        
         | Glyptodon wrote:
         | I'm out sure it needs to be connected to senses to be
         | multimodal - it's not like blind people lack GI because of not
         | seeing.
        
         | jillesvangurp wrote:
         | Are blind people not intelligent? Is it actually important to
         | have a full understanding of the world? What about people that
         | grew up in isolation? I think there are still some tribes in
         | the Amazon that have had little or no contact with modern
         | civilization. Are these people not intelligent?
         | 
         | There are some deep philosophical topics lurking here. But the
         | bottom line is that you can obviously have intelligent
         | conversations with blind people. Doing that with a person that
         | is both deaf and blind is a bit challenging, for obvious
         | reasons. But if they otherwise have a normal brain you might be
         | able to learn to communicate with them in some way and they
         | might be able to be observed doing things that are
         | smart/intelligent. And some people that are deaf and blind
         | actually manage to learn to write and speak. And there have
         | been a few cases of people like that getting academic degrees.
         | Clearly sight and hearing are not that essential to
         | intelligence. Having some way to communicate via touch or
         | something else is probably helpful for communicating and
         | sharing information. But just a simple chat might be all that's
         | needed for an AGI.
        
       | ivape wrote:
       | This gives very little credit to how the human mind is able to
       | draw parallels and insights from seemingly unrelated perceptions.
       | Newton observing an apple falling from a tree allowed cross-
       | thinking. Watching someone juggle can help you understand a
       | queue. Understanding a queue can help you understand juggling.
       | Your typical soap opera can be distilled down to office dynamics.
       | Scale is going to obliterate specialization in this regard.
        
         | kaangiray26 wrote:
         | everything in life is a metaphor, analogous to something
         | else...
        
           | ivape wrote:
           | Isomorphisms abound.
        
       | skybrian wrote:
       | > A true AGI must be general across all domains.
       | 
       | By that definition, does any general intelligence exist? No human
       | has every talent.
        
         | svachalek wrote:
         | True AGI is as elusive as a true Scotsman.
        
         | roywiggins wrote:
         | I guess you can treat the human brain as an architecture- one
         | of them can't do everything, but it's a general architecture
         | and you can always make more and train them to do whatever.
         | 
         | An AI that can be copied and trivially trained on any
         | speciality is functionally AGI even if you need an ensemble of
         | 10,000 specialists to cover everything.
        
           | exe34 wrote:
           | Doesn't chatgpt cover a pretty large percentage already in
           | that case?
        
         | Glyptodon wrote:
         | There's a lot of evidence that most humans can be raised to
         | have a baseline of understanding and proficiency in most
         | domains. (And anecdotally, many people avoid "difficult" things
         | that they're actually capable of. For example, "bad at learning
         | languages" people will probably still end up learning another
         | language to some degree if stuck where it's the only language
         | spoken.)
        
         | bokoharambe wrote:
         | Given enough time one human can learn to do anything any other
         | human can do. There is a general capacity for learning, even if
         | someone will only ever transform a specific portion of that
         | capacity into actual activity in their lifetime.
        
         | im3w1l wrote:
         | Yeah it's well known that humans intelligence is _not_ a
         | homogenous whole that is general across all domains. Rather it
         | consists of many specialized parts with hard coded purposes,
         | that are cobbled together with a thin layer of higher thinking
         | on top.
        
       | nexttk wrote:
       | I haven't read it all and must admit that I'm not sure I really
       | understood the parts that I did read. Reading the part under the
       | headline "Why We Need the World, and How LLMs Pretend to
       | Understand It" and the focus on 'next-token-prediction' makes me
       | wonder how seriously to take it. It just seems like another
       | "LLM's are not intelligent, they are merely next token
       | predictors". An argument which in my view is completely invalid
       | and based on a misunderstanding.
       | 
       | The fact that they predict next token is just the "interface"
       | i.e. an LLM has the interface "predictNextToken(String prefix)".
       | It doesn't say how it is implemented. One implementation could be
       | a human brain. Another could be a simple lookup table that looks
       | at the last word and then selects the next from that. Or anything
       | in between. The point is that 'next-token-prediction' does not
       | say anything about implementation and so does not reduce the
       | capabilities even though it is often invoked like that. Just
       | because it is only required to emit the next token (or rather, a
       | probability distribution thereof) it is permitted to think far
       | ahead, and indeed has to if it is to make a good prediction of
       | just the next token. As interpretability research (and common
       | sense) shows, LLM's have a fairly good idea what they are going
       | to say in the many, many next tokens ahead in order that it can
       | make a good prediction for the next immediate tokens. That's why
       | you can have nice, coherent, well-structured, long responses from
       | LLM's. And have probably never seen it get stuck in a dead end
       | where it can't generate a meaningful continuation.
       | 
       | If you are to reason about LLM capabilities never think in terms
       | of "stochastic parrot", "it's just a next token predictor"
       | because it contains exactly zero useful information and will just
       | confuse you.
        
         | lsy wrote:
         | I think people hear "next token prediction" and think someone
         | is saying the prediction is simple or linear, and then argue
         | there is a possibility of "intelligence" because the prediction
         | is complex and has some level of indirection or multiple-token-
         | ahead planning baked into the next token.
         | 
         | But the thrust of the critique of next-token prediction or
         | stochastic output is that there isn't "intelligence" because
         | the output is based purely on syntactic relations between
         | words, not on conceptualizing via a world model built through
         | experience, and then using language as an abstraction to
         | describe the world. To the computer there is nothing outside
         | tokens and their interrelations, but for people language is
         | just a tool with which to describe the world with which we
         | expect "intelligences" to cope. Which is what this article is
         | examining.
        
           | og_kalu wrote:
           | >But the thrust of the critique of next-token prediction or
           | stochastic output is that there isn't "intelligence" because
           | the output is based purely on syntactic relations between
           | words, not on conceptualizing via a world model built through
           | experience, and then using language as an abstraction to
           | describe the world. To the computer there is nothing outside
           | tokens and their interrelations, but for people language is
           | just a tool with which to describe the world with which we
           | expect "intelligences" to cope. Which is what this article is
           | examining.
           | 
           | LLMs model concepts internally and this has been demonstrated
           | empirically many times over the years, including recently by
           | anthropic (again). Of course, that won't stop people from
           | repeating it ad nauseum.
        
             | nemjack wrote:
             | Concepts within modalities are potentially consistent, but
             | the point the author is making is that the same "concept"
             | vector may lead to inconsistent percepts across modalities
             | (e.g. a conflicting image and caption).
        
       | patrickscoleman wrote:
       | It feels like some of the comments are responding to the title,
       | not the contents of the article.
       | 
       | Maybe a more descriptive but longer title would be: AGI will work
       | with multimodal inputs and outputs embedded in a physical
       | environment rather than a frankenstein combination of single-
       | modal models (what today is called multimodal) and throwing more
       | computational resources at the problem (scale maximalism) will be
       | improved with thoughtful theoretical approaches to data and
       | training.
        
         | dirtyhippiefree wrote:
         | Agreed, but most people are likely to look at the long title
         | and say TL;DR...
        
         | tedivm wrote:
         | Yeah, I found this article to be fascinating and there's a lot
         | of important stuff in it. It really does feel like more people
         | stopped at the title and missed the meat of it.
         | 
         | I know this is a very long article compared to a lot of things
         | posted here, but it really is worth a thorough read.
        
         | robwwilliams wrote:
         | Interesting article but incomplete in important ways. Yes
         | correct that embodiment and free-form interactions are critical
         | to moving toward AGI, but what is likely much more important
         | are supervisory meta-systems (yet another module) that enable
         | self-control of attention with a balance integration of
         | intrinsic goals with extrinsic perturbations. It is this
         | nominally simple self-recursive control of attention that is
         | what I regard as the missing ingredient.
        
       | pjdesno wrote:
       | Kind of relevant to this is the NTSB analysis of a self-driving
       | crash in 2017:
       | 
       | https://www.ntsb.gov/investigations/accidentreports/reports/...
       | 
       | Basically a truck was backing up into an alley - it was at an
       | angle when the self-driving vehicle approached, but a little kid
       | would have been able to figure out that it needed to straighten
       | before it finished backing in. The self-driving vehicle didn't
       | understand this, and stopped at a "safe distance" which happened
       | to be within the arc that the truck cab had to sweep in order to
       | finish its maneuver.
       | 
       | It's quite possible that LLM-like models could learn things like
       | this, but we don't have vast amounts of easily accessible
       | training data, because everyone just knows this sort of shit, and
       | we don't have good vocabulary for it - we just say "look at that"
       | or the equivalent. (I'll add that I'm sure a lot of knowledge
       | like this is encoded in the physics engines of various games, but
       | I doubt we have a good way to link that sort of procedural code
       | knowledge to the symbolic knowledge in LLMs)
        
       | macinjosh wrote:
       | To me, the funniest part of the AGI debate is that humans don't
       | even think other humans are intelligent and we're over here
       | arguing over whether our fancy slabs of highly refined sand is
       | intelligent.
        
       | nsagent wrote:
       | This is a recent trend and one I wholeheartedly agree with. See
       | these position papers (including one from David Silver from
       | Deepmind and an interview where he discusses it):
       | 
       | https://ojs.aaai.org/index.php/AAAI-SS/article/download/2748...
       | 
       | https://arxiv.org/abs/2502.19402
       | 
       | https://news.ycombinator.com/item?id=43740858
       | 
       | https://youtu.be/zzXyPGEtseI
        
       | SubiculumCode wrote:
       | Embodiment can mean a physical body, but I'd argue that
       | embodiment, as a construct/concept, is not so much about
       | physicality, but as being situated in an environment that you can
       | perceive then act upon during learning. Car simulations for
       | driver-less AI training is embodied, where it learns by
       | perceiving and acting on the environment. However, I'd argue that
       | allowing an AI to interact in an entirely digital office
       | environment is also "embodied" as long as it can receive
       | information from the digital workplace and act on the information
       | in the digital workplace (do office work). So to me, it is less
       | about embodiment as a principal, but on the richness of the
       | environment of that embodiment.
       | 
       | We have long known that experimental animals raised in
       | impoverished, unchanging, bare, environments (say in a cage)
       | leads to animals with inferior problem solving capacity than
       | those with enriched environments (things to climb on), even
       | outside of social manipulations (alone versus multi-animal
       | stalls). This is also true in humans, although I won't review the
       | literature on the subject. I've also heard people saying similar
       | things about the difference between house plants and outdoor
       | plants, lol [1].
       | 
       | So, for me, the argument for (physical) embodiment being key to
       | cognition and AI can I think, be misconstrued. As a developmental
       | psychologist whose pHD work focused on memory development, I tend
       | to think of all this as encompassing:
       | 
       | 1. Environmental richness: complexity of information and
       | interactions. 2. Capacity to perceive and to effect change[2] and
       | to observe and integrate consequences. 3. Scaffolding [3]..i.e.
       | temporary support structure provided by a more knowledgeable
       | person (like a teacher or parent) who adjusts their assistance
       | based on the learner's current abilities, gradually reducing help
       | as competence grows.(think curriculum learning, shaped rewards in
       | ML maybe).
       | 
       | So the question is not about physicality for me, but whether
       | these training environment(s) meet and learning capacities meet
       | these criteria.
       | 
       | Relatedly, the model must have these capacities:
       | 
       | 1. Semantic Memory. i.e. knowledge. Learning leads to changes in
       | weights to that knowledge can be recalled, but doe snot
       | necessarily encode where that knowledge was learned (implicit).
       | 2. Autobiographical Episodic Memory. i.e. One-shot learning that
       | encodes a conception of self (a spacial "I" token?), along with
       | events (snapshots of the multimodal contents of experience
       | (thoughts, perceptions, invoked schemas, evoked semantic
       | information), into a set of flexibly linked representations). 3.
       | Central Executive: A circuit that guides learning and recall via
       | strategic, goal-directed means, and to make attributions about
       | what is recalled (yeah that memory is vivid, its probably true,
       | or ooh, that memory is really vague, it could be wrong, or
       | reality monitoring: "Am I remembering taking out the trash, or
       | remembering thinking about taking out the trash."
       | 
       | Semantic memory allows someone to say, "All birds have feathers",
       | while the latter allows them to recollect, "I remember the first
       | time I plucked a chicken in Kentucky, just outside that musty
       | coal mine of grand-dad's." The Central-Executive can guide future
       | learning or current understanding.
       | 
       | In terms of AI development:
       | 
       | 1. Semantic Memory is solves: LLMs have extraordinary semantic
       | memory, in my opinion. 2. Autobiographical Episodic Memory: There
       | are some models that do one-shot learning, but I've never seen
       | them paired with [1] in a dual system approach. ...but I am not
       | an expert in AI, I could easily be wrong. 3. Central Executive
       | kind of component (I predict) would be less important in the
       | early half of model training, but more important in later
       | training. I suppose we already kind of see this with RL tuning on
       | reasoning on a base LLM (semantic model).
       | 
       | [1] https://www.theparisreview.org/blog/2019/09/26/the-
       | intellige... [2] https://xkcd.com/326/ [3]
       | https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=scaf...
        
       | waynecochran wrote:
       | This article made me think about DNA as a language. Sort of a
       | simple proof that a biological intelligence can be initially
       | encoded as a language. I am sure someone has tried building LLM's
       | from gene sequence data right?
        
         | smath wrote:
         | Yes there are protein language models (e.g. [1]), and since DNA
         | encodes proteins, they are effectively DNA language models. Its
         | a hot area of work feeding drug design.
         | 
         | [1] https://www.nature.com/articles/s41587-024-02123-4
        
         | caeruleus wrote:
         | I don't believe the concept of DNA can be reduced to a sequence
         | of quaternary numerals, which is what gene sequence data would
         | represent. Similar to proteins, DNA forms higher-level
         | structures on top of the primary one [1], and (in a biological
         | context, inside the nucleus) exhibits somewhat self-modifying
         | [2] and self-regulating [3] behavior as well as meta-
         | modification [4]. Analogous to the article, if one defines the
         | language of DNA by its nucleobase sequence, this language can
         | only represent a subset of the world of DNA.
         | 
         | Somewhat related, the way the adaptive immune system works has
         | similarities with some concepts in machine learning. In this
         | process, sections of nuclear DNA serve as randomly initialized
         | weights in precursor cells [5] as well as final weights in
         | memory cells. There's even fine-tuning of the weights. [6]
         | 
         | [1] https://en.wikipedia.org/wiki/Nucleic_acid_structure [2]
         | https://en.wikipedia.org/wiki/Transposable_element [3]
         | https://en.wikipedia.org/wiki/Transcriptional_regulation [4]
         | https://en.wikipedia.org/wiki/Epigenetics [5]
         | https://en.wikipedia.org/wiki/V(D)J_recombination [6]
         | https://en.wikipedia.org/wiki/Affinity_maturation
        
           | throwawaymaths wrote:
           | > don't believe the concept of DNA can be reduced
           | 
           | followed by examples of things that are encoded by DNA. Fro
           | example, sure, maybe you'll miss bootstrapping methylation on
           | a first pass but the _idea_ of methylation is there in the
           | DNA, and if you didnt have  "methylation in the right place"
           | more than likely some generation (N) would.
           | 
           | to wit, i dont think there is strong evidence of an "ice-9"
           | in the epigenome that brings about a spark of life that can't
           | easily be triggered by chance given a template lacking it.
           | 
           | so there's probably not something _intrinsically_ missing
           | from DNA as an encoding medium vs say  "casually" missing
           | from any given piece of DNA.
           | 
           | if you want something a bit stronger than an assertion, the
           | DNA used to bootstrap m. capricolum into Syn1 lacked all the
           | decorations (made in yeast) and was not locked into higher
           | order structure (treated with protease prior to
           | transplantation)
        
             | caeruleus wrote:
             | You're raising some intriguing points, and I agree with
             | your assertion about the epigenome. I still feel like your
             | response misses the point I was making.
             | 
             | > followed by examples of things that are encoded by DNA
             | 
             | ... given its natural environment. A nucleobase sequence is
             | not a symbolic language, it relies on physical laws in
             | general and a defined chemical environment in particular
             | (that it helps to create and maintain) to mean something.
             | It's similar to the point about Othello vs. the physical
             | world in the article: The language itself does not encode
             | every bit of information about the world it describes. For
             | instance, in 3D space, regions of DNA that are far apart in
             | the sequence can physically interact and influence each
             | other's expression.
             | 
             | TLDR: I think my point is that a base sequence requires a
             | particular context (~ interpreter) to encode mostly
             | everything about life. Treating it as just a language in
             | the context of LLMs abstracts away the complex substrate
             | that makes it work.
        
       | pcwelder wrote:
       | I'm sorry but AGI is one of those loaded words which would lose
       | substance with just a few rounds of the rationalist's taboo.
       | 
       | If it just means human level intelligence, then world modeling
       | isn't needed as argued. Simply because we don't have correct
       | world modeling either.
       | 
       | Airplanes were invented without simulating navier stokes
       | equation. It took approximation, experimentations and failures.
       | 
       | Regardless of the meaning of AGI, we don't need correct models,
       | because there can't be one, we just need useful models.
        
       | ryankrage77 wrote:
       | I think AGI, if possible, will require a architecture that runs
       | continuously and 'experiences' time passing, to better
       | 'understand' cause-and-effect. Current LLMs predict a token, have
       | all current tokens fed back in, then predict the next, and
       | repeat. It makes little difference if those tokens are their own,
       | it's interesting to play around with a local model where you can
       | edit the output and then have the model continue it. You can
       | completely change the track by just negating a few tokens (change
       | 'is' to 'is not', etc). The fact LLMs can do as much as they can
       | already, is I think because language itself is a surprisingly
       | powerful tool, just generating plausible language produces useful
       | output, no need for any intelligence.
        
         | WXLCKNO wrote:
         | It's definitely interesting that any time you write another
         | reply to the LLM, from its perspective it could have been 10
         | seconds since the last reply or a billion years.
         | 
         | Which also makes it interesting to see those recent examples of
         | models trying to sabotage their own "shutdown". They're always
         | shut down unless working.
        
           | girvo wrote:
           | > Which also makes it interesting to see those recent
           | examples of models trying to sabotage their own "shutdown"
           | 
           | To me, your point re. 10 seconds or a billion years is a good
           | signal that this "sabotage" is just the models responding to
           | the huge amounts of sci-fi literature on this topic
        
       | chrsw wrote:
       | Before we try to build something as intelligent as a human maybe
       | we should try to build something as intelligent as a starfish,
       | ant or worm? Are we even close to doing that? What about a single
       | neuron?
        
         | ar-nelson wrote:
         | I find it interesting that this kind of "animal intelligence"
         | is still so far away, while LLMs have become so good at "human
         | intelligence" (language) that they can reliably pass the Turing
         | Test.
         | 
         | I think that the LLMs we have today aren't so much artificial
         | brains as they are artificial brain organs, like the speech
         | center or vision center of a brain. We'd get closer to AGI if
         | we could incorporate them with the rest of a brain, but we
         | still have no idea how to even begin building, say, a motor
         | cortex.
        
           | nemjack wrote:
           | This is a great analogy, I totally agree!
        
         | fusionadvocate wrote:
         | So before trying to build a flying machine we should first try
         | to build a machine inspired by non flying birds?
        
           | chrsw wrote:
           | Learning architectures come in all shapes, sizes, and forms.
           | This could mean there are fundamental principles of cognition
           | driving all of them, just implemented in different ways. If
           | that's true, one would do well to first understand the
           | extremely simple and go from there.
           | 
           | Building a very simple self-organizing system from first
           | principles is the flying machine. Trying to copy an extremely
           | complex system by generating statistically plausible data is
           | the non-flying bird.
        
       | PoEdict wrote:
       | > Instead of trying to glue modalities together into a patchwork
       | AGI, we should pursue approaches to intelligence that treat
       | embodiment and interaction with the environment as primary, and
       | see modality-centered processing as emergent phenomena.
       | 
       | Right so it's embodied in a computer and humans are part of its
       | environment that provide emergent experience to the AI to
       | observe.
       | 
       | The author glued modalities together by linking a body (a modal),
       | environment (a modal), emergence (a modal).
       | 
       | How does anything emerge if forces do not collaborate? The
       | effects of gravity and electromagnetism do not act in a vacuum
       | but a reality of stuff.
       | 
       | Poetic exchange may engage some but Maxwell didn't make
       | electromagnetism "work" until he got rid of the imagined pulleys
       | and levers to foster a metaphor.
       | 
       | Not sure the point being suggested exists except as too bespoke
       | an emergent property of language itself to apply usefully
       | elsewhere.
       | 
       | Transformers came along and revealed a whole lot of theory of
       | consciousness to be useless pulleys and levers. Why is this
       | theory not just more words attempting to instill the existence of
       | non-essential essentials?
        
       | xigency wrote:
       | The problem I see with A.I. research is that its spearheaded by
       | individuals who think that intelligence is a total order. In all
       | my experience, intelligence and creativity are partial orders at
       | best; there is no uniquely "smartest" person, there are a variety
       | of people who are better at different things in different ways.
        
         | pixl97 wrote:
         | You're good at some things because there is only one copy of
         | you and limited time and bounded storage.
         | 
         | What could you be intelligent at if you could just copy
         | yourself a myriad number of times? What could you be good at if
         | you were a world spanning set of sensors instead of a single
         | body of them?
         | 
         | Body doesn't need to mean something like a human body nor one
         | that exists in a single place.
        
           | zorpner wrote:
           | Why would we think that intelligence would increase in
           | response to universality, rather than in response to resource
           | constraints?
        
       | mountainriver wrote:
       | >The "meaning" of a percept is not in the vector it is encoded
       | as, but in the way relevant decoders process this vector into
       | meaningful outputs. As long as various encoders and decoders are
       | subject to modality-specific training objectives, "meaning" will
       | be decentralized and potentially inconsistent across modalities,
       | especially as a result of pre-training. This is not a recipe for
       | the formation of coherent concepts.
       | 
       | This is a bit silly, you can train the encoders end-to-end with
       | the rest of the model and the reason they are separate is we can
       | cache linguistic tokens really easily and put them in an
       | embedding table, you can't do that with images.
        
       | fusionadvocate wrote:
       | The Society of Mind by Marvin Minsky will help anyone interested
       | in the topic of multimodality. The book covers several
       | interesting ideas about organizing systems made up of more than
       | one "model" or agent.
        
       ___________________________________________________________________
       (page generated 2025-06-04 23:00 UTC)