[HN Gopher] Why AI systems don't learn - On autonomous learning ...
       ___________________________________________________________________
        
       Why AI systems don't learn - On autonomous learning from cognitive
       science
        
       Author : aanet
       Score  : 3 points
       Date   : 2026-03-17 21:42 UTC (1 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | aanet wrote:
       | by Emmanuel Dupoux, Yann LeCun, Jitendra Malik
       | 
       | "he proposed framework integrates learning from observation
       | (System A) and learning from active behavior (System B) while
       | flexibly switching between these learning modes as a function of
       | internally generated meta-control signals (System M). We discuss
       | how this could be built by taking inspiration on how organisms
       | adapt to real-world, dynamic environments across evolutionary and
       | developmental timescales. "
        
         | dasil003 wrote:
         | If this was done well in a way that was productive for
         | corporate work, I suspect the AI would engage in Machievelian
         | maneuvering and deception that would make typical sociopathic
         | CEOs look like Mister Rogers in comparison. And I'm not sure
         | our legal and social structures have the capacity to absorb
         | that without very very bad things happening.
        
       | beernet wrote:
       | The paper's critique of the 'data wall' and language-centrism is
       | spot on. We've been treating AI training like an assembly line
       | where the machine is passive, and then we wonder why it fails in
       | non-stationary environments. It's the ultimate 'padded room'
       | architecture: the model is isolated from reality and relies on
       | human-curated data to even function.
       | 
       | The proposed System M (Meta-control) is a nice theoretical fix,
       | but the implementation is where the wheels usually come off.
       | Integrating observation (A) and action (B) sounds great until the
       | agent starts hallucinating its own feedback loops. Unless we can
       | move away from this 'outsourced learning' where humans have to
       | fix every domain mismatch, we're just building increasingly
       | expensive parrots. I'm skeptical if 'bilevel optimization' is
       | enough to bridge that gap or if we're just adding another layer
       | of complexity to a fundamentally limited transformer
       | architecture.
        
       ___________________________________________________________________
       (page generated 2026-03-17 23:01 UTC)