[HN Gopher] John Carmack on the Similarity of Human Learning and...
       ___________________________________________________________________
        
       John Carmack on the Similarity of Human Learning and LLMs Training
        
       Author : amrrs
       Score  : 39 points
       Date   : 2023-04-07 21:29 UTC (1 hours ago)
        
 (HTM) web link (twitter.com)
 (TXT) w3m dump (twitter.com)
        
       | ftxbro wrote:
       | The context is that John Carmack got $20M to try to make AGI by
       | 2030, based on non-LLM methods. In this tweet he is boosting the
       | idea that there is a gigantic learning gap between the
       | statistical efficiency of human and LLM learning, even though
       | there may not be such a huge gap in standardized test results of
       | LLMs and humans. This means that he might yet find a cool idea
       | that could justify the time and money he is spending to find
       | secret knowledge of the ancients from like pre-1990s AI
       | publications that didn't have GPU access. Maybe it will happen!
       | 
       | https://the-decoder.com/john-carmacks-general-artificial-int...
        
         | rajman187 wrote:
         | Carmack is no doubt brilliant and it shows through Oculus as a
         | product. But AGI is not about algorithm optimization the way 3D
         | graphics were in the 90s. I am not sure the approach he takes
         | (let's find a way to write this in assembly) can apply here.
         | He's said before that AGI is going to be a few thousands of
         | lines of code, ultimately [1], and it's hard to believe that's
         | really the case.
         | 
         | That said, I did get a chance to speak with him one on one last
         | year and he really emphasized a few things; the need for being
         | product oriented and giving customers what they want rather
         | than chasing cool engineering (uber)solutions; not being
         | careless with resources just because we have more power with
         | modern hardware (he poked fun at React where you spin up a new
         | thread just for an interactive button); and being aware of the
         | inefficiencies brought on by infinite resources (# of engineers
         | and/or funding, which make you think less critically about
         | timelines and delivering within bounded means)
         | 
         | [1] https://youtu.be/I845O57ZSy4
        
           | carabiner wrote:
           | The guy also tried to build a rocket to reach space via
           | "first principles." It failed, utterly.
        
           | phkahler wrote:
           | >> But AGI is not about algorithm optimization the way 3D
           | graphics were in the 90s
           | 
           | Wow, you think his ability is straight performance
           | optimisation? A big part of his early fame was from doing
           | research to find good algorithms to achieve his goals. He
           | also had that stint in real time control systems... flying
           | hovering rockets before spaceX even existed.
           | 
           | He's a problem solver with a strong ability to sift through
           | possible solutions for what actually works, and quite capable
           | of devising his own solutions when is research comes up
           | empty.
           | 
           | But yeah, he can write assembler too.
        
         | coffeebeqn wrote:
         | It's definitely Good to not put all our eggs in the LLM basket.
         | Even openAI says it's only a part of the solution - but it has
         | scaled more than most anyone expected. Will that continue
         | forever and we keep getting more impressive emergent behaviors
         | from it? I doubt that but I could certainly be wrong.
         | 
         | It still seems to be missing any sense of what is True, not
         | sure if that's possible to embed in that model or if we'll just
         | have some human feedback hacks and eventually get a better
         | model
        
           | gitfan86 wrote:
           | That debate is becoming irrelevant. Sure it would be great if
           | GTP-6 cost 50k to train instead of 100M. But even if costs
           | 100M it will produce many billions of dollars of value, so
           | the original investment is still relatively small.
        
             | scrollaway wrote:
             | Yeah but the difference between those two numbers is one is
             | accessible to individuals (albeit wealthy ones), the other
             | is accessible only to large-ish companies.
        
           | ftxbro wrote:
           | > Will that continue forever and we keep getting more
           | impressive emergent behaviors from it? I doubt that
           | 
           | I feel like putting the word 'forever' ruins the point. It's
           | the most extreme strawman.
           | 
           | It's amazing how the scaling has unlocked the emergent
           | behaviors! When I look at the scaling graphs, I see that the
           | ability to reduce 'perplexity' is continuing with scale and
           | capital investment with no sign of slowing yet (it will slow
           | eventually). I also see that reducing perplexity is
           | continually unlocking new emergent behaviors. So I would
           | guess that scaling will probably unlock so many more new
           | emergent behaviors before it eventually plateaus!
        
           | llamaLord wrote:
           | IMO, the capabilities have not improved that much between
           | GPT-3 vs GPT-4, the primary difference is just in the size of
           | the context the model can handle at any one time (Max
           | tokens).
           | 
           | While I definitely get "longer" and slightly more "in-depth"
           | answers from GPT-4 vs 3, it already feels like that
           | capability growth curve is starting to plateau.
        
             | vidarh wrote:
             | The distinction may appear subtle in many cases, but in
             | others GPT4 blows 3 away. E.g. recognizing when it doesn't
             | know and admitting it's hypothesising is something I've run
             | into with 4 but not even 3.5.
             | 
             | The capability curve _will_ necessarily appear to plateau
             | when the starting point is as good as it is now. The
             | improvements we recognize will be subtler. Halving the
             | remaining error rate will look less impressive for each
             | step.
        
             | ftxbro wrote:
             | > IMO, the capabilities have not improved that much between
             | GPT-3 vs GPT-4
             | 
             | I strongly disagree. Anyone who wants to look for themself
             | can see the GPT 4 technical report.
             | 
             | https://arxiv.org/pdf/2303.08774.pdf
        
               | CamperBob2 wrote:
               | Page 9 pretty much stopped me dead in my tracks. There's
               | no shortage of humans, even technically-savvy ones, who
               | wouldn't answer that question correctly.
        
         | ly3xqhl8g9 wrote:
         | In a way LLMs are the worst thing that could have happened in
         | the search for artificial intelligence, instead of the harsh
         | winter from the 1990s we are entering now a climate changed
         | winter: warm, products to market race winter.
         | 
         | Funnily enough, the Geoffrey Hinton of today is probably some
         | symbols researcher, shouting in the desert, like Hinton shouted
         | in the 1980s, that we need more than matmul.
         | 
         | One interesting, biology-inspired mechanism, would be _quorum
         | sensing_ [1]: a basal cognition-like, decision-making function
         | in which decentralized systems (bacteria, cells) start building
         | functionality (sensing /decision) from the bottom up. There is
         | no hint for this sort of mechanism in our current artificial
         | 'neural networks'. Not that it should, but our cells and in
         | general cells use these kind of 'tricks' to solve problems in
         | all kinds of spaces (transcriptomics, morphogenetics, etc.)
         | without requiring ridiculous amounts of energy, time, or other
         | resources.
         | 
         | [1] https://en.wikipedia.org/wiki/Quorum_sensing
        
           | joe_the_user wrote:
           | LLMs may be terrible for the search for AGI but that may be
           | good thing. If, mind you if, LLMs are always going to semi-
           | smart "parrots", then society will get a picture of what an
           | AI might without the AI being able to skynet-style takeover
           | people are maybe justifiably worried.
           | 
           | Now, as far as whether LLM or offshoots can get to AGI, I
           | know plenty of good arguments for them not being able to do
           | that. I don't think the claim that they're just accumulating
           | more abilities in each iteration in an inexplicable way is
           | true. But I've lived long enough to know that you should
           | never get too cocky when one is "arguing with success". So
           | maybe.
           | 
           | Moreover, given that LLM programming is basically just bucket
           | chemistry, if an LLM can the ability to competently pursue
           | long term goals, it seems like it will have a good chance of
           | some of its goals being random cruft that will make it quite
           | dangerous.
        
         | cpeterso wrote:
         | Is there any information available about the non-LLM methods
         | Carmack's AGI startup is researching?
        
       | breatheoften wrote:
       | Who is commenting on this and what will they say??!
       | 
       | 1-2 trillion tokens of training for gptx class models
       | 
       | It would take 20 years for a human to read that many tokens at 8
       | hrs of reading per day (although i'm pretty unclear on whether
       | that's unique token sequences or randomly selected token pairs --
       | feels like a big difference between the two!!)
       | 
       | We are all here now -- who thinks we are anywhere of note
       | together?
       | 
       | Personally I think that models that do not experience incremental
       | change remain an entirely separate class of intelligence from
       | those which evolve continuously -- but I am open to being
       | persuaded otherwise (certainly I never thought making gpt2 bigger
       | would be as impressive as gpt4) ...
        
         | meindnoch wrote:
         | >1-2 trillion tokens of training for gptx class models
         | 
         | >It would take 20 years for a human to read that many tokens at
         | 8 hrs of reading per day
         | 
         | Normal reading is ~200 words per minute.
         | 
         | That's 12,000 words per hour.
         | 
         | That's 288,000 words per day (reading for 24 hours straight).
         | 
         | That's 105,120,000 words per year.
         | 
         | That's 2,102,400,000 words per 20 years.
         | 
         | A trillion is 1,000,000,000,000.
        
           | tonightstoast wrote:
           | Plus almost no one is reading academic papers at 200 words
           | per minute. So for the level of material it could mean that
           | number is halved per 20 years.
        
         | lostmsu wrote:
         | Video from a pair of eyes is many terabytes per year though. If
         | you compress it with best codecs. Raw video would be closer to
         | 1PB per year.
         | 
         | I think people overestimate the data efficiency of the
         | biological neural networks w.r.t. transformer models.
        
           | actionfromafar wrote:
           | The _power effiency_ is pretty brutal though, in favour of
           | bio-gelpacks.
        
             | lostmsu wrote:
             | How can you be sure about power efficiency, if you don't
             | know how data efficiency compares? It is low power. But so
             | is a phone running llama.cpp
        
           | Thiez wrote:
           | People with bad eyesight are not (to the best of my
           | knowledge) less intelligent than those with perfect vision.
           | We could probably supply video at a much lower resolution
           | without significantly affecting the outcome.
        
         | GordonS wrote:
         | And, as humans, we don't only get "tokens" from what we read,
         | but from our entire sensorium.
        
       | intalentive wrote:
       | Humans are not exposed to "billions of words". They are exposed
       | to continuous signals, and somehow extract meaningful
       | representations, like "words" and "objects", out of them.
       | 
       | LLMs are exposed to human-curated data. Let's see an LLM curate
       | its own data out of nothing but experience of raw continuous
       | signals.
       | 
       | Sure, GPT trained on the internet is a nice way to condense and
       | retrieve human-curated data. But it's not going to give us a Lt.
       | Data or C-3P0 that can learn and adapt in real time.
        
         | pbw wrote:
         | > Humans are not exposed to "billions of words" ... continuous
         | signals
         | 
         | Yes humans need to learn to extract words from signals, but
         | that does not mean it's not correct to count the number of
         | words they've been exposed to in their lifetime. Your comment
         | is like saying we can't count the number of hamburgers someone
         | has eaten, because they actually eat myofibrillar proteins, not
         | burgers.
        
           | d0mine wrote:
           | Though millions seems more realistic (thousands of words per
           | day)
        
         | andsoitis wrote:
         | what do you think of midjourney's ability to describe a user-
         | provided image in text? https://the-decoder.com/midjourney-new-
         | image-tool-works-in-r...
        
         | gunshai wrote:
         | But humans are also trained on human-curated data so much so we
         | put fairly hefty price tags on that curation.
         | 
         | https://miro.medium.com/v2/resize:fit:1056/0*E1eNateTiDThGcY...
        
         | echelon wrote:
         | We've only just started. Now more capital and more minds will
         | be pouring into solving this.
         | 
         | Buckle up.
        
         | Sparkyte wrote:
         | Also LLM are not AI just parts to AI you need a bunch of
         | interfacing and injection components. The sum of which makes up
         | AI.
         | 
         | I see the progress of AI stagnating while people board the next
         | equivalent of Crypto craze because someone want to financially
         | profit off something not fully seeing realization.
         | 
         | However it is not without merit which Crypto was completely
         | without merit and something to show for itself.
         | 
         | As flawed and imperfect ChatGPT is it clearly mirrors its
         | creators and that is a compliment.
        
         | Femtodjy wrote:
         | Why not?
         | 
         | Humans learn based on human curated data too.
         | 
         | School is not natural. Our whole env is neither. Babies can't
         | survive.
         | 
         | I would even go so far to say that the potential model a LLM
         | would create internally might not be that far away of that of a
         | human.
         | 
         | And segment anything was just announced. The performance of
         | zero shot systems is tremendous.
         | 
         | It's not far fetched to assume that chatgpt combined with
         | segment anything together would allow it to create an even more
         | accurate model of the world.
        
           | bluefirebrand wrote:
           | > Humans learn based on human curated data too.
           | 
           | > School is not natural. Our whole env is neither. Babies
           | can't survive
           | 
           | A human goes to school to learn.
           | 
           | Humans in general learned based on experience of the world
           | around them. We invented language, no one taught it to us. We
           | learned to make fire, forge tools, cook food, practice
           | medicine, etc on our own.
           | 
           | School is just how we pass down that learning.
        
             | dwaltrip wrote:
             | Humanity invented those things.
             | 
             | Individual humans were taught those things, either directly
             | or indirectly by observing / listening to others.
             | 
             | We do learn through our experience, of course. But most of
             | what we learn is from others in one form or another.
        
             | pixl97 wrote:
             | If I took you and pitched you on an island as an
             | exceptionally young child with no further training, most
             | likely you would die. Even more likely you wouldn't have
             | fire and tools. A high proportion of our behaviors can be
             | traced as a continuous learning chain going back eons.
        
         | macrolocal wrote:
         | * * *
        
         | com2kid wrote:
         | > Humans are not exposed to "billions of words". They are
         | exposed to continuous signals, and somehow extract meaningful
         | representations, like "words" and "objects", out of them.
         | 
         | Plenty of studies showing long term academic achievement
         | differences based on number of unique words babies are exposed
         | to.
         | 
         | And I promise you, having a baby/toddler is _really_ damn close
         | to doing data labeling. Reading picture books, you are
         | basically labeling objects. Walking around the grocery store,
         | you are labeling objects. All the time, again and again, and
         | the same object will get labeled repeatedly, and sometimes it
         | will even be fact checked. A toddler will point to something
         | that they know is an orange, ask  "apple?" and you had sure as
         | hell better reply "orange".
        
       | cma wrote:
       | LeCun's reply:
       | 
       | > more like a thousand times more. > Between 1 and 2 trillion
       | tokens. > It would take a person 22,000 years to read through 1
       | trillion words at normal speed for 8 hours a day.
       | 
       | It's kind of surprising to me that around 1000 humans could
       | feasibly read the whole internet (as scraped for LLMs) between
       | each other in only 22 years. The internet feels so much more
       | massive than that. Though of course lots of the stuff on arxiv
       | etc. is so dense there is no way you go through it at normal
       | reading speed.
        
       | brokencode wrote:
       | I wonder if it only seems like our brains need relatively little
       | training because we have millions of years of training encoded in
       | our DNA.
       | 
       | No other animal can learn human language or logic to the extent
       | that humans can, no matter how much you train them. But humans
       | learn language easily in the first few years of life.
       | 
       | It's almost as if the human brain is preprogrammed with the
       | general concepts of all languages, and it just needs to be fine
       | tuned with specific vocabulary and grammar rules.
        
       | lairv wrote:
       | LLMs weights are initialized from a normal distribution (or
       | whatever distribution is now SOTA) while some animals can walk
       | the second they are born
        
       | cypress66 wrote:
       | GPT4 probably has a breadth of knowledge 1000x wider than a
       | human, so the fact that it uses more tokens seems reasonable.
        
       | roenxi wrote:
       | It seems likely that there is a big missing component of visual
       | data. How many Tb of visual data (effectively video) do babies
       | get exposed to?
       | 
       | That is what sets up their mental model of the world, and
       | language is fit over the top of that. It is hard to assess neural
       | net efficiency with that difference in place.
        
         | rvnx wrote:
         | It is going to amazing once LLMs get access to learn from real-
         | life visual information (e.g. daily life videos without
         | montage, or movies), and not just text.
        
       | musesum wrote:
       | funny, due to an aversion to twitter, I wanted to see if Carmac
       | cross-posted on Mastodon, which yielded a link to another HN
       | thread:
       | https://cyberfeed.io/article/0d53a16b9c8bcc0a1d2dc5f77dac225...
        
       | physPop wrote:
       | Carmack is a smart guy but don't confuse his expertise in other
       | areas with machine learning. Smells of hubris...
        
         | SquareWheel wrote:
         | He's been heavily researching machine learning for the last few
         | years, and formed an AGI startup last year.
         | 
         | He's not the world's leading researcher or anything, but he's
         | far from green in this space.
        
         | rvnx wrote:
         | Perhaps he feels that the Metaverse thing he sold to Zuck isn't
         | going anywhere, and now tries to ride a new wave.
        
       | m3kw9 wrote:
       | But we relate the words in a 3d world, llms relate words only to
       | each other
        
       | kklisura wrote:
       | Q: If we need to feed additional data to the network, do we train
       | it with existing data + new data or we can just train it on a new
       | data? I'm asking because if we need to train it with both old and
       | new data every time we have something new, that's not similar to
       | human learning.
        
         | IanCal wrote:
         | It's extremely similar to training new humans.
        
       | wkdneidbwf wrote:
       | surprises me a human would be exposed to a billion words. i never
       | thought about it, but a billion is such a large number
       | 
       | edit: oh, does he mean a billion total rather than a billion
       | unique words? haha i r smat
        
       ___________________________________________________________________
       (page generated 2023-04-07 23:02 UTC)