[HN Gopher] John Carmack on the Similarity of Human Learning and...
___________________________________________________________________
John Carmack on the Similarity of Human Learning and LLMs Training
Author : amrrs
Score : 39 points
Date : 2023-04-07 21:29 UTC (1 hours ago)
(HTM) web link (twitter.com)
(TXT) w3m dump (twitter.com)
| ftxbro wrote:
| The context is that John Carmack got $20M to try to make AGI by
| 2030, based on non-LLM methods. In this tweet he is boosting the
| idea that there is a gigantic learning gap between the
| statistical efficiency of human and LLM learning, even though
| there may not be such a huge gap in standardized test results of
| LLMs and humans. This means that he might yet find a cool idea
| that could justify the time and money he is spending to find
| secret knowledge of the ancients from like pre-1990s AI
| publications that didn't have GPU access. Maybe it will happen!
|
| https://the-decoder.com/john-carmacks-general-artificial-int...
| rajman187 wrote:
| Carmack is no doubt brilliant and it shows through Oculus as a
| product. But AGI is not about algorithm optimization the way 3D
| graphics were in the 90s. I am not sure the approach he takes
| (let's find a way to write this in assembly) can apply here.
| He's said before that AGI is going to be a few thousands of
| lines of code, ultimately [1], and it's hard to believe that's
| really the case.
|
| That said, I did get a chance to speak with him one on one last
| year and he really emphasized a few things; the need for being
| product oriented and giving customers what they want rather
| than chasing cool engineering (uber)solutions; not being
| careless with resources just because we have more power with
| modern hardware (he poked fun at React where you spin up a new
| thread just for an interactive button); and being aware of the
| inefficiencies brought on by infinite resources (# of engineers
| and/or funding, which make you think less critically about
| timelines and delivering within bounded means)
|
| [1] https://youtu.be/I845O57ZSy4
| carabiner wrote:
| The guy also tried to build a rocket to reach space via
| "first principles." It failed, utterly.
| phkahler wrote:
| >> But AGI is not about algorithm optimization the way 3D
| graphics were in the 90s
|
| Wow, you think his ability is straight performance
| optimisation? A big part of his early fame was from doing
| research to find good algorithms to achieve his goals. He
| also had that stint in real time control systems... flying
| hovering rockets before spaceX even existed.
|
| He's a problem solver with a strong ability to sift through
| possible solutions for what actually works, and quite capable
| of devising his own solutions when is research comes up
| empty.
|
| But yeah, he can write assembler too.
| coffeebeqn wrote:
| It's definitely Good to not put all our eggs in the LLM basket.
| Even openAI says it's only a part of the solution - but it has
| scaled more than most anyone expected. Will that continue
| forever and we keep getting more impressive emergent behaviors
| from it? I doubt that but I could certainly be wrong.
|
| It still seems to be missing any sense of what is True, not
| sure if that's possible to embed in that model or if we'll just
| have some human feedback hacks and eventually get a better
| model
| gitfan86 wrote:
| That debate is becoming irrelevant. Sure it would be great if
| GTP-6 cost 50k to train instead of 100M. But even if costs
| 100M it will produce many billions of dollars of value, so
| the original investment is still relatively small.
| scrollaway wrote:
| Yeah but the difference between those two numbers is one is
| accessible to individuals (albeit wealthy ones), the other
| is accessible only to large-ish companies.
| ftxbro wrote:
| > Will that continue forever and we keep getting more
| impressive emergent behaviors from it? I doubt that
|
| I feel like putting the word 'forever' ruins the point. It's
| the most extreme strawman.
|
| It's amazing how the scaling has unlocked the emergent
| behaviors! When I look at the scaling graphs, I see that the
| ability to reduce 'perplexity' is continuing with scale and
| capital investment with no sign of slowing yet (it will slow
| eventually). I also see that reducing perplexity is
| continually unlocking new emergent behaviors. So I would
| guess that scaling will probably unlock so many more new
| emergent behaviors before it eventually plateaus!
| llamaLord wrote:
| IMO, the capabilities have not improved that much between
| GPT-3 vs GPT-4, the primary difference is just in the size of
| the context the model can handle at any one time (Max
| tokens).
|
| While I definitely get "longer" and slightly more "in-depth"
| answers from GPT-4 vs 3, it already feels like that
| capability growth curve is starting to plateau.
| vidarh wrote:
| The distinction may appear subtle in many cases, but in
| others GPT4 blows 3 away. E.g. recognizing when it doesn't
| know and admitting it's hypothesising is something I've run
| into with 4 but not even 3.5.
|
| The capability curve _will_ necessarily appear to plateau
| when the starting point is as good as it is now. The
| improvements we recognize will be subtler. Halving the
| remaining error rate will look less impressive for each
| step.
| ftxbro wrote:
| > IMO, the capabilities have not improved that much between
| GPT-3 vs GPT-4
|
| I strongly disagree. Anyone who wants to look for themself
| can see the GPT 4 technical report.
|
| https://arxiv.org/pdf/2303.08774.pdf
| CamperBob2 wrote:
| Page 9 pretty much stopped me dead in my tracks. There's
| no shortage of humans, even technically-savvy ones, who
| wouldn't answer that question correctly.
| ly3xqhl8g9 wrote:
| In a way LLMs are the worst thing that could have happened in
| the search for artificial intelligence, instead of the harsh
| winter from the 1990s we are entering now a climate changed
| winter: warm, products to market race winter.
|
| Funnily enough, the Geoffrey Hinton of today is probably some
| symbols researcher, shouting in the desert, like Hinton shouted
| in the 1980s, that we need more than matmul.
|
| One interesting, biology-inspired mechanism, would be _quorum
| sensing_ [1]: a basal cognition-like, decision-making function
| in which decentralized systems (bacteria, cells) start building
| functionality (sensing /decision) from the bottom up. There is
| no hint for this sort of mechanism in our current artificial
| 'neural networks'. Not that it should, but our cells and in
| general cells use these kind of 'tricks' to solve problems in
| all kinds of spaces (transcriptomics, morphogenetics, etc.)
| without requiring ridiculous amounts of energy, time, or other
| resources.
|
| [1] https://en.wikipedia.org/wiki/Quorum_sensing
| joe_the_user wrote:
| LLMs may be terrible for the search for AGI but that may be
| good thing. If, mind you if, LLMs are always going to semi-
| smart "parrots", then society will get a picture of what an
| AI might without the AI being able to skynet-style takeover
| people are maybe justifiably worried.
|
| Now, as far as whether LLM or offshoots can get to AGI, I
| know plenty of good arguments for them not being able to do
| that. I don't think the claim that they're just accumulating
| more abilities in each iteration in an inexplicable way is
| true. But I've lived long enough to know that you should
| never get too cocky when one is "arguing with success". So
| maybe.
|
| Moreover, given that LLM programming is basically just bucket
| chemistry, if an LLM can the ability to competently pursue
| long term goals, it seems like it will have a good chance of
| some of its goals being random cruft that will make it quite
| dangerous.
| cpeterso wrote:
| Is there any information available about the non-LLM methods
| Carmack's AGI startup is researching?
| breatheoften wrote:
| Who is commenting on this and what will they say??!
|
| 1-2 trillion tokens of training for gptx class models
|
| It would take 20 years for a human to read that many tokens at 8
| hrs of reading per day (although i'm pretty unclear on whether
| that's unique token sequences or randomly selected token pairs --
| feels like a big difference between the two!!)
|
| We are all here now -- who thinks we are anywhere of note
| together?
|
| Personally I think that models that do not experience incremental
| change remain an entirely separate class of intelligence from
| those which evolve continuously -- but I am open to being
| persuaded otherwise (certainly I never thought making gpt2 bigger
| would be as impressive as gpt4) ...
| meindnoch wrote:
| >1-2 trillion tokens of training for gptx class models
|
| >It would take 20 years for a human to read that many tokens at
| 8 hrs of reading per day
|
| Normal reading is ~200 words per minute.
|
| That's 12,000 words per hour.
|
| That's 288,000 words per day (reading for 24 hours straight).
|
| That's 105,120,000 words per year.
|
| That's 2,102,400,000 words per 20 years.
|
| A trillion is 1,000,000,000,000.
| tonightstoast wrote:
| Plus almost no one is reading academic papers at 200 words
| per minute. So for the level of material it could mean that
| number is halved per 20 years.
| lostmsu wrote:
| Video from a pair of eyes is many terabytes per year though. If
| you compress it with best codecs. Raw video would be closer to
| 1PB per year.
|
| I think people overestimate the data efficiency of the
| biological neural networks w.r.t. transformer models.
| actionfromafar wrote:
| The _power effiency_ is pretty brutal though, in favour of
| bio-gelpacks.
| lostmsu wrote:
| How can you be sure about power efficiency, if you don't
| know how data efficiency compares? It is low power. But so
| is a phone running llama.cpp
| Thiez wrote:
| People with bad eyesight are not (to the best of my
| knowledge) less intelligent than those with perfect vision.
| We could probably supply video at a much lower resolution
| without significantly affecting the outcome.
| GordonS wrote:
| And, as humans, we don't only get "tokens" from what we read,
| but from our entire sensorium.
| intalentive wrote:
| Humans are not exposed to "billions of words". They are exposed
| to continuous signals, and somehow extract meaningful
| representations, like "words" and "objects", out of them.
|
| LLMs are exposed to human-curated data. Let's see an LLM curate
| its own data out of nothing but experience of raw continuous
| signals.
|
| Sure, GPT trained on the internet is a nice way to condense and
| retrieve human-curated data. But it's not going to give us a Lt.
| Data or C-3P0 that can learn and adapt in real time.
| pbw wrote:
| > Humans are not exposed to "billions of words" ... continuous
| signals
|
| Yes humans need to learn to extract words from signals, but
| that does not mean it's not correct to count the number of
| words they've been exposed to in their lifetime. Your comment
| is like saying we can't count the number of hamburgers someone
| has eaten, because they actually eat myofibrillar proteins, not
| burgers.
| d0mine wrote:
| Though millions seems more realistic (thousands of words per
| day)
| andsoitis wrote:
| what do you think of midjourney's ability to describe a user-
| provided image in text? https://the-decoder.com/midjourney-new-
| image-tool-works-in-r...
| gunshai wrote:
| But humans are also trained on human-curated data so much so we
| put fairly hefty price tags on that curation.
|
| https://miro.medium.com/v2/resize:fit:1056/0*E1eNateTiDThGcY...
| echelon wrote:
| We've only just started. Now more capital and more minds will
| be pouring into solving this.
|
| Buckle up.
| Sparkyte wrote:
| Also LLM are not AI just parts to AI you need a bunch of
| interfacing and injection components. The sum of which makes up
| AI.
|
| I see the progress of AI stagnating while people board the next
| equivalent of Crypto craze because someone want to financially
| profit off something not fully seeing realization.
|
| However it is not without merit which Crypto was completely
| without merit and something to show for itself.
|
| As flawed and imperfect ChatGPT is it clearly mirrors its
| creators and that is a compliment.
| Femtodjy wrote:
| Why not?
|
| Humans learn based on human curated data too.
|
| School is not natural. Our whole env is neither. Babies can't
| survive.
|
| I would even go so far to say that the potential model a LLM
| would create internally might not be that far away of that of a
| human.
|
| And segment anything was just announced. The performance of
| zero shot systems is tremendous.
|
| It's not far fetched to assume that chatgpt combined with
| segment anything together would allow it to create an even more
| accurate model of the world.
| bluefirebrand wrote:
| > Humans learn based on human curated data too.
|
| > School is not natural. Our whole env is neither. Babies
| can't survive
|
| A human goes to school to learn.
|
| Humans in general learned based on experience of the world
| around them. We invented language, no one taught it to us. We
| learned to make fire, forge tools, cook food, practice
| medicine, etc on our own.
|
| School is just how we pass down that learning.
| dwaltrip wrote:
| Humanity invented those things.
|
| Individual humans were taught those things, either directly
| or indirectly by observing / listening to others.
|
| We do learn through our experience, of course. But most of
| what we learn is from others in one form or another.
| pixl97 wrote:
| If I took you and pitched you on an island as an
| exceptionally young child with no further training, most
| likely you would die. Even more likely you wouldn't have
| fire and tools. A high proportion of our behaviors can be
| traced as a continuous learning chain going back eons.
| macrolocal wrote:
| * * *
| com2kid wrote:
| > Humans are not exposed to "billions of words". They are
| exposed to continuous signals, and somehow extract meaningful
| representations, like "words" and "objects", out of them.
|
| Plenty of studies showing long term academic achievement
| differences based on number of unique words babies are exposed
| to.
|
| And I promise you, having a baby/toddler is _really_ damn close
| to doing data labeling. Reading picture books, you are
| basically labeling objects. Walking around the grocery store,
| you are labeling objects. All the time, again and again, and
| the same object will get labeled repeatedly, and sometimes it
| will even be fact checked. A toddler will point to something
| that they know is an orange, ask "apple?" and you had sure as
| hell better reply "orange".
| cma wrote:
| LeCun's reply:
|
| > more like a thousand times more. > Between 1 and 2 trillion
| tokens. > It would take a person 22,000 years to read through 1
| trillion words at normal speed for 8 hours a day.
|
| It's kind of surprising to me that around 1000 humans could
| feasibly read the whole internet (as scraped for LLMs) between
| each other in only 22 years. The internet feels so much more
| massive than that. Though of course lots of the stuff on arxiv
| etc. is so dense there is no way you go through it at normal
| reading speed.
| brokencode wrote:
| I wonder if it only seems like our brains need relatively little
| training because we have millions of years of training encoded in
| our DNA.
|
| No other animal can learn human language or logic to the extent
| that humans can, no matter how much you train them. But humans
| learn language easily in the first few years of life.
|
| It's almost as if the human brain is preprogrammed with the
| general concepts of all languages, and it just needs to be fine
| tuned with specific vocabulary and grammar rules.
| lairv wrote:
| LLMs weights are initialized from a normal distribution (or
| whatever distribution is now SOTA) while some animals can walk
| the second they are born
| cypress66 wrote:
| GPT4 probably has a breadth of knowledge 1000x wider than a
| human, so the fact that it uses more tokens seems reasonable.
| roenxi wrote:
| It seems likely that there is a big missing component of visual
| data. How many Tb of visual data (effectively video) do babies
| get exposed to?
|
| That is what sets up their mental model of the world, and
| language is fit over the top of that. It is hard to assess neural
| net efficiency with that difference in place.
| rvnx wrote:
| It is going to amazing once LLMs get access to learn from real-
| life visual information (e.g. daily life videos without
| montage, or movies), and not just text.
| musesum wrote:
| funny, due to an aversion to twitter, I wanted to see if Carmac
| cross-posted on Mastodon, which yielded a link to another HN
| thread:
| https://cyberfeed.io/article/0d53a16b9c8bcc0a1d2dc5f77dac225...
| physPop wrote:
| Carmack is a smart guy but don't confuse his expertise in other
| areas with machine learning. Smells of hubris...
| SquareWheel wrote:
| He's been heavily researching machine learning for the last few
| years, and formed an AGI startup last year.
|
| He's not the world's leading researcher or anything, but he's
| far from green in this space.
| rvnx wrote:
| Perhaps he feels that the Metaverse thing he sold to Zuck isn't
| going anywhere, and now tries to ride a new wave.
| m3kw9 wrote:
| But we relate the words in a 3d world, llms relate words only to
| each other
| kklisura wrote:
| Q: If we need to feed additional data to the network, do we train
| it with existing data + new data or we can just train it on a new
| data? I'm asking because if we need to train it with both old and
| new data every time we have something new, that's not similar to
| human learning.
| IanCal wrote:
| It's extremely similar to training new humans.
| wkdneidbwf wrote:
| surprises me a human would be exposed to a billion words. i never
| thought about it, but a billion is such a large number
|
| edit: oh, does he mean a billion total rather than a billion
| unique words? haha i r smat
___________________________________________________________________
(page generated 2023-04-07 23:02 UTC)