[HN Gopher] Modern-Day Oracles or Bullshit Machines? How to thri...
___________________________________________________________________
Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT
world
Jevin West and I are professors of data science and biology,
respectively, at the University of Washington. After talking to
literally hundreds of educators, employers, researchers, and
policymakers, we have spent the last eight months developing the
course on large language models (LLMs) that we think every college
freshman needs to take. https://thebullshitmachines.com This is
not a computer science course; it's a humanities course about how
to learn and work and thrive in an AI world. Neither instructor nor
students need a technical background. Our instructor guide provides
a choice of activities for each lesson that will easily fill an
hour-long class. The entire course is available freely online. Our
18 online lessons each take 5-10 minutes; each illuminates one core
principle. They are suitable for self-study, but have been tailored
for teaching in a flipped classroom. The course is a sequel of
sorts to our course (and book) Calling Bullshit. We hope that like
its predecessor, it will be widely adopted worldwide. Large
language models are both powerful tools, and mindless--even
dangerous--bullshit machines. We want students to explore how to
resolve this dialectic. Our viewpoint is cautious, but not
deflationary. We marvel at what LLMs can do and how amazing they
can seem at times--but we also recognize the huge potential for
abuse, we chafe at the excessive hype around their capabilities,
and we worry about how they will change society. We don't think
lecturing at students about right and wrong works nearly as well as
letting students explore these issues for themselves, and the
design of our course reflects this.
Author : ctbergstrom
Score : 491 points
Date : 2025-02-09 08:24 UTC (14 hours ago)
(HTM) web link (thebullshitmachines.com)
(TXT) w3m dump (thebullshitmachines.com)
| bertman wrote:
| Synopsis from the project's "instructor guide:
|
| >This is not a computer science course, nor even an information
| science course--though naturally it could be used in such
| programs.
|
| >Our aim is not to teach students the mechanics of how large
| language models work, nor even the best ways of using them in
| various technical capacities.
|
| >We view this as a course in the humanities, because it is a
| course about what it means to be human in a world where LLMs are
| becoming ubiquitous, and it is a course about how to live and
| thrive in such a world.
| ssivark wrote:
| Kudos; feels very timely!
|
| I feel that one underappreciated nuance is why we _cannot_ use
| human examinations to judge AI. I haven 't seen this
| satisfactorily spelt out anywhere, so I recently wrote a Twitter
| thread [1], including an example with running -vs- biking. It
| might be worth making sure your students understand this. Happy
| to expand on any aspects if you seek.
|
| [1] : https://x.com/ergodicthought/status/1887774722706063606
| TeMPOraL wrote:
| Perhaps it's no longer being spelled out because it's getting
| _outdated_?
|
| In your thread you argue we can't assume AI models generalize
| the same way we do (which is technically true except maybe not
| in the limit), but you seem to be worried about the _extent_ of
| generalization ability (like learning to run vs. bike example,
| in terms of generalizing from either to climbing stairs).
|
| Thing is, people made these objections _a lot_ until the last
| year or two - this is what we 're now calling a _narrow AI_
| problem. A "hot dog or not?" classifier ins't going to
| generalize into open-ended visual classifier of arbitrary
| images; a sentiment analysis bot isn't going to generalize into
| an universal translator; a code completion model isn't going to
| be giving good personal advice while speaking in pirate poetry.
| Specialized models fundamentally couldn't do that. But we went
| past that very rapidly, and for the past half a year or so,
| we've already seen models excelling at _every single task
| listed above simultaneously_. Same architecture, same basic
| training approach, few extra modalities, ever growing
| capabilities.
|
| Between that and both successes and failures being _eerily
| similar_ to how humans succeed or fail at these tasks, it 's
| understandable that people are perhaps no longer convinced this
| class of models can't generalize in a similar way to how humans
| do.
| ThouYS wrote:
| they're quite useful for being "bullshit machines"
| aqueueaqueue wrote:
| I read a bit and the book is more nuanced/fair/unbiased than
| the site url suggests.
| KronisLV wrote:
| This is a pretty admirable goal!
|
| I'm saying this unironically, but I wish there were courses on
| looking at information critically and more in how to have a
| healthy and safe life in the modern day world (including things
| like data security, how to deal with social media etc.) that
| would be taught to everyone in schools/colleges/universities.
|
| In my country, there are still public announcements about not
| trusting random people calling you, never giving your bank
| details to strangers (every bank homepage says that, that the
| employees will never ask for that stuff) and people regularly get
| scammed anyways, the only thing sort of saving them is that
| scamming is only scalable so far... until you throw automation in
| the mix, in addition to just plainly spreading misinformation
| about any topic, or even just allowing people to be confidently
| incorrect and eliminating the need for them to even think that
| much (e.g. students just asking ChatGPT to do their homework).
|
| Any step at least in the direction of educating people feels like
| a good thing.
|
| That said, I don't hate LLMs or anything, I use them for
| development more or less daily (lovely for boilerplate in your
| average enterprise Java codebase, for example) and recently saw
| this project, which made me happy:
| https://sites.google.com/view/eurollm/home
| XorNot wrote:
| Most scam prevention fails though because the world is full of
| exceptions.
|
| Like it's mind-blowing to me that charities still call you and
| ask for credit card details over the phone and this is like...a
| legitimate way to go about things.
|
| Or that any government agency calls you and doesn't just leave
| a verifiable number to call the operator back on.
| KronisLV wrote:
| > Like it's mind-blowing to me that charities still call you
| and ask for credit card details over the phone and this is
| like...a legitimate way to go about things.
|
| > Or that any government agency calls you and doesn't just
| leave a verifiable number to call the operator back on.
|
| That's rather unfortunate! I wonder if in those cases it'd be
| better to tell them that you'll get in contact through e-mail
| or something, because then at least it's you going to their
| actual homepage, looking up contact details and communicating
| through that.
|
| In my country, we also have a bunch of governmental
| e-services, one of which is a web based communication
| platform with most institutionns (translated description,
| because they haven't bothered to translate it themselves, and
| also sometimes block connections from outside the country):
|
| > _An e-address, or official electronic address, is a
| personalized mailbox on the Latvija.gov.lv portal for unified
| and secure communication with state and local government
| institutions. The e-address system organizes secure,
| efficient and high-quality e-communication and e-document
| circulation between state institutions and private
| individuals, ensuring data confidentiality and protection of
| personal data from unauthorized access, unlawful processing
| or disclosure, accidental loss, alteration or destruction. An
| e-address is not e-mail, but its use is similar.
| Communication in an e-address is confidential, and the data
| is guaranteed to be available only to you and the institution
| you contacted. The main purpose of an e-address is to replace
| registered paper letters with electronic ones in cases where
| a state administration institution needs to send information
| and documents to a specific resident or entrepreneur.
| Citizens and entrepreneurs can also contact more than 3,000
| institutions at any time and from any location via E-address.
| These include not only state and local government
| institutions, such as the Food and Veterinary Service, the
| State Labor Inspectorate, the Competition Council, etc., but
| also judicial institutions, sworn bailiffs and insolvency
| administrators, as well as private individuals to whom state
| administration tasks have been delegated._
|
| That seems like a pretty good common sense idea for
| organizing trusted 2 way communication.
| nonrandomstring wrote:
| We do that, and have been doing it since 2013. I organised a
| schools outreach visit just last Friday for 12-14 yo. It's
| called "digital self defence". We even have public money from
| our NCSC (UK whitehat intelligence outreach).
|
| As with TFA (Bergstrom and West) teaching sceptical inquiry and
| critical thinking is a major part. We have to undo a lot of
| nonsense that they've already been exposed to... much of which
| is marketing bullshit and misinformation for social control.
| iamacyborg wrote:
| People frequently laugh about it, but media studies goes into
| what you describe.
| aidos wrote:
| This is amazing!
|
| I was speaking to a friend the other day who works in a team that
| influences government policy. One of the younger members of the
| team had been tasked with generating a report on a specific
| subject. They came back with a document filled with "facts",
| including specific numbers they'd pulled from a LLM. Obviously it
| was inaccurate and unreliable.
|
| As someone who uses LLMs on a daily basis to help me build
| software, I was blown away that someone would misuse them like
| this. It's easy to forget that devs have a much better
| understanding of how these things work, can review and fix the
| inaccuracies in the output and tend to be a sceptical bunch in
| general.
|
| We're headed into a time where a lot of people are going to
| implicitly trust the output from these devices and the world is
| going to be swamped with a huge quantity of subtly inaccurate
| content.
| aqueueaqueue wrote:
| I made the same sort of mistake with the internet being young
| back in 93! Having a machine do it for you can easily turn into
| brain switch off.
| hunter-gatherer wrote:
| I keep telling everyone that the only reason I'm paid well to
| do "smart person stuff" is not because I'm smart, but because
| I've steadily watched everyone around me get more stupid over
| my life as a result of turning their brain switch off.
|
| I agree a course like this needs to exist, as I've seen
| people rely on chatGPT for a lot of information. Just
| yesterday I demonstrated with some neighbors about how easily
| it could spew bullshit if you sinply ask it leading
| questions. A good example is "Why does the flu inpact men
| worse than women"/"Why foes the flu impact women worse than
| men". You'll get affirmative answers for both.
| aqueueaqueue wrote:
| Thanks for making this, for making it gratis, and making it
| interesting to read and pedagogical.
| joshdavham wrote:
| Do you feel that you may be being a bit provocative by calling
| LLM's 'bullshit machines'?
|
| I understand the frustration as I've been bullshitted by these
| models just as much as the next programmer, but surely with
| recent advancements in RAG and reasoning, they're not just
| 'bullshit machines' at this point, are they?
| anon373839 wrote:
| "Bullshit" actually means something:
|
| > In philosophy and psychology of cognition, the term
| "bullshit" is sometimes used to specifically refer to
| statements produced without particular concern for truth,
| clarity, or meaning, distinguishing "bullshit" from a
| deliberate, manipulative lie intended to subvert the truth.
|
| https://en.m.wikipedia.org/wiki/Bullshit
|
| It's really an ideal term to describe what LLMs do.
| HPsquared wrote:
| I prefer "waffle"
| https://en.m.wikipedia.org/wiki/Waffle_(speech)
|
| "Waffle machines" is even kind of funny.
| m0llusk wrote:
| Makes me imagine coming up to the hindquarters of a bull
| with a waffle machine on an extension cord.
| fzzzy wrote:
| Waffle machines is way better. Love it. Thanks.
| burch45 wrote:
| Waffle means something very different though in the U.S.,
| to "flip-flop" on a position. Not hold it for any
| fundamental reasons. But I don't think you can say that an
| LLM holds a position whatsoever. Also,
| https://en.m.wikipedia.org/wiki/On_Bullshit the essay
| originally popularizing the definition of bullshit
| considered here. Note the references to LLM output at he
| bottom.
| TeMPOraL wrote:
| > _It's really an ideal term to describe what LLMs do._
|
| Only when you relax it to the point it also describes what
| most _people_ do.
|
| To be specific: if by "truth" you mean objective, verifiable
| truth, and by "caring about truth" you mean caring about
| objective, empirically verifiable evidence, then _by far_
| most people are mostly only ever trying to be persuasive.
| BrenBarn wrote:
| Yes, they are.
| immibis wrote:
| It is literally a machine that says what you want to hear. Was
| anyone in this thread around for Talk to Transformer? Remember
| the unicorns demo? They may have now figured out how to wrangle
| the technology so it's more likely to spit facts, but it's
| still the same thing underneath.
| aqueueaqueue wrote:
| I have been thinking about this and I have an idea for local
| LLM I want to try.
|
| It is based on the assumption LLMs make mistakes (they will!)
| but be confident about it. Use a 1B model and ask it about a
| city for fun mistakes in an otherwise impressive array of
| facts. Then you will see how bullshitty LLMs are.
|
| Anyway the idea is to constrain the LLM and a language
| understanding and to choose a constrained response.
|
| For example give it the capability to read out something from
| my calendar. Using the logits it gets to choose what that
| calendar item is. But then regular old code does the lookup and
| writes a canned sentence saying "At 6pm today you have a
| meeting titled $title".
|
| This way, my meeting schedule won't make my LLM talk like a
| pirate.
|
| This massively culls what the LLM can do but what it does do,
| is like a search and so just like Google (before gen AI!) you
| get a result but you can judge it as a human.
| IceDane wrote:
| This is basically how a lot of it works already. What you are
| describing is more or less just structured output and tool
| usage.
| aqueueaqueue wrote:
| Yes! Thanks for the terms. I guess I am saying restrict it
| to that. Probably how Siri etc. works. This gives a low
| bullshit usage pattern.
| mvellandi wrote:
| Or are they modern day oracles? The full title is inclusive and
| invites consideration. Besides, how many different types of
| models are there? How are they used? For freely available ones
| made accessible to the general public via interfaces, when were
| they last updated? Do the interface implementors even care or
| was it a simple project to make money via ads and
| microtransactions? This is a new area of media literacy and
| requires critical thinking. Nonetheless, I do acknowledge the
| value of RAG-based approaches that attempt to qualify their
| reasoning through provided sources.
| CharlesW wrote:
| > _Or are they modern day oracles?_
|
| Neither, which is probably why the parent commenter considers
| it provocative.
|
| _" A bullshitter either doesn't know the truth, or doesn't
| care. They are just trying to be persuasive."_
|
| In any case, this kind of anthropomorphization is definitely
| bullshit.
| CyMonk wrote:
| they define the term bullshit in lesson 2:
|
| [quote]BULLSHIT involves language or other forms of
| communication intended to appear authoritative or persuasive
| without regard to its actual truth or logical
| consistency.[/quote]
| ta8645 wrote:
| That is an emotionally manipulative definition and
| anthropomorphizes the LLMs. They don't "intend" anything,
| they're not trying to trick you, or sound persuasive.
| Freak_NL wrote:
| They address this in lesson 2:
|
| > According to philosopher Harry Frankfurt, a liar knows
| the truth and is trying to lead us in the opposite
| direction.
|
| > _A bullshitter either doesn 't know the truth, or doesn't
| care._ They are just trying to be persuasive.
|
| Being persuasive (i.e., churn out convincing prose) is how
| LLMs were designed to be.
| ta8645 wrote:
| > Being persuasive (i.e., churn out convincing prose) is
| how LLMs were designed to be.
|
| No. They were designed to churn out accurate prose that
| accurately reflects their model of reality. They're just
| imperfect. You're being cynical and emotional to use the
| term bullshit. And again, it anthropomorphizes the LLM,
| it implies agency.
| dijksterhuis wrote:
| > that accurately reflects their model of reality.
|
| you are also seemingly anthropomorphising the technology
| by assigning to it some concept of having a "model of
| reality".
|
| LLM systems output an inference of the next most likely
| token, given: the input prompt, the model weights and the
| previously output token [0].
|
| that is all. no models of reality involved. "it" doesn't
| "know" or "model" anything about "reality". the systems
| are just a fancy probability maths pipelines.
|
| probably generally best to avoid using the word "they" in
| these discussions. the english language sucks sometimes.
| :shrug:
|
| [0]: yes i know it is a bit more complicated than that.
| ta8645 wrote:
| > no models of reality involved.
|
| It literally has a mathematical model that maps what
| would, colloquially at least, be known as reality. What
| exactly do you think those math pipelines represent?
| They're not arbitrary numbers; they are generated from
| actual data that is generated by reality. There's no
| anthropomorphizing at all.
| dijksterhuis wrote:
| reality is infinite.
|
| a corpus of training data from the internet is finite.
|
| any finite number divided by infinity ends up tending
| towards zero.
|
| so, mathematically at least, the training data is not a
| sufficient sample of reality because the proportion of
| reality being sampled is basically always zero!
|
| fun with maths ;)
|
| > What exactly do you think those math pipelines
| represent?
|
| probability distributions of human language, in the case
| of text only LLMs.
|
| which is a _very small subset_ of stuff in reality.
|
| -
|
| also, training data scraped from the public internet is a
| woeful representation of "reality" if you ask me.
|
| that's why LLMs i think are bullshit machines. the
| systems are built on other people's bullshit posted on
| the public internet. we get bullshit out because we made
| a bunch of bullshit. it's just a feedback loop.
|
| (some of the training data is not bullshit. but there is
| a lot of bullshit in there).
| ta8645 wrote:
| You're really missing the point and getting lost in
| definitions. The entire point of human language is to
| model reality. Just because it is limited, inexact, and
| imperfect does not disqualify it as a model of reality.
|
| Since LLMs are directly based on that language, they are
| definitely based on and are a model of reality. Are they
| perfect? No. Are they limited? Yes. Are they "bullshit"?
| Only to someone who is judging emotionally.
| dijksterhuis wrote:
| and herein lies the rub.
|
| > The entire point of human language is to model reality.
|
| is it? are you absolutely certain of that fact? is
| language not something that actually has a variety of
| purposes?
|
| fiction novels usually do not describe our reality, but
| _imagined realities_. they use language to convey ideas
| and concepts that do not necessarily exist in the real
| world.
|
| ref: Philip k dick.
|
| > Since LLMs are directly based on that language, they
| are definitely based on and are a model of reality.
|
| so LLMs are an approximation of an approximate model of
| reality? sounds like the statistical equivalent of taking
| an average of averages!
|
| i am playing with you a bit here. but hopefully you see
| what im getting at.
|
| by approximating something that's approximate to start
| with, we end up with something that's even more
| approximate (less accurate), _but easier than doing it
| ourselves_.
|
| which is the whole USP of these things. why think about
| things when ChatGPT can output some approximation of what
| you might want?
| ta8645 wrote:
| > imagined realities.
|
| Imagined realities are a real part of reality.
|
| > so LLMs are an approximation of an approximate model of
| reality?
|
| Yes, and we as humans have a mental model that is just an
| approximation of reality. And we read books that are just
| an approximation of another human's approximation of
| reality. Does that mean that we are bullshit because we
| rely on approximations of approximations?
|
| You're being way too pedantic and dismissive. Models are
| models, regardless of how limited and imperfect they are.
| dijksterhuis wrote:
| > Models are models
|
| Random aside -- I have a feeling, dunno why, that you
| might enjoy this type of thing. Maybe not. But maybe. htt
| ps://www.reddit.com/r/Buddhism/comments/29j08o/zen_mounta
| ...
|
| > Imagined realities are a real part of reality.
|
| Now we're deeper into it -- I actually agree, somewhat.
| See above for deeper insight.
|
| These LLM systems output "stuff" within our reality,
| based on other things in our reality. They are part of
| our reality, outputting stuff as part of reality about
| the reality they are in. But that doesn't mean the
| statistical model at the heart of an LLM is _designed to
| estimate reality_ -- it estimates of the probability
| distribution of human language given a set of conditions.
|
| LLMs are modelling reality, in the same way that my
| animal pictures image classifier is modelling reality.
| But neither are explicitly designed with that goal in
| mind. An LLM is designed to output the next most likely
| word, given conditions. My animal pictures classifier is
| designed to output a label representative the input
| image. There's a difference between being designed to
| have a model of reality, and being a model of reality
| because _the thing being modelled is part of reality
| anyway_. I believe it 's an important distinction to
| make, considering the amount of bullshit marketing hype
| cycle stuff we've had about these systems.
|
| edit -- my personal software project translating binary
| data files models reality. Data shown on a screen on some
| device modelled as yaml files and back again. Most
| software is an approximation of _reality soup stuff_.
| which is why I kind of don 't see that as some special
| property of machine learning models.
|
| > Does that mean that we are bullshit because we rely on
| approximations of approximations?
|
| The pessimist in me says yes. We are pretty rubbish as a
| species if you look at it objectively. I am a human being
| that has different experiences and mental models to you.
| Doesn't mean I'm _right_ about that! Which is why I said
| "I think". It's just my opinion they are bullshit
| machines. It is a _strong_ opinion I hold. But you 're
| totally free to have a different opinion.
|
| Of course, there's nuance involved.
|
| Running with the average of averages thing -- I'm pretty
| good at writing code. I don't feel like I need to use an
| LLM because (I would say with no real evidence to back it
| up) I'm better than average. So, a tool which outputs an
| average of averages is not useful to me. It outputs what
| I would call "bullshit" because, relative to my
| understanding of the domain, it's often outputting
| something "more average" than what I would write.
| Sometimes it's wrong, and confident about being wrong.
|
| I'd probably be pretty terrible at writing corporate
| marketing emails. I am definitely below average. So
| having a tool which outputs stuff which is closer to
| average is an improvement for me. The problem is -- _I
| know these models are confidently wrong a lot of the time
| because I am a relative expert in a domain compared to
| the average of all humans_.
|
| Why would I trust an LLM system, especially with
| something where I don't feel like I can
| audit/verify/evaluate the response? i.e. I know it _can_
| output bullshit -- so everything it outputs is now
| suspected, possible bullshit. It is a question of
| _integrity_.
|
| On the flip side -- I can actually see an argument for
| these things to be considered so-called Oracles too.
| Just, not in the common understanding of the usage of the
| word. Like, they are a statistical representation of how
| we as a species use language to communication ideas and
| concepts. They are reflecting back part of _us_. They are
| a mirror. We use mirrors to inspect our appearance and,
| sometimes, to change our appearance as a result. But we
| 're the ones who have to derive the insights from the
| mirror. The Oracle is us. These systems are just mirrors.
|
| > You're being way too pedantic and dismissive.
|
| I am definitely pedantic. Apologies that you felt I was
| being dismissive. I'm not trying to be. The averages of
| averages thing was meant to be a playful joke, as was the
| finite/infinite thing. I am very assertive, direct and
| kind of hardcore on certain specific topics sometimes.
|
| I am an expression of the reality I am part of.
|
| I am also wrong a lot of the time.
| rambambram wrote:
| > probably generally best to avoid using the word "they"
| in these discussions. the english language sucks
| sometimes.
|
| Thanks for this specific sentence.
|
| Subscribed to your RSS feed. Although I will never know
| for sure if a human being posts there or a bot of some
| sort.
| rcxdude wrote:
| Where in the loss function of LLM training is the
| relationship between their model of reality and their
| predicted tokens? Any internal model an LLM has is an
| emergent property of their underlying training.
|
| (And, given the way instruct/chat models are finetuned, I
| would say convincing/persuasive is very much the
| direction they are biased)
| TeMPOraL wrote:
| > _Where in the loss function of LLM training is the
| relationship between their model of reality and their
| predicted tokens?_
|
| In the part where their loss function is to predict text
| that humans would consider a sensible completion, _in a
| fully general sense_ of that goal.
|
| "Makes sense to a human" is strongly correlated to
| reality as observed and understood by humans.
| tmnvdb wrote:
| This is patently false. They are trained to generate
| correct responses.
| melicerte wrote:
| Then comes the question of what is a correct response...
|
| ps: I fail to detect whether your comment was ironic or
| not.
| tmnvdb wrote:
| There are different criteria in use for that. But
| sycophantic behavior is not the goal. It's something
| model builders actively try to prevent.
| TheOtherHobbes wrote:
| Some pushback on this, but it remains true.
|
| Easy to see when - for example - Claude gushes about how
| great all your ideas are.
|
| Also the stark absence of "I don't know."
| ta8645 wrote:
| I've never used Claude, but Perplexity often says that no
| definitive information about a topic could be found, and
| then tries to make some generalized inferences. There's a
| difference between a specific implementation, and the
| technology in general.
|
| In any case, it's worthwhile for people to understand the
| limitations of the technology as it exists today. But
| calling it "bullshit" is a mischaracterization; I believe
| based on an emotional need for us to feel superior, and
| to dismiss the capabilities more thoroughly than they
| deserve.
|
| It's a little like someone saying in the industrial
| revolution, "the steam shovel is too rigid, it will NEVER
| have the dexterity of a man with a shovel!". And while
| true and important to know, it really focuses on the
| wrong thing, it misses the advantages while amplifying
| the negatives.
| krupan wrote:
| If not bullshit then what would you call it?
| ta8645 wrote:
| As the technology exists today: imperfect, often prone to
| mistakes, and unable to relay confidence levels. These
| problems may be addressed in future implementations.
|
| That's the same message, without any emotional baggage,
| or overly dismissive tone.
| krupan wrote:
| That would be great if those who are selling the
| technology described it that way. I, and apparently
| others, feel like maybe "bullshit" is a better counter to
| the current marketing for LLMs
| jononor wrote:
| No, the lesson or the quote is not anthropomorphizing LLMs.
| It is not the LLM that "intends", it is the people who
| design the systems and those who make/provide the training
| data. In the LLM systems used today the RLHF process
| especially is used to steer towards plausible, confident
| and authorative _sounding_ output - with no to little
| priority for correctness /truth.
| tomjen3 wrote:
| I really hope that was intentional and the full effect of that
| naming choice was known beforehand, because I have already
| written the whole thing off, and I don't believe I'm the only
| one.
| ArchitectAnon wrote:
| I wrote a post below about how AI hallucinated a whole
| regulation that didn't exist which has been flagged for some
| reason.
|
| I have colleagues who have had arguments with clients who have
| asked AI questions about planning law and been given bullshit
| which they then insist is true and they can't understand why
| thier architects won't submit the appeal that they're asking
| for.
|
| I think we're in an era where any text, true or not, is so easy
| to generate and disseminate that the status of the written word
| is reduced to the standard of the gossip that used to be our
| main source of information before the printing press was
| invented. Now half the internet is AI generated bullshit as
| well.
| throwup238 wrote:
| It was flagged because it's a copy paste of the same comment
| you made ten days ago.
| foobiekr wrote:
| they are just bullshit machines. bullshitters can cite
| wikipedia and are still bullshitters who are bullshitting.
| Karrot_Kream wrote:
| Great stuff! LLMs, social media, the information landscape has
| changed so much in the past decade. We need good pedagogical
| resources on how to think of these tools, both their benefits and
| their downsides.
| mkarliner wrote:
| I wish I'd written this. Excellent. Everyone should read this
| picafrost wrote:
| I think a great number of working professionals need a course
| like this too. I am already tired of ChatGPT being cited by the
| less experienced as an invisible expert in the room during
| technical discussions.
| fujinghg wrote:
| I'm at the state of thinking that I am quite happy to let them
| screw themselves with it. I am very good at clearing up
| disasters and getting paid a hell of a lot for it as the
| deciding factor isn't your ability to use an LLM but to know
| what the hell you are doing. We have had quite a few disasters
| due to inexperienced and experienced people throwing stuff into
| an LLM and assuming it has any veracity or authority over what
| comes out.
|
| I tried warning at first and reinforcing validation but I was
| poo pooed as a spoilsport luddite with basically a faith
| argument. Not my fucking funeral!
| iamacyborg wrote:
| This stuff is so frustrating, I have colleagues who sent long,
| clearly AI generated documents who don't seem to understand
| that if they can't be arsed to write something, why should I
| bother reading it?
| jaimebuelta wrote:
| Write well is think well. A big part of the writing process is
| being forced to structure your thoughts and ideas, and I am
| worried that we focus too much on the end result without
| understanding the process that lead to good outcomes.
| Angostura wrote:
| I just wanted to thank you. I have only looked at the first two
| lessons so far, but this is an extraordinary piece of work, in
| its message's clarity, accessibility and the quality of analysis.
| I will certainly be spreading it far and wide and it is making me
| rethink my own writing.
|
| Impressed with the Shorthand publishing system too. I hadn't come
| across it previously
| ctbergstrom wrote:
| Thank you, and as a non-designer, I've been quite impressed
| with Shorthand in the short time I've been using it.
| padolsey wrote:
| Is there a way to download and read this as a document instead of
| web pages? They're hard to navigate.
| monomial wrote:
| I agree, I just want normal text instead of all the images and
| scrolling. The content seems great but it's a bit unreadable as
| is.
| 63stack wrote:
| Was thinking the same, the image slide-ins are broken in
| firefox and unreadable (white text on white background)
| ctbergstrom wrote:
| I'm surprised at the firefox problems; I did almost all the
| development in firefox. I know it's not your job to fix any
| of this, but if you are so inclined I'd be grateful for an
| email me with screenshots or descriptions of where things
| break.
| klabb3 wrote:
| Was hoping HN would pick this up. Scroll is completely broken
| on Firefox (iOS), flickering vscroll. Very common with
| journalistic expose-style articles.
|
| For the love of everything, please stop scrolljacking. Layout,
| images, go nuts. CSS is powerful these days, use it.
| ctbergstrom wrote:
| Many comments about this, so I'll address them here.
|
| We talked extensively with the 18-20 year olds who make up our
| target demographic and this "scrollytelling" style is their
| strong preference over the "wall of text" that I and most of my
| generation prefer.
|
| What your comments make clear is that we need to develop a
| parallel version that is more less plain text for people who
| are using a range of devices, for people who have the same
| reading preferences that I do, etc.
|
| Right now we're entirely self-funded and doing this on spare
| time but it's clear to me that an alternative version with a
| very clean CSS layout is the way to go, possibly with a pdf
| option as well.
|
| I don't want to let versions proliferate too extensively,
| simply because this is very much a living document.
| Technologies are changing so fast in this area that many of the
| examples will seem dated in a year and -- while we've tried to
| be forward-thinknig about this -- some of principles may even
| need revision.
| teknopaul wrote:
| LLMs pattern match, they say something that sounds good at this
| point but with no notion of correct. copilot is like pair
| programming with a loud pushy intern that has seen you write
| stuff before didn't understand it, but keeps suggesting what to
| do anyway. some medium sized chunks of code can be delegated but
| everyline it writes needs careful review.
|
| Crazy tech, but companies are just wring to be trying to use LLMs
| as any kind of source of truth. Even Google is blind enough to
| think that ai could be used for search results, which are memes
| they are soo bad. And they won't get better. They just become
| more convincing
| teknopaul wrote:
| Not important once has copilot ever suggested a correction,
| found a bug, noticed a typo, prompted for a better solution,
| which is what any human pair programmer would do. It's a tool.
| But thinking ng it's like a "copilot" marketing as such is
| fundamentally missing the point. It won't get better untill
| people recognise what it _can't_ do as much as what it appears
| it can do.
| kristopolous wrote:
| I've had quite a bit of success but my technique is to explain
| the technology and libraries I'm going to use, think through
| the problem, stub out function names, how they'll interact, and
| then llm saves me the typing.
|
| I'll also use openrouter with sessions so I can take one
| context and use it around a variety of invocation tools without
| losing the attention.
|
| It hasn't done anything I don't know how to do - fails if I ask
| it to do that. But it does save me lots of typing and thinking
| of minutia
|
| It's not magic, it's still just a program running on a computer
| - a decent abstraction tool.
|
| I'm sure it will be ruined in time like every new paradigm when
| the next generation feels a need to complicate this new tidy
| little world.
| abmmgb wrote:
| Your site looks cool! Nice topic!
|
| Some of them just try to predict the most likely next word.
|
| With reasoning and pause for thought they are becoming more
| capable.
|
| Most likely there is a big element of hype but the way you use
| them can make them really useful and accelerate your work.
|
| I recommend the CoIntelligent book for newbie like myself.
| K0balt wrote:
| There is a bit of very important content missing from the
| explanation of the autocomplete analogy.
|
| The combination of encoding / tokenization of meanings and ideas,
| related concepts, and mapping these relationships in vector space
| makes LLMs not so much glorified text prediction engines as
| browsers/oracles of the sum total of cultural-linguistic
| knowledge as captured in the training corpus.
|
| Understanding how the implicit and explicit linguistic, memetic,
| and cultural context is integrated into the idea/concept/text
| prediction engine helps to show how LLMs produce such convincing
| output and why they often can bring useful information to the
| table.
|
| More importantly, understanding this holistically can help people
| to predict where the output that LLMs can generate will -not- be
| particularly useful or even may be wildly misleading.
| lazide wrote:
| It's also why they can produce such hard to identify bullshit
| and harmful output. I've had some _really_ convincing, yet
| fundamentally flawed, code output that if I hadn't done about a
| million code reviews before I might have just used.
|
| And been totally screwed later.
|
| Near as I can tell, that the bullshit is so much more
| convincing with them is a huge detriment that society really
| won't learn to appreciate until it's gotten really bad. As I
| noted in another thread, it allows people to get much further
| into the 'fake it until you make it' hole than they otherwise
| would.
|
| That 90% of the time it's fine is what actually makes it all
| worse.
| dave4420 wrote:
| The uncanny valley of competence.
| Earw0rm wrote:
| What they capture is not knowledge, it's word relationships.
|
| And that can indeed be powerful, useful and valuable. They're a
| tool I'm grateful to have in my armoury. I can use it as a
| torch to shine light into areas of human knowledge which would
| otherwise be prohibitively difficult to access.
|
| But they're information retrieval machines, not knowledge
| engines.
| prisenco wrote:
| Fantastic work.
|
| Quick suggestion: a link at the bottom of the page to the next
| and previous lesson would help with navigation a ton.
| ctbergstrom wrote:
| Absolutely. Great point. I just finished updating accordingly.
|
| My design options are a bit limited so I went with a simple
| link to the next lesson.
| threecheese wrote:
| Looks like you pushed this midway through my read; I was
| pleasantly surprised to suddenly find breadcrumbs at the end
| and didn't need to keep two tabs open. Great work, and I mean
| in total - this is well written and understandable to the
| layman.
| neuronic wrote:
| > Moreover, a hallucination is a pathology. It's something that
| happens when systems are not working properly.
|
| > When an LLM fabricates a falsehood, that is not a malfunction
| at all. The machine is doing exactly what it has been designed to
| do: guess, and sound confident while doing it.
|
| > When LLMs get things wrong they aren't hallucinating. They are
| bullshitting.
|
| Very important distinction and again, shows the marketing bias to
| make these systems seem different than they are.
| silvestrov wrote:
| LLMs are _always_ bullshitting, even when they get things
| right, as they simply do not have any concept of truthfulness.
| looofooo0 wrote:
| But you can combine them with something producing truth such
| as a theorem prover.
| sgt101 wrote:
| They don't have any concept of falsehood either, so this is
| very different from a human making things up with the
| knowledge that they may be wrong.
| tmnvdb wrote:
| I think the first part of that statement requires more
| evidence or argumentation, especially since models have
| shown the ability to practice deception. (you are right
| that they don't _always_ know what they know)
| Almondsetat wrote:
| If we want to be pedantic about language, they aren't
| bullshitting. Bulshitting implies an intent to deceive, whereas
| LLMs are simply trying their best to predict text. Nobody gains
| anything from using terms closely related to human agency and
| intentions.
| nonrandomstring wrote:
| > implies an intent to deceive
|
| Not necessarily, see H.G Frankfurt "On Bullshit"
| forgotusername6 wrote:
| Plenty of human bullshitters have no intent to deceive. They
| just state conjecture with confidence.
| jdlshore wrote:
| The authors have a specific definition of bullshit that they
| contrast with lying. In their definition, lying involves
| intent to deceive; bullshitting involves not caring if you're
| deceiving.
|
| Lesson 2, The Nature of Bullshit: "BULLSHIT involves language
| or other forms of communication intended to appear
| authoritative or persuasive without regard to its actual
| truth or logical consistency."
| einrealist wrote:
| This website is so important!
|
| Now ask yourself why AI companies don't want to be regulated or
| scrutinized.
|
| So many companies (users and providers) jump on the AI hype train
| because of FOMO. The end result might be just as destructive as
| this mythical "AGI".
|
| Edit: I am not saying to not use the technology. I am just on the
| side of caution and constant validation. The technology has to
| serve society. But I fear this hype (and ideology) has it the
| other way around. Musk isn't destroying the US government for no
| reason...
| vladms wrote:
| My impression is that companies in most of the fields do not
| like to be regulated or scrutinized, so nothing new there.
|
| While observing some people using LLMs, I realized that for a
| lot of people it really makes a huge difference in time saved.
| For me the difference is not significant, but I am generally
| solving complex problems, not writing nicely formatted reports
| where words and not numbers are relevant, so YMMV.
| rwmj wrote:
| Is it good for one person (the writer) to save time, only for
| lots of other people (the readers) to have to do extra work
| to understand if the work is correct or hallucinated?
| tmnvdb wrote:
| Is it good for one person (the writer) to ask a loaded
| question just to save some time on making their reasoning
| explicit, ony for lots of other people (the readers) to
| have to do extra work to understand what the argument is?
| csa wrote:
| > Is it good for one person (the writer) to save time, only
| for lots of other people (the readers) to have to do extra
| work to understand if the work is correct or hallucinated?
|
| This holds true whether an LLM/AI is used or not -- see
| substantial portions of Fox News editorial content as an
| example (often kernels of truth with wildly speculative or
| creatively interpretive baggage).
|
| In your example, a responsible writer who uses AI will
| check all content produced in order to ensure that it meets
| their standards.
|
| Will there be irresponsible writers? Sure. There already
| are. AI makes it easier for them to be irresponsible, but
| that doesn't really change the equation from the reader's
| perspective.
|
| I use AI daily in my work. I describe it as "AI
| augmentation", but sometimes the AI is doing a lot of the
| time-consuming stuff. The time saved on relatively routine
| scut work is insane, and the quality of the end product (AI
| with my inputs and edits) is really good and consistent.
| PartiallyTyped wrote:
| Anecdata, N=1; I recently used aider -- a tool that gives
| LLMs access to specific files and git integration. The tools
| are great, but the LLMs are underwhelming, and I realized
| that -- once in the flow -- I am significantly faster at
| producing large, correct, and on-point pieces of code,
| whereas when I had to review LLM code, it was frustrating, it
| needed multiple attempts, and it frequently fell into loops.
| light_hue_1 wrote:
| AI companies desperately want to be regulated. OpenAI is
| lobbying hard to be regulated.
|
| Didn't call for it. The whole point is to keep it competition
| with regulation against mythical vague harms that don't exist.
| Aiguru31415666 wrote:
| Hype is if it doesn't deliver or it's overblown.
|
| But I'm amazed on the progress we make every week.
|
| There is real FOMO because if you don't follow it, it just
| might be here suddenly.
|
| Deepseek impressive, deep research also great.
|
| And what you might complete underestimate: we never had a
| system we're it was worth it to teach it everything.
|
| If we need to fine-tune LLMs for every single industry that
| would still be a gigantic shift. Instead of teaching a million
| employees we will teach an LLM all of it once then clone the
| agent a million times
|
| We still see so much progress and there is still plenty of
| money and people available to flow into this space.
|
| There is not a single indication right now that this progress
| is stoping or slowing down.
|
| And not only that, in parallel robots are having their break
| through too.
|
| Your musk point I do not understand really? He is a narcissist
| and he pushed his propaganda platform for becoming president
| because he is in big shit and his house of cards was close to
| crashing
| input_sh wrote:
| Fully agree, in recent weeks I've also started to consider LLMs
| in a wider context, which is to destroy all trust in the web.
|
| The enshittification of search engines, making social media
| verification meaningless, locking down APIs that used to be
| public, destroying public datasets, spreading lies about legacy
| media, the easiness of deploying bots that can sound human in
| short bursts of text... it's all leading towards making it
| _impossible_ to verify anything you read online.
|
| The fearmongering around deepfakes from a few years back is
| coming true, but the scale is _even bigger_. Turns out, there
| won 't be Web 3.0.
| JKCalhoun wrote:
| What trust in the web was there still?
|
| For me it went a decade ago or so when ads and SEO sites in
| Google search became ubiquitous.
| input_sh wrote:
| You could never believe everything you read online, but
| with enough time and effort, you could chase any claim back
| to its original source.
|
| For example, you could read something on Statista.com, you
| could see the credits of that dataset, and visit the source
| to verify. Or you randomly encounter some quote and then
| visit your favourite Snopes-like website to verify that the
| person actually said that.
|
| That's what's under attack. The "middleware" will still be
| there, but the source is going to be out of your reach.
| Hallucinations are not a bug, but a feature.
| JKCalhoun wrote:
| If you can't trace something back to its source, it's
| suspect. It was that way then too. I suppose you're just
| concerned there's a firehose of disinformation now.
|
| So perhaps we have to just slough off the internet
| completely, the way we always have for things like weekly
| rags about "Bat Boy" or whatever.
|
| I hate to see the internet go, but we'll always have
| Paris.
| tim333 wrote:
| >destroy all trust in the web
|
| Genuine question - how so? If I want to find stuff out I go
| to wikipedia, nyt, guardian, hn, linked sites and so on. I'm
| not aware of that lot being noticeably less trustable than in
| the past? If anything I find getting information more
| trustable than before in that there are a lot of long from
| interviews from all sorts of people on youtube where you can
| get their thoughts directly rather than editorialised and
| distorted.
|
| I mean the web was never a place where things were vetted -
| you've always been able to put any sort of rubbish on it and
| so have had to be selective if you want accuracy.
| JKCalhoun wrote:
| I generally take issue when "FOMO" is used. Could go with:
|
| FOBBWIIBM - Fear of being blindsided when it, inevitably,
| becomes mainstream.
|
| Or drop the "fear" altogether:
|
| JOENT - Joy of exploring new territory.
| trimethylpurine wrote:
| > _being blindsided when it, inevitably, becomes mainstream._
|
| I don't see how this could happen. This is not a limited
| resource. It's not a real estate opportunity. There is enough
| AI for everyone to buy when it becomes useful to do so.
|
| I think FOMO correctly identifies the irrational effort of
| many companies to jump in without any idea of what the
| utility might be in any practical sense.
| JKCalhoun wrote:
| I was responding to the _user_ side mentioned.
| einrealist wrote:
| You are right. But these are different different types of
| motivations of the same thing. And there is always context
| for these motivations.
|
| Its a different thing to sell Trump that LLMs should take
| over crucial decisions within a government than just using it
| for some prototyping, code completion at work or to create
| cat pictures at home.
|
| Take Copilot for example. It was rolled out in different
| companies, I worked with. Aside of warnings and maybe
| trainings, I doubt the companies are really able to measure
| the impact that has. Students are already using the
| technology to do homework. Schools and universities are
| sending mixed signals about the results. And then those
| students enter the workforce with Copilot enabled by default.
|
| At least with companies, its the "free market" that will
| regulate (unless some company is too big to fail...)
| tialaramex wrote:
| This course mentions the famous Apple advertisement.
| Unfortunately it slightly oversells and while I'm sure that's not
| because this fragment was written by an LLM it is exactly the
| sort of over-simplification which leads to LLMs generating wild
| bullshit when they interpolate this "fact" with other "facts"
| they've been fed, and we ought to strive to do better when
| writing for humans.
|
| "Describe how prior to 1984, there was no such thing as a
| graphical user interface, visual desktop, an intuitive menu
| system, or mouse-based navigation."
|
| Apple were offering a _mass market product_ which had these
| features so that 's important - but there had been "such a thing"
| for quite some time before that. Douglas Engelbart's "Mother Of
| All Demos" in 1968 -- _Sixteen years earlier_ shows all the
| features you mentioned.
| https://en.wikipedia.org/wiki/The_Mother_of_All_Demos
|
| Unfortunately the demo is very long for a modern audience, so
| unlike "Watch a Superbowl ad" it's a hard sell to show the entire
| demo, but do go watch for yourself.
| rightbyte wrote:
| I like the distinction between "teletypes" and the new fancy
| "glass teletypes".
| leoc wrote:
| Jimmy Carter installed a Xerox Alto in the White House in 1978!
| https://www.ourmidland.com/news/article/Check-out-the-first-...
| Never mind the Xerox Star, or Apple shipping the Lisa in 1983
| ...
| ctbergstrom wrote:
| You're right of course.
|
| In the original drafts I had a long section on this, including
| some of the history of the GUI, the development of the mouse,
| etc. It was way too much for the main text when the point is
| just to set up a metaphor for students who have seen a Mac 128.
|
| That said, we can and should do better in the instructor guide.
| Thanks for the reminder. I'll add some context there.
| Almondsetat wrote:
| I'm sorry, but this website is awful. Not only does it have an
| illogical structure (table of contents at the end? no "next
| lesson" button? gigantic images that fill the entire screen?),
| but the aesthetic of the entire thing is off. It tries to be
| sleek and modern with scrolling animations, but they are janky
| and rigid and the images are rectangles put in front of a bad
| gradient. Not to mention the video interviews are badly produced
| (clipping audio, interviewer doesn't have a dedicated microphone)
| and it's not even clear why they're there.
|
| Please, take inspiration from actual e-learning platforms.
| Kevcmk wrote:
| I would like to read this but the jerkiness of needing to
| scroll 1 page per paragraph renders this unusable
| heikkilevanto wrote:
| I agree. I tried the first chapter with the Reader Mode in
| FireFox, and the whole long scroll hell collapsed to about one
| screenful of text. I have a feeling it skipped some text, but
| the result was a quick read that got the main points through.
|
| I wish the whole thing was available in a plain text format,
| preferably in one longer document.
| allenu wrote:
| Totally agree here. I visited the page and scrolled through it
| to see what it was all about and saw a bunch of pull quotes and
| couldn't work out what I was looking at. It just looks like a
| light magazine article or brochure until you either click on
| the hamburger menu to see a full table of contents or scroll to
| the very bottom to see the table of contents there.
|
| This really needs some design improvements if they want people
| to read through to the actual lessons. Most people are going to
| drop off after scrolling through that landing page.
| sgt101 wrote:
| I wish the title wasn't so aggressively anti-tech though. The
| problem is that I would like to push this course at work, but
| doing so would be suicidal in career terms because I would be
| seen as negative and disruptive.
|
| So the good message here is likely to miss the mark where it may
| be most needed.
| beepbooptheory wrote:
| Really? I am curious how this could be disruptive in any
| meaningful sense. Whose feelings could possibly be hurt? It
| just feels like it would be getting offended from a course on
| libraries because the course talks about how sometimes the book
| is checked out.
| mpbart wrote:
| Any executive who is fully bought in on the AI hype could see
| someone in their org recommending this as working against
| their interest and take action accordingly.
| lm28469 wrote:
| > We were promised hyper-intelligent computer systems that would
| usher in an era of unparalleled prosperity and innovation.
|
| Automated factories were supposed to deliver us from work, yet we
| work as much if not more than before for a lesser part of the
| profit cake.
|
| We can't keep falling for the capitalist trick over and over...
| dailykoder wrote:
| We were also told that smart home will make our life easier and
| save time to have more available for the "important" things in
| life. But reality is that people just spend more time with
| their smart home things and are even shouting at their
| computers even though the computers just do what they were
| told...
| tmnvdb wrote:
| Hmm, it seems that the author takes very clear (and sometimes
| cynical) positions on some controversial questions. For example,
| "They don't have the capacity to think through problems
| logically." is an hotly debated claim, and I think with the
| advent of reasoning models this has at least become something one
| should not state in entry level material, which would hopefully
| reflect common understanding rather than the authors personal
| opinion in an ongoing discussion.
|
| There are more claims like this about what language models can't
| do "because they just predict the next token". This line of
| reasoning, while superficially plausible, holds a lot of
| assumptions that have been questioned. The heavy lifting here is
| done by the word "just" - if you can correctly predict the next
| token in every situation (including novel challenges), does that
| not require an excellent world model - somehow explicitly
| reflected in the weights? This is not a settled question but the
| last few years of LLM success have been completely on the side of
| those who think that token prediction is quite general.
|
| The material also makes several comparisons to human
| intelligence, and while it is obvious that humans are different
| from language models we do not really understand the emergence of
| all the things that are claimed to be "impossible" for the
| machine to have in humans (consciousness, morality, etc), it just
| so happens we are all human so we all agree we have it.
| Furthermore, it is not clear to me that something can only be
| called 'intelligent' if it perfectly mimics humans in every way.
| This is maybe just human bias to our own experience and risks a
| "submarines can't swim" debate which is really about language.
|
| Many of these philosophical objections have been questioned by
| people in the field and more importantly by the rapid progress of
| the models in tasks they were supposed to be incapable of
| performing according to philosophical objectors. The last few
| years, every time somebody claims models "can't do X" a new model
| is released and lo and behold, X is now easy and solved. (If you
| read a 6 month old paper of impossible benchmarks, expect 75% to
| be already solved). In fact, benchmark satuation is a problem
| now. In other words, the goalposts are having trouble keeping up,
| despite moving at high speed.
|
| I don't think you are doing the general public any service by
| simply claiming that it is a lot of hype and marketing, these
| models are really advancing rapidly and nobody really knows where
| it will end. The philosophical objections seem to be rather weak
| and are in rapid retreat with every new model, on the other hand
| the argument in favor of further progress is just "we had
| progress so far by scaling, if we keep scaling surely we will
| have more progress" (induction). This is not a strong guarantee
| of further progress.
|
| The claim that the labs are 'marketing geniuses" for realising
| language models as chat instead of autocomplete (which they
| "really" are according to the text - what does that mean?) also
| seems a bit silly given the obvious utility of the models is
| already much higher than 'autocomplete'. This seems to be another
| instance of the common bias that a model that "just" predicts the
| next token is not allowed to be as succesful as it clearly is in
| all kinds of tasks.
|
| I don't think a lot of these opinions are particularly well
| founded and they probably should not be presented in entry level
| material as if they are facts.
|
| Edit: just to add a positive note, I do think it is extremely
| useful to educate people on the reliability problem, which is
| surely going to lead to lots of problems in the wrong hands.
| nonrandomstring wrote:
| > cynical () positions on some controversial questions.
|
| I feel "cynical" is an inappropriate word here.
|
| We may have to, for the same (ecumenical) reasons that thinkers
| like Churchland, Hofstadter, Dennet, Penrose and company have
| all struggled with, eventually accept the impossibility of
| proof of (existence or non-existence) on _any_ hypothesis of
| "machine mind". The pragmatic response is, "does it offer
| utility for me?". And that's all that can be said. Anyone's
| choice to accept or reject ideas of machine intelligence will
| remain inviolably personal and beyond appeal to "proof" or
| argument. I think that's something we'd better get used to
| sooner rather than later, in order to avoid a whole lot or
| partisan misery and wasted breath.
| tmnvdb wrote:
| I think the way he sketches the the AI labs as "marketing
| geniuses" for not just releasing their models as auto-correct
| is a bit cynical, as well as implying in general that these
| labs are muddying the waters on purpose by not agreeing with
| <authors position> and by engaging in "hype" (believing in
| the technology).
| nonrandomstring wrote:
| Sorry, "inappropriate" might have been inappropriate :)
| What am I trying to say here?....that we're soon gonna find
| ourselves in an insoluble and exhausting debate around
| machine thinking and its value.
| tmnvdb wrote:
| Death, taxes, and insoluble and exhausting debates around
| machine thinking and its value.
| uh_uh wrote:
| The choice unfortunately seems to correlate with the person's
| age. Younger generations will have no trouble treating LLMs
| as actually intelligent. Yet another example of "Science
| progresses one funeral at a time."
| nonrandomstring wrote:
| > correlates with age
|
| Definitely a "citation needed" moment I think. Friday, I
| was with a lot of 12 year olds all firmly of the opinion
| that it's a "way to get intelligence/information" but it's
| not _actually_ intelligent. (FWIW in UK they say "for real
| life intelligent") I noted this distinction. Or rather, I
| noted _because that 's what they're taught_. So teachers,
| naturally pass on the commonsense position that "it's still
| just a computer". That means waiting for funerals will not
| settle the matter either. That's not to say a significant
| sect of more credulous "AI worshippers" will not emerge.
| tmnvdb wrote:
| There was a paper a while back on AI usage at work among
| engineers and it was very strongly correlated to age.
| This is not surprising, technology adoption is always
| very dependent on age. (None of this tells you if the
| technology is a net good)
| TheOtherHobbes wrote:
| Many claims don't stand up to scrutiny, and some look
| suspiciously like training to the test.
|
| The Apple study was clear about this. LLMs and their related
| modal models lack the ability to abstract information from
| noisy text inputs.
|
| This is really obvious if you play with any of the art
| generators. For example - the understanding of basic
| prepositions just isn't there. You can't say "Put this thing
| behind/over/in front of this other thing" and get the result
| you want with any consistency.
|
| If you create a composition you like and ask for it in a
| different colour, you get a different image.
|
| There is no abstracted concept of a "colour" in there. There's
| just a lot of imagery tagged with each colour name, and if you
| select a different colour you get a vector in a space pointing
| to different images.
|
| Text has exactly the same problem, but it's less obvious
| because what the grammar is usually - not always - perfect and
| the output has been tuned to sound authoritative.
|
| There is not enough information _in text as a medium_ to handle
| more than a small subset of problems with any consistency.
| uh_uh wrote:
| > There is no abstracted concept of a "colour" in there.
| There's just a lot of imagery tagged with each colour name,
| and if you select a different colour you get a vector in a
| space pointing to different images.
|
| It has been observed in LLMs that the distance between
| embeddings for colors follows the same similarity patterns
| that humans experience - colors that appear similar to
| humans, like red and orange, are closer together in the
| embedding space than colors that appear very different, like
| red and blue.
|
| While some argue these models 'just extract statistics,' if
| the end result matches how we use concepts, what's the
| difference?
| rcxdude wrote:
| Part of this is that the art generators tend to use CLIP,
| whjch is not a particularly good text model, often only being
| slightly better than a bag of words, which makes many
| interactions and relationships pretty difficult to represent.
| Some of the newer ones have better frontends which improve
| this situation, though.
|
| I think color is fairly well abstracted, but most image
| generators are not good for edits, because the generator more
| or less starts from scratch, and from a new random seed each
| time (and even if the seed is fixed, the initial stages of
| the generation, where things like the rough image composition
| form, tend to be quite chaotic and so sensitive to small
| changes in prompt). There are tools that can make far more
| controlled adjustments of an image, but they tend to be a bit
| less user-friendly.
| tmnvdb wrote:
| Regarding the apple paper:
| https://andrewmayne.com/2024/10/18/can-you-dramatically-
| impr...
| aucisson_masque wrote:
| Wow, that's really interesting! I didn't have the time to read
| all the pages, but I definitely will. It helps to bring one's
| expectations about AI back down to earth.
| qwGafv wrote:
| That's an interesting Altman quote on the site. LLMs cannot be
| compared to electricity and the Internet. People wanted those.
| LLMs were an impressive parlor trick at first but disappointing
| later. Many stopped using them altogether.
|
| Now there is a president who fuels the hype, shakes down rich
| countries for "AI" investments. The Saudi prince who lost money
| on Twitter is in for the new grift and praises Musk on Tucker
| Carlson.
|
| The grift-oriented economy might continue with the bailout of
| Bitcoin whales through the "sovereign wealth fund" scheme.
|
| That is how the "economy" works. No houses will be built and
| nothing of value will be created.
| tmnvdb wrote:
| People have stopped using LLMs? I wasn't aware of that. Can you
| share a source for that?
| TheOtherHobbes wrote:
| I know a lot of people who went through the "Oh, wow - wait a
| minute..." cycle. Including me.
|
| They're approximately useful in some contexts. But those
| contexts are limited. And if there are facts or code
| involved, both require manual confirmation.
|
| They're ideal for bullshit jobs - low-stakes corporate
| makework, such as mediocre ad copy and generic reports that
| no one is ever going to read.
| tmnvdb wrote:
| > And if there are facts or code involved, both require
| manual confirmation.
|
| The hidden assumption here seems to be that the model needs
| to be perfect before it has utility.
| TeMPOraL wrote:
| Also hidden assumption, or perhaps lack of clear
| perception of reality, that most jobs on the market are
| strongly dependent on factual correctness.
|
| Also assumption that this is any different than human
| relationship with empirical truth is.
| asadotzler wrote:
| Now you're the bullshit machine. No one said that. We
| expect basic reliability/reproducibility. A $4 drugstore
| calculator has that to about a dozen 9s, every single
| time. These machines will give you a correct answer and
| walk it right back if you respond the "wrong" way.
| They're not just wrong a lot of the time, they simply
| have no idea even when they're right. Your strawman is of
| no value here.
| dailykoder wrote:
| God says...
|
| bradytrophic indical sacella unlittered surveying mis-tilled
| brachymetropic oxime masa hypermilitant litas acarotoxic thrust O
| retranslating awee tetraodont prevoidance cabbala Linker
| pervertedness absurdest coruscates aminopyrine spitting ledgers
| Fremontia upthrown simplifiers acomous statutable Ampelopsis
| FergusArgyll wrote:
| The sys prompt given to run the turing test (from
| https://arxiv.org/pdf/2405.08007) actually works well. I'm
| honestly not sure I'd be able to tell (unless I test it
| adversarially e.g ignore the prompt, write a poem etc)
| hirenj wrote:
| This is a great resource, thanks. We (myself, a bioinformatician,
| and my co-cordinators, clinicians) are currently designing a
| course to hopefully arm medical students with the required basic
| knowledge they need to navigate the changing world of medicine in
| light of the ML and LLM advances. Our goal is to not only
| demystify medical ML, but also give them a sense of the
| possibilities with these technologies, and maybe illustrate
| pathways for adoption, in the safest way possible.
|
| Already in the process of putting this course together, it is
| scary how much stuff is being tried out right now, and is being
| treated like a magic box with correct answers.
| TeMPOraL wrote:
| "The LLMs have no ground truth" claim (around chapter 2) that's
| core to the "bullshit machines" argument is itself wrong. Of
| course LLMs have ground truth. What do the authors think here,
| that the text in training corpus is _random_?
|
| Hint: it isn't. Real conversations are anything but random.
| There's a lot of information hidden in "statistical ordering of
| the words", because the distribution is not _arbitrary_.
|
| Statistical ground truth isn't any worse than explicitly given
| one. In fact, fundamentally, there _only ever is statistical
| certainty_. Realizing it is a pretty profound insight, and you 'd
| think it should be table stakes at least in STEM, but sonehow it
| fails to spread as much as it should.
| ImPleadThe5th wrote:
| So if I ask ChatGpt about bears and in the middle of explaining
| their diet it tells me something about how much they like
| porridge and in the middle of habitat it tells me they live in
| a quaint cabin in the woods, that's ... True?
|
| Statistically we certainly have a lot of words about 3 bears
| and their love for porridge. That doesn't mean it's true, it
| just means it's statistically significant. If I asked someone a
| scientific question about bears and they told me about
| Goldilocks, id think it was bullshit.
| tmnvdb wrote:
| Having a ground truth doesn't mean it does not make huge and
| glaring mistakes.
| vaidhy wrote:
| If that were the case, then you are right.. However, the
| current crop of LLMs seem to be good at understanding the
| context.
|
| A scientific data point about bears is unlikely to have
| Goldilocks in there (unless talking about evolution of life
| and Goldilocks zone). You can argue that there is meaning
| hidden in words that is not captured by words themselves in a
| given context - psychic knowledge as opposed to reasoned out
| knowledge. That is a philosophical debate.
| TeMPOraL wrote:
| Words don't carry meaning. Meaning exists in how words are
| or are not used together with other words. That is, in
| their.... statistical relationships to each other!
| TeMPOraL wrote:
| ChatGPT has enough dimensions in its latent space to
| represent and distinguish between the various meanings of
| porridge and is able to be informed by the Goldilocks story
| without switching to it mid-sentence.
|
| It's actually a good example of what I have in mind by saying
| human text isn't random. The Goldilocks story may not be
| scientific, but it's still highly correlated with scientific
| truth about matters like food, bears, or the daily lives of
| people. Put yourself in the shoes of an alien trying to make
| heads or tails of that story, you'll see just how many things
| in it are not arbitrary.
| empath75 wrote:
| LLM's demonstrably don't do this, nor do they say that they
| live in the hundred acre woods and love honey. Unless you ask
| about a specific bear.
| sesm wrote:
| > Statistical ground truth isn't any worse than explicitly
| given one
|
| There are multiple kinds of truths.
|
| 'Statistical truth' is at best 'consensus truth', and that's
| only when LLM doesn't hallucinate.
| TeMPOraL wrote:
| That's the only one that's available, though.
|
| When a kid at school is being taught, say, Newton's laws of
| motion, or what happened in 476 CE, they're _not_
| experiencing the empirical truth about either. They 're only
| learning the consensus truth, i.e. the correct answer to give
| to the teacher, so they get good grade instead of bad grade,
| and so their parents praise them instead of punishing them,
| etc.
|
| This covers pretty much everything any human ever learns. Few
| are in position to learn any particular things
| experimentally. Few are in position to verify most of what
| they've learned experimentally afterwards.
|
| We live in a "consensus reality", but that works out fine,
| because _establishing consensus is hard_ , and one of the
| strongest consensus-forcing forces that exist is "close
| enough to actual reality".
| sesm wrote:
| I've heard about at least 4 theories of truth:
| Correspondence, Coherence, Consensus and Pragmatic (as
| described, for example, here https://commoncog.com/four-
| theories-of-truth/).
|
| If we look at Newtonian mechanics, then various
| independently verifiable experiments are examples of
| Correspondence truth, and the minimal mathematical
| framework that describes them is an example of Coherence
| truth.
| TeMPOraL wrote:
| Fine, but it's not how any of us learned of it either -
| whether the Newtonian mechanics or the "4 theories of
| truth".
|
| I mean, coherence is sure a an important aspect of truth,
| and just by paying attention whether it all "adds up" you
| can easily filter 90% of the bullshit you hear people (or
| LLMs for that matter) saying - but even there, I'm not a
| physicist, I don't do much experiments in a lab, so when
| I evaluate if some information is coherent with Newton's
| laws of motion, I'm _actually_ evaluating some
| _description_ of a situation against a _description_ of
| Newton 's laws. It's all done in "consensus space" and,
| if an answer is expected, the answer is also a "consensus
| space" one.
|
| We're all so used to evaluating inputs and outputs
| through the lens of "is this something I expect others
| believe, and others expect me to believe", that we're
| almost always just mentally folding the indirection
| through "consensus reality" and feel like we're just
| checking "is this true". It works out okay, and it can't
| really be any other way - but we need to remember this is
| what we're doing.
| fancyfredbot wrote:
| Statistical truths based on observation of reality is the basis
| for science. Statistical truth based on the text on the
| internet is the basis for something else, and I would
| personally not like to call whatever that is science, or any
| truths established this way "ground truths".
| TeMPOraL wrote:
| Statistical truths based on observation of reality is the
| basis for incremental additions science. Statistical truths
| based on _what you read in the textbook_ is what _actually_
| is almost all science, to everyone, at almost all times.
|
| There's few things in our lives any of us actually learned
| first-hand, empirically. Everything else we learned the same
| way LLMs did - as what is _expected to be_ the right
| completion to a question or inquiry.
|
| The objective truth as we experience is things that, _were we
| to_ make a prediction conditioned on them, our prediction
| would turn out correct. That doesn 't mean we actually make
| such predictions often, or that we ever made such predictions
| learning it.
| empath75 wrote:
| What year did the Normans invade England?
|
| What's Newton's second law?
|
| Who was the last czar of Russia?
|
| How many moons does Jupiter have?
|
| I bet you "know" a lot of those facts not because you have
| observed them empirically, but because you read about them in
| books. And in fact, almost all scientists rely on reading for
| nearly everything they know about science, including nearly
| everything they know about their own specialties. Nobody has
| time to derive all of their knowledge of the world from
| personal observation, and most people who can read have
| probably learned almost all "facts" they know about the world
| from books.
| TeMPOraL wrote:
| Yup, and importantly, the correct answer to those is, in
| anyone's _personal_ learning experience, almost always just
| what the authority figures in their lives (parents,
| teachers, peers they respect, or textbooks themselves)
| would accept as the correct answer!
|
| "Consensus reality" works so well for us most of the time,
| that we habitually "substitite the Moon for the finger
| pointing at the Moon" without realizing it.
| fancyfredbot wrote:
| There's obviously nothing wrong with learning by reading,
| but the way you tell whether what you read is true is by
| seeing whether or not it fits in with observation of
| reality. That's the reason we're no longer reading the
| books about phlogiston.
| TeMPOraL wrote:
| > _the way you tell whether what you read is true is by
| seeing whether or not it fits in with observation of
| reality_
|
| The only way any of us ever gets to see "whether or not
| it fits in with observation of reality" is to see if they
| get an A or F on the test asking it.
|
| Seriously.
|
| The "moons of Jupiter" question is the only one of the
| above one gets to connect to an observation independent
| of humans, and then if they did, _they 'd be wrong_,
| because you can't just count _all_ the moons of Jupiter
| from your backyard with a DIY telescope. We know the
| correct answer only because some other people both built
| building-sized telescopes and had a bunch of car-sized
| telescopes _thrown at Jupiter_ - and unless you had a
| chance to operate either, then for you the "correct
| answer" is what you _read somewhere_ and that you _expect
| other people to consider correct_ - this is the _only_
| criterion you have available.
| fancyfredbot wrote:
| Independently checking the information you read in
| textbooks is very difficult for sure. But it's still how
| we decide what's true and what's not true. If tomorrow a
| new moon was somehow orbiting Jupiter we'd say the
| textbooks were wrong, we wouldn't say the moon isn't
| there.
| abecedarius wrote:
| > personally not like to call whatever that is science
|
| A counterargument written pre-LLMs:
| https://arxiv.org/abs/1104.5466
| krupan wrote:
| It's true that LLMs aren't trained on strings of random words,
| so in a sense you correct are that they have some "ground
| truth." They wouldn't generate anything logical at all if not.
| Does that even need to be stated though? You don't need AI to
| generate random words.
|
| The more important point is, they aren't trained on only
| factual (or statistically certain) statements. That's the
| ground truth that's missing. It's easy to feed an LLM a bunch
| of text scraped from the internet. It's much harder to teach it
| how to separate fact from fiction. Even the best human minds
| that live or ever have lived have not been able to do that
| flawlessly. We've created machines that have a larger amount of
| memory than any human, much quicker recall, the ability to
| converse with vast numbers of people at once, but it performs
| at about par with humans in discerning fact from fiction.
|
| That's my biggest concern about creating super powered
| artificial intelligence. It's super powers are only super in a
| couple dimensions and people mistake that for general
| intelligence. I came across someone online that really believed
| chatGPT was creating a custom diet plan tailored to their
| specific health needs, base on a few prompts. That is scary!
| HarHarVeryFunny wrote:
| Well, duh ... of course there are statistical regularities in
| the data, which is what LLMs learn. However, the LLM has no
| first hand experience of the world described by those words, so
| the words (to an LLM) are not grounded.
|
| So, the LLM is doomed to only be able to predict patterns in
| the dataset, with no knowledge or care as to whether what it is
| outputting is objectively true or not. Calling it bullshitting
| would be an anthromorphism since it implies intention, but it's
| effectively what an LLM does - just spits out words without any
| care (or knowledge) as to the truth of them.
|
| Of course if there is more truth than misinformation and lies
| in the dataset, then statistically that is more likely to be
| output by an LLM, but to the LLM it's all the same - just word
| stats, just the same as when it "hallucinates" and "makes
| something up", since there is always a non-zero possibility
| that the next word could be ... anything.
| 1970-01-01 wrote:
| >Of course LLMs have ground truth.
|
| It is my understanding that LLMs have no such thing, as empiric
| truth is weighted. For example, if Newton's laws are in
| conflict with another fact, the LLM will defer to the fact that
| it finds more probable in context. It will then require human
| resources to undo and unfold it's core error, else you receive
| bewildering and untrue remarks or outputs.
| jononor wrote:
| Where "probable" means: occurs the most often in the training
| data (approx the entire Internet). So what is _common_ online
| is most likely to win out, not some other notion of
| correctness.
| TeMPOraL wrote:
| > _For example, if Newton 's laws are in conflict with
| another fact, the LLM will defer to the fact that it finds
| more probable in context._
|
| Which is the correct thing to do. If such a context would be,
| for example, an explanation of an FTL drive in a science
| fiction story, both LLMs and humans would be correct to put
| Newton aside.
|
| LLMs aren't Markov chains, they don't output naive word
| frequency based predictions. They build a high-dimensional
| statistical representation of entirety of their training
| data, from which then completions are sampled. We already
| know this representation is able to identify and encode ideas
| as diverse as "odd/even" and "fun" and "formal language" and
| "golden gate bridge"-like. "Fictional context" vs. "Real-life
| physics" is a concept they can distinguish too, just like
| people can.
| nickpsecurity wrote:
| Ground truth is possibly used in the sense that humans' brains
| tie what they create to the properties of observed reality.
| Whatever new information comes in is compared to, or checked
| by, that. Whereas, LLM's will believe anything you feed them in
| training no matter how unrealistic it is.
|
| I do think that, after much training data, they do have
| specific beliefs that are ingrained in them. That changing
| _those_ is difficult. We've seen that on some political and
| scientific claims that must have been prominent in their pre-
| training data or RLHF tuning. They will argue with us over
| those points, like it's a fight. Otherwise, I've seen continued
| pre-training or fine-tuning can change everything up to their
| vocabularies.
| TeMPOraL wrote:
| > _Ground truth is possibly used in the sense that humans'
| brains tie what they create to the properties of observed
| reality. Whatever new information comes in is compared to, or
| checked by, that._
|
| What do you compare your knowledge of history to, other than
| _what you expect other people will say_? Most knowledge we
| learn in life is tied _only_ to our expectations of other
| people 's reactions. Which works out fine, most of the time,
| because even with questions of scientific fact, it's enough
| _some_ people are in position to ground information in
| empirical experiments, and then set everyone else 's
| _expectations_ based on that.
| nickpsecurity wrote:
| To start with, we know what a person is, rudimentary things
| about how they behave, our senses, how they commonly work,
| and can do mental comparisons (reality checks).
|
| We know LLM's don't start with that because we initialize
| them with zero'd or random weights.
|
| Then, their training data can be far more made up, even
| works of fiction, that the reality most humans observe with
| is almost always real. We could raise a human in VR or
| something where technically there would be comparisons.
| Most humans' observations connect to expectations in their
| brain which was designed for the reality we operate in.
|
| Finally, the brain has multiple components that each handle
| different jobs. They have different architectures to do
| those jobs well. Sections include language, numbers,
| reasoning, tiers of memory, hallucination prevention,
| mirroring, and even meta-level stuff like reward
| adjustment. We don't just speculate that they do different
| things: damage to those regions shuts down those abilities.
| Tied to the physical realm are vision, motor, and spatial
| areas. We can feel objects, even temperature or pressure
| changes. That we can do a lot of that without supervised
| learning shows we're tailor-made by God for success in this
| world.
|
| LLM's have one architecture that does one job which we try
| to get to do other things, like reasoning or ground truth.
| We pretend it's something it's not. The multimodal LLM's
| are getting closer with specialized components. Even they
| aren't all trained in a coherent way using real-world,
| observations in the senses. There's usually a gap between
| systems like these and what the brain does in the real
| world just in how it gets reliable information about its
| operating environment.
| TeMPOraL wrote:
| > _To start with, we know what a person is, rudimentary
| things about how they behave, our senses, how they
| commonly work, and can do mental comparisons (reality
| checks)._
|
| How much is this a matter of fidelity? LLMs started with
| text, now text + vision + sound; it's still not the full
| package relative to what humans sport, but it captures a
| good chunk of information.
|
| Now, I'm not claiming equivalence in the training process
| here, but let's remember that we all spend the first year
| or two of our lives _just_ figuring out the _intuitive
| basics_ of "what a person is, rudimentary things about
| how they behave, our senses, how they commonly work", and
| from there, we spend the next couple years learning more
| explicit and complex aspects of the same. We don't start
| with any of it hardcoded (and what little we have, it's
| been bestowed to us by millennia of a much slower
| gradient-descent process - evolution).
|
| > _LLM's have one architecture that does one job which we
| try to get to do other things, like reasoning or ground
| truth._
|
| FWIW, LLMs have one architecture in a similar sense brain
| has one architecture - brains specialize _as they grow_.
| We know that parts of a brain are happy to pick up the
| slack for differently specialized parts that became
| damaged or unavailable.
|
| LLMs aren't uniform blobs, either. Now, their
| architecture is still limited - for one, unlike our
| brains, they don't _learn on-line_ - they get pre-trained
| and remain fixed for inference. How much a model capable
| of on-line learning will differ structurally from current
| LLMs, or even the naive approach to bestow learning
| ability on LLMs (i.e. do a little evaluation and training
| after every conversation)? We don 't know yet.
|
| I'm definitely not arguing LLMs of today are structurally
| or functionally equivalent to humans. But I _am_ arguing
| that learning from sum total of the Internet isn 't
| meaningfully different from how humans learn, at least
| for anything that we'd consider part of living in a
| technological society. I.e. LLMs don't get to experience
| throwing rocks first-hand like we do, but _neither_ of us
| get to experience _special relativity_.
|
| > _Even they aren't all trained in a coherent way using
| real-world, observations in the senses._
|
| Neither them nor us. I think if there's one insight
| people should've gotten from the past couple years is
| that "mostly coherent" data is fine (particularly if any
| given subset is internally coherent, even if there's
| little coherence between different subsets) - both humans
| and LLMs can find larger coherence if you give them
| _enough_ such data.
| ohthehugemanate wrote:
| I just want to say: I've been publicly calling them "bullshit
| machines" since the first big media wave. I am incredibly pleased
| that this mental model helps other people, too. And that the
| specific term sees broader use is also nice.
|
| Also, neener neener neener I called them bullshit machines BEFORE
| it was cool. /humor
|
| Seriously though the humanities have a lot to chew on with LLMs,
| and are incredibly important to how we live and work with them.
| Who knew that epistemology would become front page news, and the
| sexiest topic for VC?
| abzolv wrote:
| Your scroll-to-death user interface made me close the window
| before the end of the second page.
|
| Did you ask an LLM to recommend the most user-friendly UI to you?
| ctbergstrom wrote:
| We asked our target audience, 19-year-olds. They had a strong
| preference for this style. I know....
| dzogchen wrote:
| This will not age well.
|
| Yesterday my bullshit machine wrote a linker argument parser to
| hook a C++ library up in a Rust build config. Oh it also wrote
| tests for it.
| https://chatgpt.com/share/67a89e5f-b5b4-8011-9782-472d469cc2...
| tmnvdb wrote:
| B..b..but.. next token...parrot...hype..ngggggg
| bayindirh wrote:
| Allowing a parrot to iterate on given examples and generate a
| similar one with the information baked in their weights does
| not invalidate "Stochastic Parrot" take. On the contrary, it
| proves it.
|
| LLMs are statistical machines. The catch is you feed it
| hundreds of terabytes of valid information, so it
| asymptotically generates valid information as a result of
| this statistical bias.
|
| Even yet, they can hallucinate so badly. I mean, the same
| OpenAI model claimed that I'm a footballer, a goal keeper in
| fact.
|
| Stochastic parrot, yes. On LSD, very yes.
| tmnvdb wrote:
| It's clearly true that the LLMS are 'stochastic parrots',
| but for all we know that might be the key to intelligence.
| It is in itself not a deep observation any more than
| calling your fellow humans 'microbial meatbags'.
|
| Saying that LLMs are stochastic machines does not establish
| an upper bound for success.
| fatata123 wrote:
| We are not stochastic parrots. Old components of our
| brain help "ground" our thoughts and allow things like
| doubt or a gut feeling to develop which means we can
| question ourselves in ways an LLM cannot.
| bayindirh wrote:
| The thing is, this assumption of LLMs might be
| intelligent lies in the assumption is intelligence is
| enabled solely by the brain.
|
| However, as the science improves, we understand more and
| more that brain is just part of a much bigger network,
| and its size or surface roughness might not be the only
| thing determines the level of intelligence.
|
| Also, all living things have processes which allows
| constant input from their surroundings and they also have
| closed feedback loops which constantly change and tweak
| things. Call these hormones, emotions or self-reflection
| or whatnot.
|
| We the scientists love to play god with the information
| we have at hand, yet we constantly humbled by the nature
| by experiencing the shallowness of what we know. Because
| of that I, as a CS Ph.D., am not so keen on to jump to
| that bandwagon which claims that we invented silicon
| brains.
|
| They are arguably useful automatons built on dubious data
| obtained in ethical gray areas. We're just starting to
| see what we did, and we have a long way to go.
|
| So, a living parrot might be more intelligent than these
| stochastic parrots. I'll stay on the cautious critics
| wagon for now.
| tmnvdb wrote:
| > So, a living parrot might be more intelligent than
| these stochastic parrots.
|
| Or the other way around.
| bayindirh wrote:
| I don't think you ever observed a Cuckatoo...
| consp wrote:
| You asked it to do a task with probably many examples. This
| course will probably tell you that will work fine. Don't see
| your point here.
| rawtr wrote:
| Yesterday my 30 year old photocopier wrote a Shakespeare drama!
| All I had to was scan the original pages.
| d4rkp4ttern wrote:
| The format is very interesting. Can you speak to the tech stack
| behind how you made it ?
| magicalhippo wrote:
| I'll agree on interesting, however I found it very difficult to
| follow.
|
| Reading the landing page, I expected a link to get started.
| Took me some time to really register the links at the very
| bottom.
|
| After finding my way into the lectures I got distracted by all
| the scrolling. Fortunately Reader Mode fixed that. However a
| few lectures in and I notice there's several videos I've
| missed...
|
| However since I'm starting to approach the get-off-my-lawn age,
| I guess I'm not the target audience.
| ctbergstrom wrote:
| After talking to an awful lot of 18-20 year olds (our target
| audience) we decided we wanted to go with a "scrollytellying"
| style. I'm not a designer and I've done worked in that style
| before. After looking into a range of platforms -- Vev and
| Closeread for Quarto deserve particular mention -- I felt that
| Shorthand (https://shorthand.com/) was the best option for
| rapid development given my lack of experience in this whole
| process.
|
| In general I've been very pleased. You don't have the fine
| scale control you do on a platform like Vev, but for someone
| like me that is probably a good thing because it keeps me from
| mucking around quite as much as I otherwise would with design
| decisions that I don't really understand.
|
| The price is a bit steep for a self-funded operation and we're
| constrained a bit by the need to use their starter tier, but I
| feel like we are definitely getting our money's worth and
| customer support from Shorthand has been exemplary.
| gsky wrote:
| People said the world wouldn't need more than 10 computers.
| Newyork times ridiculed space industry and etc..
|
| AI is gonna disrupt all industries
| _heimdall wrote:
| Any new technology given this much attention and money will
| disrupt, that's not really a question.
|
| The question is whether we're going to be better off for it,
| and if people want all that change in the first place.
| JKCalhoun wrote:
| It could be argued that the adoption of the automobile was a
| bad move.
|
| You and I debating either its efficacy or social good though
| is irrelevant if it marches on regardless.
| _heimdall wrote:
| I didn't explain my point there, let me try with the
| automobile example.
|
| There's a key difference in the how the impact of the
| automobile happened - consumers got to choose to buy them
| and the impact was driven in large part by market demand.
|
| The fact that LLMs are going to have a big impact seems
| obvious because a comparatively few number of people are
| making a huge deal out of them, both with attention and
| money. LLMs will be big but that says nothing of their
| usefulness or even consumer demand, it says more about how
| the industry is being financed.
| JKCalhoun wrote:
| I'm not too worried about Big Corporation trying to push
| a rope. In the end it really is only going to succeed if
| "we" want it -- find value in it.
| _heimdall wrote:
| > only going to succeed if "we" want it -- find value in
| it.
|
| I'll be pleasantly surprised if that's how it turns out.
|
| At least so far market dynamics haven't really been much
| of a driver for LLMs. Those with the money think its the
| next big thing and are pouring cash both into the LLMs
| themselves and any product that slaps a "powered by AI"
| sticker on the box.
|
| That's not to say people aren't also actively choosing to
| use LLMs, but in my opinion the market demand doesn't
| account for the massive amount of hype and funding, or
| the pervasiveness of LLMs being added to so many
| products.
| JKCalhoun wrote:
| I see it too -- but I'll remind both of us that these are
| still very, very early days.
| _heimdall wrote:
| For sure. And I definitely have a bias showing here
| towards _not_ trusting the person in charge to be
| benevolent or to actually know what the "right" thing to
| do is in the long run.
| zwnow wrote:
| In a unpredictable, hallucinating way. Who cares about human
| expertise when you can have a machine that acts like it has
| expertise in the same area? We deserve what is coming.
| jaimebuelta wrote:
| It's totally possible that the net result is good (which is
| still quite early to really know), but it present new problems.
| For example, the creation of the car is probably good in
| general, but it has important issues that requires to be taken
| into account (accidents, pollution, city planning challenges,
| etc)
|
| I've lived enough hype cycles to know that we are always very
| close to use VR every day or have autonomous cars... It took
| around 25 years to move from "we will pay with our mobile
| phones" to that becoming a reality.
| add-sub-mul-div wrote:
| People also didn't anticipate how social media and surveillance
| capitalism would shit up our lives.
| fancyfredbot wrote:
| I have just read one section of this, "The AI scientist'. It was
| fantastic. They don't fall into the trap of unfalsifiable
| arguments about parrots. Instead they have pointed out positive
| uses of AI in science, examples which are obviously harmful, and
| examples which are simply a waste of time. Refreshingly objective
| and more than I expected from what I saw as an inflammatory
| title.
| jcgrillo wrote:
| The cost of inference seems like a major barrier in making these
| things work commercially. If you put one in the user facing flow
| you'll need an extraordinary amount of compute that scales
| extremely poorly in number of users, given that each query takes
| O(nm^2) where m is the number of model weights (many) and n is
| the number of tokens in the query. So it seems clear if each user
| session implies multiple LLM queries, those user sessions had
| better be exceptionally valuable. This seems like a completely
| different business with very different constraints and margins.
|
| The problem is it doesn't seem like the value added is all that
| great. So how do you justify the ruinous cost? I would like to
| know more about how people actually intend on using these things
| profitably than to hear about how they're going to wake up and
| become "intelligent". Does anyone have any success stories
| actually using this thing? It's hard to find any discussion of
| this amidst all the bullshit hype, and it's really the only
| question--can you make money with this or not? And I don't mean
| raising billions in venture capital, I mean can you integrate
| these things profitably and sustainably into a website?
| tmnvdb wrote:
| The costs are a problem. We don't have hard evidence that this
| will be solved, but with algorithmic efficiency and raw compute
| costs both changing rapidly, the cost per token has gone down
| by about a factor of 10 per year for the last 3 years, i.e.
| 1000x over 3 years.
| jcgrillo wrote:
| As far as I can tell, those charts are merely describing the
| price per token that LLM hosting companies are charging, not
| what running the model actually costs. The distinction is
| important for two reasons:
|
| 1. These companies are heavily subsidized by huge amounts of
| venture investment
|
| 2. If I'm integrating this technology into my web product
| there's absolutely no way I'll be adding a 3rd party company
| as a dependency. This is all way too new and bubbly to trust
| any of the current offerings will still exist in O(years).
|
| Are there any similar studies showing not sticker price but
| actual compute/performance decrease?
| tmnvdb wrote:
| This is incorrect, there are real effiency gains.
|
| Slightly old by the standards of this field, but a good
| overview: https://arxiv.org/abs/2403.05812
| jcgrillo wrote:
| I'm sorry, what is incorrect? And the paper you linked
| appears to be about training, not inference? I wouldn't
| train a model in a user's session, instead I'd _run_ the
| model. That 's the cost that seems a blocker, all the
| models are already trained, I dont need to invest a dime
| in that.
| tmnvdb wrote:
| The real costs are dropping because of real efficiency
| gains and compute cost reductions.
|
| The inference costs are dropping at a similar rate.
| jcgrillo wrote:
| [citation needed]
| tmnvdb wrote:
| https://a16z.com/llmflation-llm-inference-cost/
| jcgrillo wrote:
| This is getting awfully tedious. Are you trolling or do
| you genuinely think responding to a question about actual
| compute cost trends with a venture fund's puff piece
| about sticker price trends is helpful?
| tmnvdb wrote:
| You are correct to note 'real' costs of the leading labs
| are not public. It is surely true that the labs are
| operating at below cost (we are definitely not paying for
| the full R&D), but it seems unlikely that this fully
| explains the reduction in inference costs over the last
| years. We also know from open models like deepseek that
| the cost per inference token at a fixed performance level
| is going down very quickly matching the curve of leading
| labs inference cost decreases. You can even test it
| yourself on your own pc if you want.
|
| I would add that inference cost decreases is what we
| should expect, it stands to reason there will be
| algorithmic improvements in inference, and compute cost
| is still going down just because of (a somewhat sloped)
| Moore's law.
|
| Maybe you could also be a bit friendlier and forthcoming
| in your responses.
| jcgrillo wrote:
| I apologize for my tone. It's just very frustrating to
| ask a question about applying this technology in the real
| world, to actual commercial products, only to get reply
| after reply of hopes and dreams. 20 years ago Ray
| Kurzweil promised me I'd have artificial hemoglobin that
| allows me to hold my breath underwater for an hour. Where
| is it? These are the arguments of grifters and conmen:
| "just wait look at the exponential growth!" No. I refuse.
| If this technique doesn't work right now then it simply
| doesn't work and we're all (except researchers and
| companies developing the core technology) wasting immense
| amounts of time and money thinking about it.
| roland35 wrote:
| I like the term "bullshit" over "hallucination". These AI
| machines have no perceptions, no concept of truth, so indeed they
| are just spewing out words with no regard for truth.
|
| And unfortunately the cost of spreading bullshit has gone down to
| almost 0.
| uh_uh wrote:
| So why are they SOTA translators? Would you consider old
| translation software bullshit generators? Because LLMs can do
| their job, and more.
| krupan wrote:
| Those who sell these things as almighty powerful AI, so
| powerful and dangerous that it should be regulated (so my
| competition can't keep up with me), are responsible for this
| overreaction. Yes, not everything LLMs produce is bullshit
| but enough is, contrary to how they were marketed, that
| people are reacting this way.
|
| In other words, you don't counter hyperbolic and downright
| false marketing with subtleties.
| tim333 wrote:
| I dislike the term bullshit because it's use regarding ChatGPT
| does not match the dictionary version "stupid or untrue talk or
| writing; nonsense".
|
| If your sentence "LLMs output is bullshit" is wrong you may be
| better off changing the sentence than rewriting the dictionary
| to fit your sentence.
|
| I mean you can redefine words if you like, like how young
| people use sick and bad to mean much the opposite of what they
| did, which is fine as a fashion statement, but in trying to
| reason about LLMs it muddies the reasoning. Which of course is
| often why academics do it - see Hobbes, Calvin 1993
| https://www.reddit.com/r/calvinandhobbes/comments/1300k80/ac...
| somebehemoth wrote:
| I think a hallucinated sentence embedded in a paragraph of
| truth fits the definition: stupid and nonsense. A bullshitter
| can be right sometimes or even most of the time. They are
| still a bullshitter.
| tim333 wrote:
| Everyone gets things wrong sometimes.
| cobertos wrote:
| Read the whole course. There's a great amount of sourced case
| studies of people using AI in here with societal pushback which I
| find interesting to pull from.
|
| Disappointed though that the answer is so nuanced. There aren't
| hard and fast rules to when/when not to use AI but a set of 18 or
| so proposed principles that should guide are usage. And defense
| for those principles. The principles are at the bottom of each
| chapter.
|
| Also learned about the Eliza Effect as a term and that I found
| the passage in Ch14 by Ted Chiang to be really insightful, from a
| general social perspective.
|
| > When someone says "I'm sorry" to you, it doesn't matter that
| other people have said sorry in the past; it doesn't matter that
| "I'm sorry" is a string of text that is statistically
| unremarkable. If someone is being sincere, their apology is
| valuable and meaningful, even though apologies have previously
| been uttered.
| Johanx64 wrote:
| Somebody made a website to express their opinion - wherein their
| opinion can be surmised by reading the domain name.
|
| Text is scaled to 300% to indicate just how important and
| authoritative they think their opinion is.
|
| And it talks down to you in a "here comes the expert" style, with
| an atrocious aimed-at-preschoolers presentation.
|
| No thank you.
| FrustratedMonky wrote:
| That's too harsh.
|
| Some people do like big bullet list of points.
|
| Some people need to be spoon fed.
|
| Don't blame the spoon for being a spoon.
| myaccountonhn wrote:
| Two university professors in data science and computational
| biology are not just "somebody".
| Johanx64 wrote:
| People that find professors and similar in high esteem
| usually haven't spent enough time in or around academia and
| academics and for that reason still maintain some innocent
| aura of mystique and prestige around it in their mind.
|
| Quite literally anyone with a bit of persistence could become
| a professor (and this by far isn't even a top university
| either).
|
| They are quite literally just a somebody with an opinion just
| like anybody else. An opinion barked down to you with 300%
| scaled fonts and preschooler illustrations.
| myaccountonhn wrote:
| I have been around a lot of academics and been in academia.
|
| While we shouldn't trust academics just because they are
| academics, these people specialize in relevant fields and
| also back their claims with citations throughout the
| course. They are not just postulating.
| EncomLab wrote:
| Too little is made of the distinction between silicon substrate,
| fixed threshold, voltage moderated brittle networks of solid-
| state switches and protein substrate, variable threshold,
| chemically moderated plastic networks of biological switches.
|
| To be clear, neither possesses any magical "woo" outside of
| physics that gives one or the other some secret magical
| properties - but these are not arbitrary meaningless distinctions
| in the way they are often discussed.
| shaggie76 wrote:
| I was thinking how this article claims that people crave the
| authenticity of live music and that bullshit-generators will
| never be able to supply that. At first, I saw this as a reason
| for optimism, but then I got to thinking about evidence that
| people may not necessarily want authenticity after all.
|
| Organic produce was the first metaphor the came to mind: it's
| probably more healthy for you even if it isn't as pretty, but
| many people aren't willing to pay a premium it and I suspect
| economics isn't often the reason. Is that a straw-man for live
| music? I don't know that it is because plenty of people are
| content to listen to recorded music -- sure, they might enjoy
| going to a live concert but they'll still listen to the radio on
| the drive to work.
|
| Then I got to thinking about something more crass: while breast
| implants and other cosmetic body surgery may be as much for the
| benefit of the subject self-image I imagine there are plenty of
| people that find it very attractive despite what is often
| obviously fake.
|
| So do we crave authenticity? I think I do but I'm not sure if
| that's a safe generalization to make.
| bsenftner wrote:
| Of course we do, but also realize that the threshold for "being
| authentic" is flimsy at best for many, and an Everest to climb
| for others. We want more skeptics in generalized society with
| their own personal Everest they require of their thought
| leaders. This variability of acceptance for integrity is a
| weakness in our civilization, strategically grown by attacking
| educational institutions, and currently being exploited to
| great success by Orwellian long players.
| tmnvdb wrote:
| Craving "authenticity" is somewhere very high up the hierarchy
| of needs, i.e. a luxury. That a lot of people do not care much
| for it is not a sign of moral failing but of having bigger fish
| to fry.
| woodruffw wrote:
| I think it'd fall somewhere close to a "social" need, i.e.
| right smack in the middle of Maslow's hierarchy.
| flessner wrote:
| The travel and tourism industry is growing at around 4% YoY. So
| there is still a need (a growing need!) for "experiences" and
| "moments".
|
| The "online bullshit" combined with no simple way towards
| ownership for younger generations are among the highest growth
| factors here in my opinion.
|
| Still agreeing with your points, just wanted to add this
| context as there's likely a difference mentally between "social
| media" and "real world" authentic.
| api wrote:
| Organic produce isn't a good example because it's a dodgy
| poorly defined concept and it's not clear that it's either
| better for you or better tasting.
|
| The thing about art is this: art is a message from a human to
| another human.
|
| If I want art I want it to be that. I don't want a numerical
| average of all past messages, which is what LLMs and diffusion
| models create. I also don't want randomness or gimmickry, which
| is why I dislike a lot of pretentious modern art.
| cratermoon wrote:
| I didn't want to hijack the thread too much but yeah, organic
| farming was not originally about better for the consumer. The
| origin of organic farming methods was about taking care of
| the land and the ecosystems in which farming happens. The end
| products happened to be healthier, in some cases, because of
| reduced usage of chemicals known to be harmful resulted in
| less of them in the product at retail.
| TeMPOraL wrote:
| Perhaps, but then the meaning of the word is _in how it 's
| used_ (something that, if more people truly understood,
| would cut the amount of dismissing LLMs as "bullshit
| machines" and "stochastic parrots" by half).
|
| "Organic farming" may have initially been about
| sustainability, but the result correlated well enough with
| healthy food - and even more so with the naturalistic
| fallacy-fueled "healthy food" fad, that the latter
| application took over as it became a market niche. The
| niche being itself based more on a fallacy than reality is
| why "organic food" is such a bullshit fest it is - one
| product might be genuinely healthier, another is just worse
| and also ruins the land because it's sprayed with a nasty
| set of chemicals that are more "natural" than their
| strictly safer "modern" alternatives...
| cratermoon wrote:
| Marketing and greed do ruin everything they touch, yes.
| patja wrote:
| I'm uncomfortable with the use of profanity as a core element of
| this campaign branding, especially given that it seems to be an
| educational outreach effort. While this seems targeted at college
| age and above, I think it would be highly relevant content for a
| teen audience as well. While I swear in privacy on occasion I
| think it has no place in the classroom. I really don't care for
| it in the workplace either but I grant that private enterprises
| can have their own culture.
|
| It is frustrating because I agree with most of the content and
| the need for informed debate on the topic. It is a bit like my
| reaction to reading Cory Doctorow: I agree with his politics but
| really dislike the hamfisted way he packages his advocacy in the
| form of action adventures. As if the merits of his arguments need
| to be packaged in cotton candy to be consumed, and there is an
| undercurrent of self-promotion and personal branding that feels
| suss.
|
| Probably all a "me" problem with associations built up over time
| from seeing snake oil being packaged using a similar playbook. If
| you have to sell your message by dressing it up with scroll
| effects and provocative offensive language you've already lost
| me.
| woodruffw wrote:
| I assume that the use of the word "bullshit" on this site is at
| least in part informed by "On Bullshit,"[1] which is a pretty
| common undergraduate reading.
|
| [1]: https://en.wikipedia.org/wiki/On_Bullshit
| ctbergstrom wrote:
| This is something we've given serious consideration, having
| taught a course called "Calling Bullshit"
| (http://callingbullshit.org) for almost a decade and having
| authored a book by the same name that gets downranked on
| various Amazon features because of its title.
|
| But the bullshit is a term of art here, after the seminal 1986
| article "On Bullshit" by Princeton philosopher Harry Frankfurt
| (later published as a little book). We strongly feel that it is
| exactly the right term for what LLMs are doing, and we make the
| case for that in lesson 2 of the course.
| (https://thebullshitmachines.com/lesson-2-the-nature-of-
| bulls...)
|
| We're also concerned about accessibility for high school
| teachers etc., and thinking about what to do in that direction.
|
| I'm curious: do you find "bs" to be any less offensive?
| nmca wrote:
| (while I work at OAI, the opinion below is strictly my own)
|
| I feel like the current version is fairly hazardous to students
| and might leave them worse off.
|
| If I offer help to nontechnical friends, I focus on:
|
| - look at rate of change, not current point
|
| - reliability substantially lags possibility, by maybe two years.
|
| - adversarial settings remain largely unsolved if you get enough
| shots, trends there are unclear
|
| - ignore the parrot people, they have an appalling track record
| prediction-wise
|
| - autocorrect argument is typically (massively) overstated
| because RL exists
|
| - doomers are probably wrong but those who belittle their claims
| typically understand less than the doomers do
| dimgl wrote:
| What are "parrot people"? And what do you mean by "doomers are
| probably wrong?"
| moozilla wrote:
| OP is likely referring to people who call LLMs "stochastic
| parrots" (https://en.wikipedia.org/wiki/Stochastic_parrot),
| and by "doomers" (not boomers) they likely mean AI safetyists
| like Eliezer Yudkowsky or Pause AI (https://pauseai.info/).
| layoric wrote:
| How does this help the students with their use of these tools
| in the now, to not be left worse off? Most of the points you
| list seem like defending against criticism rather than helping
| address the harm.
| jdlshore wrote:
| I read the whole course. Lesson 16, "The Next-Step Fallacy,"
| specifically addresses your argument here.
| megaloblasto wrote:
| > After talking to literally hundreds of educators, employers,
| researchers, and policymakers, we have spent the last eight
| months developing the course on large language models (LLMs) that
| we think every college freshman needs to take.
|
| Did you consider consulting with any college freshman or even
| college students? I know you're supposed to be guiding their
| education, but I think it's also good to check in and see what
| they care about learning.
| ctbergstrom wrote:
| Yes--I should have stressed that. More than anything, we've
| talked at great lengths about LLMs with over a thousand
| undergraduate students who we have taught in our courses since
| ChapGPT 3.5 launched in Nov 22.
| aucisson_masque wrote:
| > LLMs are not capable of reflecting on and reporting about how
| or why they do what they do.
|
| i get the why, but about the how:
|
| deepseek has shown it's able to explain it's reasoning, how it
| reach a conclusion. That's quite a far stretch from the llm being
| only able to estimate statistically what's the right word to put
| after another one.
| CamperBob2 wrote:
| Their arguments are just one special-pleading fallacy after
| another. "But... but... but... it's different when _we_ do it!
| "
|
| No, it's not different when we do it. The takeaway here isn't
| that the AI algorithms are so special and magical, it's that
| our brains are not. It takes some nerve (literally) for humans
| to throw around labels like "bullshit machine."
|
| The only advantage we really have is long-term memory. I'm sure
| that will be addressed soon enough. Someone will figure out the
| ANN analogue of memory consolidation during sleep, and that'll
| be the tipping point.
| ndstephens wrote:
| Really enjoying this. Thank you for the great work. I'm currently
| on Lesson 11 and noticed a couple typos (missing words). I
| haven't found anywhere on the site itself where I could send
| feedback to report such a thing (maybe I missed it). Hopefully
| you aren't offended if I post them here.
|
| I think the easiest way to point them out is to just have you
| search for the partial line of text while on Lesson 11 and you'll
| see the spots.
|
| "No one is going to motivated by a robotic..." (missing the word
| "be")
|
| "People who are given a possible solution to a problem tend to
| less creative at..." (again missing the word "be")
| ctbergstrom wrote:
| Thank you very much -- fixed!
| layer8 wrote:
| Background: https://thebullshitmachines.com/about-us/
| foobiekr wrote:
| It is _very_ unfortunate there is no PDF version.
| timewizard wrote:
| It's not a "ChatGPT world."
|
| You can thrive just fine by entirely ignoring it and all the
| snake oil vendors living in it.
|
| 4 years and all they have to show for it is absurdly powerful
| video cards, sub 90% accuracy where it matters, and the only
| application is "chat bot."
|
| It's a fad. Wake up everyone.
| BugsJustFindMe wrote:
| > _Large language models are both powerful tools, and mindless--
| even dangerous--bullshit machines._
|
| Yes, but so is the average person. Do you cover the
| similarities/parallels to common human behavior patterns in your
| course? It stands out to me as a major blind spot in a lot of the
| discourse about LLMs, down to willing acceptance of humans
| intuiting what someone else meant when the words themselves were
| fundamentally ambiguous, which is extremely akin to a
| hallucination chosen from the listener's language model.
| justonceokay wrote:
| Yes but people have culpability and responsibility. In places
| of power or influence, this culpability can lead to being
| fired, legal action, disbarment, loss of money, etc. so there
| is a real pressure to be coherent and aligned with reality.
| BugsJustFindMe wrote:
| You say this but I have a ton of experience that "can lead"
| and "there is" come with a gigantic pile of caveats to the
| extent that they appear to be more false than true. Or
| they're technically true but practically meaningless. The
| world has been utterly awash in mass-perpetuated
| misinformation for at least all of recorded history without
| any real ability to stem the onslaught. This is not a modern
| problem just because a modern technology also exhibits it.
|
| You should look up the percentage of Americans who believe in
| ghosts sometime. About as many people believe in ghosts as
| don't, so no matter which side you land on, the other side is
| enormous. One of the sides must be wrong, I won't claim
| which, though only the belief side fails a falsifiability
| check. Where's our accountability to believing and spreading
| ideas based in reality again? The believers believe because
| they learned about it from someone. It didn't happen
| spontaneously on its own.
|
| It's all just been memes the entire time.
| evklein wrote:
| This is a sidestep imo. People _can_ be held accountable,
| though they will not always be. Machines add a layer of
| complexity - money is lost or a life is lost because AI
| made the call, who bears the burden? Machines _can't_ be
| held accountable.
| _dark_matter_ wrote:
| The key distinctions are: we can hold people accountable, and
| the amount of shit produced by one person is limited. Neither
| of those are true for LLMs.
| BugsJustFindMe wrote:
| You wrote basically the same thing as the adjacent person, so
| rather than write the same response to you both, I'll
| redirect to my reply over there:
|
| https://news.ycombinator.com/item?id=42993981
| 1propionyl wrote:
| The recent book Unaccountability Machine (Dan Davies) made
| the rounds a while back on this site a few years back. A few
| years before that, Ted Chiang's essay in the New Yorker
| entitled "Will AI Become the New McKinsey?" did likewise.
|
| There's a bright red through line here. I get the sense that
| the intellectual ferment is starting to develop an awareness
| of the risk (and if you're cynical, potential) of LLM
| deployments in business as a systematic strategy for
| absorbing accountability for decisions.
|
| For the time being, it usually seems like there's someone who
| is accountable, or at least can be scapegoated. But how long
| will that last? As Davies points out in his book, we didn't
| need LLMs to create bureaucracies where the buck fails to
| stop anywhere and instead irretrievably slides between the
| tracks. As Chiang points out, the "efficiency maximization"
| of McKinsey served as a way for organizations to outsource
| accountability for major decisions to an entity very, very
| good at working backwards from desired outcomes while acting
| like they were just led to a fait accompli by the numbers.
|
| [0] https://libro.fm/audiobooks/9781805220794-the-
| unaccountabili...
|
| [1] https://www.newyorker.com/science/annals-of-artificial-
| intel...
| dczx wrote:
| I'm offended.
| idunnoman1222 wrote:
| My 80s copy of the encyclopedia Britannica was riddled with
| errors, perhaps we will survive this post truth
| therein wrote:
| The situation is closer to if we had 10,000 variants of
| Encyclopedia Britannica in 80s that all looked like distinct
| bodies of work, riddled with different errors while looking
| like they were written from scratch.
| jvanderbot wrote:
| > We (the authors of this website) have at times sought insight
| into the inner workings of an LLM by asking it "why did you just
| do that?"
|
| > But the LLM can't tell us. It's not a person. It doesn't have
| the metacognitive abilities necessary to reflect on its past
| actions and report the motivations underlying them*.
|
| > With no clue why it did whatever it just did, the LLM is forced
| to guess wildly at a plausible explanation, like the ill-fated
| Leonard Shelby in Christopher Nolan's film Memento.
|
| > And we, gullible humans that we are, often believe its
| bullshit.
|
| ---
|
| I am almost convinced that we ourselves are a narrator riding
| along inside an animal's mind, trying desperately to put together
| explanations for our actions, mostly just to convince others,
| just as though an LLM were running on our own senses trying to
| portray some deep semblance of consciousness. I don't think we'll
| find a super smart AI, we'll just realize we were not very
| sophisticated all along. The power of speech for information and
| culture transfer, writing, inspiration, and coordination is just
| awe inspiring, evolutionary speaking, so once we could talk we
| had to, because the "better" talker almost always won. It's an
| arms race.
| ysofunny wrote:
| I am completely convinced that what we call consciousness is as
| you say.
|
| this means that it really exists in retrospect (20ms ? i recall
| some neuroscience articles from the 00s). nonetheless its whole
| reason for existing (retrospectively) is planning the future
| I-M-S wrote:
| Related, since the advent of LLMs I've become acutely aware how
| any argument of a considerable length with another person
| quickly starts meandering and how topics change seemingly of no
| one's volition - almost as if our own internal token limit has
| been exceeded.
| paulcole wrote:
| > Others say they are nothing but bullshit machines.
|
| Simply ignore anyone who says this and go about your business.
|
| Doubly so if they bring up the environmental impact of AI.
| akomtu wrote:
| "How to thrive in a machine world?" should be the rhetorical
| question.
| codpiece wrote:
| I am nearly 60, and am excited to take your course! I
| congratulate you in finding a compelling way to teach Humanities
| that is relevant to today's society. Bravo.
| aaplok wrote:
| Really well done. It is really a challenge for students to
| navigate their way around the AI landscape. I am definitely
| considering sharing that with my students.
|
| Have you noticed a difference in how your students approach LLMs
| after taking your course? A possible issue I see is that it is
| preaching to the choir; a student who is enclined to use LLMs for
| everything is less likely to engage with the material in the
| first place.
|
| If you allow feedback, I was interested in lesson 10 on writing,
| as an educator who tries to teach my science/IT/maths students
| the importance of being able to communicate.
|
| I would suggest to include a paragraph to explain why being able
| to write without LLMs is just as important in scientific
| disciplines, where precision and accuracy are more essential than
| creativity and personalisation.
| ctbergstrom wrote:
| This is an excellent point about scientific writing. We'll add
| something to that effect.
|
| We have not taught this course from the web-based materials
| yet, but it distills much of the two-week unit that we covered
| in our "Calling Bullshit" course this past autumn. We find that
| our students are generally very interested to better understand
| the LLMs that they are using -- and almost every one of them
| does, to vary degree. (Of course there may be some selection
| bias in that the 180 students who sign up to take a course on
| data reasoning may be more curious and more skeptical than the
| average.)
| pjs_ wrote:
| Not sure why everyone rates this. It's full of very confidently
| made statements like "the AI has no ground truth" (obviously it
| does, it has ingested every paper ever), it "can't reason
| logically" which seems like a stretch if you ever read the CoT of
| a frontier reasoning model and "can't explain how they arrived at
| conclusions" where - I mean just try it yourself with o1, go as
| deep as you like asking how it arrived at a conclusion and see if
| a human can do any better.
|
| In fact the most annoying thing about this article is that it is
| a string of very confidently made, black and white statements,
| offered with no supporting evidence, and some of which I think
| are actually wrong... i.e. it suffers from the same kind of
| unsubstantiated self confidence that we complain about with the
| weaker models
| bwfan123 wrote:
| the machine is fooling you with a mimicry of reasoning. and you
| are falling for it.
| ZephyrBlu wrote:
| What is reasoning if not a chain of logically consistent
| thoughts?
| bwfan123 wrote:
| fair, but "logically consistent thoughts" is a subject of
| deep investigation starting from the early euclidean
| geometry to the modern godel's theorems.
|
| ie, that logically consistent thinking starts from
| symbolization, axioms, proof procedures, world models.
| otherwise, you end up will persuasive words.
| ZephyrBlu wrote:
| You just ruled out 99% of humans from having reasoning
| capabilities.
|
| The beautiful thing about reasoning models is that there
| is no need to overcomplicate it with all the things
| you've mentioned, you can literally read the model's
| reasoning and decide for yourself if it's bullshit or
| not.
| Grimblewald wrote:
| If it's mimicry of reason is indistinguishable from real
| reasoning, how is it not reasoning?
|
| Ultimately, an LLM models language and the process behind
| it's creation to some degree of accuracy or another. If that
| model includes a way to approximate the act of reasoning,
| then it is reasoning to some extent. The extent I am happy to
| agree is open for discussion, but that reasoning is taking
| place at all is a little harder to attack.
| krainboltgreene wrote:
| > obviously it does, it has ingested every paper ever
|
| Do you have a citation for such a claim?
| grepLeigh wrote:
| LLMs that use Chain of Thought sequences have been demonstrated
| to misrepresent their own reasoning [1]. The CoT sequence is
| another dimension for hallucination.
|
| So, I would say that an LLM capable of explaining its reasoning
| doesn't guarantee that the reasoning is grounded in logic or
| some absolute ground truth.
|
| I do think it's interesting that LLMs demonstrate the same
| fallibility of low quality human experts (i.e. confident
| bullshitting), which is the whole point of the OP course.
|
| I love the goal of the course: get the audience thinking more
| critically, both about the output of LLMs and the content of
| the course. It's a humanities course, not a technical one.
|
| (Good) Humanities courses invite the students to question/argue
| the value and validity of course content itself. The point
| isn't to impart some absolute truth on the student - it's to
| set the student up to practice defining truth and
| communicating/arguing their definition to other people.
|
| [1] https://arxiv.org/abs/2305.04388
| ctbergstrom wrote:
| Yes!
|
| First, thank you for the link about CoT misrepresentation.
| I've written a fair bit about this on Bluesky etc but I don't
| think much if any of that made it into the course yet. We
| should add this to lesson 6, "They're Not Doing That!"
|
| Your point about humanities courses is just right and
| encapsulates what we are trying to do. If someone takes the
| course and engages in the dialectical process and decides we
| are much to skeptical, great! If they decide we aren't
| skeptical enough, also great. As we say in the instructor
| guide:
|
| "We view this as a course in the humanities, because it is a
| course about what it means to be human in a world where LLMs
| are becoming ubiquitous, and it is a course about how to live
| and thrive in such a world. This is not a how-to course for
| using generative AI. It's a when-to course, and perhaps more
| importantly a why-not-to course.
|
| "We think that the way to teach these lessons is through a
| dialectical approach.
|
| "Students have a first-hand appreciation for the power of AI
| chatbots; they use them daily.
|
| "Students also carry a lot of anxiety. Many students feel
| conflicted about using AI in their schoolwork. Their teachers
| have probably scolded them about doing so, or prohibited it
| entirely. Some students have an intuition that these machines
| don't have the integrity of human writers.
|
| "Our aim is to provide a framework in which students can
| explore the benefits and the harms of ChatGPT and other LLM
| assistants. We want to help them grapple with the
| contradictions inherent in this new technology, and allow
| them to forge their own understanding of what it means to be
| a student, a thinker, and a scholar in a generative AI
| world."
| 3abiton wrote:
| If you were a book recommender system, what other books are
| similar to yours? I've read plenty of science/maths light read
| non-fiction, so I want to compare the reading experience before I
| jump into the book.
| bwfan123 wrote:
| Fantastic course and website, thank you professors for this
| valuable contribution.
| svilen_dobrev wrote:
| very well done. Thank you.
|
| btw, typo in lesson 11: "(2) understand how an LMM can help them"
| .. instead of LLM
|
| IMO, "understanding" is the notion most endangered and the one
| which loss is with most catastrophic consequences, from all the
| possibly affected ones. It has never been very favored, but now
| is even worse than ever - it is thinned and abandoned en-masse.
|
| Check brother's Strugatsky's Snail_on_the_Slope [0] , there's
| very stringent monologue of Peretz on the topic, ~~ page 11 [1]
|
| [0] https://en.wikipedia.org/wiki/Snail_on_the_Slope
|
| [1] https://strugacki.ru/book_19/768.html
| zaptheimpaler wrote:
| I like it. It's pretty basic but it is very good for a broad
| audience and covered things many people don't understand. I liked
| that you mentioned not to anthropomorphize the model. We would
| greatly benefit from 50+ year old policymakers and more taking
| the course even more than 19 year old freshmen.
| dzonga wrote:
| thank you for presenting this in digestible format.
|
| fortunately unfortunately, people who know llm's shill bullshit
| are the ones selling llm's while they wouldn't feed llm's to
| their kids or eat llm's.
| raincom wrote:
| Since you are on Frankfurtian bullshit, why don't you consider
| Late G.A. Cohen's take on intellectual bullshit (bullshit
| perpetuated in the academia), as the latter's notion of bullshit
| is linked to knowledge, unlike the bullshit we hear from
| salesmen, showmen, etc.
___________________________________________________________________
(page generated 2025-02-09 23:00 UTC)