[HN Gopher] 1960s chatbot ELIZA beat OpenAI's GPT-3.5 in a recen...
___________________________________________________________________
1960s chatbot ELIZA beat OpenAI's GPT-3.5 in a recent Turing test
study
Author : layer8
Score : 120 points
Date : 2023-12-03 10:56 UTC (12 hours ago)
(HTM) web link (arstechnica.com)
(TXT) w3m dump (arstechnica.com)
| gumballindie wrote:
| Time for a new test as the turning test while fun is out of date.
| cykros wrote:
| Voight-Kampff.
| TechRemarker wrote:
| " GPT-3.5, the base model behind the free version of ChatGPT, has
| been conditioned by OpenAI specifically not to present itself as
| a human, which may partially account for its poor performance."
| So 3.5 isn't bad at passing it compared to a 1960s test, it just
| has been design specifically not to, contrary to the implications
| of the title.
| taspeotis wrote:
| > some interrogators reported thinking that ELIZA was "too bad"
| to be a current AI model, and therefore was more likely to be a
| human intentionally being uncooperative
| masswerk wrote:
| Given that ELIZA doesn't really work outside its supposed
| social setup, which makes some of its behavior acceptable
| and/or intelligible, it would be hard to believe that anyone
| would have been actually fooled by this. (Still, the
| performance of ELIZA in the bounds of its intended setup is
| remarkable, especially for its small rule set. But there is
| little doubt that this is insufficient for a sustained free
| conversation.) - Let's call it a "parasitic win". ;-)
| ithkuil wrote:
| It clearly shows that the Turing test is not only about the bot
| but also about the humans
| rvense wrote:
| You can think of a language as an psychological object, a
| thing that can be described by a grammar and a dictionary.
| That's basically where linguistics started, and it's part of
| it. But some things about language are really best understood
| and described by looking at it as a thing that people do
| together, a situated, interactice activity.
|
| And I think that part of the reason that both Eliza and GPT
| work so well is that the observer really, really wants it to,
| because that's a crucial part of how this activity is
| structured, as described by Grice and others:
|
| https://en.m.wikipedia.org/wiki/Cooperative_principle
| trehalose wrote:
| Yep. It demonstrates that the Turing test is strongly
| influenced by the culture it's administered in. One bot can
| influence _other_ bots ' performance on the test just by
| existing and being familiar to the public.
| gyudin wrote:
| Another meaningless study written for the sake of writing
| something. Didn't even bother to prepare a proper role for GPT-4.
|
| And the whole article can be shortened down to "GPT 3.5 has been
| conditioned by OpenAI specifically not to present itself as a
| human." and it does it with success :)
| RandomLensman wrote:
| GPT4 might not be great at the Turing test either, it seems
| (https://arxiv.org/abs/2310.20216 and that is the simple non-
| adverserial version of the test).
| hfhdjdks wrote:
| "are you a human?" "No, I'm a model trained by OpenAI"
|
| Nearly fooled me
| mewpmewp2 wrote:
| > "No, I'm a model trained by OpenAI"
|
| That's exactly what a human would say.
| Kim_Bruning wrote:
| > Hello, I am Eliza.
|
| * Are you a human
|
| > Would you prefer if I were not a human?
|
| Eliza wins this round! %-P
| throwuwu wrote:
| What a waste of time. Why wouldn't they put in a tiny bit of
| effort to use an uncensored model like a dolphin 2.0 finetune?
| Not serious work at all.
| transfire wrote:
| How does that make you feel?
| anthk wrote:
| If you use Linux/BSD you have Eliza under Alt-x (or Esc-x) doctor
| for Emacs.
|
| Also, under Perl set up cpanminus if you don't want to play with
| Emacs: curl -L https://cpanmin.us | perl -
| App::cpanminus ~/perl5/bin/cpanm -n local::lib
| echo 'eval $(perl -I ~/perl5/lib/perl5/ -Mlocal::lib)' >>
| ~/.profile echo 'export PATH=$HOME/perl5/bin:$PATH'
| >> ~/.profile
|
| Logout and login again. Install rivescript:
| cpanm -n RiveSript
|
| Try Eliza under RiveScript:
| ~/perl5/bin/rivescript ~/perl5/lib/perl5/RiveScript/demo/
|
| If you want to create your own chatbot in RiveScript, the
| language it's pretty easy and you can bind Perl functions, code
| and who knows to your bot, so if you ask for the weather, for
| instance, you could call an HTTP module for Perl to call wttr.in
| for instance. Happy Hacking.
| 3836293648 wrote:
| I thought ELIZA war lost and only clones remained?
| homarp wrote:
| https://news.ycombinator.com/item?id=27321906
|
| it was eventually found
| toss1 wrote:
| >> raises further questions about using the Turing test to
| evaluate AI model performance.
|
| Aside from the fact that ELIZA was tuned to present as human and
| GPT3.5 was not, the validity of the basic Turing Test still seems
| a key issue. Fooling ordinary humans is a common occurrence, as
| Richard Feynman pointed out, "... and the easiest one to fool is
| ourselves.".
| usrbinbash wrote:
| The turing test, as has been often discussed, is pointless.
| Reason being that the tests results are determined at least as
| much, if not more, by the interrogator and the human interrogee
| as they are by the machine.
|
| It doesn't test "Can a machine think" so much as it tests "Can a
| human be tricked". There is a reason why ML research mostly
| ignores the Turing "Test".
|
| But, alas, its name has a nice ring to it, it is known (by name
| at least) to pop culture, and easy enough to understand for
| everyone, and so it gets far more attention than it deserves.
| Pretty much like Asimovs three "laws" of robotics :D
| n2d4 wrote:
| That doesn't mean it's pointless. For most practical
| (commercial?) purposes, "can a human be tricked" is a much more
| important property than "can it think". If a decent portion of
| humans can't distinguish the AI anymore, with trickery or not,
| THAT opens up space for plenty of applications, regardless of
| conscience or not. Think of support bots that sound like
| humans, sales, email drafts, etc.
|
| Like, most people would say that animals can think, but clearly
| it can't pass the Turing test. Would you rather have
| unconscious GPT-4 or a conscious pig "working" at your job?
| usrbinbash wrote:
| > For most practical (commercial?) purposes, "can a human be
| tricked" is a much more important property than "can it
| think".
|
| I disagree.
|
| In my experience human customers don't care much about
| whether its obvious when they are communicating with a robot.
| In fact, in several use cases it's considered good practice,
| and may even be legally required, to tell them when they are.
|
| What they DO care about, is whether their requests,
| inquiries, complaints, etc. are handled quickly and
| efficiently.
|
| We don't create value for our customers by giving them good
| chatterbots. We create value by giving them highly functional
| automated systems.
| n2d4 wrote:
| I mean, if it can't handle requests like that, then it'll
| also not pass the Turing test (since asking the AI to do
| help with simple support request would get it to fail). I
| doubt that "thinking" systems are the only ones that can
| handle a support ticket (GPT-4 does quite alright on a lot
| of them already), at least if you require consciousness as
| a requirement for "thinking".
|
| Maybe we just disagree on what we mean when we say "think"?
| usrbinbash wrote:
| > then it'll also not pass the Turing test
|
| Chatterbots, including the 60s-era ELIZA, which are
| mostly useless for this kind of task, have repeatedly
| been shown to be able to trick people into chosing the
| machine during the Turing "Test", so I'm not sure how you
| come by that conclusion.
|
| > at least if you require consciousness as a requirement
| for "thinking".
|
| I don't require consciousness, or an abstract and not-
| even-clearly-defined "thinking" capability from an
| advanced ML model. I require fitness for purpose,
| whatever that purpose is, and in most use cases I ever
| encountered, "trick the humen into believing I'm a human"
| isn't one of them.
|
| > Maybe we just disagree on what we mean when we say
| "think"?
|
| I don't think we really disagree, I think the problem is
| that "thinking", "consciousness", "intelligence" and a
| lot of other terms that often get thrown around when
| talking about AI, are really not well defined in
| technical terms.
| thaumasiotes wrote:
| > Like, most people would say that animals can think, but
| clearly [they] can't pass the Turing test.
|
| This is not actually clear; people commonly feel that they
| are having meaningful conversations with animals.
|
| The answer to "can a human be tricked" is always yes,
| regardless of the particular context.
| FrustratedMonky wrote:
| It gets to the point of the 'philosophical zombie'. That If an
| AI can completely mimic a human, then some internal subjective
| 'experiences' must be occurring that are similar to a humans
| internal subjective experience of reality. Call it
| consciousness or whatever.
|
| Of course it can't be proven one way or the other, we can't
| prove humans are conscious either.
|
| Humans think they are conscious because they have an internal
| sense of self. How do you determine if a machine has that? If
| it can mimic a human was one hypothetical test.
|
| There are more robust versions of Turing Test than this one, so
| can't toss it out as an idea.
| Ukv wrote:
| > It doesn't test "Can a machine think" so much as it tests
| "Can a human be tricked".
|
| But importantly, Turing didn't just mean it'd be tested on
| ability to make small talk. He chose the format because he
| considered text to be a sufficiently general interface for
| testing any particular intelligent behavior, and gave examples
| such as feeding in chess moves.
|
| If it can perform all such intelligent behavior at human level
| (as far as human interrogators can tell), then that's a
| significant result and - depending on what philosophy of mind
| you subscribe to - may or may not be sufficient to demonstrate
| that it can "think".
| usrbinbash wrote:
| > (as far as human interrogators can tell),
|
| And there's the problem again.
|
| The "Test" isn't evaluating a machine. It evaluates an
| interrogators response _to that_ machine.
|
| The core problem of answering the question "can machines
| think" isn't evaluated by Turings method. The "Test" doesn't
| test what it is supposed to test.
|
| And that has nothing to do with "what philosophy of mind you
| subscribe to", that is simply a fundamental flaw in the tests
| design, and the reason why it's pretty much pointless; it
| defines "thinking" as "the ability to trick humans". Well,
| there are optical illusions that can trick people into seeing
| things that don't exist.
| Ukv wrote:
| > And that has nothing to do with "what philosophy of mind
| you subscribe to",
|
| There are definitely those that would reason "Using the
| tools and expertise available to us, this behaves
| indistinguishably from a human - therefore we should
| believe it is intelligent, like humans are".
|
| It's not the same question, and Turing is clear that it's a
| surrogate question due to the ambiguity of the original
| question, but it's not unrelated either.
| seydor wrote:
| ChatGPT, make me a wild clickbait claim about yourself
| epistasis wrote:
| People go wild for new code integrating with ChatGPT, but us
| eMacs users have been living in the future since first picking up
| the editor.
|
| M-x doctor can fix all your woes.
| facialwipe wrote:
| Emacs and eMacs are two very different things.
| tonyarkles wrote:
| Lol and iOS autocorrect definitely prefers the latter.
| epistasis wrote:
| I had forgotten that eMacs even exist. But there must have
| been another bad update to iOS's autocorrection wieghts
| because I definitely did not have an eMacs problem the past
| few years while writing comments on my phone...
| tonyarkles wrote:
| Ahhhh interesting, I have had that problem forever but
| maybe it's because I often shrug it off and let it keep
| the auto-correct as-is... further reinforcing to it that
| that's what I meant?
|
| On the other hand, it did offer (correctly) to use the
| word "fucking" the other day for me instead of trying to
| change that to "ducking", so maybe we're making a bit of
| progress!
| CoastalCoder wrote:
| There's a great Corecursive podcast on the history of ELIZA: [0]
|
| [0] https://corecursive.com/eliza-with-jeff-shrager/
| Tiberium wrote:
| Is there no repository with all the prompts they've used for
| their study? In a research like this prompts are the main part,
| so it's kind of weird for me that they only show one prompt in
| the paper.
| hedora wrote:
| There are some nice tidbits in the article:
|
| > The study also showed that participants' education and
| familiarity with large language models (LLMs) did not
| significantly predict their success in detecting AI.
|
| Also, humans only score 61%, and the best chat gpt prompt gets
| 41% (but 49% in the graph?).
|
| I'm definitely having trouble distinguishing between content farm
| crap written by humans and chatgpt output. They're usually of
| comparable accuracy.
|
| I'd be curious to see an extended study where they had a dozen
| human participants or so, and got statistically significant
| percentages for each individual. Is the worst human below 41%? Is
| the best close to 100%?
| bbor wrote:
| This is a great example of how everyone misunderstands the Turing
| test.
___________________________________________________________________
(page generated 2023-12-03 23:02 UTC)