[HN Gopher] 1960s chatbot ELIZA beat OpenAI's GPT-3.5 in a recen...
       ___________________________________________________________________
        
       1960s chatbot ELIZA beat OpenAI's GPT-3.5 in a recent Turing test
       study
        
       Author : layer8
       Score  : 120 points
       Date   : 2023-12-03 10:56 UTC (12 hours ago)
        
 (HTM) web link (arstechnica.com)
 (TXT) w3m dump (arstechnica.com)
        
       | gumballindie wrote:
       | Time for a new test as the turning test while fun is out of date.
        
         | cykros wrote:
         | Voight-Kampff.
        
       | TechRemarker wrote:
       | " GPT-3.5, the base model behind the free version of ChatGPT, has
       | been conditioned by OpenAI specifically not to present itself as
       | a human, which may partially account for its poor performance."
       | So 3.5 isn't bad at passing it compared to a 1960s test, it just
       | has been design specifically not to, contrary to the implications
       | of the title.
        
       | taspeotis wrote:
       | > some interrogators reported thinking that ELIZA was "too bad"
       | to be a current AI model, and therefore was more likely to be a
       | human intentionally being uncooperative
        
         | masswerk wrote:
         | Given that ELIZA doesn't really work outside its supposed
         | social setup, which makes some of its behavior acceptable
         | and/or intelligible, it would be hard to believe that anyone
         | would have been actually fooled by this. (Still, the
         | performance of ELIZA in the bounds of its intended setup is
         | remarkable, especially for its small rule set. But there is
         | little doubt that this is insufficient for a sustained free
         | conversation.) - Let's call it a "parasitic win". ;-)
        
         | ithkuil wrote:
         | It clearly shows that the Turing test is not only about the bot
         | but also about the humans
        
           | rvense wrote:
           | You can think of a language as an psychological object, a
           | thing that can be described by a grammar and a dictionary.
           | That's basically where linguistics started, and it's part of
           | it. But some things about language are really best understood
           | and described by looking at it as a thing that people do
           | together, a situated, interactice activity.
           | 
           | And I think that part of the reason that both Eliza and GPT
           | work so well is that the observer really, really wants it to,
           | because that's a crucial part of how this activity is
           | structured, as described by Grice and others:
           | 
           | https://en.m.wikipedia.org/wiki/Cooperative_principle
        
           | trehalose wrote:
           | Yep. It demonstrates that the Turing test is strongly
           | influenced by the culture it's administered in. One bot can
           | influence _other_ bots ' performance on the test just by
           | existing and being familiar to the public.
        
       | gyudin wrote:
       | Another meaningless study written for the sake of writing
       | something. Didn't even bother to prepare a proper role for GPT-4.
       | 
       | And the whole article can be shortened down to "GPT 3.5 has been
       | conditioned by OpenAI specifically not to present itself as a
       | human." and it does it with success :)
        
         | RandomLensman wrote:
         | GPT4 might not be great at the Turing test either, it seems
         | (https://arxiv.org/abs/2310.20216 and that is the simple non-
         | adverserial version of the test).
        
         | hfhdjdks wrote:
         | "are you a human?" "No, I'm a model trained by OpenAI"
         | 
         | Nearly fooled me
        
           | mewpmewp2 wrote:
           | > "No, I'm a model trained by OpenAI"
           | 
           | That's exactly what a human would say.
        
           | Kim_Bruning wrote:
           | > Hello, I am Eliza.
           | 
           | * Are you a human
           | 
           | > Would you prefer if I were not a human?
           | 
           | Eliza wins this round! %-P
        
       | throwuwu wrote:
       | What a waste of time. Why wouldn't they put in a tiny bit of
       | effort to use an uncensored model like a dolphin 2.0 finetune?
       | Not serious work at all.
        
       | transfire wrote:
       | How does that make you feel?
        
       | anthk wrote:
       | If you use Linux/BSD you have Eliza under Alt-x (or Esc-x) doctor
       | for Emacs.
       | 
       | Also, under Perl set up cpanminus if you don't want to play with
       | Emacs:                   curl -L https://cpanmin.us | perl -
       | App::cpanminus              ~/perl5/bin/cpanm -n local::lib
       | echo  'eval $(perl -I ~/perl5/lib/perl5/ -Mlocal::lib)' >>
       | ~/.profile              echo 'export PATH=$HOME/perl5/bin:$PATH'
       | >> ~/.profile
       | 
       | Logout and login again. Install rivescript:
       | cpanm -n RiveSript
       | 
       | Try Eliza under RiveScript:
       | ~/perl5/bin/rivescript ~/perl5/lib/perl5/RiveScript/demo/
       | 
       | If you want to create your own chatbot in RiveScript, the
       | language it's pretty easy and you can bind Perl functions, code
       | and who knows to your bot, so if you ask for the weather, for
       | instance, you could call an HTTP module for Perl to call wttr.in
       | for instance. Happy Hacking.
        
       | 3836293648 wrote:
       | I thought ELIZA war lost and only clones remained?
        
         | homarp wrote:
         | https://news.ycombinator.com/item?id=27321906
         | 
         | it was eventually found
        
       | toss1 wrote:
       | >> raises further questions about using the Turing test to
       | evaluate AI model performance.
       | 
       | Aside from the fact that ELIZA was tuned to present as human and
       | GPT3.5 was not, the validity of the basic Turing Test still seems
       | a key issue. Fooling ordinary humans is a common occurrence, as
       | Richard Feynman pointed out, "... and the easiest one to fool is
       | ourselves.".
        
       | usrbinbash wrote:
       | The turing test, as has been often discussed, is pointless.
       | Reason being that the tests results are determined at least as
       | much, if not more, by the interrogator and the human interrogee
       | as they are by the machine.
       | 
       | It doesn't test "Can a machine think" so much as it tests "Can a
       | human be tricked". There is a reason why ML research mostly
       | ignores the Turing "Test".
       | 
       | But, alas, its name has a nice ring to it, it is known (by name
       | at least) to pop culture, and easy enough to understand for
       | everyone, and so it gets far more attention than it deserves.
       | Pretty much like Asimovs three "laws" of robotics :D
        
         | n2d4 wrote:
         | That doesn't mean it's pointless. For most practical
         | (commercial?) purposes, "can a human be tricked" is a much more
         | important property than "can it think". If a decent portion of
         | humans can't distinguish the AI anymore, with trickery or not,
         | THAT opens up space for plenty of applications, regardless of
         | conscience or not. Think of support bots that sound like
         | humans, sales, email drafts, etc.
         | 
         | Like, most people would say that animals can think, but clearly
         | it can't pass the Turing test. Would you rather have
         | unconscious GPT-4 or a conscious pig "working" at your job?
        
           | usrbinbash wrote:
           | > For most practical (commercial?) purposes, "can a human be
           | tricked" is a much more important property than "can it
           | think".
           | 
           | I disagree.
           | 
           | In my experience human customers don't care much about
           | whether its obvious when they are communicating with a robot.
           | In fact, in several use cases it's considered good practice,
           | and may even be legally required, to tell them when they are.
           | 
           | What they DO care about, is whether their requests,
           | inquiries, complaints, etc. are handled quickly and
           | efficiently.
           | 
           | We don't create value for our customers by giving them good
           | chatterbots. We create value by giving them highly functional
           | automated systems.
        
             | n2d4 wrote:
             | I mean, if it can't handle requests like that, then it'll
             | also not pass the Turing test (since asking the AI to do
             | help with simple support request would get it to fail). I
             | doubt that "thinking" systems are the only ones that can
             | handle a support ticket (GPT-4 does quite alright on a lot
             | of them already), at least if you require consciousness as
             | a requirement for "thinking".
             | 
             | Maybe we just disagree on what we mean when we say "think"?
        
               | usrbinbash wrote:
               | > then it'll also not pass the Turing test
               | 
               | Chatterbots, including the 60s-era ELIZA, which are
               | mostly useless for this kind of task, have repeatedly
               | been shown to be able to trick people into chosing the
               | machine during the Turing "Test", so I'm not sure how you
               | come by that conclusion.
               | 
               | > at least if you require consciousness as a requirement
               | for "thinking".
               | 
               | I don't require consciousness, or an abstract and not-
               | even-clearly-defined "thinking" capability from an
               | advanced ML model. I require fitness for purpose,
               | whatever that purpose is, and in most use cases I ever
               | encountered, "trick the humen into believing I'm a human"
               | isn't one of them.
               | 
               | > Maybe we just disagree on what we mean when we say
               | "think"?
               | 
               | I don't think we really disagree, I think the problem is
               | that "thinking", "consciousness", "intelligence" and a
               | lot of other terms that often get thrown around when
               | talking about AI, are really not well defined in
               | technical terms.
        
           | thaumasiotes wrote:
           | > Like, most people would say that animals can think, but
           | clearly [they] can't pass the Turing test.
           | 
           | This is not actually clear; people commonly feel that they
           | are having meaningful conversations with animals.
           | 
           | The answer to "can a human be tricked" is always yes,
           | regardless of the particular context.
        
         | FrustratedMonky wrote:
         | It gets to the point of the 'philosophical zombie'. That If an
         | AI can completely mimic a human, then some internal subjective
         | 'experiences' must be occurring that are similar to a humans
         | internal subjective experience of reality. Call it
         | consciousness or whatever.
         | 
         | Of course it can't be proven one way or the other, we can't
         | prove humans are conscious either.
         | 
         | Humans think they are conscious because they have an internal
         | sense of self. How do you determine if a machine has that? If
         | it can mimic a human was one hypothetical test.
         | 
         | There are more robust versions of Turing Test than this one, so
         | can't toss it out as an idea.
        
         | Ukv wrote:
         | > It doesn't test "Can a machine think" so much as it tests
         | "Can a human be tricked".
         | 
         | But importantly, Turing didn't just mean it'd be tested on
         | ability to make small talk. He chose the format because he
         | considered text to be a sufficiently general interface for
         | testing any particular intelligent behavior, and gave examples
         | such as feeding in chess moves.
         | 
         | If it can perform all such intelligent behavior at human level
         | (as far as human interrogators can tell), then that's a
         | significant result and - depending on what philosophy of mind
         | you subscribe to - may or may not be sufficient to demonstrate
         | that it can "think".
        
           | usrbinbash wrote:
           | > (as far as human interrogators can tell),
           | 
           | And there's the problem again.
           | 
           | The "Test" isn't evaluating a machine. It evaluates an
           | interrogators response _to that_ machine.
           | 
           | The core problem of answering the question "can machines
           | think" isn't evaluated by Turings method. The "Test" doesn't
           | test what it is supposed to test.
           | 
           | And that has nothing to do with "what philosophy of mind you
           | subscribe to", that is simply a fundamental flaw in the tests
           | design, and the reason why it's pretty much pointless; it
           | defines "thinking" as "the ability to trick humans". Well,
           | there are optical illusions that can trick people into seeing
           | things that don't exist.
        
             | Ukv wrote:
             | > And that has nothing to do with "what philosophy of mind
             | you subscribe to",
             | 
             | There are definitely those that would reason "Using the
             | tools and expertise available to us, this behaves
             | indistinguishably from a human - therefore we should
             | believe it is intelligent, like humans are".
             | 
             | It's not the same question, and Turing is clear that it's a
             | surrogate question due to the ambiguity of the original
             | question, but it's not unrelated either.
        
       | seydor wrote:
       | ChatGPT, make me a wild clickbait claim about yourself
        
       | epistasis wrote:
       | People go wild for new code integrating with ChatGPT, but us
       | eMacs users have been living in the future since first picking up
       | the editor.
       | 
       | M-x doctor can fix all your woes.
        
         | facialwipe wrote:
         | Emacs and eMacs are two very different things.
        
           | tonyarkles wrote:
           | Lol and iOS autocorrect definitely prefers the latter.
        
             | epistasis wrote:
             | I had forgotten that eMacs even exist. But there must have
             | been another bad update to iOS's autocorrection wieghts
             | because I definitely did not have an eMacs problem the past
             | few years while writing comments on my phone...
        
               | tonyarkles wrote:
               | Ahhhh interesting, I have had that problem forever but
               | maybe it's because I often shrug it off and let it keep
               | the auto-correct as-is... further reinforcing to it that
               | that's what I meant?
               | 
               | On the other hand, it did offer (correctly) to use the
               | word "fucking" the other day for me instead of trying to
               | change that to "ducking", so maybe we're making a bit of
               | progress!
        
       | CoastalCoder wrote:
       | There's a great Corecursive podcast on the history of ELIZA: [0]
       | 
       | [0] https://corecursive.com/eliza-with-jeff-shrager/
        
       | Tiberium wrote:
       | Is there no repository with all the prompts they've used for
       | their study? In a research like this prompts are the main part,
       | so it's kind of weird for me that they only show one prompt in
       | the paper.
        
       | hedora wrote:
       | There are some nice tidbits in the article:
       | 
       | > The study also showed that participants' education and
       | familiarity with large language models (LLMs) did not
       | significantly predict their success in detecting AI.
       | 
       | Also, humans only score 61%, and the best chat gpt prompt gets
       | 41% (but 49% in the graph?).
       | 
       | I'm definitely having trouble distinguishing between content farm
       | crap written by humans and chatgpt output. They're usually of
       | comparable accuracy.
       | 
       | I'd be curious to see an extended study where they had a dozen
       | human participants or so, and got statistically significant
       | percentages for each individual. Is the worst human below 41%? Is
       | the best close to 100%?
        
       | bbor wrote:
       | This is a great example of how everyone misunderstands the Turing
       | test.
        
       ___________________________________________________________________
       (page generated 2023-12-03 23:02 UTC)