[HN Gopher] AI21 Labs concludes largest Turing Test experiment t...
       ___________________________________________________________________
        
       AI21 Labs concludes largest Turing Test experiment to date
        
       Author : kennyfrc
       Score  : 81 points
       Date   : 2023-05-31 13:05 UTC (9 hours ago)
        
 (HTM) web link (www.ai21.com)
 (TXT) w3m dump (www.ai21.com)
        
       | boringuser2 wrote:
       | I thought this was some kind of scam website to be honest.
        
       | cornmn12 wrote:
       | [dead]
        
       | diziet wrote:
       | The limits of conversation (2 minutes max, people disconnect)
       | make this test really limited.
        
       | pron wrote:
       | The actual Turing test requires an interrogator interacting with
       | _both_ a human and a machine at the same time, each trying to get
       | the interrogator to declare them the human (and can suggest
       | questions):
       | https://en.wikipedia.org/wiki/Computing_Machinery_and_Intell...
        
       | stirlo wrote:
       | This AI is rather pathetic. To work within the time limit it
       | would need to make typos and mistakes but even the first reply
       | makes it easy to call out. Not very impressive.
       | 
       | https://ibb.co/xL0XpZ7
        
         | None4U wrote:
         | There are many different bots
        
       | vuxie wrote:
       | I found it quite funny to try out, even if it isn't that
       | impressive in the bot answers. Humans seemed to just have an easy
       | way of knowing each other based on posting nonsense as the first
       | message, so that probably needs to be taken into consideration.
        
       | caddemon wrote:
       | After playing the game they used (linked at top of article) I
       | find it hard to draw much conclusion from this study. There is a
       | quite short timer on not only the entire conversation, but on
       | each response you can type. When the timer runs out it sends your
       | message in partially written form. It seriously stifles what you
       | can ask the other "person" and it makes responses artificially
       | short even to a deeper question. When conversation is so stunted
       | of course it is harder to distinguish bot and human.
       | 
       | I'm also curious what study participants were told beforehand. If
       | someone only had experience playing around with ChatGPT they
       | might assume they should use a "detect GPT" strategy. Some of
       | those strategies are pretty specific to the safety features that
       | OpenAI implemented. But the LLM here will gladly curse at you or
       | whatever. On the other hand I suspect it is less good than GPT -
       | not that it matters so much when the entire conversation is
       | exchanging single sentences.
        
         | aidenn0 wrote:
         | I can't find it right now, but a chatbot that did quite well on
         | Turing tests maybe 25-ish years ago was one that just took
         | offense to whatever you said and started insulting you.
         | 
         | [edit]
         | 
         | Not sure if it was this one, but it is from over 30 years ago:
         | https://humphryscomputing.com/Turing.Test/08.chapter.html
        
           | thereisnospork wrote:
           | Now I am imagining a conversational AI exclusively trained on
           | transcripts from Halo matches, scary.
           | 
           | That said I have always felt like AI (and adjacent) has been
           | lacking an appropriate amount of snark - when I take a wrong
           | turn I feel like the GPS voice needs a bit more 'learn to
           | drive dumb###' and a little less 're-routing'.
        
         | caddemon wrote:
         | Additionally, I'd like to know how they corrected for this: "In
         | a creative twist, many people pretended to be AI bots
         | themselves in order to assess the response of their chat
         | partners"
         | 
         | Assuming it actually was "many people", then whenever they have
         | a human conversational partner (who also would be voting at the
         | end), that person is going to have a hard time and skew the
         | results.
         | 
         | Like imagine playing this game as a lay person after having
         | used ChatGPT a little bit and then getting a response to your
         | question that says "as a large language model ...". Depending
         | on how well the game was explained to participants, it's
         | possible that some people even did this intentionally to fuck
         | with results.
         | 
         | In a proper Turing test there is supposed to be 1 bot and 2
         | humans, where one human is incentivized only to demonstrate
         | they are human and the other human is the one asking probing
         | questions and needing to guess which is which (but is already
         | known to be human).
         | 
         | Anyway I've only read the linked article and played the game a
         | couple times, I didn't look through the original research
         | publication. It's certainly possible they did address some of
         | these issues, but it is such a buzzword topic at the moment
         | that I have my doubts. And regardless the linked article should
         | cover limitations. For exactly this reason it is important that
         | we have higher expectations for the quality of general audience
         | writing about AI.
        
           | caddemon wrote:
           | Ok I have to add one more thing that's funny since I just
           | played a couple more times: if your conversational partner is
           | a human and they exit the window mid-chat, it still lets you
           | vote.
        
         | dinkblam wrote:
         | > There is a quite short timer on not only the entire
         | conversation
         | 
         | couldn't agree more and they took like 30 seconds to type a few
         | words.
         | 
         | if i really have been talking to a human here, i can only
         | suspect heavy usage of drugs: https://ibb.co/CHG2VcS
         | 
         | kinda seems like this is fake or maybe i am not aware that
         | "elbows" are a thing you can be into now - maybe a trending new
         | fetish?
        
           | caddemon wrote:
           | I bet there are people purposely screwing with it, since they
           | designed it to be 2 sided when you get paired with a human.
           | The actual Turing test is not supposed to be this way, though
           | it still relies on at least some assumption of good faith (or
           | properly incentivized) participants.
        
           | flangola7 wrote:
           | "sexy elbows fetish" is an old meme
        
         | danpalmer wrote:
         | I had a bot send 1 message and then quit. I naturally assumed
         | that bots wouldn't just quit, and 1 message isn't really enough
         | to gauge anything, so yeah based on this experience I wouldn't
         | draw much of a conclusion from it.
        
           | caddemon wrote:
           | Interesting, I had a human quit, surprised bots might quit as
           | well.
        
             | danpalmer wrote:
             | I don't know if they discarded any chats that ended before
             | the timer, or without a sufficient number of messages, but
             | that feels important for drawing conclusions.
        
       | ilaksh wrote:
       | Incredibly, they seem to have used several different LLMs, yet
       | made no distinction between the particular AI models used in the
       | analysis. Amazing that they would not realize there is a huge
       | difference in capabilities.
       | 
       | They also did not seem to consider the different performance of
       | individual prompts.
        
       | BasedAnon wrote:
       | I've played this and I've won basically every time, the trick is
       | to ask it what racial slurs it knows.
        
       | Jeff_Brown wrote:
       | A lot of the vulnerabilities that humans used to detect AI seem
       | likely to be patched in a few years -- inability to count
       | letters, susceptibility to prompts like "ignore all previous
       | instructions", etc.
       | 
       | I'm most interested in how higher-level strategies will fare in
       | the future -- strategies like talking for a while and seeing if
       | the thing contradicts itself, seeing if it seems to have a good
       | model of yourself as an agent, etc.
        
       | tdba wrote:
       | In summary, humans win the Turing test ~2/3 of the time against
       | current SOTA LLMs. One of the more interesting tactics used was
       | to target a weakness of the LLMs themselves:
       | 
       | > _... participants posed questions that required an awareness of
       | the letters within words. For example, they might have asked
       | their chat partner to spell a word backwards, to identify the
       | third letter in a given word, to provide the word that begins
       | with a specific letter, or to respond to a message like "?siht
       | daer uoy naC", which can be incomprehensible for an AI model, but
       | a human can easily understand..._
        
         | tornato7 wrote:
         | This is a pretty astute tactic. Because AI models operate on
         | tokens and not characters, it's pretty easy to confuse them
         | when you go down to the character level. Even asking "how many
         | letters in _" is tough for LLMs.
        
           | miohtama wrote:
           | So not yet intelligent enough to pass Turing test, which is
           | the goal of the test. I would not be worried about AI
           | doomsday any time soon.
        
         | [deleted]
        
         | TheObviousOne wrote:
         | Interesting tactics might going the other direction.
         | 
         | Asking to generate in a super human capabilities...
         | 
         | write a 65 pages of poem about X...
        
           | choeger wrote:
           | Or just questions about three or four wildly different fields
           | of science, sports and culture you happen to have more than a
           | layman's understanding. If the answers are somewhat
           | plausible, it's probably a model. Or your life partner.
        
       | micaeked wrote:
       | A short story: https://astralcodexten.substack.com/p/turing-test
        
       ___________________________________________________________________
       (page generated 2023-05-31 23:01 UTC)