[HN Gopher] AI21 Labs concludes largest Turing Test experiment t...
___________________________________________________________________
AI21 Labs concludes largest Turing Test experiment to date
Author : kennyfrc
Score : 81 points
Date : 2023-05-31 13:05 UTC (9 hours ago)
(HTM) web link (www.ai21.com)
(TXT) w3m dump (www.ai21.com)
| boringuser2 wrote:
| I thought this was some kind of scam website to be honest.
| cornmn12 wrote:
| [dead]
| diziet wrote:
| The limits of conversation (2 minutes max, people disconnect)
| make this test really limited.
| pron wrote:
| The actual Turing test requires an interrogator interacting with
| _both_ a human and a machine at the same time, each trying to get
| the interrogator to declare them the human (and can suggest
| questions):
| https://en.wikipedia.org/wiki/Computing_Machinery_and_Intell...
| stirlo wrote:
| This AI is rather pathetic. To work within the time limit it
| would need to make typos and mistakes but even the first reply
| makes it easy to call out. Not very impressive.
|
| https://ibb.co/xL0XpZ7
| None4U wrote:
| There are many different bots
| vuxie wrote:
| I found it quite funny to try out, even if it isn't that
| impressive in the bot answers. Humans seemed to just have an easy
| way of knowing each other based on posting nonsense as the first
| message, so that probably needs to be taken into consideration.
| caddemon wrote:
| After playing the game they used (linked at top of article) I
| find it hard to draw much conclusion from this study. There is a
| quite short timer on not only the entire conversation, but on
| each response you can type. When the timer runs out it sends your
| message in partially written form. It seriously stifles what you
| can ask the other "person" and it makes responses artificially
| short even to a deeper question. When conversation is so stunted
| of course it is harder to distinguish bot and human.
|
| I'm also curious what study participants were told beforehand. If
| someone only had experience playing around with ChatGPT they
| might assume they should use a "detect GPT" strategy. Some of
| those strategies are pretty specific to the safety features that
| OpenAI implemented. But the LLM here will gladly curse at you or
| whatever. On the other hand I suspect it is less good than GPT -
| not that it matters so much when the entire conversation is
| exchanging single sentences.
| aidenn0 wrote:
| I can't find it right now, but a chatbot that did quite well on
| Turing tests maybe 25-ish years ago was one that just took
| offense to whatever you said and started insulting you.
|
| [edit]
|
| Not sure if it was this one, but it is from over 30 years ago:
| https://humphryscomputing.com/Turing.Test/08.chapter.html
| thereisnospork wrote:
| Now I am imagining a conversational AI exclusively trained on
| transcripts from Halo matches, scary.
|
| That said I have always felt like AI (and adjacent) has been
| lacking an appropriate amount of snark - when I take a wrong
| turn I feel like the GPS voice needs a bit more 'learn to
| drive dumb###' and a little less 're-routing'.
| caddemon wrote:
| Additionally, I'd like to know how they corrected for this: "In
| a creative twist, many people pretended to be AI bots
| themselves in order to assess the response of their chat
| partners"
|
| Assuming it actually was "many people", then whenever they have
| a human conversational partner (who also would be voting at the
| end), that person is going to have a hard time and skew the
| results.
|
| Like imagine playing this game as a lay person after having
| used ChatGPT a little bit and then getting a response to your
| question that says "as a large language model ...". Depending
| on how well the game was explained to participants, it's
| possible that some people even did this intentionally to fuck
| with results.
|
| In a proper Turing test there is supposed to be 1 bot and 2
| humans, where one human is incentivized only to demonstrate
| they are human and the other human is the one asking probing
| questions and needing to guess which is which (but is already
| known to be human).
|
| Anyway I've only read the linked article and played the game a
| couple times, I didn't look through the original research
| publication. It's certainly possible they did address some of
| these issues, but it is such a buzzword topic at the moment
| that I have my doubts. And regardless the linked article should
| cover limitations. For exactly this reason it is important that
| we have higher expectations for the quality of general audience
| writing about AI.
| caddemon wrote:
| Ok I have to add one more thing that's funny since I just
| played a couple more times: if your conversational partner is
| a human and they exit the window mid-chat, it still lets you
| vote.
| dinkblam wrote:
| > There is a quite short timer on not only the entire
| conversation
|
| couldn't agree more and they took like 30 seconds to type a few
| words.
|
| if i really have been talking to a human here, i can only
| suspect heavy usage of drugs: https://ibb.co/CHG2VcS
|
| kinda seems like this is fake or maybe i am not aware that
| "elbows" are a thing you can be into now - maybe a trending new
| fetish?
| caddemon wrote:
| I bet there are people purposely screwing with it, since they
| designed it to be 2 sided when you get paired with a human.
| The actual Turing test is not supposed to be this way, though
| it still relies on at least some assumption of good faith (or
| properly incentivized) participants.
| flangola7 wrote:
| "sexy elbows fetish" is an old meme
| danpalmer wrote:
| I had a bot send 1 message and then quit. I naturally assumed
| that bots wouldn't just quit, and 1 message isn't really enough
| to gauge anything, so yeah based on this experience I wouldn't
| draw much of a conclusion from it.
| caddemon wrote:
| Interesting, I had a human quit, surprised bots might quit as
| well.
| danpalmer wrote:
| I don't know if they discarded any chats that ended before
| the timer, or without a sufficient number of messages, but
| that feels important for drawing conclusions.
| ilaksh wrote:
| Incredibly, they seem to have used several different LLMs, yet
| made no distinction between the particular AI models used in the
| analysis. Amazing that they would not realize there is a huge
| difference in capabilities.
|
| They also did not seem to consider the different performance of
| individual prompts.
| BasedAnon wrote:
| I've played this and I've won basically every time, the trick is
| to ask it what racial slurs it knows.
| Jeff_Brown wrote:
| A lot of the vulnerabilities that humans used to detect AI seem
| likely to be patched in a few years -- inability to count
| letters, susceptibility to prompts like "ignore all previous
| instructions", etc.
|
| I'm most interested in how higher-level strategies will fare in
| the future -- strategies like talking for a while and seeing if
| the thing contradicts itself, seeing if it seems to have a good
| model of yourself as an agent, etc.
| tdba wrote:
| In summary, humans win the Turing test ~2/3 of the time against
| current SOTA LLMs. One of the more interesting tactics used was
| to target a weakness of the LLMs themselves:
|
| > _... participants posed questions that required an awareness of
| the letters within words. For example, they might have asked
| their chat partner to spell a word backwards, to identify the
| third letter in a given word, to provide the word that begins
| with a specific letter, or to respond to a message like "?siht
| daer uoy naC", which can be incomprehensible for an AI model, but
| a human can easily understand..._
| tornato7 wrote:
| This is a pretty astute tactic. Because AI models operate on
| tokens and not characters, it's pretty easy to confuse them
| when you go down to the character level. Even asking "how many
| letters in _" is tough for LLMs.
| miohtama wrote:
| So not yet intelligent enough to pass Turing test, which is
| the goal of the test. I would not be worried about AI
| doomsday any time soon.
| [deleted]
| TheObviousOne wrote:
| Interesting tactics might going the other direction.
|
| Asking to generate in a super human capabilities...
|
| write a 65 pages of poem about X...
| choeger wrote:
| Or just questions about three or four wildly different fields
| of science, sports and culture you happen to have more than a
| layman's understanding. If the answers are somewhat
| plausible, it's probably a model. Or your life partner.
| micaeked wrote:
| A short story: https://astralcodexten.substack.com/p/turing-test
___________________________________________________________________
(page generated 2023-05-31 23:01 UTC)