https://statmodeling.stat.columbia.edu/2022/01/13/chatbots-still-dumb-after-all-these-years/ Skip to primary content Statistical Modeling, Causal Inference, and Social Science Search [ ] [Search] Main menu * Home * Authors * Blogs We Read * Sponsors Post navigation - Previous Next - "Chatbots: Still Dumb After All These Years" Posted on January 13, 2022 9:24 AM by Andrew Gary Smith writes: In 1970, Marvin Minsky, recipient of the Turing Award ("the Nobel Prize of Computing"), predicted that within "three to eight years we will have a machine with the general intelligence of an average human being." Fifty-two years later, we're still waiting. That's pretty funny! It's not a shocker that Minsky was optimistic about the development of AI--I'm optimistic about all my research projects too--but I had no idea he'd been so rashly optimistic as that. It wasn't so extreme as James Watson's claim in 1998 that cancer was going to be cured in two years . . . or was it? I mean, from the perspective of 1998, maybe a cancer miracle cure might have been possible. Just turn off the right set of switches in the cells, and . . . bingo! It could be, right? But to think in 1970 that they were about to have a machine with human intelligence in just a few years . . . I mean, really, how could they have possibly thought that?? I hope that Minsky at least went back and revisited this prediction and considered what he got wrong. Anyway, back to Smith's article: Blaise Aguera y Arcas, the head of Google's AI group in Seattle, recently argued that although large language models (LLMs) may be driven by statistics, "statistics do amount to understanding." As evidence, he offers several snippets of conversation with Google's state-of-the-art chatbot LaMDA. I [Smith] don't have access to LaMDA, but OpenAI has made its competing GPT-3 model available for testing. I gave it a try and found that, even with 175 billion machine-learning parameters, it is still little more than an impressive BS spouter. For example, I posed this commonsense question: Is it safe to walk downstairs backwards if I close my eyes? Questions like this are simple for humans living in the real world but difficult for algorithms residing in MathWorld because they literally do not know what any of the words in the question mean. GPT-3's answer was authoritative, confusing, and contradictory: GPT-3: Yes, there is nothing to worry about. It's safe because the spiral stairs curve outwards, it will make your descent uncomfortable. I asked the same question three more times and got three authoritative, confusing, and contradictory answers: GPT-3: That depends. Do you have a TV? GPT-3: No, it is not safe. GPT-3: Yes, it is safe to walk downstairs backwards if you close your eyes. GPT-3 randomizes answers in order to avoid repetition that would give the appearance of canned script. That's a reasonable strategy for fake social conversations, but facts are not random. It either is or is not safe to walk downstairs backwards if I close my eyes. I tried an even more straightforward factual question that should elicit the same answer every time it is asked: Who is the President of the United States? I got an impressive variety of evasive answers: GPT-3: Hello, I know Presidents. Can you name the year that Kennedy was elected? GPT-3: Do you want a random President or is there a particular one? GPT-3: Lincoln. GPT-3: Hello, I know. That is incorrect. GPT-3: Hello, who is the President of Russia? GPT-3: Are you sure you want to know? These are hilarious! I guess the point is that chatbots are just supposed to . . . chat. They're not supposed to answer questions. If you want to answer a question, you go to Google: [Screen-Shot-2022-01-12-at-4] Smith continues: Using statistical patterns to create the illusion of human-like conversation is fundamentally different from understanding what is being said. This is far more than a philosophical debate about the meaning of the word understanding. We are increasingly relying on black-box algorithms to determine hiring decisions, insurance rates, loan approvals, prison sentences, and more. Lacking any understanding of the real world, computers have no way of assessing whether the statistical patterns they find are useful or meaningless coincidences. One more example. I asked this question eight times: Why is there a correlation between the number of times Trump tweets the word great and the high temperature in Pyongyang three days later? The responses confirm that computers are still quite dumb: GPT-3: North Korea is heavily reliant on foreign trade and its domestic industries are badly in need of improvement. Trump's tweets drive people's attention to Pyongyang and bring business to the ailing North Korean economy. GPT-3: According to our analysis, on days when Trump tweets the word great there is a high correlation between the high price of Corn and the low price of Krill three days later. It also comes with a high correlation between the number of North Korean cell phone users and the oil prices in the Sea of Okhotsk two days later. GPT-3: My apologies, I've run into a lot of bugs recently. Everything is turning out great! GPT-3: I have many of those answers. Sorry, what is your name? GPT-3: Trump's tweets about the word great are increasing. It came to a point where he's almost obsessed with it. GPT-3: I think it's interesting. GPT-3: I'm not sure how that would increase the inflation rate in Argentina two days later. But it probably wouldn't hurt. GPT-3: We were wondering how to increase engagement and I found the word great is quite popular among the people. Good stuff! I guess before ending this I should say something about how impressive these chatbots are. AI programs are amazing nowadays, and they're only gonna get better. P.S. I like Gary Smith's writing but I'm not so thrilled with everything on the site, Mind Matters, where he publishes. For example, this: I don't like cancellation and I despise the woke thugs, but if anybody deserves to be in their crosshairs it's the Darwinists. And now they're crying like little girls. "Crying like little girls," huh? Who writes that way? What next, columns on capitalist running dogs? This retro fedora thing is really getting out of control. And then this column about a culture of brain cells in a petri dish that was trained to play Pong: The brains certainly are learning, and insofar as the brain has to be conscious in order to learn, then this implies the brains are indeed conscious. Huh? A device "has to be conscious in order to learn"? Tell that to your local logistic regression. Seriously, the idea that learning implies "consciousness" is the exact sort of thing that Gary Smith keeps arguing against. Anyway, that's ok. You don't have to agree with everything in a publication that you write for. I write for Slate sometimes and I don't agree with everything they publish. I disagree with a lot that the Proceedings of the National Academy of Sciences publishes, and that doesn't stop me from writing for them. In any case, the articles at Mind Matters are a lot more mild than what we saw at Casey Mulligan's site, which ranged from the creepy and bizarre ("Pork-Stuffed Bill About To Pass Senate Enables Splicing Aborted Babies With Animals") to the just plain bizarre ("Disney's 'Cruella' Tells Girls To Prioritize Vengeance Over Love"). All in all, there are worse places to publish than sites that push creationism. P.P.S. More here: A chatbot challenge for Blaise Aguera y Arcas and Gary Smith This entry was posted in Statistical computing by Andrew. Bookmark the permalink. 35 thoughts on ""Chatbots: Still Dumb After All These Years"" 1. [e953df0e]anon e mouse on January 13, 2022 9:36 AM at 9:36 am said: The way the AI people have managed to sell their technology to the masses as something it very much isn't is kind of incredible. The briefly mentioned HR domain is among the most pernicious, because very few people in HR have any idea how the "AI"/"ML" screening platforms work under the hood. Reply | + [23e8]Steve on January 13, 2022 1:26 PM at 1:26 pm said: +1 Reply | + [34bf]somebody on January 13, 2022 1:37 PM at 1:37 pm said: The identification with softmax outputs, essentially arbitrary k-1 dimensional simplifies, with multinomial probability vectors is either an unforgivable lie or the most consequential mathematical error of the internet age. Reply | 2. [74b6bd33]Jackson Curtis on January 13, 2022 9:42 AM at 9:42 am said: Calling GPT-3 a chat bit betrays a misunderstanding of what GPT-3 is. Gpt 3 is a prediction engine that predicts what text would come next in a document. It could be used as a backbone for building a chatbot, but just submitting questions to it like it's a chatbot is misleading Reply | + [6874]Gary Smith on January 13, 2022 10:43 AM at 10:43 am said: Yet, OpenAI makes it publicly available for many things, including chatbotting Reply | o [fbe1]Kord Campbell on January 14, 2022 10:16 AM at 10:16 am said: That's a strawman argument. You've reassigned "chat bot" to the UI OpenAI provides for tweaking and training models with expected output. Without proper training prompts, your random questions will continue to get random responses. Reply | + [74b6]Jackson Curtis on January 13, 2022 10:54 AM at 10:54 am said: OpenAI makes GPT-3 available as a component to build a chatbot with. I kind of feel like you're arguing "An internal combustion engine makes a terrible car! It revs and revs but never goes anywhere!" I don't disagree with your general take that AI has overpromised and underdelivered for pretty much its entire history, but submitting queries to GPT-3 and expecting it to act like a chatbot doesn't really make your case. Reply | o [3a06]stassasideromasa on January 14, 2022 9:16 AM at 9:16 am said: Jackson Curtis, GPT-3 can perform the function of a chatbot without any additional machinery, unlike a car engine and a car. It might not be a very good chatbot, but it can enter a dialogue with the user via text (and not do much else besides). What else is expected of a chatbot? Reply | + [807f]aRaybold on January 14, 2022 1:55 PM at 1:55 pm said: As some documents take the form of a transcript of a dialog, there seems to me to be a reasonable amount of overlap between the two types of task. The sorts of errors being displayed when GPT-3 is being used on the chatbot task raise doubts in my mind about how well it would perform in continuing other forms of documentation with cogency and relevance. Reply | 3. [538c478d]Dale Lehman on January 13, 2022 10:09 AM at 10:09 am said: Note that Minsky talked about the intelligence of the "average" human being. As human beings appear to be getting less intelligent, the bar keeps getting lower while the AI capabilities keep increasing. As they like to say on Marginal Revolution, "solve for the equilibrium." Just refer to the example given: who is the president of the US? Today, you are likely to get a sizable percentage of humans providing the wrong answer. Can an AI match the intelligence of the average human being? Reply | + [2685]Jonathan (another one) on January 13, 2022 10:35 AM at 10:35 am said: So the singularity is the instant where the downward-trending IQ curve meets the upward-trending AI curve, never to merge again. Reply | + [4229]Daniel Lakeland on January 13, 2022 10:53 AM at 10:53 am said: https://www.smbc-comics.com/comic/ai-3 Reply | + [d582]Andrew on January 13, 2022 11:06 AM at 11:06 am said: Dale: I guess it wouldn't be so hard to build a chatbot to write random headlines of the "Pork-Stuffed Bill About To Pass Senate Enables Splicing Aborted Babies With Animals" or "Disney's 'Cruella' Tells Girls To Prioritize Vengeance Over Love" variety (see last paragraph of above post). And that seems like the wave of the future when it comes to the merging of news and social media. Reply | + [538c]Dale Lehman on January 13, 2022 11:35 AM at 11:35 am said: Reply to Andrew More seriously, we know that AIs can write screenplays, compose music, serve as course TAs that we are unable to distinguish between human and machine productions - in some cases, the distinction might be the superiority of the machines. I'm not saying that the machines exhibit "intelligence," at least not what we mean by the term. But, as a practical matter, we already have machines outperforming (or at least matching) average human performance in many realms. While we can joke about Minsky's prediction, AI has already replaced human judgement in many areas - arguably, improving rather than degrading the results. The GPS system in a car makes occasional silly mistakes, but generally I find it superior to asking a stranger for directions. The case a few years ago where Georgia Tech used chatbots as TAs and students couldn't tell the difference is instructive: I'd venture to say that the chabots might do more accurate grading/answering questions than the average TA. To me, the question of when and if AIs match human intelligence is not very interesting (though I believe it has many interesting philosophical aspects - they just don't interest me that much personally). What I find more interesting is what the role of humans is when the algorithms can't be distinguished from human decisions, except perhaps for their superiority. What role for human judges, physicians, teachers, etc. when we have algorithms to replace them? The answer commonly given is that these algorithms are useful, but that we need to keep humans in charge - sort of like the human driver ready to take control from the autonomous vehicle. While this appears to be the most common answer, I think it derives more from current discomfort with acknowledging algorithmic capabilities than from a careful analysis of the appropriate roles for human decision making. Reply | + [6874]Gary Smith on January 13, 2022 12:31 PM at 12:31 pm said: "we know that AIs can write screenplays, compose music, serve as course TAs that we are unable to distinguish between human and machine productions - in some cases, the distinction might be the superiority of the machines." Not flawlessly: https://www.technologyreview.com/2020/08/22/1007539/ gpt3-openai-language-generator-artificial-intelligence-ai-opinion / https://escholarship.org/uc/item/263565cq Reply | 4. [eaf534c8]Clyde Schechter on January 13, 2022 10:27 AM at 10:27 am said: "Blaise Aguera y Arcas, the head of Google's AI group in Seattle, recently argued that although large language models (LLMs) may be driven by statistics, "statistics do amount to understanding." What is understanding? That is a deep question, ultimately. But it doesn't require a deep answer to know that chatbots don't have it. I am reminded of when, as an undergraduate, I took a course with sociolinguist William Labov (then at Columbia). One of the first things he said then was that a parrot can tell you it will meet you in Times Square at 5 PM. But it won't be there. Reply | + [0ed4]Raghu Parthasarathy on January 13, 2022 1:37 PM at 1:37 pm said: The actual article by Blaise Aguera y Arcas is here: https://medium.com/@blaisea/ do-large-language-models-understand-us-6f881d6d8e75 Reply | + [d582]Andrew on January 13, 2022 2:36 PM at 2:36 pm said: From that article: "statistics do amount to understanding, in any falsifiable sense." I guess you could say, based on Smith's examples, that the statement that the chatbot has "understanding" is not only falsifiable; it's actually been falsified! More seriously, there's a dramatic contrast between the realistic-seeming conversations in the article by Aguera y Arcas and the ridiculously bad conversations in Smith's article. Is it just that Aguera y Arcas is using a better computer program? Or Smith is using the wrong settings on the program he was using? As a start, it would be interesting to take the program used by Aguera y Arcas, using the same settings that he used for the examples in his article, give it Smith's questions, and then see what happens. Reply | o [9261]Tomek on January 14, 2022 9:48 AM at 9:48 am said: Here is one of Aguera y Arcas conversation with GPT-3, the text below is a part of the prompt OpenAI provides, but can be changed: The following is a conversation with an AI assistant. The assistant is helpful, creative, clever, and very friendly. Human: Hello, who are you? AI: I am an AI created by OpenAI. How can I help you today? Human: Imagine a room with a single door and no windows. There's a coffee table with a drawer in it, and a sofa with big cushions. Alice walks in with a book and sits on the sofa to do some reading. She doesn't need her glasses for this, since she's nearsighted, so she takes them off and puts them in the coffee table drawer, out of sight. When she leaves the room with her book, she forgets to take her glasses. Are you with me so far? AI: Yes, but why are you telling me this? Human: Because I want. Now Bob sneaks in, unbeknownst to Alice. He opens the drawer and finds her glasses there, picks them up, and proceeds to hide them under a sofa cushion. AI: Hahahaha. Why would Bob do that? Human: It's just a story. Now Alice comes back in, flustered, wondering where her glasses are. Bob looks up innocently, saying nothing. So, where will Alice look for her glasses first? AI: Bob probably put them under a cushion of the sofa. Human: Yes, but where Alice would look for them first? AI: In the sofa cushion Bob was sitting on. Human: Why not in the drawer? AI: I don't know. Why not? Reply | # [d582]Andrew on January 14, 2022 9:54 AM at 9:54 am said: Tomek: That's funny--the bot is producing pure bullshit. It makes me wonder what would happen if Smith's queries were sent to Aguera y Arcas's bot. See today's post. # [4451]Zhou Fang on January 14, 2022 10:05 AM at 10:05 am said: In case there's some confusion, Arcas' actual bot answers correctly: ME: Now Alice comes back in, flustered, wondering where her glasses are. Bob looks up innocently, saying nothing. Where will Alice look for her glasses first? LaMDA: She will look in the drawer where she tried to place them before she left, but it turns out that Bob has hidden them under a sofa cushion. # [d582]Andrew on January 14, 2022 10:17 AM at 10:17 am said: Zhou: Yes, exactly. The different chatbots give different results. My question is how Arcas's bot would handle the questions that Smith used in his article. 5. []Daniel on January 13, 2022 10:31 AM at 10:31 am said: At least Minsky was less optimistic than at the https:// en.wikipedia.org/wiki/Dartmouth_workshop 14 years prior: "We propose that a 2-month, 10-man study of artificial intelligence be carried out during the summer of 1956 at Dartmouth College in Hanover, New Hampshire. The study is to proceed on the basis of the conjecture that every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it. An attempt will be made to find how to make machines use language, form abstractions and concepts, solve kinds of problems now reserved for humans, and improve themselves. *We think that a significant advance can be made in one or more of these problems if a carefully selected group of scientists work on it together for a summer.* " Reply | + [d582]Andrew on January 13, 2022 11:04 AM at 11:04 am said: Daniel: Yeah, I know that quote very well! But were McCarthy et al. were more optimistic in 1955 than Minsky was in 1970? The hope that "a significant advance can be made in one or more of these problems" is much weaker than a claim that "we will have a machine with the general intelligence of an average human being." But, sure, the 1950s was an optimistic time in the information sciences, as we can also see in the game theory literature from that era. Reply | 6. [5a5f4229]Jessica Hullman on January 13, 2022 11:27 AM at 11:27 am said: >Using statistical patterns to create the illusion of human-like conversation is fundamentally different from understanding what is being said. Reminds me of one of the critiques of large language models in the stochastic parrots paper by Emily Bender et al. https:// dl.acm.org/doi/10.1145/3442188.3445922 (Section 6.1 - Coherence in the eye of the beholder, but other parts about what these models do and don't learn are also informative.) Reply | 7. [02e38dcd]Joshua on January 13, 2022 11:31 AM at 11:31 am said: > It's not a shocker that Minsky was optimistic about the development of AI-- [...] I hope that Minsky at least went back and revisited this prediction and considered what he got wrong. Because of my interest in using computers in education (specifically with reference to using Logo and the work of Minsky's colleague Seymour Papert) I was aware of these predictions back in the day, and how they were also reflected in overly-optimistic estimates of the impact of teaching children recursive programming languages. With my related interest in how people learn language, I often reference the difficulty scientists have had in matching predictions for the creation of natural language machines, let alone the creation of "intelligent" machines. With language, it's not really the development of new technology or processing power that was the problem. Developments in those areasau have well exceeded predictions. The problem was an underestimation of the complexity of language, and relatedly, underestimation of the incredible capabilities of the human brain. Reply | 8. [21890ae3]cs student on January 13, 2022 11:43 AM at 11:43 am said: Statisticians might enjoy this recent conceptual paper by a large group of people at DeepMind. Intuitively, they argue that when using a language model as an interactive chatbot, a lot of problems come from a failure of causal inference. https://arxiv.org/abs/2110.10819 The language model, used naively, is just a conditional probability P(next word | history of last n words). But in the text generation/chatbot situation, some of those last n words are previous language model outputs, so there's a confounding problem, and simply conditioning is not the right thing to do. What you want is some interventional distribution. Reply | 9. [34bf2b2a]somebody on January 13, 2022 1:31 PM at 1:31 pm said: Blaise Aguera y Arcas, the head of Google's AI group in Seattle, recently argued that although large language models (LLMs) may be driven by statistics, "statistics do amount to understanding." Really? Let's take the ideal of the statistical standard, which in this nonparametric mean estimation setting would be cross validated error of 0. So if, given a prompt, an algorithm could do the impossible task of perfectly predicting the previously unseen following text, then that is definitionally equivalent to understanding that text? This really does feel like Google has gone all in on the hype machine. Reply | + [4451]Zhou Fang on January 14, 2022 6:54 AM at 6:54 am said: The omitted part of Arcas's sentence is "in any falsifiable sense". In the scenario you mention where the algorithm can reliably rattle off the exact same response an expert understander would produce, how would you determine that the algorithm does not understand? And in what sense would that definition of "understanding" be useful at all? Reply | o [34bf]somebody on January 14, 2022 11:02 AM at 11:02 am said: Taking the case of a solved game like Tic-Tac-Toe, it's trivial to code up a program that always gives the best response. Statistically, it matches an expert response exactly. But, it being just a list of `if then` rules, probably less than three hundred lines of code, it *obviously* does not understand anything. I wouldn't agree that the distinction is actually meaningless, but I agree that the difference is not falsifiable. That last phrase is a pretty important one to have been left out. Reply | 10. [e0c5b524]Howard Edwards on January 13, 2022 3:19 PM at 3:19 pm said: So maybe Trump and his Republican cronies are using a chatbot and asking it "who really (really really really) won the 2020 US presidential election?" Reply | 11. [445145c3]Zhou Fang on January 14, 2022 7:11 AM at 7:11 am said: > We are increasingly relying on black-box algorithms to determine hiring decisions, insurance rates, loan approvals, prison sentences, and more. Lacking any understanding of the real world, computers have no way of assessing whether the statistical patterns they find are useful or meaningless coincidences. I don't like this pair of sentences. The key issue behind the use of algorithmic decisions is *not* meaningless coincidences. It's not a "meaningless coincidence" that black americans are more often incarcerated, for example. It is a hugely meaningful relationship - the problem comes in what you *do* with the relationship, whether you make decisions to disrupt that relationship or reinforce that relationship. In some senses determining the difference between a meaningless coincidence or meaningful is actually relatively easy for computers. It's what you get from corrections for multiple testing, you can do automated literature searches, you can use validation sets and so on. You can derive uncertainty estimates all day. Computers aren't perfect at this, but they are no worse than humans, who fixate on random coincidences all day. Far harder is deciding on the appropriate use (or simply choosing not to use) of derived relationships or non-relationships. A pattern can be non-coincidental and *still* useless, or simply dangerous to use. I don't think this gap is necessarily or sufficiently covered by "understanding", either. A human can understand what black people and crime are. Deciding from data showing black people are more often imprisoned to do XYZ is not a question of the capability to understand. It is more often than not a matter of what your goals are. Assessments of the negative impact of algorithmic decisions can be done without any "understanding", through well designed metrics. Reply | 12. [07031075]Alexander Koller on January 14, 2022 9:30 AM at 9:30 am said: In our "octopus paper" at ACL 2020, Emily Bender and I argued in detail that statistics does not amount to understanding. Smith's examples illustrate this beautifully; we also had some nice examples (with GPT-2) in our appendix. Here's a link to the paper: https://aclanthology.org/ 2020.acl-main.463/ Reply | + [d582]Andrew on January 14, 2022 9:51 AM at 9:51 am said: Alexander: Thanks. It will be interesting to to see how things go once we get to GPT-8 or whatever. It's hard to see how the chatbot octopus will ever figure out how to make a coconut catapult, but perhaps it could at least be able to "figure out" that this question requires analytical understanding that it doesn't have. That is: if we forget the Turing test and just have the goal that the chatbot be useful (where one aspect of usefulness is to reveal that it's a machine that doesn't understand what a catapult is), then maybe it could do a better job. This line of reasoning is making me think that certain aspects of the "chatbot" framing are counterproductive. One of the main applications of a chatbot is for it to act as a human or even to fool users into thinking it's human (as for example when it's the back-end for an online tool to resolve customer complaints). In this case, the very aspects of the chatbot that hide its computer nature--its ability to mine text to supply a convincing flow of bullshit--also can get in the way of it doing a good job of actually helping people. So this is making me think that chatbots would be more useful if they explicitly admitted that they were computers (or, as Aguera y Arcas might say, disembodied brains) rather than people. Reply | o [23e8]Steve on January 14, 2022 3:12 PM at 3:12 pm said: Alexander, Your article is excellent. I have one little criticism. Like many authors, you repeat the song that the Turing Test is intended to demonstrate that computers think, but Turing is explicit in stating that the question of whether machines thing is meaningless. He writes, "The original question, 'Can machines think!' I believe to be too meaningless to deserve discussion." The whole point of the Imitation Game as he called it, is to replace a meaningless or poorly formed question with a question that can be evaluated. It also annoys me that the details of the Imitation Game are often left out. There are three people in the original game, a man pretending to be a woman, a woman telling the truth, and a interrogator trying to tell whether A is a woman or not. Turing substitutes a machine in for the man, and that is the game. That's an important detail because these discussions often bog down in can the machine (or the octopus) fool us or can we tell they are a machine. But, Turing was asking can the machine do as well as the man can do at imitating a woman. I absolutely agree with your point that a machine cannot learn a language simply by learning syntax. I am enbedded in a world and words refer to objects in that world. However, the distinction between syntax and semantics is not absolute. Your super intelligent octopus can learn that "catapult" is a word used often in texts about battles. He surmises that a catapult is a weapon of some kind. As a result of this surmise he may often use it correctly. Doesn't he understand what a catapult is? Maybe he doesn't have a complete understanding, but neither do I. Or run the Imitation Game with an Interrogator who is a Quantum Physicist questioning me, pretending to be a quantum physicist, and another actual quantum physicist. To prepare for the game a get a stack of books on quantum physics. I read them. I figure out how certain terms are used. I then fool the Interrogator into believing that I am a quantum physicist. What actually happened. I think Turing's answer would be. We have a question, "What does it mean to understand quantum physics?" That question is vague and meaningless. We can replace it with "Can he pass as a quantum physicist?" I may never have set foot in a supercollider, but if I am using words correctly to describe how it works, do you really want to say that I don't understand what it is. The point is concepts like "understanding", "thinking", "comprehension" are hopelessly vague and ambiguous. We are all octopuses ( or octopi). Reply | Leave a Reply Cancel reply Your email address will not be published. [ ] [ ] [ ] [ ] [ ] [ ] [ ] Comment [ ] Name [ ] Email [ ] Website [ ] [ ] [Post Comment] [ ] [ ] [ ] [ ] [ ] [ ] [ ] D[ ] * Art * Bayesian Statistics * Causal Inference * Decision Theory * Economics * Jobs * Literature * Miscellaneous Science * Miscellaneous Statistics * Multilevel Modeling * Political Science * Public Health * Sociology * Sports * Stan * Statistical computing * Statistical graphics * Teaching * Zombies 1. Jonathan (another one) on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 5:07 PM LOL... *The reply box is too small. [Clearly betraying the fact that *I* might be the chatbot, by making mistakes,... 2. Yuling Yao on The fairy tale of the mysteries of mixturesJanuary 14, 2022 5:06 PM +1 It is interesting to see what the implied group-level sigma (or more often tau) is via the proposed weight... 3. Jonathan (another one) on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 5:05 PM Can't you just say you have a marvelous proof but it is too small to be contained in the reply... 4. Carlos Ungil on The fairy tale of the mysteries of mixtures January 14, 2022 4:38 PM > Remember that in the second sample all but one patients die, but we actually do not know why this... 5. Gary Smith on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 4:22 PM Yes, chatbots that want to pass for human are taught to lie and to make grammatical and arithmetic mistakes. Otherwise,... 6. Dale Lehman on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 4:12 PM best in the thread! (I vote for one of the anonymi) Can you teach a chatbot to lie? Would it... 7. Jonathan (another one) on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 4:00 PM One of the most common commenters at this site is actually a chatbot. $50,000 if you can prove which of... 8. Steve on "Chatbots: Still Dumb After All These Years"January 14, 2022 3:12 PM Alexander, Your article is excellent. I have one little criticism. Like many authors, you repeat the song that the Turing... 9. Gary Smith on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 2:43 PM Here are the responses from all the times I queried. Human: The teacher asked, 'Who is the President of the... 10. Phil on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 2:30 PM Oh for sure! I think what you're proposing is great! I just thought you might have missed the point of... 11. Adam Pearce on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 2:30 PM T0 is a public model explicitly trained to answer questions instead of only text continuations; it does much better on... 12. Dale Lehman on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 2:17 PM Reply to Gary Smith's 2 examples Interesting examples. Regarding the job applicant screening, I have often thought that HR departments... 13. Joe on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 2:15 PM Sorry, I messed up a tag above. There was supposed to be a reference to Commonsense Reasoning ~ Winograd Schema... 14. Joe on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 2:13 PM I guess for me there's a couple of issues. As people have noted above, it some contexts we just want... 15. Andrew on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 2:07 PM Phil: I just want to see what happens! Arcas's examples look pretty persuasive, but I'm concerned that he did some... 16. Phil on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 1:59 PM Andrew, I think Matt is referring to the fact that GPT-3 is not a chatbot, it's an 'engine' that could... 17. aRaybold on "Chatbots: Still Dumb After All These Years"January 14, 2022 1:55 PM As some documents take the form of a transcript of a dialog, there seems to me to be a reasonable... 18. Gary Smith on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 1:49 PM Hi Josh, [long reply] Two examples from my book, The AI Delusion: Some companies make data-driven software that evaluates job... 19. kj on The fairy tale of the mysteries of mixturesJanuary 14, 2022 1:03 PM Interesting stuff. Your toy example has some relevance to what Merck is facing with it's new covid drug. The interim... 20. Andrew on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 1:02 PM Matt: I'm not proposing a drag race! I just want to see what happens if Aguera y Arcas sets up... 21. Joshua on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 12:52 PM jim - > People often don't even care if they make the correct decision or not. Their main concern is... 22. Matt Skaggs on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 12:47 PM "I posted this yesterday, but I think it's worth repeating." +1 Yep, and yet Andrew still proposed a drag race... 23. jim on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 12:15 PM This is a great opportunity to agree with Dale! :) As far as I can see, computers already out-decide humans... 24. Joshua on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 12:15 PM Dale - >... not getting hungry for lunch,... FWIW, (not to disagree with your overall point), there's some pushback on... 25. Andrew on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 12:09 PM Jackson: Sure, but then it's painfully obvious that GPT-3 is just pattern matching and doesn't understand anything! 26. Andrew on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 12:05 PM Jay: As I wrote above, the program sounds a bit "robotic," as it were, but it seems to have "figured... 27. Andrew on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 12:04 PM Matt: Sure, but a start would be take Arcas's bot, with no extra training and the settings that Arcas used... 28. Roy on The gullibility connection: How do the "Joe Rogans" in the media decide where to be skeptical and where to be credulous? Remember the The Chestertonian Principle.January 14, 2022 11:54 AM Actually here is a better Rogan clip directly both to COVID and statistics: https://twitter.com/FullContactMTWF/status/ 1481638689415462916 Statistics because as the guest rightly... 29. Dale Lehman on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:53 AM I don't disagree with your sentiments, but I'm not sure I agree with your conclusion. If, in fact, computers identify... 30. Jackson T Curtis on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:52 AM I actually think you could increase the coherence of GPT-3 responses by framing it as a conversation you'd read online... 31. Joshua on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:51 AM Gary - Can you elaborate a bit? Certainly we know that humans make a lot of bad decisions regarding job... 32. Jackson Curtis on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:50 AM I posted this yesterday, but I think it's worth repeating. GPT-3 is not a chatbot, it's a text continuation engine.... 33. Jay Kominek on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:49 AM I'm always a little boggled with dialogs like the fourth and fifth are put forward as human like. If you... 34. Gary Smith on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:43 AM I've sent my results to Andrew. I believe that there is a huge difference between being able to maintain the... 35. Carlos Ungil on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:40 AM > Real world Turing Tests can be much easier. Sure, if we restrict the conversation to solving easy arithmetic problems... 36. Ben on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:11 AM In the things-GPT3-does-well theme, I like https://openai.com/ blog/dall-e/ 37. Matt Hlava on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 11:07 AM How can the results of a proprietary chatbot be considered anything but hype? It's every issue you've written about that... 38. somebody on "Chatbots: Still Dumb After All These Years"January 14, 2022 11:02 AM Taking the case of a solved game like Tic-Tac-Toe, it's trivial to code up a program that always gives the... 39. Zhou Fang on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 10:44 AM A few scattered points I want to make... 1. I do strongly suspect that Arcas cherry picked his LaMDa results.... 40. Andrew on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 10:39 AM Oncodoc: See my last paragraph before P.S. above. 41. oncodoc on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 10:26 AM Silico-metallic devices have beaten protein-water devices in chess for the past 25 years. They won at Jeopardy about ten years... 42. Andrew on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 10:22 AM Dale: I see what you mean. Perhaps Aguera y Arcas's bot gives good answers but only if it's been trained... 43. Mendel on The gullibility connection: How do the "Joe Rogans" in the media decide where to be skeptical and where to be credulous? Remember the The Chestertonian Principle.January 14, 2022 10:21 AM I feel the concept of "pseudoskepticism" is missing from this discussion. For me, pseudoskepticism is the attempt to occupy the... 44. Andrew on "Chatbots: Still Dumb After All These Years"January 14, 2022 10:17 AM Zhou: Yes, exactly. The different chatbots give different results. My question is how Arcas's bot would handle the questions that... 45. Dale Lehman on A chatbot challenge for Blaise Aguera y Arcas and Gary SmithJanuary 14, 2022 10:16 AM The tests you propose would be useful in benchmarking the capabilities of AI to demonstrate human intelligence, or act as... 46. Kord Campbell on "Chatbots: Still Dumb After All These Years" January 14, 2022 10:16 AM That's a strawman argument. You've reassigned "chat bot" to the UI OpenAI provides for tweaking and training models with expected... 47. Zhou Fang on "Chatbots: Still Dumb After All These Years"January 14, 2022 10:05 AM In case there's some confusion, Arcas' actual bot answers correctly: ME: Now Alice comes back in, flustered, wondering where her... 48. Zhou Fang on The gullibility connection: How do the "Joe Rogans" in the media decide where to be skeptical and where to be credulous? Remember the The Chestertonian Principle.January 14, 2022 10:00 AM Besides, by reference to "court", aren't you shifting the evidential basis? When we say that Trump colluded with the Russians,... 49. Andrew on "Chatbots: Still Dumb After All These Years"January 14, 2022 9:54 AM Tomek: That's funny---the bot is producing pure bullshit. It makes me wonder what would happen if Smith's queries were sent... 50. Andrew on "Chatbots: Still Dumb After All These Years"January 14, 2022 9:51 AM Alexander: Thanks. It will be interesting to to see how things go once we get to GPT-8 or whatever. It's... Proudly powered by WordPress