[HN Gopher] GPT4 gets a 0 on Steven Landsbrug's undergrad econ exam
       ___________________________________________________________________
        
       GPT4 gets a 0 on Steven Landsbrug's undergrad econ exam
        
       Author : DantesKite
       Score  : 33 points
       Date   : 2023-04-12 17:57 UTC (5 hours ago)
        
 (HTM) web link (twitter.com)
 (TXT) w3m dump (twitter.com)
        
       | SilasX wrote:
       | This should really link directly to Landsburg's post rather a
       | Twitter promotion of it:
       | 
       | https://www.thebigquestions.com/2023/04/05/gpt-4-fails-econo...
       | 
       | Edit: Also, the post says GPT-4 scored 4/90, a little better than
       | 0.
        
         | hgsgm wrote:
         | Landsburg's test has intentionally misleading questions with
         | ambiguous wording, including two he admits had mistakes per his
         | own standard for himself. He then graded exceptionally harshly.
         | I don't believe his students average 75% on the rubric he used
         | for ChatGPT.
         | 
         | I believe he made his mind up before he started the experiment.
        
           | Our_Benefactors wrote:
           | This was my impression as well. When reading his nitpicky
           | reasons a question was scored _zero_ out of nine (the same as
           | leaving it blank!) made me think that sitting this course
           | would have me hating his guts by the end of it.
        
             | carlmr wrote:
             | It reminds me of one of my professors, actually also in
             | economics. He would grade in a similar way.
             | 
             | Basically he had some very extreme views, based on terribly
             | simplified economic models, and would want you paraphrasing
             | what he said, ideally quoting.
             | 
             | The closer you could match his wording and extreme views
             | the better you scored.
        
           | usaar333 wrote:
           | Ya, these questions seem too hard for an undergrad test or at
           | least feel like they were only written in an age where you
           | want to be adversarial to LLMs.
           | 
           | Don't get me wrong - I can give simple tests that LLMs fail
           | and a college kid could easily pass -- but I don't think I'd
           | come up with them unless I wanted to be adversarial to LLMs.
           | 
           | Example simple question that GPT4/Bard/etc. fail hard:
           | 
           | > The moose is behind the cat. The dog is in front of the
           | moose. The moose is in front of the rat. The rat is behind
           | the cat. The elk is in front of the dog.
           | 
           | What's a possible animal ordering?
        
       | stuff4ben wrote:
       | I'd expect most economists couldn't do a good job either. It's
       | black-box magic, suitable for wizards.
        
         | phendrenad2 wrote:
         | "Most economists couldn't pass an undergrad econ exam" is
         | probably untrue, right?
        
           | [deleted]
        
       | hackrnusr wrote:
       | Because it couldn't explain corporate executive compensation
       | levels?
        
       | Ancalagon wrote:
       | Hard to contrast this with the report of gpt-4 getting a B in
       | quantum computing yesterday
        
         | usaar333 wrote:
         | Aaronson's test seems to require less "fluid intelligence" to
         | get right. Someone who has read the book and just pattern
         | matches basically can do well on the test. A lot of tests work
         | this way.
         | 
         | Steve's test requires.. more fluid intelligence. Question 2 is
         | simply not something you see on the internet in any common
         | form, but the if you do lots of deduction of the implications
         | -- ya, you get it.
         | 
         | I've noticed similar weird fails with GPT-4. It seems like a
         | genius recalling factual information and even (to some degree)
         | translating common word problems into math. It starts failing
         | hard though when you give it things way outside its domain
         | (often children could outperform it since humans seem to not
         | need so many shots of novel knowledge to understand things)
        
           | TheCoelacanth wrote:
           | That's putting it nicely. Aaronson's test was more focused on
           | shamelessly parroting the professor's ideological views which
           | were strongly hinted at in the question.
           | 
           | GPT-4 indeed was better than the real students at that task
           | since it lacks a sense of shame.
           | 
           | This one seems more focused on actual application of economic
           | principles.
        
       | maister wrote:
       | Anyone had a look at those questions?
       | 
       | > Question 2: In a country where everyone is identical, 100
       | people wait in line each day to buy raspberries at a controlled
       | price. The government has decided to hand out free coffee to the
       | people standing in line. The coffee costs the government $1 per
       | cup, but the people in line value that coffee at only 75 cents
       | per cup. What is the social cost of providing the coffee?
       | 
       | > Grading Remarks: This answer completely misses the key fact
       | that free coffee will cause the line to get longer (in fact it
       | must cause the line to get longer, given the stated assumption
       | that everyone is identical, hence initially indifferent between
       | standing in line and not standing in line). In fact (unless one
       | assumes a very small population), the line must grow until the
       | extra waiting time completely dissipates the value of the free
       | coffee; thus the social cost of providing the coffee is $100.
       | 
       | Wait what? So the line must get longer, *given the stated
       | assumption that everyone is identical*. Well what about the given
       | stated assumption *that 100 people wait in line each day*? One of
       | the stated assumptions can be dropped just like that, while the
       | other is treated like a law of nature? Also where is the logic in
       | "everyone is identical, hence initially indifferent between
       | standing in line and not standing in line"? Am I missing
       | something?
        
         | fabbari wrote:
         | Think about it these terms: originally there are 100 people in
         | the line because the 101st will not queue since the time to get
         | to the beginning of the line is more valuable than the
         | raspberries at the controlled price. If nothing changes there
         | will not be more people in that queue.
         | 
         | When the government gives free coffee to the people in line now
         | the perceived value is the value of the raspberries at the
         | controlled price plus the $ .75 of the coffee, making it worth
         | to queue more - until the queue gets to the length at which the
         | waiting time exceeds the value of the added coffee.
        
           | skeaker wrote:
           | If your first paragraph was actually in the question then I
           | would agree, but in the actual question, the reason there are
           | exactly 100 people in line is that the question says there
           | are 100 people in line. Phrasing certain parameters in your
           | question as though they are absolutes and then expecting the
           | person answering the question to change those parameters is
           | like the worst kind of trick question. It's like when a
           | toddler asks you if a ninja or a pirate would win in a fight,
           | then says the pirate would win because he secretly has super
           | strength. It's cute when a toddler does it because they
           | innocently don't understand why their question wouldn't
           | logically convey that hidden parameter to the listener, but
           | inexcusable for a professor.
        
         | mumrik wrote:
         | Before the coffee, the equilibrium was for the line to be 100
         | persons long.
         | 
         | After the introduction of the coffee, the equilibrium changes
         | so that people are willing to wait in a longer line than 100
         | people, although the individuals themselves don't have any
         | preferences (like A only standing in line if it's less than 20
         | people and so on).
         | 
         | I think not much thought was given to explain the answers in
         | depth since the main point was to show where GPTs reasoning was
         | lacking.
        
           | carlmr wrote:
           | >After the introduction of the coffee, the equilibrium
           | changes so that people are willing to wait in a longer line
           | than 100 people
           | 
           | I do agree with OP that this has to be stated as part of the
           | question. Otherwise you can't expect anyone to give the
           | stated answer.
           | 
           | Basically any answer is correct.
        
           | maister wrote:
           | > I think not much thought was given to explain the answers
           | in depth since the main point was to show where GPTs
           | reasoning was lacking.
           | 
           | Funnily enough it seems as if the prof has also fallen for a
           | reasoning error then. How can both be equally true? The line
           | gets longer than 100 people and the social cost of the coffee
           | is exactly 100 dollar (1 Dollar for every person).
        
           | adrianmonk wrote:
           | The point is that it's a badly-worded question because there
           | are two ways to interpret it.
           | 
           | It's clear (after reading the score's explanation) what the
           | intended interpretation is, which is the one you're
           | describing.
           | 
           | The other interpretation is that the length of the line is
           | set by the government based on how many people they choose to
           | to allow to buy raspberries at a controlled price. Once 100
           | people are in line, they won't let a 101st person get in
           | line, so it is always 100 people. In this interpretation, a
           | cup of coffee that would normally be worth $1.00 is only
           | valued at $0.75 because what the people are there for is
           | raspberries, not coffee. Some of them may not even like
           | coffee that much.
        
       | sverona wrote:
       | Honestly this doesn't surprise me. I don't have GPT-4 access, but
       | I fed ChatGPT a simple calculus question as a sanity check.
       | 
       | Question: Calculate the integral of x arctan x dx. Show your
       | work.
       | 
       | Answer: We can solve this integral using integration by parts:
       | 
       | Let u = arctan x and dv = x dx, then we have:
       | 
       | du/dx = 1/(1+x^2) v = (1/2)x^2
       | 
       | Using the integration by parts formula, we get:
       | 
       | [?] x arctan x dx = uv - [?] v du/dx dx = (1/2)x^2 arctan x -
       | (1/2)[?] x^2/(1+x^2) dx
       | 
       | To solve the second integral, we can use substitution. Letting u
       | = 1+x^2, then du/dx = 2x, and dx = du/(2x). Substituting into the
       | integral, we get:
       | 
       | [?] x^2/(1+x^2) dx = (1/2) [?] du/u = (1/2) ln|u| + C = (1/2)
       | ln|1+x^2| + C
       | 
       | Putting it all together, we get:
       | 
       | [?] x arctan x dx = (1/2)x^2 arctan x - (1/2) [(1/2) ln|1+x^2|] +
       | C = (1/2)x^2 arctan x - (1/4) ln|1+x^2| + C
       | 
       | Therefore, the integral of x arctan x dx is (1/2)x^2 arctan x -
       | (1/4) ln|1+x^2| + C, where C is the constant of integration.
       | 
       | ---
       | 
       | This is wrong, although it gets pretty damn close (see if you can
       | spot the mistake.) I was an exceptionally kind grader when I
       | taught and Landsburg seems like an exceptionally picky one. But
       | either way ChatGPT wouldn't pass calculus 2, either.
        
         | MuffinFlavored wrote:
         | I think the real power will be when LLMs are trained to
         | formulate the response "math detected, let me attach a low
         | score to anything other than 'pipe user input into wolfram
         | alpha and then use output'"
         | 
         | Same for "let me call my XYZ plugin to do XYZ specialized task"
         | to fill the gaps where LLMs fall short (most places when it
         | comes to accuracy?). once ChatGPT can access the web/do basic
         | curl/wget requests, then maybe visually inspect an image or a
         | PDF, it's game on for "ok, AI is serious now".
        
         | hgsgm wrote:
         | Did it lose an x in the middle?
        
         | JCharante wrote:
         | Would you consider it fair if ChatGPT answered using the
         | WolframAlpha plugin?
         | 
         | https://user-images.githubusercontent.com/13973198/231566582...
        
           | sverona wrote:
           | Would you let your students use Mathematica on such an exam?
        
             | skeaker wrote:
             | If it was built into their consciousness, then yeah
        
           | hgsgm wrote:
           | Can you try it again without WA?
           | 
           | I'm curious how often ChatGPT makes mistakes because its
           | temperature makes it sometimes randomly give wrong answers
           | _intentionally_ , to be generatively creative.
        
       ___________________________________________________________________
       (page generated 2023-04-12 23:03 UTC)