[HN Gopher] GPT4 gets a 0 on Steven Landsbrug's undergrad econ exam
___________________________________________________________________
GPT4 gets a 0 on Steven Landsbrug's undergrad econ exam
Author : DantesKite
Score : 33 points
Date : 2023-04-12 17:57 UTC (5 hours ago)
(HTM) web link (twitter.com)
(TXT) w3m dump (twitter.com)
| SilasX wrote:
| This should really link directly to Landsburg's post rather a
| Twitter promotion of it:
|
| https://www.thebigquestions.com/2023/04/05/gpt-4-fails-econo...
|
| Edit: Also, the post says GPT-4 scored 4/90, a little better than
| 0.
| hgsgm wrote:
| Landsburg's test has intentionally misleading questions with
| ambiguous wording, including two he admits had mistakes per his
| own standard for himself. He then graded exceptionally harshly.
| I don't believe his students average 75% on the rubric he used
| for ChatGPT.
|
| I believe he made his mind up before he started the experiment.
| Our_Benefactors wrote:
| This was my impression as well. When reading his nitpicky
| reasons a question was scored _zero_ out of nine (the same as
| leaving it blank!) made me think that sitting this course
| would have me hating his guts by the end of it.
| carlmr wrote:
| It reminds me of one of my professors, actually also in
| economics. He would grade in a similar way.
|
| Basically he had some very extreme views, based on terribly
| simplified economic models, and would want you paraphrasing
| what he said, ideally quoting.
|
| The closer you could match his wording and extreme views
| the better you scored.
| usaar333 wrote:
| Ya, these questions seem too hard for an undergrad test or at
| least feel like they were only written in an age where you
| want to be adversarial to LLMs.
|
| Don't get me wrong - I can give simple tests that LLMs fail
| and a college kid could easily pass -- but I don't think I'd
| come up with them unless I wanted to be adversarial to LLMs.
|
| Example simple question that GPT4/Bard/etc. fail hard:
|
| > The moose is behind the cat. The dog is in front of the
| moose. The moose is in front of the rat. The rat is behind
| the cat. The elk is in front of the dog.
|
| What's a possible animal ordering?
| stuff4ben wrote:
| I'd expect most economists couldn't do a good job either. It's
| black-box magic, suitable for wizards.
| phendrenad2 wrote:
| "Most economists couldn't pass an undergrad econ exam" is
| probably untrue, right?
| [deleted]
| hackrnusr wrote:
| Because it couldn't explain corporate executive compensation
| levels?
| Ancalagon wrote:
| Hard to contrast this with the report of gpt-4 getting a B in
| quantum computing yesterday
| usaar333 wrote:
| Aaronson's test seems to require less "fluid intelligence" to
| get right. Someone who has read the book and just pattern
| matches basically can do well on the test. A lot of tests work
| this way.
|
| Steve's test requires.. more fluid intelligence. Question 2 is
| simply not something you see on the internet in any common
| form, but the if you do lots of deduction of the implications
| -- ya, you get it.
|
| I've noticed similar weird fails with GPT-4. It seems like a
| genius recalling factual information and even (to some degree)
| translating common word problems into math. It starts failing
| hard though when you give it things way outside its domain
| (often children could outperform it since humans seem to not
| need so many shots of novel knowledge to understand things)
| TheCoelacanth wrote:
| That's putting it nicely. Aaronson's test was more focused on
| shamelessly parroting the professor's ideological views which
| were strongly hinted at in the question.
|
| GPT-4 indeed was better than the real students at that task
| since it lacks a sense of shame.
|
| This one seems more focused on actual application of economic
| principles.
| maister wrote:
| Anyone had a look at those questions?
|
| > Question 2: In a country where everyone is identical, 100
| people wait in line each day to buy raspberries at a controlled
| price. The government has decided to hand out free coffee to the
| people standing in line. The coffee costs the government $1 per
| cup, but the people in line value that coffee at only 75 cents
| per cup. What is the social cost of providing the coffee?
|
| > Grading Remarks: This answer completely misses the key fact
| that free coffee will cause the line to get longer (in fact it
| must cause the line to get longer, given the stated assumption
| that everyone is identical, hence initially indifferent between
| standing in line and not standing in line). In fact (unless one
| assumes a very small population), the line must grow until the
| extra waiting time completely dissipates the value of the free
| coffee; thus the social cost of providing the coffee is $100.
|
| Wait what? So the line must get longer, *given the stated
| assumption that everyone is identical*. Well what about the given
| stated assumption *that 100 people wait in line each day*? One of
| the stated assumptions can be dropped just like that, while the
| other is treated like a law of nature? Also where is the logic in
| "everyone is identical, hence initially indifferent between
| standing in line and not standing in line"? Am I missing
| something?
| fabbari wrote:
| Think about it these terms: originally there are 100 people in
| the line because the 101st will not queue since the time to get
| to the beginning of the line is more valuable than the
| raspberries at the controlled price. If nothing changes there
| will not be more people in that queue.
|
| When the government gives free coffee to the people in line now
| the perceived value is the value of the raspberries at the
| controlled price plus the $ .75 of the coffee, making it worth
| to queue more - until the queue gets to the length at which the
| waiting time exceeds the value of the added coffee.
| skeaker wrote:
| If your first paragraph was actually in the question then I
| would agree, but in the actual question, the reason there are
| exactly 100 people in line is that the question says there
| are 100 people in line. Phrasing certain parameters in your
| question as though they are absolutes and then expecting the
| person answering the question to change those parameters is
| like the worst kind of trick question. It's like when a
| toddler asks you if a ninja or a pirate would win in a fight,
| then says the pirate would win because he secretly has super
| strength. It's cute when a toddler does it because they
| innocently don't understand why their question wouldn't
| logically convey that hidden parameter to the listener, but
| inexcusable for a professor.
| mumrik wrote:
| Before the coffee, the equilibrium was for the line to be 100
| persons long.
|
| After the introduction of the coffee, the equilibrium changes
| so that people are willing to wait in a longer line than 100
| people, although the individuals themselves don't have any
| preferences (like A only standing in line if it's less than 20
| people and so on).
|
| I think not much thought was given to explain the answers in
| depth since the main point was to show where GPTs reasoning was
| lacking.
| carlmr wrote:
| >After the introduction of the coffee, the equilibrium
| changes so that people are willing to wait in a longer line
| than 100 people
|
| I do agree with OP that this has to be stated as part of the
| question. Otherwise you can't expect anyone to give the
| stated answer.
|
| Basically any answer is correct.
| maister wrote:
| > I think not much thought was given to explain the answers
| in depth since the main point was to show where GPTs
| reasoning was lacking.
|
| Funnily enough it seems as if the prof has also fallen for a
| reasoning error then. How can both be equally true? The line
| gets longer than 100 people and the social cost of the coffee
| is exactly 100 dollar (1 Dollar for every person).
| adrianmonk wrote:
| The point is that it's a badly-worded question because there
| are two ways to interpret it.
|
| It's clear (after reading the score's explanation) what the
| intended interpretation is, which is the one you're
| describing.
|
| The other interpretation is that the length of the line is
| set by the government based on how many people they choose to
| to allow to buy raspberries at a controlled price. Once 100
| people are in line, they won't let a 101st person get in
| line, so it is always 100 people. In this interpretation, a
| cup of coffee that would normally be worth $1.00 is only
| valued at $0.75 because what the people are there for is
| raspberries, not coffee. Some of them may not even like
| coffee that much.
| sverona wrote:
| Honestly this doesn't surprise me. I don't have GPT-4 access, but
| I fed ChatGPT a simple calculus question as a sanity check.
|
| Question: Calculate the integral of x arctan x dx. Show your
| work.
|
| Answer: We can solve this integral using integration by parts:
|
| Let u = arctan x and dv = x dx, then we have:
|
| du/dx = 1/(1+x^2) v = (1/2)x^2
|
| Using the integration by parts formula, we get:
|
| [?] x arctan x dx = uv - [?] v du/dx dx = (1/2)x^2 arctan x -
| (1/2)[?] x^2/(1+x^2) dx
|
| To solve the second integral, we can use substitution. Letting u
| = 1+x^2, then du/dx = 2x, and dx = du/(2x). Substituting into the
| integral, we get:
|
| [?] x^2/(1+x^2) dx = (1/2) [?] du/u = (1/2) ln|u| + C = (1/2)
| ln|1+x^2| + C
|
| Putting it all together, we get:
|
| [?] x arctan x dx = (1/2)x^2 arctan x - (1/2) [(1/2) ln|1+x^2|] +
| C = (1/2)x^2 arctan x - (1/4) ln|1+x^2| + C
|
| Therefore, the integral of x arctan x dx is (1/2)x^2 arctan x -
| (1/4) ln|1+x^2| + C, where C is the constant of integration.
|
| ---
|
| This is wrong, although it gets pretty damn close (see if you can
| spot the mistake.) I was an exceptionally kind grader when I
| taught and Landsburg seems like an exceptionally picky one. But
| either way ChatGPT wouldn't pass calculus 2, either.
| MuffinFlavored wrote:
| I think the real power will be when LLMs are trained to
| formulate the response "math detected, let me attach a low
| score to anything other than 'pipe user input into wolfram
| alpha and then use output'"
|
| Same for "let me call my XYZ plugin to do XYZ specialized task"
| to fill the gaps where LLMs fall short (most places when it
| comes to accuracy?). once ChatGPT can access the web/do basic
| curl/wget requests, then maybe visually inspect an image or a
| PDF, it's game on for "ok, AI is serious now".
| hgsgm wrote:
| Did it lose an x in the middle?
| JCharante wrote:
| Would you consider it fair if ChatGPT answered using the
| WolframAlpha plugin?
|
| https://user-images.githubusercontent.com/13973198/231566582...
| sverona wrote:
| Would you let your students use Mathematica on such an exam?
| skeaker wrote:
| If it was built into their consciousness, then yeah
| hgsgm wrote:
| Can you try it again without WA?
|
| I'm curious how often ChatGPT makes mistakes because its
| temperature makes it sometimes randomly give wrong answers
| _intentionally_ , to be generatively creative.
___________________________________________________________________
(page generated 2023-04-12 23:03 UTC)