[HN Gopher] Superhuman performance of an LLM on the reasoning ta...
___________________________________________________________________
Superhuman performance of an LLM on the reasoning tasks of a
physician
Author : amichail
Score : 21 points
Date : 2025-05-29 20:44 UTC (2 hours ago)
(HTM) web link (arxiv.org)
(TXT) w3m dump (arxiv.org)
| spwa4 wrote:
| I don't understand what the aim is here. LLMs have disadvantages
| compared to human doctors that make them a really, really bad
| option.
|
| 1) they can't take measurements themselves.
|
| 2) they don't adapt on the job. Illnesses do. In other words, if
| there is a contagious health emergency, an LLM would see the
| patients ... and ignore the emergency.
|
| 3) they are very bad at figuring out if a patient is lying to
| them (which is a required skill: combined with 2, people would
| figure out how to get the LLM to prescribe them morfine and ...)
|
| 4) they are generally socially problematic. A big part of being a
| doctor is gently convincing a patient their slightly painful toe
| does not in fact justify a diagnosis of bone cancer ... WITHOUT
| doing tests (that would be unethical, as there's zero chance of
| those tests yielding positive results)
|
| 5) they will not adapt to people. LLMs will not adapt, people
| will. This means patients will exploit LLMs to achieve a whole
| bunch of aims (like getting drugs, getting days off, getting free
| hospital stays, ...) and it doesn't matter how good LLMs are. An
| adaptive system vs a non-adaptive system ... it's a matter of
| time.
|
| 6) they are not themselves patients. This is a fundamental
| problem: it will be very hard for an LLM to collect new
| information about "the human condition" and new problems it may
| generate. There's many examples of this, from patients drinking
| radium solution (it lights up in the dark, so surely, it must
| give extra energy, right? Even sexual energy, right?) to rivers
| or ponds that turn out to have serious diseases lurking around.
| Meaning a doctor needs to be able to make the decision to go
| after problems in society when society finds a new,
| catastrophically dumb, way to hurt itself.
|
| Now you might say "but they would still be good in the developing
| world, wouldn't they?". Yes, but as the tuberculosis vaccine
| efforts sadly showed: the developing world is developing
| partially because they invest nothing whatsoever in (poor)
| people's health. Nothing. Zero. Rien. Which means, making health
| services cheaper (e.g. providing a cheap tuberculosis vaccine)
| ... has the problem that it does not increase the value of zero.
| They won't pay for healthcare ... and they won't pay for cheaper
| healthcare. And while Bill Gates ad the US government do pay for
| a bit of this, they're not sustainable solutions. If, however,
| you train a local with basic medical skills, there's a lot they
| can do for free, which _actually_ helps.
| timschmidt wrote:
| 3 and 4 can be highly problematic behaviors in doctors.
| Patients who have real medical issues are often ignored,
| scolded, or otherwise denied treatment because of a doctor's
| perception.
| gpt5 wrote:
| There are two very interesting results here:
|
| 1. ChatGPT O1 significantly outperformed any combination of
| Doctor + Resources (Median score of 86% vs 34%-42% of doctors).
| Hence superhuman results (at least compared against average
| physicians)
|
| 2. ChatGPT + Doctor performs worse than just ChatGPT alone.
|
| This means that the situation is getting similar to Chess - where
| adding Magnus Carlsen as a helper to Stockfish (a strong open
| source chess enginer)) could only make Stockfish worse.
| inopinatus wrote:
| Thoroughly proves that with careful prompt engineering, you too
| can ask for more funding for your next paper.
___________________________________________________________________
(page generated 2025-05-29 23:00 UTC)