[HN Gopher] Superhuman performance of an LLM on the reasoning ta...
       ___________________________________________________________________
        
       Superhuman performance of an LLM on the reasoning tasks of a
       physician
        
       Author : amichail
       Score  : 21 points
       Date   : 2025-05-29 20:44 UTC (2 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | spwa4 wrote:
       | I don't understand what the aim is here. LLMs have disadvantages
       | compared to human doctors that make them a really, really bad
       | option.
       | 
       | 1) they can't take measurements themselves.
       | 
       | 2) they don't adapt on the job. Illnesses do. In other words, if
       | there is a contagious health emergency, an LLM would see the
       | patients ... and ignore the emergency.
       | 
       | 3) they are very bad at figuring out if a patient is lying to
       | them (which is a required skill: combined with 2, people would
       | figure out how to get the LLM to prescribe them morfine and ...)
       | 
       | 4) they are generally socially problematic. A big part of being a
       | doctor is gently convincing a patient their slightly painful toe
       | does not in fact justify a diagnosis of bone cancer ... WITHOUT
       | doing tests (that would be unethical, as there's zero chance of
       | those tests yielding positive results)
       | 
       | 5) they will not adapt to people. LLMs will not adapt, people
       | will. This means patients will exploit LLMs to achieve a whole
       | bunch of aims (like getting drugs, getting days off, getting free
       | hospital stays, ...) and it doesn't matter how good LLMs are. An
       | adaptive system vs a non-adaptive system ... it's a matter of
       | time.
       | 
       | 6) they are not themselves patients. This is a fundamental
       | problem: it will be very hard for an LLM to collect new
       | information about "the human condition" and new problems it may
       | generate. There's many examples of this, from patients drinking
       | radium solution (it lights up in the dark, so surely, it must
       | give extra energy, right? Even sexual energy, right?) to rivers
       | or ponds that turn out to have serious diseases lurking around.
       | Meaning a doctor needs to be able to make the decision to go
       | after problems in society when society finds a new,
       | catastrophically dumb, way to hurt itself.
       | 
       | Now you might say "but they would still be good in the developing
       | world, wouldn't they?". Yes, but as the tuberculosis vaccine
       | efforts sadly showed: the developing world is developing
       | partially because they invest nothing whatsoever in (poor)
       | people's health. Nothing. Zero. Rien. Which means, making health
       | services cheaper (e.g. providing a cheap tuberculosis vaccine)
       | ... has the problem that it does not increase the value of zero.
       | They won't pay for healthcare ... and they won't pay for cheaper
       | healthcare. And while Bill Gates ad the US government do pay for
       | a bit of this, they're not sustainable solutions. If, however,
       | you train a local with basic medical skills, there's a lot they
       | can do for free, which _actually_ helps.
        
         | timschmidt wrote:
         | 3 and 4 can be highly problematic behaviors in doctors.
         | Patients who have real medical issues are often ignored,
         | scolded, or otherwise denied treatment because of a doctor's
         | perception.
        
       | gpt5 wrote:
       | There are two very interesting results here:
       | 
       | 1. ChatGPT O1 significantly outperformed any combination of
       | Doctor + Resources (Median score of 86% vs 34%-42% of doctors).
       | Hence superhuman results (at least compared against average
       | physicians)
       | 
       | 2. ChatGPT + Doctor performs worse than just ChatGPT alone.
       | 
       | This means that the situation is getting similar to Chess - where
       | adding Magnus Carlsen as a helper to Stockfish (a strong open
       | source chess enginer)) could only make Stockfish worse.
        
       | inopinatus wrote:
       | Thoroughly proves that with careful prompt engineering, you too
       | can ask for more funding for your next paper.
        
       ___________________________________________________________________
       (page generated 2025-05-29 23:00 UTC)