[HN Gopher] The revenge of the data scientist
       ___________________________________________________________________
        
       The revenge of the data scientist
        
       Author : hamelsmu
       Score  : 63 points
       Date   : 2026-03-28 19:12 UTC (4 days ago)
        
 (HTM) web link (hamel.dev)
 (TXT) w3m dump (hamel.dev)
        
       | jamesblonde wrote:
       | I say this quite a lot to data scientists who are now building
       | agents _:
       | 
       | 1. think of the context data as training data for your requests
       | (the LLM performs in-context learning based on your provided
       | context data)
       | 
       | 2. think of evals as test data to evaluate the performance of
       | your agents. Collect them from agent traces and label them
       | manually. If you want to "train" a LLM to act as a judge to label
       | traces, then again, you will need lots of good quality examples
       | (training data) as the LLM-as-a-Judge does in-context learning as
       | well.
       | 
       | _From my book - https://www.amazon.com/Building-Machine-Learning-
       | Systems-Fea...
        
         | pbronez wrote:
         | Yup, agree. "Evaluations" = Tests
         | 
         | Gets pretty meta when you're evaluating a model which needs to
         | evaluate the output of another agent... gotta pin things down
         | to ground truth somewhere.
        
       | maxwg wrote:
       | I can see cases like the recently mentioned pg_textsearch
       | (https://news.ycombinator.com/item?id=47589856) being perfect
       | cases for this kind of development style succeeding - where you
       | have the clear test cases, benchmarks, etc you can meet.
       | 
       | Though for greenfield development, writing the test cases (like
       | the spec) is equally as hard, if not harder than writing the
       | code.
       | 
       | I also observe that LLMs tend to find themselves trapped in local
       | minima. Once the codebase architecture has been solidified, very
       | rarely will it consider larger refactors. In some ways - very
       | similar to overfitting in ML
        
       | djoldman wrote:
       | > The bulk of the work is setting up experiments to test how well
       | the AI generalizes to unseen data, debugging stochastic systems,
       | and designing good metrics.
       | 
       | In my experience, this is missing a big part of the work:
       | confirming what the data actually is, sometimes despite what
       | people think it is.
        
       | uduni wrote:
       | So true... I get more mileage from just watching an agent work
       | than building sophisticated LLM-as-judge workflows
        
       | convexly wrote:
       | I mean it is a similar loop. Define what good looks like, measure
       | how far off you are, iterate. I would say though that the people
       | who've been doing that for years just have a head start that
       | prompt engineers don't.
        
       | Flashtoo wrote:
       | These are good practices to keep in mind when setting up GenAI
       | solutions, but I'm not convinced that this part of the job will
       | allow "data scientist" as a profession to thrive. Here's my
       | pessimistic take.
       | 
       | Data scientists were appreciated largely because of their ability
       | to create models that unlock business value. Model creation was a
       | dark magic that you needed strong mathematical skills to perform
       | - or at least that's the image, even if in reality you just slap
       | XGBoost on a problem and call it a day. Data scientists were
       | enablers and value creators.
       | 
       | With GenAI, value creation is apparently done by the LLM provider
       | and whoever in your company calls the API, which could really be
       | any engineering team. Coaxing the right behavior out of the LLM
       | is a bit of black magic in itself, but it's not something that
       | requires deep mathematical knowledge. Knowing how gradients are
       | calculated in a decoder-only transformer doesn't really help you
       | make the LLM follow instructions. In fact, all your business
       | stakeholders are constantly prompting chatbots themselves, so
       | even if you provide some expertise here they will just see you as
       | someone doing the same thing they do when they summarize an
       | email.
       | 
       | So that leaves the part the OP discusses: evaluation and
       | monitoring. These are not sexy tasks and from the point of view
       | of business stakeholders they are not the primary value add. In
       | fact, they are barriers that get in the way of taking the POC
       | someone slapped together in Copilot (it works!) and putting that
       | solution in production. It's not even strictly necessary if you
       | just want to move fast and break things. Appreciation for this
       | kind of work is most present in large risk-averse companies, but
       | even there it can be tricky to convince management that this is a
       | job that needs to be done by a highly paid statistician with a
       | graduate degree.
       | 
       | What's the way forward? Convince management that people with the
       | job title "data scientist" should be allowed to gatekeep building
       | LLM solutions? Maybe I'm overestimating how good the average AI-
       | aware software engineer is at this stuff, but I don't see the
       | professional moat.
        
         | redhale wrote:
         | I agree with your take.
         | 
         | I don't really see why evals are assumed to be exclusively in
         | the domain of data scientists. In my experience SWEs-turned-AI
         | Engineers are much better suited to building agents. Some
         | struggle more than others, but "evals as automated tests" is,
         | imo, so obvious a mental model, and can be so well adapted to
         | by good SWEs, that data scientists have no real role on many
         | "agent" projects.
         | 
         | I'm not saying this is good or bad, just that it's what I'm
         | observing in practice.
         | 
         | For context, I'm a SWE-turned-AI Engineer, so I may be biased
         | :)
        
         | cdavid wrote:
         | I agree. It is difficult to convince leadership to do this work
         | at all ("it works on my example, ship it"), and in my
         | experience most DS don't even want to do it.
         | 
         | One of the key value is that it forces some thinking about what
         | is the task you want to solve in the first way. In many cases,
         | it is difficult if not impossible to do it, which implies the
         | underlying product should not be built at all. But nobody wants
         | to hear that.
         | 
         | Doing eval only makes sense if making the product better
         | impacts something the business cares about, which is very
         | difficult to do in practice.
        
       | daemonk wrote:
       | I have a data science/engineering background. From my
       | perspective, using AI is like mining the solution space for
       | optimality. The solution space is the combinatorics of the
       | billions of parameters and their cardinalities. You try to narrow
       | down the search space with your prompt and hopefully guide your
       | mining with more semantic-based heuristics towards your optimal
       | solution.
       | 
       | You might hit a local maxima or go down a blind path. I tend to
       | completely start my code base from scratch every week. I would
       | make things more generic, remove unnecessary complexity, or add
       | new features. And hope that can move me past the local maxima.
        
       | __mharrison__ wrote:
       | I just spent yesterday applying Kaparthy's autoresearch on an ML
       | problem.
       | 
       | I teach ML for a living and was amazed with what the tokens gave
       | back to me after many rounds of experiments. If Kaggle was still
       | a thing, AI would generally beat it.
       | 
       | The challenge I've seen is that most data science/ml modeling
       | work is quite weak. Folks don't even know the basic tools well.
       | Not sure if giving AI to them will really open up many doors to
       | them.
       | 
       | As always experts love minions of juniors doing their deeds. Non-
       | experts get to wade through slop.
        
         | twelfthnight wrote:
         | I agree AI could probably do a decent job on Kaggle problems.
         | Of course, almost no DS job is building models with well-
         | defined objectives and perfect data. The DS and MLE folks I
         | work with mostly spend their time reframing ill-posed product
         | requests into ML systems that can be maintained and improved
         | with feedback loops.
         | 
         | A _huge_ part of a DS is saying "No" to bad ideas posed by non-
         | experts. The issue with LLMs is all they ever say is "Yes" and
         | "Wow, that's such a great idea!"
        
       ___________________________________________________________________
       (page generated 2026-04-01 23:00 UTC)