[HN Gopher] The revenge of the data scientist
___________________________________________________________________
The revenge of the data scientist
Author : hamelsmu
Score : 63 points
Date : 2026-03-28 19:12 UTC (4 days ago)
(HTM) web link (hamel.dev)
(TXT) w3m dump (hamel.dev)
| jamesblonde wrote:
| I say this quite a lot to data scientists who are now building
| agents _:
|
| 1. think of the context data as training data for your requests
| (the LLM performs in-context learning based on your provided
| context data)
|
| 2. think of evals as test data to evaluate the performance of
| your agents. Collect them from agent traces and label them
| manually. If you want to "train" a LLM to act as a judge to label
| traces, then again, you will need lots of good quality examples
| (training data) as the LLM-as-a-Judge does in-context learning as
| well.
|
| _From my book - https://www.amazon.com/Building-Machine-Learning-
| Systems-Fea...
| pbronez wrote:
| Yup, agree. "Evaluations" = Tests
|
| Gets pretty meta when you're evaluating a model which needs to
| evaluate the output of another agent... gotta pin things down
| to ground truth somewhere.
| maxwg wrote:
| I can see cases like the recently mentioned pg_textsearch
| (https://news.ycombinator.com/item?id=47589856) being perfect
| cases for this kind of development style succeeding - where you
| have the clear test cases, benchmarks, etc you can meet.
|
| Though for greenfield development, writing the test cases (like
| the spec) is equally as hard, if not harder than writing the
| code.
|
| I also observe that LLMs tend to find themselves trapped in local
| minima. Once the codebase architecture has been solidified, very
| rarely will it consider larger refactors. In some ways - very
| similar to overfitting in ML
| djoldman wrote:
| > The bulk of the work is setting up experiments to test how well
| the AI generalizes to unseen data, debugging stochastic systems,
| and designing good metrics.
|
| In my experience, this is missing a big part of the work:
| confirming what the data actually is, sometimes despite what
| people think it is.
| uduni wrote:
| So true... I get more mileage from just watching an agent work
| than building sophisticated LLM-as-judge workflows
| convexly wrote:
| I mean it is a similar loop. Define what good looks like, measure
| how far off you are, iterate. I would say though that the people
| who've been doing that for years just have a head start that
| prompt engineers don't.
| Flashtoo wrote:
| These are good practices to keep in mind when setting up GenAI
| solutions, but I'm not convinced that this part of the job will
| allow "data scientist" as a profession to thrive. Here's my
| pessimistic take.
|
| Data scientists were appreciated largely because of their ability
| to create models that unlock business value. Model creation was a
| dark magic that you needed strong mathematical skills to perform
| - or at least that's the image, even if in reality you just slap
| XGBoost on a problem and call it a day. Data scientists were
| enablers and value creators.
|
| With GenAI, value creation is apparently done by the LLM provider
| and whoever in your company calls the API, which could really be
| any engineering team. Coaxing the right behavior out of the LLM
| is a bit of black magic in itself, but it's not something that
| requires deep mathematical knowledge. Knowing how gradients are
| calculated in a decoder-only transformer doesn't really help you
| make the LLM follow instructions. In fact, all your business
| stakeholders are constantly prompting chatbots themselves, so
| even if you provide some expertise here they will just see you as
| someone doing the same thing they do when they summarize an
| email.
|
| So that leaves the part the OP discusses: evaluation and
| monitoring. These are not sexy tasks and from the point of view
| of business stakeholders they are not the primary value add. In
| fact, they are barriers that get in the way of taking the POC
| someone slapped together in Copilot (it works!) and putting that
| solution in production. It's not even strictly necessary if you
| just want to move fast and break things. Appreciation for this
| kind of work is most present in large risk-averse companies, but
| even there it can be tricky to convince management that this is a
| job that needs to be done by a highly paid statistician with a
| graduate degree.
|
| What's the way forward? Convince management that people with the
| job title "data scientist" should be allowed to gatekeep building
| LLM solutions? Maybe I'm overestimating how good the average AI-
| aware software engineer is at this stuff, but I don't see the
| professional moat.
| redhale wrote:
| I agree with your take.
|
| I don't really see why evals are assumed to be exclusively in
| the domain of data scientists. In my experience SWEs-turned-AI
| Engineers are much better suited to building agents. Some
| struggle more than others, but "evals as automated tests" is,
| imo, so obvious a mental model, and can be so well adapted to
| by good SWEs, that data scientists have no real role on many
| "agent" projects.
|
| I'm not saying this is good or bad, just that it's what I'm
| observing in practice.
|
| For context, I'm a SWE-turned-AI Engineer, so I may be biased
| :)
| cdavid wrote:
| I agree. It is difficult to convince leadership to do this work
| at all ("it works on my example, ship it"), and in my
| experience most DS don't even want to do it.
|
| One of the key value is that it forces some thinking about what
| is the task you want to solve in the first way. In many cases,
| it is difficult if not impossible to do it, which implies the
| underlying product should not be built at all. But nobody wants
| to hear that.
|
| Doing eval only makes sense if making the product better
| impacts something the business cares about, which is very
| difficult to do in practice.
| daemonk wrote:
| I have a data science/engineering background. From my
| perspective, using AI is like mining the solution space for
| optimality. The solution space is the combinatorics of the
| billions of parameters and their cardinalities. You try to narrow
| down the search space with your prompt and hopefully guide your
| mining with more semantic-based heuristics towards your optimal
| solution.
|
| You might hit a local maxima or go down a blind path. I tend to
| completely start my code base from scratch every week. I would
| make things more generic, remove unnecessary complexity, or add
| new features. And hope that can move me past the local maxima.
| __mharrison__ wrote:
| I just spent yesterday applying Kaparthy's autoresearch on an ML
| problem.
|
| I teach ML for a living and was amazed with what the tokens gave
| back to me after many rounds of experiments. If Kaggle was still
| a thing, AI would generally beat it.
|
| The challenge I've seen is that most data science/ml modeling
| work is quite weak. Folks don't even know the basic tools well.
| Not sure if giving AI to them will really open up many doors to
| them.
|
| As always experts love minions of juniors doing their deeds. Non-
| experts get to wade through slop.
| twelfthnight wrote:
| I agree AI could probably do a decent job on Kaggle problems.
| Of course, almost no DS job is building models with well-
| defined objectives and perfect data. The DS and MLE folks I
| work with mostly spend their time reframing ill-posed product
| requests into ML systems that can be maintained and improved
| with feedback loops.
|
| A _huge_ part of a DS is saying "No" to bad ideas posed by non-
| experts. The issue with LLMs is all they ever say is "Yes" and
| "Wow, that's such a great idea!"
___________________________________________________________________
(page generated 2026-04-01 23:00 UTC)