[HN Gopher] Taming LLMs: Using Executable Oracles to Prevent Bad...
___________________________________________________________________
Taming LLMs: Using Executable Oracles to Prevent Bad Code
Author : mad44
Score : 27 points
Date : 2026-03-26 17:51 UTC (5 hours ago)
(HTM) web link (john.regehr.org)
(TXT) w3m dump (john.regehr.org)
| dktoao wrote:
| "Our goal should be to give an LLM coding agent zero degrees of
| freedom"
|
| Wouldn't that just be called inventing a new language with all
| the overhead of the languages we already have? Are we getting to
| the point where getting LLMs to be productive and also write good
| code is going to require so much overhead and additional
| procedures and tools that we might as well write the code
| ourselves. Hmmm...
| seanw444 wrote:
| Yeah, precision LLM coding is kind of an oxymoron. English
| language -> codebase is essentially lossily-compressed logic by
| definition. The less lossy the compression becomes, the more
| you probably approach re-inventing programming languages. Which
| then means that in order to use LLMs to code, you're accepting
| some degree of imprecision.
| virgilp wrote:
| Actually, no. We always needed good checks - that's why you
| have techniques like automated canary analysis, extensive
| testing, checking for coverage - these are forms of "executable
| oracles". If you wanted to be able to do continuous deployment
| - you had to be very thorough in your validation.
|
| LLMs just take this to the extreme. You can no longer rely on
| human code reviews (well you can but you give away all the LLM
| advantages) so then if you take out "human judgement" *from
| validation*[1], you have to resort to very sophisticated
| automated validation. This is it - it's not about "inventing a
| new language", it's about being much more thorough (and
| innovative, and efficient) in the validation process.
|
| [1] never from design, or specification - you shouldn't
| outsource that to AI, I don't think we're close to an AI that
| can do that even moderately effective without human help.
| nitwit005 wrote:
| If the LLM generates code exactly matching a specification,
| the specification becomes a conventional programing language.
| The LLM is just transforming from one language to another.
| sanxiyn wrote:
| Yes, but a programming language with a proverbial
| sufficiently smart compiler. That is very useful.
| Quekid5 wrote:
| Try writing an exhaustive spec for anything non-trivial
| and you might see the problem.
| amelius wrote:
| Zero degrees of freedom is a step too far.
|
| What you want is correctness preserving transformations. Add to
| this some metrics such as code size, execution speed.
| voxaai wrote:
| ran into this with creative generation. for code, formal
| constraints work great. but when the quality criteria cant be
| typed (feels right for this audience, sounds like infrastructure
| not a toy) constraints made things worse. what worked was
| competing generators with different objectives, then rank against
| the brief. the variance from competition was more useful than the
| precision from constraints.
| CrazyStat wrote:
| Wow, a sudden change in writing style that's not at all
| intended to disguise the fact that you're an llm!
| RS-232 wrote:
| Has anyone had success using 2 agents, with one as the creator
| and one as an adversarial "reviewer"? Is the output usually
| better or worse?
| esafak wrote:
| This is routine. We have Gemini (which is not our coding model)
| review our PRs and it genuinely catches mistakes. Even using
| the same model as the creator, without its context to bias it,
| would probably catch many mistakes.
| sanxiyn wrote:
| That works well. Anthropic wrote a writeup on it.
|
| https://www.anthropic.com/engineering/harness-design-long-ru...
| mapontosevenths wrote:
| This is how its meant to be done. Usually with the reviewer
| being the stronger model.
|
| That said, with both the test driven development this post
| describes and the reviewer model (its best to do both) you have
| to provide an escape hatch or out for the model. If you let the
| model get inescapably stuck with an impossible test or
| constraints it will just start deleting tests or rewriting the
| entire codebase in rust or something.
|
| My escape hatch is "expert advice". I let the weak LLM phone a
| friend when its stuck and ask a smarter LLM for assistance. Its
| since stopped going crazy and replacing all my tests with
| gibberish... mostly.
| ReptileMan wrote:
| Now is Haskell's time to shine.
___________________________________________________________________
(page generated 2026-03-26 23:00 UTC)