[HN Gopher] Taming LLMs: Using Executable Oracles to Prevent Bad...
       ___________________________________________________________________
        
       Taming LLMs: Using Executable Oracles to Prevent Bad Code
        
       Author : mad44
       Score  : 27 points
       Date   : 2026-03-26 17:51 UTC (5 hours ago)
        
 (HTM) web link (john.regehr.org)
 (TXT) w3m dump (john.regehr.org)
        
       | dktoao wrote:
       | "Our goal should be to give an LLM coding agent zero degrees of
       | freedom"
       | 
       | Wouldn't that just be called inventing a new language with all
       | the overhead of the languages we already have? Are we getting to
       | the point where getting LLMs to be productive and also write good
       | code is going to require so much overhead and additional
       | procedures and tools that we might as well write the code
       | ourselves. Hmmm...
        
         | seanw444 wrote:
         | Yeah, precision LLM coding is kind of an oxymoron. English
         | language -> codebase is essentially lossily-compressed logic by
         | definition. The less lossy the compression becomes, the more
         | you probably approach re-inventing programming languages. Which
         | then means that in order to use LLMs to code, you're accepting
         | some degree of imprecision.
        
         | virgilp wrote:
         | Actually, no. We always needed good checks - that's why you
         | have techniques like automated canary analysis, extensive
         | testing, checking for coverage - these are forms of "executable
         | oracles". If you wanted to be able to do continuous deployment
         | - you had to be very thorough in your validation.
         | 
         | LLMs just take this to the extreme. You can no longer rely on
         | human code reviews (well you can but you give away all the LLM
         | advantages) so then if you take out "human judgement" *from
         | validation*[1], you have to resort to very sophisticated
         | automated validation. This is it - it's not about "inventing a
         | new language", it's about being much more thorough (and
         | innovative, and efficient) in the validation process.
         | 
         | [1] never from design, or specification - you shouldn't
         | outsource that to AI, I don't think we're close to an AI that
         | can do that even moderately effective without human help.
        
           | nitwit005 wrote:
           | If the LLM generates code exactly matching a specification,
           | the specification becomes a conventional programing language.
           | The LLM is just transforming from one language to another.
        
             | sanxiyn wrote:
             | Yes, but a programming language with a proverbial
             | sufficiently smart compiler. That is very useful.
        
               | Quekid5 wrote:
               | Try writing an exhaustive spec for anything non-trivial
               | and you might see the problem.
        
         | amelius wrote:
         | Zero degrees of freedom is a step too far.
         | 
         | What you want is correctness preserving transformations. Add to
         | this some metrics such as code size, execution speed.
        
       | voxaai wrote:
       | ran into this with creative generation. for code, formal
       | constraints work great. but when the quality criteria cant be
       | typed (feels right for this audience, sounds like infrastructure
       | not a toy) constraints made things worse. what worked was
       | competing generators with different objectives, then rank against
       | the brief. the variance from competition was more useful than the
       | precision from constraints.
        
         | CrazyStat wrote:
         | Wow, a sudden change in writing style that's not at all
         | intended to disguise the fact that you're an llm!
        
       | RS-232 wrote:
       | Has anyone had success using 2 agents, with one as the creator
       | and one as an adversarial "reviewer"? Is the output usually
       | better or worse?
        
         | esafak wrote:
         | This is routine. We have Gemini (which is not our coding model)
         | review our PRs and it genuinely catches mistakes. Even using
         | the same model as the creator, without its context to bias it,
         | would probably catch many mistakes.
        
         | sanxiyn wrote:
         | That works well. Anthropic wrote a writeup on it.
         | 
         | https://www.anthropic.com/engineering/harness-design-long-ru...
        
         | mapontosevenths wrote:
         | This is how its meant to be done. Usually with the reviewer
         | being the stronger model.
         | 
         | That said, with both the test driven development this post
         | describes and the reviewer model (its best to do both) you have
         | to provide an escape hatch or out for the model. If you let the
         | model get inescapably stuck with an impossible test or
         | constraints it will just start deleting tests or rewriting the
         | entire codebase in rust or something.
         | 
         | My escape hatch is "expert advice". I let the weak LLM phone a
         | friend when its stuck and ask a smarter LLM for assistance. Its
         | since stopped going crazy and replacing all my tests with
         | gibberish... mostly.
        
       | ReptileMan wrote:
       | Now is Haskell's time to shine.
        
       ___________________________________________________________________
       (page generated 2026-03-26 23:00 UTC)