[HN Gopher] Launch HN: Sentrial (YC W26) - Catch AI agent failur...
       ___________________________________________________________________
        
       Launch HN: Sentrial (YC W26) - Catch AI agent failures before your
       users do
        
       Hey HN! We're Neel and Anay, and we're building Sentrial
       (https://sentrial.com). It's production monitoring for AI products.
       We automatically detect failure patterns: loops, hallucinations,
       tool misuse, and user frustrations the moment they happen. When
       issues surface, Sentrial diagnoses the root cause by analyzing
       conversation patterns, model outputs, and tool interactions, then
       recommends specific fixes.  Here's a demo if you're interested:
       https://www.youtube.com/watch?v=cc4DWrJF7hk. When agents fail,
       choose wrong tools, or blow cost budgets, there's no way to know
       why - usually just logs and guesswork. As agents move from demos to
       production with real SLAs and real users, this is not sustainable.
       Neel and I lived this, building agents at SenseHQ and Accenture
       where we found that debugging agents was often harder than actually
       building them. Agents are untrustworthy in prod because there's no
       good infrastructure to verify what they're actually doing.  In
       practice this looks like: - A support agent that began
       misclassifying refund requests as product questions, which meant
       customers never reached the refund flow. - A document drafting
       agent that would occasionally hallucinate missing sections when
       parsing long specs, producing confident but incorrect outputs.
       There's no stack trace or 500 error and you only figure this out
       when a customer is angry.  We both realized teams were flying blind
       in production, and that agent native monitoring was going to be
       foundational infrastructure for every serious AI product. We
       started Sentrial as a verification layer designed to take care of
       this.  How it works: You wrap your client with our SDK in only a
       couple of lines. From there, we detect drift for you: - Wrong tool
       invocations - Misunderstood intents - Hallucinations - Quality
       regressions over time. You see it on our platform before a customer
       files a ticket.  There's a quick mcp set up, just give claude code:
       claude mcp add --transport http Sentrial
       https://www.sentrial.com/docs/mcp  We have a free tier (14 days, no
       credit card required). We'd love any feedback from anyone running
       agents whether they be for personal use or within a professional
       setting.  We'll be around in the comments!
        
       Author : anayrshukla
       Score  : 22 points
       Date   : 2026-03-11 16:24 UTC (6 hours ago)
        
 (HTM) web link (www.sentrial.com)
 (TXT) w3m dump (www.sentrial.com)
        
       | rajit wrote:
       | How do you identify "wrong tool" invocations (how is the "wrong
       | tool" defined)?
        
         | anayrshukla wrote:
         | Good question. We don't define "wrong tool" in some universal
         | way, because that really depends on the workflow.
         | 
         | What we do in practice is let the team mark a few tool calls as
         | right or wrong in context, then use that to learn the pattern
         | for that agent. From there, we can flag similar cases
         | automatically by looking at the convo state, the tool chosen,
         | the arguments, and what happened next.
         | 
         | So we're learning what "correct" looks like for your workflow
         | and then catching repeats of the same kind of mistake.
        
       | BoorishBears wrote:
       | I know your homepage isn't your business, but I'm bet Claude
       | could fix the janky horizontal overflow on mobile in a prompt.
       | Makes for a very distracting read
        
         | anayrshukla wrote:
         | Will fix ASAP.
        
           | _joel wrote:
           | There's some serious irony in this thread.
        
           | lpellis wrote:
           | The github link is also going to a 404.
           | 
           | I built a tool to check for these issues, was curious if it
           | would find it all, but yes.
           | 
           | https://pagewatch.ai/s-bm6jq1qs6y1x/b560hmfx/dashboard/previ.
           | ..
        
         | claudeomusic wrote:
         | Agreed - fix fast. No way to take a tool seriously about taking
         | care of production that has such a blatant production issue
        
       | mzelling wrote:
       | The landing page design reminds me of Perplexity's ad campaigns.
       | It's a clean look. I'd find your product more enticing if you
       | framed your offerings more around evaluation + automatic
       | optimization of production agents. There's real value there. The
       | current selling points -- trace sessions, track tool calls,
       | measure token usage, and calculate costs -- seem easily
       | implementable at home with a bit of vibe coding.
        
       | ZekiAI2026 wrote:
       | Interesting gap to explore: Sentrial catches drift and anomalies
       | -- failures that happen by accident. What's the defense against
       | failures that happen by design?
       | 
       | Prompt injection is the clearest example: an attacker embeds
       | instructions in content your agent processes. The agent does
       | exactly what it's told. No wrong tool invocations, no
       | hallucinations in the traditional sense -- just an agent
       | successfully executing injected instructions. From a monitoring
       | perspective it looks like normal operation.
       | 
       | Same with adversarial inputs crafted to stay inside your learned
       | "correct" patterns: tool calls are right, arguments are
       | plausible, outputs pass quality checks. The manipulation is in
       | what the agent was pointed at, not in how it behaved.
       | 
       | Curious whether your anomaly detection has a layer for
       | adversarial intent vs. operational drift, or whether that's
       | explicitly out of scope for now.
        
       ___________________________________________________________________
       (page generated 2026-03-11 23:01 UTC)