[HN Gopher] From OpenAPI spec to MCP: How we built Xata's MCP se...
       ___________________________________________________________________
        
       From OpenAPI spec to MCP: How we built Xata's MCP server
        
       Author : tudorg
       Score  : 36 points
       Date   : 2025-05-25 08:41 UTC (2 days ago)
        
 (HTM) web link (xata.io)
 (TXT) w3m dump (xata.io)
        
       | _pdp_ wrote:
       | I mean there are 2 other posts related to data exfiltration
       | attacks against MCP severs on the main page of HN at the time of
       | this comment - at this point I think you want to involve a
       | security person to make sure it is not vulnerable to stupid
       | things.
        
         | Atotalnoob wrote:
         | The MCP attacks are really just due to bad token scoping.
         | 
         | If you allow Y to do X, if an attacker takes control of Y, of
         | course they can do X.
        
           | wild_egg wrote:
           | Can you elaborate on "bad token scoping"?
           | 
           | I don't think your XY phrasing fully describes the GitHub MCP
           | exploit and curious if you think that's somehow a "token
           | scoping" issue.
        
             | fkyoureadthedoc wrote:
             | I'm unaware of the GitHub MCP "exploit", but given the
             | overall state of LLM/MCP security FUD, there's probably
             | some self promotion blog post from a security company about
             | an LLM doing something stupid with GitHub data that the
             | owner of the LLM using system didn't intend.
             | 
             | For example, let's say I create an application that lets
             | you chat with my open source repo. I set up my LLM with a
             | GitHub tool. I don't want to think about oauth and getting
             | a token from the end user, so I give it a PAT that I
             | generated from my account. I'm even more lazy so I just
             | used a PAT I already had laying around, and it
             | unfortunately had read/write access to SSH keys. The user
             | can add their ssh key to my account and do malicious
             | things.
             | 
             | Oh no, MCP is super vulnerable, please buy my LLM security
             | product.
             | 
             | If you give the LLM a tool, and you give the LLM input from
             | a user, the user has access to that tool. That shrimple.
        
               | wild_egg wrote:
               | https://news.ycombinator.com/item?id=44097390
               | 
               | Also currently on the front page. It's mainly that this
               | tool hits the trifecta of having privileged access,
               | untrusted inputs, and ability to exfiltrate. Most tools
               | only do 1-2 of those so attacks need to be more
               | sophisticated to coordinate that.
        
         | truemotive wrote:
         | GitLab Duo got hit with an oopsie, "AI agent runs with same
         | privilege to site content as the authenticated user" kinda
         | oopsie where you could just exfiltrate private repo information
         | via a pixel gif.
         | 
         | I knew it would get bad, but this bad already? I yearn for
         | rigor haha
        
       | alooPotato wrote:
       | i really dont get why we cant just feed the openapi spec to the
       | LLM instead of having this intermediate MCP representation. Don't
       | really buy the whole 'the api docs will overwhelm an LLM" - that
       | hasn't been my experience.
        
         | wild_egg wrote:
         | I haven't looked at MCP payloads properly to compare but often
         | the raw OpenAPI spec is overly verbose and eats context space
         | pretty quick.
         | 
         | Really trivial to have the LLM first filter it down to the
         | sections it cares about and then condense those sections
         | though.
         | 
         | Wrap that process in a small tool and give that to the LLM
         | along with a `fetch` tool that handles credentials based on
         | URLs and agent capabilities explode pretty rapidly.
        
         | crystal_revenge wrote:
         | I see this question frequently related to MCP, but I'm guessing
         | these questions come from people who haven't built a lot of
         | products using LLMs?
         | 
         | Even if you're LLM could learn the openai spec, you still have
         | to figure out how to concretely receive a response back. This
         | is necessary for virtually any application build using an LLM
         | and requires support for far, far more use cases than just
         | calling an API.
         | 
         | Consider the following use case: - You need to include some
         | relevant contextual data from a local RAG system. - There are
         | local functions that you want the model to be able to call -
         | The API example you describe - You need to access data from a
         | database
         | 
         | In all of these cases, if you have experience working with
         | LLMs, you've implemented some ad hoc template solution to pass
         | the context into the model. You might have writing something
         | like "Here is the info relevant to this task {{info}}" or
         | "These are the tools you can use {{tools}}", but in each case
         | you've had to craft a prompting solution specific to one
         | problem.
         | 
         | MCP solves this by making a generic interface to sending a wide
         | range of information to the model to make use of. While the
         | hype can be a bit much, it's a pretty good (minus the lack of
         | foresight around security) and obvious solution to this current
         | problem in AI Engineering.
        
           | otabdeveloper4 wrote:
           | Just ask the model to respond with JSON. Give it a template
           | example response.
           | 
           | You don't need a spec.
           | 
           | For sending prompts to the LLM you will absolutely need to
           | hand-craft custom prompts anyways, as each model responds
           | slightly different.
        
           | wild_egg wrote:
           | > you still have to figure out how to concretely receive a
           | response back
           | 
           | Isn't that handled by whatever Tool API you're using? There's
           | usually a `function_call_output` or `tool_result` message
           | type. I haven't had a need for a separate protocol just to
           | send responses.
        
         | truemotive wrote:
         | If you're working from OpenAPI, ideally you want to be able to
         | process _any_ , potentially full of shit formatting spec file.
         | I find that half the integrations I run into have some old
         | weird version of Swagger, and the rest work like hell to stay
         | up to date with the 3.x spec track.
         | 
         | I agree, I wish, it will be a solved problem eventually. Just
         | feeding a complex data model like that to the paper shredder
         | that is the LLM, for making decisions about whether DELETE or
         | POST is used is just asking for trouble.
        
       | lmeyerov wrote:
       | Slightly different experience here
       | 
       | We have been adding MCP remote server to louie.ai, think a
       | semantic layer over DBs for automating investigations, analytics,
       | and viz over operational systems. MCP is nice so people can now
       | use from Slack, VS Code, CLI, etc, without us building every
       | single integration when they want to use it outside of our AI
       | notebooks. And same starting point of openAPI spec, and even
       | better, fastapi standard web framework for the REST layer.
       | 
       | Using frameworks has been good. However, for chat ergonomics, we
       | find we are defining custom tools, as talking directly to REST
       | APIs is better than nothing, but that doesn't mean it's good. The
       | tool layer isn't that fancy, but getting the ergonomics right
       | matters, at least in our experience. Most of our time has been on
       | security and ergonomics. (And for fun, we had an experiment of
       | vibe coding this while hitting enterprise-level quality goals.)
        
       ___________________________________________________________________
       (page generated 2025-05-27 23:01 UTC)