[HN Gopher] From OpenAPI spec to MCP: How we built Xata's MCP se...
___________________________________________________________________
From OpenAPI spec to MCP: How we built Xata's MCP server
Author : tudorg
Score : 36 points
Date : 2025-05-25 08:41 UTC (2 days ago)
(HTM) web link (xata.io)
(TXT) w3m dump (xata.io)
| _pdp_ wrote:
| I mean there are 2 other posts related to data exfiltration
| attacks against MCP severs on the main page of HN at the time of
| this comment - at this point I think you want to involve a
| security person to make sure it is not vulnerable to stupid
| things.
| Atotalnoob wrote:
| The MCP attacks are really just due to bad token scoping.
|
| If you allow Y to do X, if an attacker takes control of Y, of
| course they can do X.
| wild_egg wrote:
| Can you elaborate on "bad token scoping"?
|
| I don't think your XY phrasing fully describes the GitHub MCP
| exploit and curious if you think that's somehow a "token
| scoping" issue.
| fkyoureadthedoc wrote:
| I'm unaware of the GitHub MCP "exploit", but given the
| overall state of LLM/MCP security FUD, there's probably
| some self promotion blog post from a security company about
| an LLM doing something stupid with GitHub data that the
| owner of the LLM using system didn't intend.
|
| For example, let's say I create an application that lets
| you chat with my open source repo. I set up my LLM with a
| GitHub tool. I don't want to think about oauth and getting
| a token from the end user, so I give it a PAT that I
| generated from my account. I'm even more lazy so I just
| used a PAT I already had laying around, and it
| unfortunately had read/write access to SSH keys. The user
| can add their ssh key to my account and do malicious
| things.
|
| Oh no, MCP is super vulnerable, please buy my LLM security
| product.
|
| If you give the LLM a tool, and you give the LLM input from
| a user, the user has access to that tool. That shrimple.
| wild_egg wrote:
| https://news.ycombinator.com/item?id=44097390
|
| Also currently on the front page. It's mainly that this
| tool hits the trifecta of having privileged access,
| untrusted inputs, and ability to exfiltrate. Most tools
| only do 1-2 of those so attacks need to be more
| sophisticated to coordinate that.
| truemotive wrote:
| GitLab Duo got hit with an oopsie, "AI agent runs with same
| privilege to site content as the authenticated user" kinda
| oopsie where you could just exfiltrate private repo information
| via a pixel gif.
|
| I knew it would get bad, but this bad already? I yearn for
| rigor haha
| alooPotato wrote:
| i really dont get why we cant just feed the openapi spec to the
| LLM instead of having this intermediate MCP representation. Don't
| really buy the whole 'the api docs will overwhelm an LLM" - that
| hasn't been my experience.
| wild_egg wrote:
| I haven't looked at MCP payloads properly to compare but often
| the raw OpenAPI spec is overly verbose and eats context space
| pretty quick.
|
| Really trivial to have the LLM first filter it down to the
| sections it cares about and then condense those sections
| though.
|
| Wrap that process in a small tool and give that to the LLM
| along with a `fetch` tool that handles credentials based on
| URLs and agent capabilities explode pretty rapidly.
| crystal_revenge wrote:
| I see this question frequently related to MCP, but I'm guessing
| these questions come from people who haven't built a lot of
| products using LLMs?
|
| Even if you're LLM could learn the openai spec, you still have
| to figure out how to concretely receive a response back. This
| is necessary for virtually any application build using an LLM
| and requires support for far, far more use cases than just
| calling an API.
|
| Consider the following use case: - You need to include some
| relevant contextual data from a local RAG system. - There are
| local functions that you want the model to be able to call -
| The API example you describe - You need to access data from a
| database
|
| In all of these cases, if you have experience working with
| LLMs, you've implemented some ad hoc template solution to pass
| the context into the model. You might have writing something
| like "Here is the info relevant to this task {{info}}" or
| "These are the tools you can use {{tools}}", but in each case
| you've had to craft a prompting solution specific to one
| problem.
|
| MCP solves this by making a generic interface to sending a wide
| range of information to the model to make use of. While the
| hype can be a bit much, it's a pretty good (minus the lack of
| foresight around security) and obvious solution to this current
| problem in AI Engineering.
| otabdeveloper4 wrote:
| Just ask the model to respond with JSON. Give it a template
| example response.
|
| You don't need a spec.
|
| For sending prompts to the LLM you will absolutely need to
| hand-craft custom prompts anyways, as each model responds
| slightly different.
| wild_egg wrote:
| > you still have to figure out how to concretely receive a
| response back
|
| Isn't that handled by whatever Tool API you're using? There's
| usually a `function_call_output` or `tool_result` message
| type. I haven't had a need for a separate protocol just to
| send responses.
| truemotive wrote:
| If you're working from OpenAPI, ideally you want to be able to
| process _any_ , potentially full of shit formatting spec file.
| I find that half the integrations I run into have some old
| weird version of Swagger, and the rest work like hell to stay
| up to date with the 3.x spec track.
|
| I agree, I wish, it will be a solved problem eventually. Just
| feeding a complex data model like that to the paper shredder
| that is the LLM, for making decisions about whether DELETE or
| POST is used is just asking for trouble.
| lmeyerov wrote:
| Slightly different experience here
|
| We have been adding MCP remote server to louie.ai, think a
| semantic layer over DBs for automating investigations, analytics,
| and viz over operational systems. MCP is nice so people can now
| use from Slack, VS Code, CLI, etc, without us building every
| single integration when they want to use it outside of our AI
| notebooks. And same starting point of openAPI spec, and even
| better, fastapi standard web framework for the REST layer.
|
| Using frameworks has been good. However, for chat ergonomics, we
| find we are defining custom tools, as talking directly to REST
| APIs is better than nothing, but that doesn't mean it's good. The
| tool layer isn't that fancy, but getting the ergonomics right
| matters, at least in our experience. Most of our time has been on
| security and ergonomics. (And for fun, we had an experiment of
| vibe coding this while hitting enterprise-level quality goals.)
___________________________________________________________________
(page generated 2025-05-27 23:01 UTC)