[HN Gopher] Show HN: sllm - Split a GPU node with other develope...
       ___________________________________________________________________
        
       Show HN: sllm - Split a GPU node with other developers, unlimited
       tokens
        
       Running DeepSeek V3 (685B) requires 8xH100 GPUs which is about
       $14k/month. Most developers only need 15-25 tok/s. sllm lets you
       join a cohort of developers sharing a dedicated node. You reserve a
       spot with your card, and nobody is charged until the cohort fills.
       Prices start at $5/mo for smaller models.  The LLMs are completely
       private (we don't log any traffic).  The API is OpenAI-compatible
       (we run vLLM), so you just swap the base URL. Currently offering a
       few models.
        
       Author : jrandolf
       Score  : 97 points
       Date   : 2026-04-04 15:18 UTC (7 hours ago)
        
 (HTM) web link (sllm.cloud)
 (TXT) w3m dump (sllm.cloud)
        
       | mmargenot wrote:
       | This is a great idea! I saw a similar (inverse) idea the other
       | day for pooling compute (https://github.com/michaelneale/mesh-
       | llm). What are you doing for compute in the backend? Are you
       | locked into a cohort from month to month?
        
       | vova_hn2 wrote:
       | 1. Is the given tok/s estimate for the total node throughput, or
       | is it what you can realistically expect to get? Or is it the
       | worst case scenario throughput if everyone starts to use it
       | simultaneously?
       | 
       | 2. What if I try to hog all resources of a node by running some
       | large data processing and making multiple queries in parallel?
       | What if I try to resell the access by charging per token?
       | 
       | Edit: sorry if this comment sounds overly critical. I think that
       | pooling money with other developers to collectively rent a server
       | for LLM inference is a really cool idea. I also thought about it,
       | but haven't found a satisfactory answer to my question number 2,
       | so I decided that it is infeasible in practice.
        
         | jrandolf wrote:
         | 1. It's an average. 2. We have sophisticated rate limiter.
        
           | poly2it wrote:
           | Does it take user time zones into account?
        
             | jrandolf wrote:
             | Yes
        
       | esafak wrote:
       | Like vast.ai and TensorDock, and presumably others.
        
       | spuz wrote:
       | It seems crazy to me that the "Join" button does not have a price
       | on it and yet clicking it simply forwards you to a Stripe page
       | again with no price information on it. How am I supposed to know
       | how much I'm about to be charged?
        
         | jrandolf wrote:
         | That was an error on our part lol. We'll update with the price.
        
       | peter_d_sherman wrote:
       | What a brilliant idea!
       | 
       | Split a "it needs to run in a datacenter because its hardware
       | requirements are so large" AI/LLM across multiple people who each
       | want shared access to that particular model.
       | 
       | Sort of like the Real Estate equivalent of subletting, or
       | splitting a larger space into smaller spaces and subletting each
       | one...
       | 
       | Or, like the Web Host equivalent of splitting a single server
       | into multiple virtual machines for shared hosting by multiple
       | other parties, or what-have-you...
       | 
       |  _I could definitely see marketplaces similar to this, popping up
       | in the future!_
       | 
       | It seems like it should make AI cheaper for everyone... that is,
       | "democratize AI"... in a "more/better/faster/cheaper" way than AI
       | has been democratized to date...
       | 
       | Anyway, it's a brilliant idea!
       | 
       | Wishing you a lot of luck with this endeavor!
        
       | kaoD wrote:
       | How is the time sharing handled? I assume if I submit a unit of
       | work it will load to VRAM and then run (sharing time? how many
       | work units can run in parallel?)
       | 
       | How large is a full context window in MiB and how long does it
       | take to load the buffer? I.e. how many seconds should I expect my
       | worst case wait time to take until I get my first token?
        
         | ninjha wrote:
         | > how many work units can run in parallel
         | 
         | not original author but batching is one very important trick to
         | make inference efficient, you can reasonably do tens to low
         | hundreds in parallel (depending on model size and gpu size)
         | with very little performance overhead
        
         | jrandolf wrote:
         | vLLM handles GPU scheduling, not sllm. The model weights stay
         | resident in VRAM permanently so there's no loading/unloading
         | per request. vLLM uses continuous batching, so incoming
         | requests are dynamically added to the running batch every
         | decode step and the GPU is always working on multiple requests
         | simultaneously. There is no "load to VRAM and run" per request;
         | it's more like joining an already-running batch.
         | 
         | TTFT is under 2 seconds average. Worst case is 10-30s.
        
           | kaoD wrote:
           | > The model weights stay resident in VRAM permanently so
           | there's no loading/unloading per request.
           | 
           | Yes, I was thinking about context buffers, which I assume are
           | not small in large models. That has to be loaded into VRAM,
           | right?
           | 
           | If I keep sending large context buffers, will that hog the
           | batches?
        
             | jrandolf wrote:
             | Not if you are the only one. We have rate limits to prevent
             | this in case, idk, you share your key with 1000 people lol.
        
       | spuz wrote:
       | Is this not a more restricted version of OpenRouter? With
       | OpenRouter you pay for credits that can be used to run any
       | commercial or open-source model and you only pay for what you
       | use.
        
         | jrandolf wrote:
         | OpenRouter is a little different. We are trying to experiment
         | with maximizing a single GPU cluster.
        
       | singpolyma3 wrote:
       | 25 t/s is barely usable. Maybe for a background runner
        
         | lelanthran wrote:
         | > 25 t/s is barely usable. Maybe for a background runner
         | 
         | That's over a 1000 words/s if you were typing. If 1000 words/s
         | is too slow for your use-case, then perhaps $5/m is just not
         | for you.
         | 
         | I kinda like the idea of paying $5/m for unlimited usage at the
         | specified speed.
         | 
         | It beats a 10x higher speed that hits daily restrictions in
         | about 2 hours, and weekly restrictions in 3 days.
        
           | singpolyma3 wrote:
           | Sure if it was just a matter of typing. But in practise it
           | means sitting and staring for minutes at nothing happening
           | with a "thinking" until something finally happens.
           | 
           | I mean my local 122b is only 20t/s so for background stuff it
           | can be used for that. But not for anything interactive IME.
        
             | lelanthran wrote:
             | > I mean my local 122b is only 20t/s so for background
             | stuff it can be used for that. But not for anything
             | interactive IME.
             | 
             | What are you running that local 122b on? I mean, this looks
             | attractive to me for $5/m running unlimited at 20t/s-25t/s,
             | but if I could buy hardware to get that running locally, I
             | don't mind doing so.
        
       | freedomben wrote:
       | This is an excellent idea, but I worry about fairness during
       | resource contention. I don't often need queries, but when I do
       | it's often big and long. I wouldn't want to eat up the whole
       | system when other users need it, but I also would want to have
       | the cluster when I need it. How do you address a case like this?
        
         | jrandolf wrote:
         | We implement rate-limiting and queuing to ensure fairness, but
         | if there are a massive amount of people with huge and long
         | queries, then there will be waits. The question is whether
         | people will do this and more often than not users will be idle.
        
           | freedomben wrote:
           | Is there any way to buy into a pool of people with similar
           | usage patterns? Maybe I'm overthinking it, but just wondering
        
             | ssl-3 wrote:
             | I think it'd be best to pool with people with different
             | patterns, not the same patterns. Perhaps it would be best
             | to pool with people in different timezones, and/or with
             | different work/sleep schedules.
             | 
             | If everyone in a pool uses it during the ~same periods and
             | sleeps during the ~same periods, then the node would
             | oscillate between contention and idle -- every day. This
             | seems largely avoidable.
             | 
             | (Or, darker: Maybe the contention/idle dichotomy is a
             | feature, not a bug. After all, when one has control of
             | $14k/month of hardware that is sitting idle reliably-enough
             | for significant periods every day, then one becomes
             | incentivized to devise a way to sell that idle time for
             | other purposes.)
        
           | mogili1 wrote:
           | Rate limit essentially is a token limit
        
             | ibejoeb wrote:
             | It depends on how it's implemented. If it's a fixed window,
             | then your absolute ceiling is tokens/windows in a month. If
             | it's a function of other usage, like a timeshare, you're
             | still paying for some price for a month and you get what
             | you get without paying more per token. There's an intrinsic
             | limit based on how many tokens the model can process on
             | that gpu in a month anyway, even if it's only you.
        
             | delusional wrote:
             | Time x capacity is also a limit. There's always a limit.
        
           | petterroea wrote:
           | To be fair this is the price you pay for sharing a GPU.
           | Probably good for stuff that doesn't need to be done "now"
           | but that you can just launch and run in the background. I bet
           | some graphs that show when the gpu is most busy could be
           | useful as well
        
         | pokstad wrote:
         | This problem sounds like an excellent opportunity. We need a
         | race to the bottom for hosting LLMs to democratize the tech and
         | lower costs. I cheer on anyone who figures this out.
        
         | cyanydeez wrote:
         | Also, cache ejection during contention qill degrade everyones
         | service.
         | 
         | I question whether they actually understand LLMs at scale.
        
           | zozbot234 wrote:
           | I suppose it's meant to be a "minimum viable" third-party
           | inference platform, where you're literally selling
           | subscription-based access (i.e. fixed price, not PAYGO by
           | token) to a single GPU cluster, and then only once enough
           | users subscribe to make it viable (which is very nice from
           | them, it works like a Kickstarter/group coupon model and
           | creates a guaranteed win-win for the users). But they could
           | easily expand to more than just the minimum cluster size,
           | which would somewhat improve efficiency. (Deepseek themselves
           | scale out their model over huge amounts of GPUs, which is how
           | they manage to price their tokens quite cheap.)
        
         | zozbot234 wrote:
         | Ultimately the most sensible way of handling this is you end up
         | with "surge pricing" for the highest-priority tokens whenever
         | the inference platform is congested, over and above the base
         | subscription (but perhaps ultimately making the subscription a
         | bit cheaper).
        
       | varunr89 wrote:
       | $40/mo for deepseek r1 seems steep compared to a pro sub on open
       | ai /claude unless you run 24x7. im not sure how sharing is making
       | this affirdable.
        
         | lelanthran wrote:
         | > $40/mo for deepseek r1 seems steep compared to a pro sub on
         | open ai /claude unless you run 24x7.
         | 
         | "Running 24x7" is what people want to do with openclaw.
        
       | Lalabadie wrote:
       | This is the most "Prompted ourselves a Shadcn UI" page I've seen
       | in a while lol
       | 
       | I dig the idea! I'm curious where the costs will land with actual
       | use.
        
         | jrandolf wrote:
         | Thanks lol. I actually like Shadcn's style. It's sad that
         | people view it as AI now.
        
       | mogili1 wrote:
       | Can you show a comparison of cost of we went per token pricing.
        
       | QuantumNomad_ wrote:
       | > How does billing work?
       | 
       | > When you join a cohort, your card is saved but not charged
       | until the cohort fills. Stripe holds your card information -- we
       | never store it. Once the cohort fills, you are charged and
       | receive an API key for the duration of the cohort.
       | 
       | Have any cohorts filled yet?
       | 
       | I'm interested in joining one, but only if it's reasonable to
       | assume that the cohort will be full within the next 7 days or so.
       | (Especially because in a little over a week I'm attending an LLM-
       | centered hackathon where we can either use AWS LLM credits
       | provided by the organizer, or we can use providers of our own
       | choosing, and I'd rather use either yours or my own hardware
       | running vLLM than the LLM offerings and APIs from AWS.)
       | 
       | I'd be pretty annoyed if I join a cohort and then it takes like 3
       | months before the cohort has filled and I can begin to use it. By
       | then I will probably have forgotten all about it and not have
       | time to make use of the API key I am paying you for.
        
         | jrandolf wrote:
         | No cohorts have been filled yet. We're still early. We are
         | seeing reservations pick up quickly, but I'd be able to give
         | you a more concrete estimate of fill velocity after about a
         | week.
         | 
         | That said, we're planning to add a 7-day window: if a cohort
         | doesn't fill within 7 days of your reservation, it cancels
         | automatically and your card is released. We don't want anyone's
         | payment method sitting in limbo indefinitely.
        
       | p_m_c wrote:
       | Do you own the GPUs or are you multiplexing on a 3rd party GPU
       | cloud?
        
         | jrandolf wrote:
         | Multiplexing on a GPU cloud.
        
       | RIMR wrote:
       | I read the FAQ, and I can't imagine this is going to work the way
       | you want it to. It fundamentally doesn't make sense as a business
       | model.
       | 
       | I can sign up for a cohort today, but there's not even a hint of
       | how long it will take the cohort to fill up. The most subscribed
       | cohort is only at 42% (and dropping), so maybe days to weeks?
       | That's a long time to wait if you have a use case to satisfy.
       | 
       | And then the cohort expires, and I have to sign up for another
       | one and play the waiting game again? Nobody wants that level of
       | unreliability.
       | 
       | Also, don't say "15-25 tok/s". That is a min-max figure, but your
       | FAQ says that this is actually a maximum. It makes no sense to
       | measure a maximum as a range, and you state no minimum so I can
       | only assume that it is 0 tok/s. If all users in the cohort use it
       | simultaneously, the best they're getting is something like 1.5
       | tok/s (probably less), which is abyssmal.
       | 
       | You mention "optimization", but I have no idea what that means.
       | It certainly doesn't mean imposing token limits, because your FAQ
       | says that won't happen. If more than 25 users are using the
       | cohort simultaneously, it is a physical impossibility to improve
       | performance to the levels you advertise without sacrificing
       | something else, like switching to a smaller model, which would
       | essentially be fraud, or adding more GPUs which will bankrupt you
       | at these margins. With 465 users per cohort, a large chunk of
       | whom will be using tools like OpenClaw, nobody will ever see the
       | performance you are offering.
       | 
       | The issue here is you are trying to offer affordable AI GPU nodes
       | without operating at a loss. The entire AI industry is operating
       | at a loss right now because of how expensive this all is. This
       | strategy literally won't work right now unless you start courting
       | VCs to invest tens to hundreds of millions of dollars so you can
       | get this off the ground by operating at a loss until hopefully
       | you turn a profit at some point in the future, but at that point
       | developers will probably be able to run these models at home
       | without your help.
        
         | jrandolf wrote:
         | Going on ChatGPT.com and using their AI for 24 hours doesn't
         | mean you are actually using their LLM for 24 hours. It's only
         | live for as long as the output is being generated. You reading,
         | waiting for tool calls, etc. don't count toward concurrency.
         | Factor in time-zones, lunch times, etc...it's more likely that
         | we'd have an underutilization problem.
         | 
         | For filling up the cohorts, I agree and we're launching for a
         | week to gather feedback.
        
       | tensor-fusion wrote:
       | Interesting direction. One adjacent pattern we've been working on
       | is a bit less about partitioning a shared node for more tokens,
       | and more about letting developers keep a local workflow while
       | attaching to an existing remote GPU via a share link / CLI / VS
       | Code path. In labs and small teams we've found the pain is often
       | not just allocation, but getting access into the everyday
       | workflow without moving code + environment into a full remote VM
       | flow. Curious whether your users mostly want higher GPU
       | utilization, or whether they also want workflow portability from
       | laptops and homelabs. I'm involved with GPUGo / TensorFusion, so
       | that's the lens I'm looking through.
        
       | scottcha wrote:
       | Pretty cool idea, but whats the stack behind this? As 15-25 tok/s
       | seems a bit low as expected SoA for most providers is around 60
       | tok/s and quality of life dramatically improves above that.
        
       | IanCal wrote:
       | Can you explain the benefits over something like openrouter?
        
         | jrandolf wrote:
         | 24/7 LLM for $10/month.
        
           | johndough wrote:
           | Isn't this a bad deal? Or is there an error in my math?
           | 
           | For $40, I'd get 20 tok/s * 2.6M seconds per month = 52M
           | tokens of DeepSeek v3.2 per month if I run it 24/7, which is
           | not realistic for most workloads.
           | 
           | On OpenRouter [1], $40 buys 105M tokens from the same model,
           | which is more than 52M tokens, and I can freely choose when
           | to use them.
           | 
           | [1]: https://openrouter.ai/deepseek/deepseek-v3.2
        
             | jrandolf wrote:
             | 20 tok/s is an average. It can be more, it can be less. If
             | you are running off-peak I'm sure you'd get some crazy
             | number.
        
       | moralestapia wrote:
       | This is great, thanks!
       | 
       | I personally would like something like this but with "regular"
       | GPU access. Some people still use them for something other than
       | LLMs ^^.
        
         | jrandolf wrote:
         | There is vast.ai!
        
           | moralestapia wrote:
           | Wow!
           | 
           | I recall hearing about them years ago.
           | 
           | Good to see they're thriving!
        
       | MuffinFlavored wrote:
       | > Running DeepSeek V3 (685B) requires 8xH100 GPUs which is about
       | $14k/month. Most developers only need 15-25 tok/s.
       | 
       | > deepseek-v3.2-685b, $40/mo/slot for ~20 tok/s, 465 slots total
       | 
       | > 465 users x 20 tok/s = 9,300 tok/s needed
       | 
       | > The node peaks at ~3,000 tok/s total. So at full capacity they
       | can really only serve:
       | 
       | > 3,000 / 20 = 150 concurrent users at 20 tok/s
       | 
       | > That's only 32% of the cohort being active simultaneously.
        
         | artificialprint wrote:
         | People work 8 hours a day presumably, I guess they are banking
         | on this idea
        
       | artificialprint wrote:
       | Didn't make sense to launch multiple 10 and 40 bucks
       | subscriptions right at the start, because now they are competing
       | with each other.
       | 
       | Also mobile version is a bit broken, but good idea and good luck!
        
         | jrandolf wrote:
         | I'm feeling it Mr. Crabs.
        
       | trvz wrote:
       | The absolute lack of any kind of legal information makes this
       | website criminal.
        
         | copperx wrote:
         | There's a big difference between non-compliant, illegal, and
         | criminal.
        
       | avereveard wrote:
       | Interesting there's a trickle of low intensity job one can always
       | get running but like glm own plan is $30/mo and something about
       | 300tps now I know that one is subsidized but still.
        
       | spencer9714 wrote:
       | Interesting concept. One thing I'm curious about if I'm in a
       | cohort for something like DeepSeek V3 and another user spins up a
       | heavy 24/7 job, how do you keep TTFT from degrading? vLLM's
       | continuous batching helps, but there's still a physical limit
       | with shared VRAM/compute. I've been grappling with this exact
       | 'noisy neighbor' issue while building Runfra. We actually ended
       | up moving toward a credit per task model on idle GPUs
       | specifically to avoid that resource contention entirely.
       | 
       | Curious how you're thinking about isolation here. Is there any
       | hard guarantee on a 'slice' of the GPU, or is it mostly just
       | handled by the vLLM scheduler?
        
       ___________________________________________________________________
       (page generated 2026-04-04 23:00 UTC)