[HN Gopher] Show HN: sllm - Split a GPU node with other develope...
___________________________________________________________________
Show HN: sllm - Split a GPU node with other developers, unlimited
tokens
Running DeepSeek V3 (685B) requires 8xH100 GPUs which is about
$14k/month. Most developers only need 15-25 tok/s. sllm lets you
join a cohort of developers sharing a dedicated node. You reserve a
spot with your card, and nobody is charged until the cohort fills.
Prices start at $5/mo for smaller models. The LLMs are completely
private (we don't log any traffic). The API is OpenAI-compatible
(we run vLLM), so you just swap the base URL. Currently offering a
few models.
Author : jrandolf
Score : 97 points
Date : 2026-04-04 15:18 UTC (7 hours ago)
(HTM) web link (sllm.cloud)
(TXT) w3m dump (sllm.cloud)
| mmargenot wrote:
| This is a great idea! I saw a similar (inverse) idea the other
| day for pooling compute (https://github.com/michaelneale/mesh-
| llm). What are you doing for compute in the backend? Are you
| locked into a cohort from month to month?
| vova_hn2 wrote:
| 1. Is the given tok/s estimate for the total node throughput, or
| is it what you can realistically expect to get? Or is it the
| worst case scenario throughput if everyone starts to use it
| simultaneously?
|
| 2. What if I try to hog all resources of a node by running some
| large data processing and making multiple queries in parallel?
| What if I try to resell the access by charging per token?
|
| Edit: sorry if this comment sounds overly critical. I think that
| pooling money with other developers to collectively rent a server
| for LLM inference is a really cool idea. I also thought about it,
| but haven't found a satisfactory answer to my question number 2,
| so I decided that it is infeasible in practice.
| jrandolf wrote:
| 1. It's an average. 2. We have sophisticated rate limiter.
| poly2it wrote:
| Does it take user time zones into account?
| jrandolf wrote:
| Yes
| esafak wrote:
| Like vast.ai and TensorDock, and presumably others.
| spuz wrote:
| It seems crazy to me that the "Join" button does not have a price
| on it and yet clicking it simply forwards you to a Stripe page
| again with no price information on it. How am I supposed to know
| how much I'm about to be charged?
| jrandolf wrote:
| That was an error on our part lol. We'll update with the price.
| peter_d_sherman wrote:
| What a brilliant idea!
|
| Split a "it needs to run in a datacenter because its hardware
| requirements are so large" AI/LLM across multiple people who each
| want shared access to that particular model.
|
| Sort of like the Real Estate equivalent of subletting, or
| splitting a larger space into smaller spaces and subletting each
| one...
|
| Or, like the Web Host equivalent of splitting a single server
| into multiple virtual machines for shared hosting by multiple
| other parties, or what-have-you...
|
| _I could definitely see marketplaces similar to this, popping up
| in the future!_
|
| It seems like it should make AI cheaper for everyone... that is,
| "democratize AI"... in a "more/better/faster/cheaper" way than AI
| has been democratized to date...
|
| Anyway, it's a brilliant idea!
|
| Wishing you a lot of luck with this endeavor!
| kaoD wrote:
| How is the time sharing handled? I assume if I submit a unit of
| work it will load to VRAM and then run (sharing time? how many
| work units can run in parallel?)
|
| How large is a full context window in MiB and how long does it
| take to load the buffer? I.e. how many seconds should I expect my
| worst case wait time to take until I get my first token?
| ninjha wrote:
| > how many work units can run in parallel
|
| not original author but batching is one very important trick to
| make inference efficient, you can reasonably do tens to low
| hundreds in parallel (depending on model size and gpu size)
| with very little performance overhead
| jrandolf wrote:
| vLLM handles GPU scheduling, not sllm. The model weights stay
| resident in VRAM permanently so there's no loading/unloading
| per request. vLLM uses continuous batching, so incoming
| requests are dynamically added to the running batch every
| decode step and the GPU is always working on multiple requests
| simultaneously. There is no "load to VRAM and run" per request;
| it's more like joining an already-running batch.
|
| TTFT is under 2 seconds average. Worst case is 10-30s.
| kaoD wrote:
| > The model weights stay resident in VRAM permanently so
| there's no loading/unloading per request.
|
| Yes, I was thinking about context buffers, which I assume are
| not small in large models. That has to be loaded into VRAM,
| right?
|
| If I keep sending large context buffers, will that hog the
| batches?
| jrandolf wrote:
| Not if you are the only one. We have rate limits to prevent
| this in case, idk, you share your key with 1000 people lol.
| spuz wrote:
| Is this not a more restricted version of OpenRouter? With
| OpenRouter you pay for credits that can be used to run any
| commercial or open-source model and you only pay for what you
| use.
| jrandolf wrote:
| OpenRouter is a little different. We are trying to experiment
| with maximizing a single GPU cluster.
| singpolyma3 wrote:
| 25 t/s is barely usable. Maybe for a background runner
| lelanthran wrote:
| > 25 t/s is barely usable. Maybe for a background runner
|
| That's over a 1000 words/s if you were typing. If 1000 words/s
| is too slow for your use-case, then perhaps $5/m is just not
| for you.
|
| I kinda like the idea of paying $5/m for unlimited usage at the
| specified speed.
|
| It beats a 10x higher speed that hits daily restrictions in
| about 2 hours, and weekly restrictions in 3 days.
| singpolyma3 wrote:
| Sure if it was just a matter of typing. But in practise it
| means sitting and staring for minutes at nothing happening
| with a "thinking" until something finally happens.
|
| I mean my local 122b is only 20t/s so for background stuff it
| can be used for that. But not for anything interactive IME.
| lelanthran wrote:
| > I mean my local 122b is only 20t/s so for background
| stuff it can be used for that. But not for anything
| interactive IME.
|
| What are you running that local 122b on? I mean, this looks
| attractive to me for $5/m running unlimited at 20t/s-25t/s,
| but if I could buy hardware to get that running locally, I
| don't mind doing so.
| freedomben wrote:
| This is an excellent idea, but I worry about fairness during
| resource contention. I don't often need queries, but when I do
| it's often big and long. I wouldn't want to eat up the whole
| system when other users need it, but I also would want to have
| the cluster when I need it. How do you address a case like this?
| jrandolf wrote:
| We implement rate-limiting and queuing to ensure fairness, but
| if there are a massive amount of people with huge and long
| queries, then there will be waits. The question is whether
| people will do this and more often than not users will be idle.
| freedomben wrote:
| Is there any way to buy into a pool of people with similar
| usage patterns? Maybe I'm overthinking it, but just wondering
| ssl-3 wrote:
| I think it'd be best to pool with people with different
| patterns, not the same patterns. Perhaps it would be best
| to pool with people in different timezones, and/or with
| different work/sleep schedules.
|
| If everyone in a pool uses it during the ~same periods and
| sleeps during the ~same periods, then the node would
| oscillate between contention and idle -- every day. This
| seems largely avoidable.
|
| (Or, darker: Maybe the contention/idle dichotomy is a
| feature, not a bug. After all, when one has control of
| $14k/month of hardware that is sitting idle reliably-enough
| for significant periods every day, then one becomes
| incentivized to devise a way to sell that idle time for
| other purposes.)
| mogili1 wrote:
| Rate limit essentially is a token limit
| ibejoeb wrote:
| It depends on how it's implemented. If it's a fixed window,
| then your absolute ceiling is tokens/windows in a month. If
| it's a function of other usage, like a timeshare, you're
| still paying for some price for a month and you get what
| you get without paying more per token. There's an intrinsic
| limit based on how many tokens the model can process on
| that gpu in a month anyway, even if it's only you.
| delusional wrote:
| Time x capacity is also a limit. There's always a limit.
| petterroea wrote:
| To be fair this is the price you pay for sharing a GPU.
| Probably good for stuff that doesn't need to be done "now"
| but that you can just launch and run in the background. I bet
| some graphs that show when the gpu is most busy could be
| useful as well
| pokstad wrote:
| This problem sounds like an excellent opportunity. We need a
| race to the bottom for hosting LLMs to democratize the tech and
| lower costs. I cheer on anyone who figures this out.
| cyanydeez wrote:
| Also, cache ejection during contention qill degrade everyones
| service.
|
| I question whether they actually understand LLMs at scale.
| zozbot234 wrote:
| I suppose it's meant to be a "minimum viable" third-party
| inference platform, where you're literally selling
| subscription-based access (i.e. fixed price, not PAYGO by
| token) to a single GPU cluster, and then only once enough
| users subscribe to make it viable (which is very nice from
| them, it works like a Kickstarter/group coupon model and
| creates a guaranteed win-win for the users). But they could
| easily expand to more than just the minimum cluster size,
| which would somewhat improve efficiency. (Deepseek themselves
| scale out their model over huge amounts of GPUs, which is how
| they manage to price their tokens quite cheap.)
| zozbot234 wrote:
| Ultimately the most sensible way of handling this is you end up
| with "surge pricing" for the highest-priority tokens whenever
| the inference platform is congested, over and above the base
| subscription (but perhaps ultimately making the subscription a
| bit cheaper).
| varunr89 wrote:
| $40/mo for deepseek r1 seems steep compared to a pro sub on open
| ai /claude unless you run 24x7. im not sure how sharing is making
| this affirdable.
| lelanthran wrote:
| > $40/mo for deepseek r1 seems steep compared to a pro sub on
| open ai /claude unless you run 24x7.
|
| "Running 24x7" is what people want to do with openclaw.
| Lalabadie wrote:
| This is the most "Prompted ourselves a Shadcn UI" page I've seen
| in a while lol
|
| I dig the idea! I'm curious where the costs will land with actual
| use.
| jrandolf wrote:
| Thanks lol. I actually like Shadcn's style. It's sad that
| people view it as AI now.
| mogili1 wrote:
| Can you show a comparison of cost of we went per token pricing.
| QuantumNomad_ wrote:
| > How does billing work?
|
| > When you join a cohort, your card is saved but not charged
| until the cohort fills. Stripe holds your card information -- we
| never store it. Once the cohort fills, you are charged and
| receive an API key for the duration of the cohort.
|
| Have any cohorts filled yet?
|
| I'm interested in joining one, but only if it's reasonable to
| assume that the cohort will be full within the next 7 days or so.
| (Especially because in a little over a week I'm attending an LLM-
| centered hackathon where we can either use AWS LLM credits
| provided by the organizer, or we can use providers of our own
| choosing, and I'd rather use either yours or my own hardware
| running vLLM than the LLM offerings and APIs from AWS.)
|
| I'd be pretty annoyed if I join a cohort and then it takes like 3
| months before the cohort has filled and I can begin to use it. By
| then I will probably have forgotten all about it and not have
| time to make use of the API key I am paying you for.
| jrandolf wrote:
| No cohorts have been filled yet. We're still early. We are
| seeing reservations pick up quickly, but I'd be able to give
| you a more concrete estimate of fill velocity after about a
| week.
|
| That said, we're planning to add a 7-day window: if a cohort
| doesn't fill within 7 days of your reservation, it cancels
| automatically and your card is released. We don't want anyone's
| payment method sitting in limbo indefinitely.
| p_m_c wrote:
| Do you own the GPUs or are you multiplexing on a 3rd party GPU
| cloud?
| jrandolf wrote:
| Multiplexing on a GPU cloud.
| RIMR wrote:
| I read the FAQ, and I can't imagine this is going to work the way
| you want it to. It fundamentally doesn't make sense as a business
| model.
|
| I can sign up for a cohort today, but there's not even a hint of
| how long it will take the cohort to fill up. The most subscribed
| cohort is only at 42% (and dropping), so maybe days to weeks?
| That's a long time to wait if you have a use case to satisfy.
|
| And then the cohort expires, and I have to sign up for another
| one and play the waiting game again? Nobody wants that level of
| unreliability.
|
| Also, don't say "15-25 tok/s". That is a min-max figure, but your
| FAQ says that this is actually a maximum. It makes no sense to
| measure a maximum as a range, and you state no minimum so I can
| only assume that it is 0 tok/s. If all users in the cohort use it
| simultaneously, the best they're getting is something like 1.5
| tok/s (probably less), which is abyssmal.
|
| You mention "optimization", but I have no idea what that means.
| It certainly doesn't mean imposing token limits, because your FAQ
| says that won't happen. If more than 25 users are using the
| cohort simultaneously, it is a physical impossibility to improve
| performance to the levels you advertise without sacrificing
| something else, like switching to a smaller model, which would
| essentially be fraud, or adding more GPUs which will bankrupt you
| at these margins. With 465 users per cohort, a large chunk of
| whom will be using tools like OpenClaw, nobody will ever see the
| performance you are offering.
|
| The issue here is you are trying to offer affordable AI GPU nodes
| without operating at a loss. The entire AI industry is operating
| at a loss right now because of how expensive this all is. This
| strategy literally won't work right now unless you start courting
| VCs to invest tens to hundreds of millions of dollars so you can
| get this off the ground by operating at a loss until hopefully
| you turn a profit at some point in the future, but at that point
| developers will probably be able to run these models at home
| without your help.
| jrandolf wrote:
| Going on ChatGPT.com and using their AI for 24 hours doesn't
| mean you are actually using their LLM for 24 hours. It's only
| live for as long as the output is being generated. You reading,
| waiting for tool calls, etc. don't count toward concurrency.
| Factor in time-zones, lunch times, etc...it's more likely that
| we'd have an underutilization problem.
|
| For filling up the cohorts, I agree and we're launching for a
| week to gather feedback.
| tensor-fusion wrote:
| Interesting direction. One adjacent pattern we've been working on
| is a bit less about partitioning a shared node for more tokens,
| and more about letting developers keep a local workflow while
| attaching to an existing remote GPU via a share link / CLI / VS
| Code path. In labs and small teams we've found the pain is often
| not just allocation, but getting access into the everyday
| workflow without moving code + environment into a full remote VM
| flow. Curious whether your users mostly want higher GPU
| utilization, or whether they also want workflow portability from
| laptops and homelabs. I'm involved with GPUGo / TensorFusion, so
| that's the lens I'm looking through.
| scottcha wrote:
| Pretty cool idea, but whats the stack behind this? As 15-25 tok/s
| seems a bit low as expected SoA for most providers is around 60
| tok/s and quality of life dramatically improves above that.
| IanCal wrote:
| Can you explain the benefits over something like openrouter?
| jrandolf wrote:
| 24/7 LLM for $10/month.
| johndough wrote:
| Isn't this a bad deal? Or is there an error in my math?
|
| For $40, I'd get 20 tok/s * 2.6M seconds per month = 52M
| tokens of DeepSeek v3.2 per month if I run it 24/7, which is
| not realistic for most workloads.
|
| On OpenRouter [1], $40 buys 105M tokens from the same model,
| which is more than 52M tokens, and I can freely choose when
| to use them.
|
| [1]: https://openrouter.ai/deepseek/deepseek-v3.2
| jrandolf wrote:
| 20 tok/s is an average. It can be more, it can be less. If
| you are running off-peak I'm sure you'd get some crazy
| number.
| moralestapia wrote:
| This is great, thanks!
|
| I personally would like something like this but with "regular"
| GPU access. Some people still use them for something other than
| LLMs ^^.
| jrandolf wrote:
| There is vast.ai!
| moralestapia wrote:
| Wow!
|
| I recall hearing about them years ago.
|
| Good to see they're thriving!
| MuffinFlavored wrote:
| > Running DeepSeek V3 (685B) requires 8xH100 GPUs which is about
| $14k/month. Most developers only need 15-25 tok/s.
|
| > deepseek-v3.2-685b, $40/mo/slot for ~20 tok/s, 465 slots total
|
| > 465 users x 20 tok/s = 9,300 tok/s needed
|
| > The node peaks at ~3,000 tok/s total. So at full capacity they
| can really only serve:
|
| > 3,000 / 20 = 150 concurrent users at 20 tok/s
|
| > That's only 32% of the cohort being active simultaneously.
| artificialprint wrote:
| People work 8 hours a day presumably, I guess they are banking
| on this idea
| artificialprint wrote:
| Didn't make sense to launch multiple 10 and 40 bucks
| subscriptions right at the start, because now they are competing
| with each other.
|
| Also mobile version is a bit broken, but good idea and good luck!
| jrandolf wrote:
| I'm feeling it Mr. Crabs.
| trvz wrote:
| The absolute lack of any kind of legal information makes this
| website criminal.
| copperx wrote:
| There's a big difference between non-compliant, illegal, and
| criminal.
| avereveard wrote:
| Interesting there's a trickle of low intensity job one can always
| get running but like glm own plan is $30/mo and something about
| 300tps now I know that one is subsidized but still.
| spencer9714 wrote:
| Interesting concept. One thing I'm curious about if I'm in a
| cohort for something like DeepSeek V3 and another user spins up a
| heavy 24/7 job, how do you keep TTFT from degrading? vLLM's
| continuous batching helps, but there's still a physical limit
| with shared VRAM/compute. I've been grappling with this exact
| 'noisy neighbor' issue while building Runfra. We actually ended
| up moving toward a credit per task model on idle GPUs
| specifically to avoid that resource contention entirely.
|
| Curious how you're thinking about isolation here. Is there any
| hard guarantee on a 'slice' of the GPU, or is it mostly just
| handled by the vLLM scheduler?
___________________________________________________________________
(page generated 2026-04-04 23:00 UTC)