Posts by AGARTHA_NOBLE@shortstacksran.ch
(DIR) Post #B6WzlCj3rqAiXweV16 by AGARTHA_NOBLE@shortstacksran.ch
0 likes, 0 repeats
@sapphire Not so sure about M1, but with that VRAM you can probably load high-precision (8 bit quant) Qwen 3.6 MoE 35B and expect pretty snappy performance, even using Ollama. Lemme look into it.
(DIR) Post #B6WzlCvp6ONdBWcguO by AGARTHA_NOBLE@shortstacksran.ch
0 likes, 0 repeats
@sapphire Yea, looks like either 6 or 8 bit quant works fine, but caveat emptor on your tokens per second. I suggest Qwen 3.6 MoE for speed, 27B Dense if you want accuracy in coding or what have you.
(DIR) Post #B6WzlD6SSqt3iVbBU8 by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@sapphire If you only want to line up one model and make that shit run fast, tinker with a vLLM config, but honestly Ollama will probably suit you super well, and you can connect OpenWeb UI to it remotely over docker pod to chat with the little virtual idiot, if that's all you want.
(DIR) Post #B6X0ZFrsSVyKlrdr3w by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@mischievoustomato @sapphire Trade-off math: do you want to run a lot of models really easily and to be able to swap between them seamlessly(ollama), or do you want to tune one model to serve a lot of people really well(vllm)? There are other options that I haven't yet touched like lmstudio, and llamacpp, but those fall into the second bucket, not the first- ollama's a wrapper around llamacpp, meaning you don't have to care about implementation so long as you've got the hardware to support it.
(DIR) Post #B6X0t9hF78gdQE9moC by AGARTHA_NOBLE@shortstacksran.ch
0 likes, 0 repeats
@sapphire Looks like it's up to about 48gb by default, which still puts you in the right range for a quantized Qwen, plus respectable KV cache. I'm using a bunch of tricks to get my cache window to about 110K tokens; you'll likely be in the same vicinity, and those models support up to 256k tokens natively.
(DIR) Post #B6X0t9u0LgtY3o7yhU by AGARTHA_NOBLE@shortstacksran.ch
0 likes, 0 repeats
@sapphire I THINK that's just soft-cap, too; there's not really anything preventing you from accessing all but 4GB with the GPU on those things except for sanity.
(DIR) Post #B6X19QiBcTyIpLMHrc by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@sapphire All the better- run the highest quantization you can make work. Definitely worth trimming KV-cache for higher quants, too; smaller models are super susceptible to quantization.
(DIR) Post #B6X1Amo6xEcde2mClE by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@mischievoustomato @sapphire Kobold's really solid for the smaller stuff, and the framework's good, but for creative work I'm starting to look towards Stabilitymatrix: https://lykos.ai/
(DIR) Post #B6ZDyP7St2ugQKAuau by AGARTHA_NOBLE@shortstacksran.ch
0 likes, 0 repeats
@sickburnbro ...No, that's not right.
(DIR) Post #B6ZDz0lSvHtXB6Asa0 by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@sapphire @sickburnbro With traditional transformer models at 1T or 2.5T, yeah, maybe, but it sounds kinda like your guy is doomering just like everyone else, doesn't know what he's talking about. Never underestimate the world's ability to get better, even at the expense of your expectations.
(DIR) Post #B6ZDzgcNJTZwz4FDF2 by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@sapphire @sickburnbro >Look at my quant. You notice something? That's right he's got a PIECE of PAPER in TENSORFLOW
(DIR) Post #B6ZE00c2yPCSzqFEVU by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@sickburnbro @petra Right on function, wrong on architecture.
(DIR) Post #B7W89AWQTkSFR4ffGa by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@Owl @MK2boogaloo @vriska @Oughtism @bronzeagebond @mangeurdenuage No I mean if we're going to talk about Christian larpers, let's take full account here. A lot of these supposed trad Christian people are going to sit with a straight face and talk about their devotion to God and then outright try to get other men horny with pornographic content under this guise of "w-we're just putting beauty out into the world uwu" like it's still not extremely gay or they'll be writing outright smut on the timeline and then two posts down will be talking about the church or something. It's phenomenal, it's such a break in consistency that it even draws my brow upward.
(DIR) Post #B7subzZgdpnoVZp596 by AGARTHA_NOBLE@shortstacksran.ch
0 likes, 0 repeats
@Hoss Finally, more monstergirl greentexts.
(DIR) Post #B7yO7m5AKa34qnmTvk by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@sapphire @Turdicus Probably a good idea to listen to the only actually married guy on Fedi tbh, he's just correct here.
(DIR) Post #B7ySUYGIHU5OSq6Q2C by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@Junes WAZOOO!!
(DIR) Post #B7yWOFsi6bdlXOKBqy by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 1 repeats
@bronze Imagine having violent south pacific sex with tan fox for babymaking purposes.
(DIR) Post #B80v1lrnuADXFSCDMu by AGARTHA_NOBLE@shortstacksran.ch
0 likes, 0 repeats
@sapphire @Groomschild @teto did someone say golbin for insert peanus?
(DIR) Post #B81CUoemJFYhWHXTBQ by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@p
(DIR) Post #B81EIiTJg8pQlrXHYO by AGARTHA_NOBLE@shortstacksran.ch
1 likes, 0 repeats
@p I have bond burgered your sister. Pray I do not bond burger her any more.