Post B8hGpIYdH2H6xqUkUK by ent@noauthority.social
 (DIR) More posts by ent@noauthority.social
 (DIR) Post #B8bvcbwxJR7n7aty2i by ent@noauthority.social
       0 likes, 0 repeats
       
       I cannot overstate my hatred for current RAM prices. Thanks to the cartel, my laptop pre-order ended up getting downgraded on the RAM module it'll ship with, and I'm NOT paying $1600 for a stick of 64GB. That is absolutely absurd
       
 (DIR) Post #B8bxN3ZN3VXg3ba4o4 by stormbringer@noauthority.social
       0 likes, 0 repeats
       
       @ent I feel your pain.
       
 (DIR) Post #B8bzPCv5Sj0xXmtLyy by Smeetoo@noauthority.social
       0 likes, 0 repeats
       
       @ent Hit some flea markets or thrift stores for an old laptop and grab the RAM
       
 (DIR) Post #B8bzqqwmkWnlVOruam by Fox@noauthority.social
       0 likes, 0 repeats
       
       @ent It's so bad, but to be fair the bestest stuff isn't required if you want to have a good time. Previous generation systems can still be built for reasonable amounts of money if you're willing to determine what it is that you really need...A barebones system with lots of ram and cores with the supporting hardware can still be had for $1000. It would run just about anything you could ask if to. Maybe not at ultra high graphics. But for everything else it's fine.
       
 (DIR) Post #B8c9pn2ylOuof4thtQ by cryptoduke@noauthority.social
       0 likes, 0 repeats
       
       @ent I blame those AI fuckheads... the ones who made it and the ones who keep using it
       
 (DIR) Post #B8cA8wHYeSs6o1qbui by ent@noauthority.social
       0 likes, 0 repeats
       
       @Smeetoo That doesn't work. There are very few sources for LPCAMM2 and it's a newer type of RAM
       
 (DIR) Post #B8cAFObaGCZ0NyIFf6 by ent@noauthority.social
       0 likes, 0 repeats
       
       @Fox The problem is to reasonably run local LLMs you need a chipset that supports unified RAM or a cluster of GPUs.
       
 (DIR) Post #B8cAXMgk7qbreZH6Mi by ent@noauthority.social
       0 likes, 0 repeats
       
       @cryptoduke To be fair, using it hastens the bubble pop. The only reason I got an AI subscription is so I can burn my monthly subscription's value in tokens at API rates within a single session and use it to keep abreast of what the frontier models are capable of. They spend 10x what I give them, so may as well take advantage.
       
 (DIR) Post #B8d1hNebPHJTDT1Ri4 by Fox@noauthority.social
       0 likes, 0 repeats
       
       @ent Well....Actually, it depends. The model size and Quantization level will determine it's size. If you have the Ram, you can run it in a CPU with 8 cores. It won't be fast, at all, but you will get answers.Some Ternary models will fit in 2GBs of Ram and capable of basic tasks. They run on a CPU just fine. But yes, in general, VRAM is king. Need at least 8GBs for something useful in the classic sense. 12-16GBs opens up another level. So yeah, it just depends. 🧐
       
 (DIR) Post #B8d6vaqJfPLTwi68Bc by ent@noauthority.social
       0 likes, 0 repeats
       
       @FoxI'm aware of quantized options and I have a GPU with 12GB VRAM. I have not been remotely impressed with the output of what can actually run on that with a decent rate of token output. One of my basic tests I've used is asking for a Limerick, and the open models I've used can't even get the structure, much less anything else in such a request, correct.
       
 (DIR) Post #B8d7jj2rRe6Drds2Eq by Fox@noauthority.social
       0 likes, 0 repeats
       
       @ent Huh...Weird. At minimum a 4Q 12B model should actually be half decent...I bet I could get an 8B to do that. I wonder if it's your settings then. What could make a difference is the models with reasoning. An 8B 4Q in a Qwen model should be half decent. Hell even a Mistral AI model works with tools as a research assistant... and that's 8B 4Q.You got me curious. Let me know if you got some absolutely failing tests and I could try them on my end of you want. Could be context... 🤔
       
 (DIR) Post #B8hGpIYdH2H6xqUkUK by ent@noauthority.social
       0 likes, 0 repeats
       
       @FoxI don't see how it would be a context issue. Asking for a Limerick is a single self-contained prompt that should be able to be fulfilled from training data. It doesn't even need to be a model optimized for use in chat mode.
       
 (DIR) Post #B8hGrPH7Gl6MNSJ6sC by ent@noauthority.social
       0 likes, 0 repeats
       
       @FoxThat's the simplest method I've used for testing open weight models, and I haven't seen a single one I can run locally get the Limerick structure correct, much less one that rhymes. I'll have to double check, but I know I've run Llama and Mistral models around the 8B parameter mark with none of them able to fulfill that.
       
 (DIR) Post #B8hij1WuT1oL8HUAMa by Fox@noauthority.social
       0 likes, 0 repeats
       
       @ent .... Are you sure? Limericks come in a few variations. Here; AABBA Style with no specific syllable count.First result: "There was a laptop so slow, It took ages to boot -- what a show!With a crash and a whine, It froze up like a pine,Now my work is all stuck in the glow. "Ministral 3 8B Reasoning Q4_K_M.Did you have some constraints about the limerick structure specifically? 🤔
       
 (DIR) Post #B8nvAgwAmFCeO4B5nc by ent@noauthority.social
       0 likes, 0 repeats
       
       @FoxThose times I tried in the past I only gave it a topic and didn't specify structure. Didn't keep records of which models I tried with, but tried a more recent Llama model and it actually got it, so that's something.It's less scientific, but another test I've used is joint story writing in a manner like setting up a D&D campaign and seeing how well it does. That usually hasn't given me great results as the machine will literally lose the plot over time.