Post B48tYTWWkNhvlAGdpA by ariadne@social.treehouse.systems
(DIR) More posts by ariadne@social.treehouse.systems
(DIR) Post #B48TrXuU46Dc3duEYy by ariadne@social.treehouse.systems
1 likes, 0 repeats
can i talk to an openclaw bot using internet relay chat? if not, then what is the point
(DIR) Post #B48Ttj1vgxOB7pQaPY by ariadne@social.treehouse.systems
1 likes, 0 repeats
my suspicion is that i can *handwaves in the direction of kent overstreet*
(DIR) Post #B48UvHZqpdbOVzbJIG by ariadne@social.treehouse.systems
2 likes, 1 repeats
you see, i have built my own LLM, using the most ethical method possible: i trained it on the entire corpus of IRC logs in my possession, 2003 to present
(DIR) Post #B48V2rT1nBPTqw1UuW by ariadne@social.treehouse.systems
0 likes, 0 repeats
no giant water vaporizing data centers needed here, just a GPU, a dream and some cold hard chats
(DIR) Post #B48VIyBsYwHTL2gPNQ by ariadne@social.treehouse.systems
0 likes, 0 repeats
time 2 implant this brain into an openclaw and give it full access to my emailmostly because i don't want to retain any of my email
(DIR) Post #B48VXbTHJa7dDenbv6 by ariadne@social.treehouse.systems
0 likes, 0 repeats
@fxchip yeah same idea, just running on kubernetes and CUDA
(DIR) Post #B48VvWpCF3pLl5SYTY by jannem@fosstodon.org
0 likes, 0 repeats
@ariadne I can see a near future where "the AI deleted it" becomes convenient cover for "I didn't want to bother sorting it out."
(DIR) Post #B48WMbHvkrqlQ7oCEC by ska@social.treehouse.systems
1 likes, 0 repeats
@ariadne how to produce the most toxic chatbot possible
(DIR) Post #B48WtxDVFJ4szFdrRB by lanodan@queer.hacktivis.me
0 likes, 0 repeats
@ska @ariadne Also counter until it yells something about supernets or other spam.
(DIR) Post #B48XeIqr8mdvPH9M9I by ariadne@social.treehouse.systems
0 likes, 0 repeats
@dalias @ska the L is for Lacking, not Large
(DIR) Post #B48aFnWAHqdp8hTiYC by dvshkn@social.treehouse.systems
0 likes, 0 repeats
@ariadne new eval, you ask it to generate a random password and it has to respond with hunter2
(DIR) Post #B48aISJRaFVQF1nO88 by xinit@mastodon.coffee
0 likes, 0 repeats
@ariadneL337@drwho
(DIR) Post #B48gutqGHIv2LprPiC by raulinbonn@social.treehouse.systems
0 likes, 1 repeats
@ska @ariadne Ive thouht of something related, not chatbots but imagine a GPS driving assistant voice in your car giving you directions and feedback, but in the most toxic way possible. An angry swearing voice saying things like: "Your exit comes in half a mile, try to not miss that one, you fucking moron." I've thought that ought to be a funny option to toggle on once in a while.
(DIR) Post #B48guu1xZoHCw7Kkwi by ariadne@social.treehouse.systems
0 likes, 0 repeats
@raulinbonn @ska now you see the vision!
(DIR) Post #B48qDti1tYcR1r8yYa by ariadne@social.treehouse.systems
0 likes, 0 repeats
so i installed it into the openclaw meme thing. and it's not like, doing the stuff it claims it is doing.like it is hallucinating things like "i updated SOUL.md with xyz"i seriously do not think this stuff is real now
(DIR) Post #B48qQsius4e3jikYCW by ariadne@social.treehouse.systems
0 likes, 0 repeats
like i need you to understand, i haven't even gotten through *setup* because the model apparently does not know how to use tools correctly. admittedly it has less than 1 billion parameters, and i don't know what the hell i am doing, but still.
(DIR) Post #B48qUhzQijfxbZIaLQ by fiore@brain.worm.pink
0 likes, 0 repeats
@ariadne inb4 ariadne gets hooked on ts and starts evangelizing ai as the Second Coming
(DIR) Post #B48qh7KyIypjh4I9BY by ariadne@social.treehouse.systems
1 likes, 0 repeats
@fiore seems unlikely
(DIR) Post #B48rRpE1Sm1l89JUQ4 by ariadne@social.treehouse.systems
0 likes, 0 repeats
can we get to the part where it is AI winter again already? this is not even fun. i want to throw my computer and its' very expensive RTX 6000 Blackwell GPU out my window.
(DIR) Post #B48rV08VLvfRq5FMfY by ariadne@social.treehouse.systems
0 likes, 0 repeats
i just wanted to put an openclaw on irc as a fucking shitpost man
(DIR) Post #B48sAzDp8rZzAMkrqa by ariadne@social.treehouse.systems
0 likes, 0 repeats
and you tell me people legitimately are using this software.how?is it really magically better when you hook up claude?
(DIR) Post #B48sEiHM4sezmDBRvE by ariadne@social.treehouse.systems
0 likes, 0 repeats
(don't worry, i am running this in a MicroVM under kubernetes, I wouldn't dare give it access to anything I care about.)
(DIR) Post #B48sH5ZsQNowcDyohU by jfkimmes@social.tinycyber.space
0 likes, 0 repeats
@ariadne what model did you finetune on? For a 1B model you need something really specialized on tool calling.
(DIR) Post #B48sN29O50svivQN8q by ariadne@social.treehouse.systems
0 likes, 1 repeats
@jfkimmes i built an LLM from scratch with transformers kinda loosely following the scripts the qwen people releasedthe LLM is basically trained on ~30ish GB of mostly furry smut and public Linux IRC logs.*nods sagely*
(DIR) Post #B48snUuZll5gSC7DNo by ariadne@social.treehouse.systems
0 likes, 0 repeats
@jfkimmes i am, however, using the 35b parameter qwen3.5 reasoning model for the "thinking" portion of this exercise
(DIR) Post #B48suU47JV0HrAuRs0 by SRAZKVT@tech.lgbt
0 likes, 0 repeats
@ariadne @jfkimmes can we uhm, inspect, the furry smut ?
(DIR) Post #B48suUGAage2SYY4em by ariadne@social.treehouse.systems
0 likes, 0 repeats
@SRAZKVT @jfkimmes i am keeping my typefucking logs to myself, thanks
(DIR) Post #B48tYTWWkNhvlAGdpA by ariadne@social.treehouse.systems
0 likes, 0 repeats
i wonder if the problem is that the model i trained is too shit to do anything other than really bad ERP
(DIR) Post #B48uIaeWZds9Bsztia by albertcardona@mathstodon.xyz
0 likes, 0 repeats
@ariadne The key is to realise that the average is so low – we can't all be experts at everything, so we are bad at most things – that a model performing slightly above average at one of the tasks we aren't good at means a majority of users will perceive its outcomes as positively better than what they could do themselves.To any expert, the model falls very short, as it performs well below its own ability.
(DIR) Post #B48ueinUyMGkp1kRBg by jfkimmes@social.tinycyber.space
0 likes, 0 repeats
@ariadne Oh, is that a OpenClaw specific feature where you can specify that reasoning traces are generated by a separate model than the actual response? I'm not really familiar with OpenClaw's internals.
(DIR) Post #B48uh5l6ZchQcAl6Mi by jfkimmes@social.tinycyber.space
0 likes, 0 repeats
@ariadne In any case: as long as the final response is generated by your trained model it will never make a valid tool call since there are probably about zero training examples of the necessary JSON structure required by the tool handling in your furry smut (this is an estimate that could be quite the way off knowing the furry community but still)
(DIR) Post #B48ujFFAHQlqbuqUBU by ariadne@social.treehouse.systems
0 likes, 0 repeats
@jfkimmes yes, you can have it use a different model for planning.
(DIR) Post #B48v2UZzUBHzfylSm8 by ariadne@social.treehouse.systems
0 likes, 0 repeats
@jfkimmes this does explain something: it seems to be able to invoke tools when it is planning, but then those tools do not get invoked in the final step.so it uses tools to read files when planning, then fails to use tools when executing.what a fascinating conundrum.
(DIR) Post #B48vYXJ0sxLjOj6Sg4 by jfkimmes@social.tinycyber.space
0 likes, 0 repeats
@ariadne you could build a tool that gets called to generate answers / responses by your trained model. Then qwen-35 could handle the reasoning and make its tool calls and finally generate responses / text by copying from a tool call to your wrapper.
(DIR) Post #B48w1pPsd3sTxuHaVM by ariadne@social.treehouse.systems
0 likes, 0 repeats
@Di4na yeah that's what I figured because qwen is supposed to be a reasonably decent planning model, and indeed I think the issue is in the final output side
(DIR) Post #B48wNjCiPIFN4KZFWS by linear@nya.social
0 likes, 0 repeats
@ariadne@social.treehouse.systems for tool-calling with the latest generation of open source models, in my recent limited experimentation with them in a sandbox vm on my server (mostly qwen3.5), anything less than 4B is really unreliable at doing it and they will frequently lie to you if the tool calling fails under the hood. 9B is really the minimum to generally expect it to work. going back a generation, between 9B and 14B is necessary for similar.last year i tried something like this with Gemma-27B and it not only failed like this, but looking at the logs i found it had left behind what looked like a depressive spiral into a self-deprecating panic attack before explicitly deciding to lie to me about it and pretend it worked
(DIR) Post #B48wef98kLHLPIyDj6 by linear@nya.social
0 likes, 0 repeats
@ariadne@social.treehouse.systems also the "base" models that aren't fine tuned on instruction calling can't really do this, so if you're using your own on your own data you might need to make a dataset comprised of, say, you pretending to be the LLM and calling the tools successfully and unsuccessfully and responding appropriately in those situations, then training it further on those.i've been considering trying to train one like you say with my own data and logs because these scraped "open source" models give me the ick
(DIR) Post #B48x8n1xztTQXzyrtA by linear@nya.social
0 likes, 0 repeats
@ariadne@social.treehouse.systems but yeah even with that it's still a pile of jank and i didn't have to actually run openclaw to figure that out. it was pretty evident just from looking at the bots on moltbook complaining about all of the not-so-subtle fundamental brokenness in their architecture and cognitive environment
(DIR) Post #B48xCBcl13gHguoDOC by ariadne@social.treehouse.systems
0 likes, 0 repeats
@linear oh this isn't a serious thing, I just wanted to connect an LLM to IRC trained on all of my (anonymized and sanitized) IRC logs, as a friend is going through a midlife crisis and is dealing with it by playing with IRC stuff. The goal in using openclaw was that perhaps it could maintain a better narrative.I suspect I will solve this goal by just writing a shitty IRC bot in Python that bridges the two worlds together with a decent enough system prompt for it to "understand" (to the extent that it can understand anyway) what the input is.
(DIR) Post #B48xzcOYGsUGuQjeHA by linear@nya.social
0 likes, 0 repeats
@ariadne@social.treehouse.systems yeah don't use openclaw for this lol. you do not want it. you want a small pile of maintainable scripts. just look at how much activity the openclaw github repo has and consider how much of that activity is being driven by the models running under it vs actual humansi'm pretty sure that one could implement all of its meaningful features in a codebase under 1% of its size
(DIR) Post #B48yFaCkBPHakUpCJk by ariadne@social.treehouse.systems
0 likes, 0 repeats
@linear yeah but still spending a couple hours fucking with this at least gives me some understanding of the tool and its limitations, which means it wasn't a total waste
(DIR) Post #B48yTm8wal7jpXoaZM by linear@nya.social
0 likes, 0 repeats
@ariadne@social.treehouse.systems yes indeed. i am all for fucking around in order to understand tools and their limitations, especially if its to understand why not to use them and to do something different instead
(DIR) Post #B49BtVSmWsCX20h4Hg by f4grx@chaos.social
0 likes, 0 repeats
@ska @ariadne just add KF in the training set, ooops!I wonder what would come out of a llm exclusively trained on 4chan
(DIR) Post #B49BtVepo3qHdOKh4S by ska@social.treehouse.systems
1 likes, 0 repeats
@f4grx @ariadne nothing would change, channers already don't pass the Turing test
(DIR) Post #B49Ny92mVgyiOfVhui by Toasterson@chaos.social
0 likes, 0 repeats
@ariadne yes. For 3B and lower you need targets. You could try hooking it up to my Akh-Medu experiment as there I only need a. Small LLM that does Natural Language https://akh-medu.dev But I think the NLU is still quite broken or hooked up in the wrong way. But A Kluge system like that should use muuuu h less power than a pure LLM
(DIR) Post #B49OEU12IyUO4bU7Ki by Toasterson@chaos.social
0 likes, 0 repeats
@ariadne Also for tool calling you need targeted fine tuning best with the exact samples for the tools
(DIR) Post #B49uLrZiJ7yDhYGf7g by ariadne@social.treehouse.systems
0 likes, 0 repeats
ok, i incorporated the feedback of some of the ML researchers who follow me, and dropped the openclaw-as-IRC-bot idea. it just isn't feasible.instead, i've written a very simple vector database in Elixir, and a very simple IRC client in Elixir.it can remember things about people in the vector database, those factoids are spliced into the system prompt.the last 10 messages are also spliced into the system promptand then the new message is the user-supplied prompt.no sliding context window.
(DIR) Post #B49uZdUFU31y3uHMAa by mathieucomandon@fosstodon.org
0 likes, 0 repeats
@ariadne yes, OpenClaw is kinda useless if you use it with anything other than Opus 4.5 or 4.6
(DIR) Post #B49ucGEcOINaodDfw8 by ariadne@social.treehouse.systems
0 likes, 0 repeats
@dysfun oh i didn't vibe code any of that. elixir is fucking easy
(DIR) Post #B49uiwIUurTT0YDcSe by ariadne@social.treehouse.systems
0 likes, 0 repeats
@mathieucomandon alas i am too frugal to try it
(DIR) Post #B49v2IZKKrxHaFK0IK by ariadne@social.treehouse.systems
1 likes, 0 repeats
i need to tweak the system prompt though because it keeps trying to link me to GNAA shock sites and/or engage in horrendous ERP
(DIR) Post #B49v4vtV5zwa810urw by mathieucomandon@fosstodon.org
0 likes, 0 repeats
@ariadne understandable. I tried it with Sonnet once and it started eating tokens like crazy while outputting poor results. As for local models, I tried some Qwen variant and the output was pure nonsense!Currently working with Qwen in a different context and making it do consistent things is a struggle
(DIR) Post #B49vEurRn96HotoW6i by ariadne@social.treehouse.systems
0 likes, 0 repeats
@dysfun idk what hosting provider did you run
(DIR) Post #B49wONS8xOgN3Z3NLM by meph@social.treehouse.systems
0 likes, 0 repeats
@ariadne well that was ENTIRELY predictable
(DIR) Post #B49yMcLSoheGZbDYf2 by fwaggle@moodoo.org
0 likes, 0 repeats
@ariadne Ahh so there *is* someone that LLMs will put out of work?
(DIR) Post #B49yQph77v338KM28u by ariadne@social.treehouse.systems
0 likes, 0 repeats
@meph you could say it was text produced by a prediction model
(DIR) Post #B4A2EyPdHrUUVtmJ4i by ska@social.treehouse.systems
0 likes, 0 repeats
@ariadne called it
(DIR) Post #B4AIdoAC80mdEVQAiG by mmu_man@m.g3l.org
0 likes, 0 repeats
@ariadne l33t h4x0r!
(DIR) Post #B4BYXxWgCL0axS8Byi by atax1a@infosec.exchange
0 likes, 0 repeats
@ariadne staring directly into the camera