IkariDev's picture
👋 Open to Work

IkariDev

IkariDev
P(doom) 10%

AI & ML interests

None yet

Recent Activity

reacted to Undi95's post with 🔥 about 8 hours ago
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek. I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use? No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note. The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match. Later, code will be checked by actually running tests. The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights. Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again. The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move. I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks. Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try. Did you already tried something like that? What was your result? I'm curious!
reacted to Undi95's post with ❤️ 25 days ago
Hi! I will get off the internet for a moment. I launched the Hanami Project because I didn't supported SillyTavern UI anymore atm. Too much options for my dead brain, still very good, but I wanted more simple, professional, phone accessible and sober front end for when I will be gone from home. I did my maximum to finish it before I go, I want you to have it, I want my work to be used (even if it's AI slop for some of you) for who care. Here's the github repo: https://github.com/Undi95/Hanami If you have any suggestion, bugs report, pull request or anything, post it, if you want to modify it, fork it, but keep the credit, and add myself haha. If you search an option, a function, you will find it. But at first, the front end will be what you expect: minimalist, but customizable, empty at first. Navigate to see all it can do. Everything is well organized. Context is full ? No worries anymore, with memory file, files access, auto compaction and smooth transition, you can continue your chat like nothing happened. (Inspired from Claude) The front end have a final option for everyone : The tools calling for action and emotion could be a bit too much for smaller model, you can, in this case, use the "Simple" option in Settings > Model > Model mode. "Simple: no tools are exposed to the model — Hanami handles memory server-side (facts are extracted during compaction) and guesses the emotion from the text. Pick this for small models, which often fail at tool calling." My last gift for myself, and for you. Cya!
View all activity

Organizations

Pygmalion's profile picture Caldera AI's profile picture OpenOrca's profile picture CyberHarem's profile picture The Waifu Research Department's profile picture M.O.F.U.'s profile picture NeverSleep's profile picture Social Post Explorers's profile picture NeverSleep - Historical's profile picture The Chaotic Neutrals's profile picture dreamgen-preview's profile picture He-He's profile picture MergeFuel's profile picture Edgerunners's profile picture kalo-team's profile picture Anthracite Core's profile picture The Smug Clovers's profile picture Nuwa's profile picture Anthracite's profile picture SillyTilly's profile picture AnimeResearchModels's profile picture AnimeResearchModels Private's profile picture