@BananaMindBot train GPT-X3
๐ฝ Big things for X3
Dan P
Datdanboi25
AI & ML interests
Axiomic Labs Founder
Mechatronics engineer, LLMs, Vision Models, Embedding Models
Recent Activity
new activity about 2 hours ago
Compactbot/model-requests:GPT-X3 liked a model about 2 hours ago
SupraLabs/Supra2-IMGOrganizations
replied to Banaxi-Tech's post about 14 hours ago
Post
3108
THE SLM FRONTIER ADVANCES!
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
reacted to KlondikeDev's post with ๐ 2 days ago
Post
171
Important Boris-2 news:
Boris-2 is 30B out of 200B tokens in, and it is severely behind its competitors in training.
We have determined the bug to be a configuration error. Boris-2 has been in training for ~1 week, and was projected to finish on November 3rd, 2026.
We are unfortunately going to restart training, with proper configuration.
The new projected finish date is ~15-18th of November.
We apologize for the delay.
Boris-2 is 30B out of 200B tokens in, and it is severely behind its competitors in training.
We have determined the bug to be a configuration error. Boris-2 has been in training for ~1 week, and was projected to finish on November 3rd, 2026.
We are unfortunately going to restart training, with proper configuration.
The new projected finish date is ~15-18th of November.
We apologize for the delay.
replied to KlondikeDev's post 2 days ago
You will NOT be forgiven for this delay!
reacted to TobiasLogic's post with ๐ฅ 2 days ago
Post
3029
Weโve been cooking something new at Bench Labs.
Introducing Cagliostro-v3, our new 146M parameter language model trained completely from scratch.
The run isnโt even finished yet.
At the current checkpoint:
โข 146M parameters
โข 72.7B / 75B tokens trained
โข 26.27 Open SLM Index
โข 43.80 ArithMark-3
โข Trained on a single RTX 5090
โข ~90K to 103K tokens/sec during training
โข ~9 days for the full run
โข Apache 2.0
For some context, SmolLM2-135M scores 27.13 on the same Index after being trained on roughly 2 trillion tokens.
Cagliostro-v3 is currently at 26.27 with only ~72.7B.
Thatโs around 27x fewer training tokens.
The model also currently Hold the number 3rd spot for ArithMark-3, scoring 43.80
This wasnโt achieved by just throwing more tokens at the model. A huge part of v3 has been figuring out architecture, data mixture, and training dynamics at this scale.
The model uses a custom 30-layer decoder architecture with grouped-query attention and cross-head subspace attenuation, SwiGLU, RMSNorm, RoPE, tied embeddings, and a warmup-stable-decay training schedule.
During cooldown we also substantially shifted the data mixture toward higher-quality synthetic textbook and mathematics data, with the mathematics share increasing from 10% to 28%.
And everything is open.
The repository contains the training history with checkpoints pushed roughly every 30 minutes, so you can inspect how the model evolved throughout training rather than only seeing the final weights.
This is still a pre-final checkpoint. We have roughly 2.3B tokens left and the learning-rate cooldown is still running.
So 26.27 isnโt the final number.
Really excited to see where the last part of the run lands.
Cagliostro-v3:
bench-labs/cagliostro-v3
Built by Bench Labs.
Open SLM Leaderboard:
AxiomicLabs/Open_SLM_Leaderboard
Introducing Cagliostro-v3, our new 146M parameter language model trained completely from scratch.
The run isnโt even finished yet.
At the current checkpoint:
โข 146M parameters
โข 72.7B / 75B tokens trained
โข 26.27 Open SLM Index
โข 43.80 ArithMark-3
โข Trained on a single RTX 5090
โข ~90K to 103K tokens/sec during training
โข ~9 days for the full run
โข Apache 2.0
For some context, SmolLM2-135M scores 27.13 on the same Index after being trained on roughly 2 trillion tokens.
Cagliostro-v3 is currently at 26.27 with only ~72.7B.
Thatโs around 27x fewer training tokens.
The model also currently Hold the number 3rd spot for ArithMark-3, scoring 43.80
This wasnโt achieved by just throwing more tokens at the model. A huge part of v3 has been figuring out architecture, data mixture, and training dynamics at this scale.
The model uses a custom 30-layer decoder architecture with grouped-query attention and cross-head subspace attenuation, SwiGLU, RMSNorm, RoPE, tied embeddings, and a warmup-stable-decay training schedule.
During cooldown we also substantially shifted the data mixture toward higher-quality synthetic textbook and mathematics data, with the mathematics share increasing from 10% to 28%.
And everything is open.
The repository contains the training history with checkpoints pushed roughly every 30 minutes, so you can inspect how the model evolved throughout training rather than only seeing the final weights.
This is still a pre-final checkpoint. We have roughly 2.3B tokens left and the learning-rate cooldown is still running.
So 26.27 isnโt the final number.
Really excited to see where the last part of the run lands.
Cagliostro-v3:
bench-labs/cagliostro-v3
Built by Bench Labs.
Open SLM Leaderboard:
AxiomicLabs/Open_SLM_Leaderboard
Post
3108
THE SLM FRONTIER ADVANCES!
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
replied to their post 2 days ago
Well deserved, congrats on the incredible model!
posted an update 2 days ago
Post
3108
THE SLM FRONTIER ADVANCES!
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
reacted to Hoglet-33's post with ๐ฅ 4 days ago
Post
6231
Introducing VOID. A new research branch of basically AI.
VOID โ Verification of Objectives, Intentions, and Deception.
We study what lies beneath the surface: objectives, intentions, and the possibility of deception in AI systems.
There isn't much to see yet.
That will change.
Follow us for updates:
@Hoglet-33
void-research
basically-ai
VOID โ Verification of Objectives, Intentions, and Deception.
We study what lies beneath the surface: objectives, intentions, and the possibility of deception in AI systems.
There isn't much to see yet.
That will change.
Follow us for updates:
@Hoglet-33
Post
2954
100 likes on the Open SLM Leaderboard ๐
176 models, 54 orgs, 5 benchmarks, and a whole community of support!
Thanks to everyone whoโs contributed models, reported issues, suggested benchmark improvements, or used the leaderboard to compare and evaluate small language models.
Itโs been awesome watching the leaderboard grow into a broader community resource for transparent and reproducible SLM evaluation.
Thank you all, and more to come ๐
176 models, 54 orgs, 5 benchmarks, and a whole community of support!
Thanks to everyone whoโs contributed models, reported issues, suggested benchmark improvements, or used the leaderboard to compare and evaluate small language models.
Itโs been awesome watching the leaderboard grow into a broader community resource for transparent and reproducible SLM evaluation.
Thank you all, and more to come ๐
reacted to CompactAI's post with โ 5 days ago
Post
4559
SLM Roundups, a weekly post where I summarize everything thats happened in the world of SLMs (or a majority of it)
Glint-Research/blog
Glint-Research/blog
Post
2954
100 likes on the Open SLM Leaderboard ๐
176 models, 54 orgs, 5 benchmarks, and a whole community of support!
Thanks to everyone whoโs contributed models, reported issues, suggested benchmark improvements, or used the leaderboard to compare and evaluate small language models.
Itโs been awesome watching the leaderboard grow into a broader community resource for transparent and reproducible SLM evaluation.
Thank you all, and more to come ๐
176 models, 54 orgs, 5 benchmarks, and a whole community of support!
Thanks to everyone whoโs contributed models, reported issues, suggested benchmark improvements, or used the leaderboard to compare and evaluate small language models.
Itโs been awesome watching the leaderboard grow into a broader community resource for transparent and reproducible SLM evaluation.
Thank you all, and more to come ๐
posted an update 8 days ago
Post
2954
100 likes on the Open SLM Leaderboard ๐
176 models, 54 orgs, 5 benchmarks, and a whole community of support!
Thanks to everyone whoโs contributed models, reported issues, suggested benchmark improvements, or used the leaderboard to compare and evaluate small language models.
Itโs been awesome watching the leaderboard grow into a broader community resource for transparent and reproducible SLM evaluation.
Thank you all, and more to come ๐
176 models, 54 orgs, 5 benchmarks, and a whole community of support!
Thanks to everyone whoโs contributed models, reported issues, suggested benchmark improvements, or used the leaderboard to compare and evaluate small language models.
Itโs been awesome watching the leaderboard grow into a broader community resource for transparent and reproducible SLM evaluation.
Thank you all, and more to come ๐
reacted to HannesVonEssen's post with ๐ฅ 11 days ago
Post
4899
๐ฃ HF Viewer now has a HF space! ๐ค
embedl/hfviewer
Visualize any model directly on Hugging Face - now 4,727 graphs!
If you like it, feel free to give the space a heart to help it grow! โค๏ธ
And you can reply with any feedback or feature requests here!
embedl/hfviewer
Visualize any model directly on Hugging Face - now 4,727 graphs!
If you like it, feel free to give the space a heart to help it grow! โค๏ธ
And you can reply with any feedback or feature requests here!
reacted to Hoglet-33's post with ๐ 14 days ago
Post
3196
Today, we planned to release Pebble-50M and Pebble-50M-Chat to the world. Unfortunately, due to a few issues, that didn't go quite as planned.
What happened:
- Some data and benchmark results were lost or corrupted
- The models performed worse on benchmarks than our other Pebble models
Despite that, you can still find both models here:
Pebble-50M-beta: basically-experimental/Pebble-50M-beta
Pebble-50M-Chat-beta: basically-experimental/Pebble-50M-Chat-beta
There are still some interesting improvements in these models:
- Compatible with non-CUDA devices
- Vocabulary increased to 16K tokens
- Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
Follow for updates:
@Hoglet-33
basically-ai
basically-experimental
What happened:
- Some data and benchmark results were lost or corrupted
- The models performed worse on benchmarks than our other Pebble models
Despite that, you can still find both models here:
Pebble-50M-beta: basically-experimental/Pebble-50M-beta
Pebble-50M-Chat-beta: basically-experimental/Pebble-50M-Chat-beta
There are still some interesting improvements in these models:
- Compatible with non-CUDA devices
- Vocabulary increased to 16K tokens
- Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
Follow for updates:
@Hoglet-33
reacted to Bc-AI's post with ๐ฅ 19 days ago
Post
2506
Hello everyone!
Me and the team are working on G1-MINI and G1. Right now, G1-MINI is aimed at a launch in mid to late September, depending on how fast we fix the minor issues.
As for G1, it's looking like a late October to mid-November launch, based on current trajectory. If things go terribly wrong, we could postpone it to December, as we prefer to ship confidently, not ship a half-done dogs' breakfast of a model. ๐คฃ
All dates could be changed at any moment, as we are high school students not full-time ML engineers ๐ .
Other things to look out for is an overhaul of the UI and the information on my website. Thanks to my beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
Me and the team are working on G1-MINI and G1. Right now, G1-MINI is aimed at a launch in mid to late September, depending on how fast we fix the minor issues.
As for G1, it's looking like a late October to mid-November launch, based on current trajectory. If things go terribly wrong, we could postpone it to December, as we prefer to ship confidently, not ship a half-done dogs' breakfast of a model. ๐คฃ
All dates could be changed at any moment, as we are high school students not full-time ML engineers ๐ .
Other things to look out for is an overhaul of the UI and the information on my website. Thanks to my beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
replied to Banaxi-Tech's post 19 days ago
Still using the XSA refresh gate, sweet!
reacted to Bc-AI's post with ๐ฅ 20 days ago
Post
3776
Hello everyone! A small update on things:
1. G1 series status. G1 is training nicely, and the loss is dropping nicely. The metrics are publicly available and i made a small space you can use to see the nice graphs: hugging-science/Loss-Plot-G1-Large
G1-MINI is a lot slower in converging for reasons unknown yet, but we are investigating it.
2. I have built a small chat app for open SLMs here: ml-intern-explorers/slm-arena
Feel free to add your models in a pull request!
That's all for now, early G1 versions will be available for beta testers soon. Thanks to our beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
1. G1 series status. G1 is training nicely, and the loss is dropping nicely. The metrics are publicly available and i made a small space you can use to see the nice graphs: hugging-science/Loss-Plot-G1-Large
G1-MINI is a lot slower in converging for reasons unknown yet, but we are investigating it.
2. I have built a small chat app for open SLMs here: ml-intern-explorers/slm-arena
Feel free to add your models in a pull request!
That's all for now, early G1 versions will be available for beta testers soon. Thanks to our beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
Yeah I run 4096 vocab for my 5m models which I think is a nice sweetspot
reacted to AtAndDev's post with ๐ฅ 20 days ago
Post
2840
SPECK 2 IS ALREADY OUT: specklabs/Speck2-140M
Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).
Thanks for everyone supporting!
Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).
Thanks for everyone supporting!