Peter Kotsch
AI & ML interests
Recent Activity
Organizations
9B variants?
A crazy thinking model
Mtp quality? π€
WINNER , Qwen 3.8 27B version(s) // new versions plus GAIN and COLD FUSION training method (and models).
Qwen 3.8-27b context usage is about 10x of 3.6-27b
DFlash draft model instead of MTP one
Interestingly, the just released 3.8 version has exactly the same architecture - meaning all the capability gains come from training improvements!
See the diff (0 changes) here!
https://hfviewer.com/compare/qwen3.6-27b-vs-qwen3.8-27b
Qwen 3.8 27B
wow
First! :)
I'm really pleased π
No DFlash?
thinking level
Yes. It is on the list.
We are still revising the 9-14B pipeline.
We do have an interm Qwen 3.5 9B from the experimental pipeline here:It matches/exceeds 27B Qwen 3.5 ; and meets in some cases 27B Qwen 3.6 performance.
There is still a lot of optimizations to do at this time.
That's insane. Thanks. I've been using that model : Qwen3.5-9B-The-Defiant-Fable-Uncnr-Heretic-NEO-MAX-MTP-IQ4_XS
and yes its very performant already. It blew all other 9B models when it comes to my very specific benchmark task. Basically created solution in one shot! Can't believe this was possible. the 3.6 9B should be able to do even better, probably could chain it for reasonable tasks inside a harness loop to outmatch bigger models.
is there any chance you will train Qwen 3.6 9B version same way too?