Post
44
Jeb update for the week. Most of what changed came out of a review.
Dipankar went back to the per-question eval records we publish, re-tempered them himself, and found the 4B and 9B were still shipping the score temperature fit to the training target. The 27B had already moved off it.
So on 2026-09-29 we refit the 4B and 9B without retraining. Score questions now use the hard fit, T 0.83 on both. Choice and noul stay where they were, and accuracy doesn't move. Across the 20 evaluation sets, question-weighted ECE goes from 0.0863 to 0.0800 on the 4B and from 0.0796 to 0.0709 on the 9B.
It isn't free. Jevals HelpSteer2 gets worse, 0.048 to 0.074 on the 4B and 0.039 to 0.082 on the 9B.
Then we ran the check he pushed us toward. HelpSteer2 has 418 validation rows that none of our reported eval sets use and that never went near the training pool. Those rows want a score temperature of 1.26 on the 4B and 1.29 on the 9B, not 0.83.
So the problem is how our calibration split drew HelpSteer2. If your traffic looks like HelpSteer2, keep the older train fit (1.20 on the 4B, 1.22 on the 9B, both in temperatures.json) or refit on your own labels. Both model cards say that now. For v3, the HelpSteer2 part of calibration comes from held-out validation data.
The community quants have also gone further than I expected. bartowski's 9B GGUF is at 3,542 downloads, and mradermacher's 4B, 9B and 27B GGUFs are at 600, 860 and 779. That's more than our own repos put together. Thank you both. And thank you Dipankar, that was exactly the kind of review I hoped for when we put the records up.
A question for anyone running Jeb on real traffic: are you using the shipped temperatures as they are, or refitting on your own labels? If you refit, I'd like to know what you ended up with.
The review thread: https://huggingface.co/posts/jbrashear/856060098865438
Dipankar went back to the per-question eval records we publish, re-tempered them himself, and found the 4B and 9B were still shipping the score temperature fit to the training target. The 27B had already moved off it.
So on 2026-09-29 we refit the 4B and 9B without retraining. Score questions now use the hard fit, T 0.83 on both. Choice and noul stay where they were, and accuracy doesn't move. Across the 20 evaluation sets, question-weighted ECE goes from 0.0863 to 0.0800 on the 4B and from 0.0796 to 0.0709 on the 9B.
It isn't free. Jevals HelpSteer2 gets worse, 0.048 to 0.074 on the 4B and 0.039 to 0.082 on the 9B.
Then we ran the check he pushed us toward. HelpSteer2 has 418 validation rows that none of our reported eval sets use and that never went near the training pool. Those rows want a score temperature of 1.26 on the 4B and 1.29 on the 9B, not 0.83.
So the problem is how our calibration split drew HelpSteer2. If your traffic looks like HelpSteer2, keep the older train fit (1.20 on the 4B, 1.22 on the 9B, both in temperatures.json) or refit on your own labels. Both model cards say that now. For v3, the HelpSteer2 part of calibration comes from held-out validation data.
The community quants have also gone further than I expected. bartowski's 9B GGUF is at 3,542 downloads, and mradermacher's 4B, 9B and 27B GGUFs are at 600, 860 and 779. That's more than our own repos put together. Thank you both. And thank you Dipankar, that was exactly the kind of review I hoped for when we put the records up.
A question for anyone running Jeb on real traffic: are you using the shipped temperatures as they are, or refitting on your own labels? If you refit, I'd like to know what you ended up with.
The review thread: https://huggingface.co/posts/jbrashear/856060098865438