Post
33
Jev-style decisions, now at 2B: Jev-Style-2B-Decision-v3
State and typed questions in (pick one, yes/no, score); a calibrated probability for every option out, in one pass.
Try it in your browser:
chaoliangUNSW/jev-style-2b
• JevBench v1.4.1, 231 public items (self-run with the official v1.4.1 harness, not an official board entry): 73.6% (170/231), the highest public accuracy among the Qwen3.5-2B-family systems on the board (decider-2b 71.0%, open-jev-zefan-2b 64.5%; the lead over decider-2b is inside the 95% CI, 67.6–78.9%). 42 of the 82 board systems score higher, almost all 4B or larger. +9.5 points over our 0.8B v3. Jev 1.13 is well ahead at 86.6%.
• tweet_topic, zero-shot (n = 1,693): 82.2% accuracy [80.4, 84.0], above Jev's 79.3%; macro-F1 is below Jev (0.678 vs 0.694).
• fin_topic, zero-shot (n = 4,117): 61.1%, +14.4 points over our 0.8B v3 (Jev: 67.0%).
• 25,600 tokens per call, no option cap, nothing truncated.
• Read once, then ask: on an Apple M1 Max (GGUF F16), the first question about a 24,501-token input took 16.2 s and a further question about the same state 0.17 s (shared machine, indicative).
Contamination checks against the training pool: 0 hits for the JevBench items, and no exact or near-duplicate overlap with the tweet_topic and fin_topic test sets. Needs the shipped runtimes (PyTorch, llama.cpp with the bundled scorer, MLX); trained on a reduced data pool (60M tokens).
Models (Apache-2.0; some training data has restrictive or unclear terms and includes OpenAI GPT and Anthropic Claude outputs, see "Training data and licences" on the main card):
chaoliangUNSW/Jev-Style-2B-Decision-v3
chaoliangUNSW/Jev-Style-2B-Decision-v3-GGUF
chaoliangUNSW/Jev-Style-2B-Decision-v3-MLX
All v3 builds: chaoliangUNSW/jev-style-decision-v3-08b-2b-6ab87f32380cbd8c03b608b9
Not affiliated with TypeSafe or Jev.
State and typed questions in (pick one, yes/no, score); a calibrated probability for every option out, in one pass.
Try it in your browser:
chaoliangUNSW/jev-style-2b
• JevBench v1.4.1, 231 public items (self-run with the official v1.4.1 harness, not an official board entry): 73.6% (170/231), the highest public accuracy among the Qwen3.5-2B-family systems on the board (decider-2b 71.0%, open-jev-zefan-2b 64.5%; the lead over decider-2b is inside the 95% CI, 67.6–78.9%). 42 of the 82 board systems score higher, almost all 4B or larger. +9.5 points over our 0.8B v3. Jev 1.13 is well ahead at 86.6%.
• tweet_topic, zero-shot (n = 1,693): 82.2% accuracy [80.4, 84.0], above Jev's 79.3%; macro-F1 is below Jev (0.678 vs 0.694).
• fin_topic, zero-shot (n = 4,117): 61.1%, +14.4 points over our 0.8B v3 (Jev: 67.0%).
• 25,600 tokens per call, no option cap, nothing truncated.
• Read once, then ask: on an Apple M1 Max (GGUF F16), the first question about a 24,501-token input took 16.2 s and a further question about the same state 0.17 s (shared machine, indicative).
Contamination checks against the training pool: 0 hits for the JevBench items, and no exact or near-duplicate overlap with the tweet_topic and fin_topic test sets. Needs the shipped runtimes (PyTorch, llama.cpp with the bundled scorer, MLX); trained on a reduced data pool (60M tokens).
Models (Apache-2.0; some training data has restrictive or unclear terms and includes OpenAI GPT and Anthropic Claude outputs, see "Training data and licences" on the main card):
chaoliangUNSW/Jev-Style-2B-Decision-v3
chaoliangUNSW/Jev-Style-2B-Decision-v3-GGUF
chaoliangUNSW/Jev-Style-2B-Decision-v3-MLX
All v3 builds: chaoliangUNSW/jev-style-decision-v3-08b-2b-6ab87f32380cbd8c03b608b9
Not affiliated with TypeSafe or Jev.