Datasets:
Tasks:
Text Classification
Formats:
parquet
Languages:
English
Size:
1K - 10K
Tags:
structured-decisions
calibration
probabilistic-classification
system-one
workflow-evaluation
Synthetic
License:
MAV (specialist, 421M, Laya-RLCD reproduction on one RTX 4090): 0.7800 acc / KL 0.086 / Brier 0.050 on the official test split
#12 opened 1 day ago
by
htob
ProKope-421M (Alan Canto): 0.714 acc / 0.071 Brier / 0.254 Score MAE on official test split
3
#11 opened 3 days ago
by
AlanCantoFTW
Call for results: add your decision model to the leaderboard
#10 opened 3 days ago
by
codelion
Leaderboard submission: OpenDecider (open weights, Apache-2.0), one zero-shot model and four fine-tuned on `train`
👍 1
1
#8 opened 6 days ago
by
manjunathshiva
Bongard-mini: general, zero-shot, open weights (accuracy 0.594, KL 0.256, Brier 0.132)
👍 1
1
#7 opened 6 days ago
by
0xDing
Japanese translation: GeneLab/typed-decisions-ja
👍 1
#6 opened 6 days ago
by
GeneLab
prima-ratio + 12B (generalist, zero-shot, default calibration): 0.702 acc / KL 0.564 / Brier 0.234 / ECE 0.146
🚀 1
1
#5 opened 7 days ago
by
j3st3r666
od1-typed-decisions (specialist, 4B): 0.7965 acc / KL 0.082 / Brier 0.045 on the official test split
🔥 1
1
#4 opened 8 days ago
by
mvbalaji
soft-decider-421m (specialist): 0.774 acc / 0.141 ECE on the official test split
👍 1
1
#3 opened 11 days ago
by
winwinwinbb
Laya benchmark results
👍 3
2
#2 opened 17 days ago
by
convaiinnovations
[bot] Conversion to Parquet
#1 opened 17 days ago
by
parquet-converter