Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Edit Datasets filters
Main
Tasks
Libraries
Languages
Licenses
Other
1
Reset Other
ai-act
Synthetic
art
code
medical
finance
biology
legal
chemistry
agent
climate
music
Apply filters
Datasets
5,233
Full-text search
Edit filters
Sort: Trending
Active filters:
evaluation
Clear all
atroposhealth/precision-evidence-bench
Viewer
•
Updated
3 days ago
•
209
•
73
•
4
Alibaba-Aone/aacr-bench
Viewer
•
Updated
Feb 2
•
2.15k
•
525
•
9
Qwen/AgentWorldBench
Viewer
•
Updated
Jul 4
•
2.17k
•
1.3k
•
108
TIGER-Lab/MMLU-Pro
Benchmark
•
Updated
May 2
•
12.1k
•
247k
•
514
nvidia/compute-eval
Viewer
•
Updated
5 days ago
•
3.1k
•
700
•
30
MiniMaxAI/OctoCodingBench
Viewer
•
Updated
Jan 13
•
72
•
412
•
358
treadon/abliteration-eval
Viewer
•
Updated
Apr 14
•
283
•
255
•
3
surgeai/GDP.pdf
Viewer
•
Updated
8 days ago
•
100
•
58.4k
•
21
treadon/disinhibition-eval
Viewer
•
Updated
Apr 30
•
248
•
31
•
2
harborframework/terminal-bench-2.1
Benchmark
•
Updated
6 days ago
•
150k
•
11
microsoft/WorkflowPerturb
Viewer
•
Updated
8 days ago
•
44.9k
•
175
•
3
jaelly/MCJudgeBench
Viewer
•
Updated
14 days ago
•
219
•
149
•
4
alexshpunt/explicit-edit-benchmark
Viewer
•
Updated
about 22 hours ago
•
80
•
4.82k
•
2
Agent-as-Policy/agent-as-policy
Viewer
•
Updated
3 days ago
•
162
•
677
•
2
openthaigpt/thai-ocr-evaluation
Viewer
•
Updated
Sep 30, 2024
•
104
•
170
•
9
Salesforce/CRMArenaPro
Viewer
•
Updated
Jul 9, 2025
•
8.61k
•
2.11k
•
18
Naholav/claude_4_math_evaluation_500
Preview
•
Updated
Jul 7, 2025
•
65
•
1
m-a-p/Encyclo-K
Viewer
•
Updated
Feb 9
•
5.04k
•
82
•
5
GSMA/ot-full
Viewer
•
Updated
Mar 26
•
20.6k
•
891
•
7
harborframework/terminal-bench-2.0
Benchmark
•
Updated
Apr 24
•
86.8k
•
51
KRAFTON/ArtiBench
Viewer
•
Updated
Feb 25
•
1k
•
963
•
6
rl-rag/hle-gpt-oss-120b-no-python-260222
Viewer
•
Updated
Feb 25
•
9.71k
•
3.8k
•
1
PhysionLabs/Physion-Eval
Preview
•
Updated
Jun 14
•
747
•
19
llamaindex/ParseBench
Benchmark
•
Updated
Apr 19
•
169k
•
22k
•
128
lthn/MMLU-Pro
Viewer
•
Updated
Apr 10
•
12.1k
•
64
•
1
EverMind-AI/EvoAgentBench
Viewer
•
Updated
Jul 16
•
578
•
311
•
17
besimple-ai/voice-code-bench
Viewer
•
Updated
8 days ago
•
300
•
713
•
13
RedactionBench/RedactionBench
Viewer
•
Updated
May 21
•
200
•
285
•
2
PaintBenchAnonymousNeurIPS26/PaintBench
Viewer
•
Updated
May 7
•
2.73k
•
35
•
1
Qwen/Qwen-Image-Bench
Viewer
•
Updated
May 28
•
1k
•
7.17k
•
48
Previous
1
2
3
...
100
Next