Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
SaylorTwift
's Collections
benchmarks
RULER Datasets Falcon-H1-3B-Base
RULER Datasets Lamma3-Instruct
RULER Datasets Qwen2.5-Instruct
RULER Datasets Qwen-3-Instruct
RULER Datasets Qwen-3
agents
Agents ressources
benchmarks
updated
18 days ago
Upvote
-
Sort: Collection
meituan-longcat/LARYBench
Updated
Apr 30
•
1.75k
•
20
llamaindex/ParseBench
Benchmark
•
Updated
Apr 19
•
169k
•
10.4k
•
106
nvidia/QCalEval
Viewer
•
Updated
Apr 13
•
243
•
1.61k
•
19
allenai/olmOCR-bench
Benchmark
•
Updated
Feb 19
•
7.15k
•
267
LongHorizonReasoning/longcot
Viewer
•
Updated
Apr 20
•
5k
•
317
•
13
mercor/apex-agents
Benchmark
•
Updated
Jun 11
•
480
•
47.1k
•
142
hsiung/MagicBench
Viewer
•
Updated
Apr 18
•
50
•
19
•
10
openlifescienceai/medmcqa
Viewer
•
Updated
Jan 4, 2024
•
193k
•
42.5k
•
232
openai/healthbench-professional
Viewer
•
Updated
Apr 22
•
525
•
1.34k
•
55
ShadenA/MathNet
Viewer
•
Updated
Jun 16
•
55.6k
•
9.94k
•
90
claw-eval/Claw-Eval
Benchmark
•
Updated
May 8
•
2.88k
•
30
FrontierCS/Frontier-CS
Viewer
•
Updated
6 days ago
•
268
•
4.35k
•
6
ByteDance-Seed/EdgeBench
Viewer
•
Updated
12 days ago
•
51
•
9.12k
•
78
Upvote
-
Sort: Collection
Share collection
View history
Collection guide
Browse collections