AI Benchmarks & Leaderboards Put things like:
AI Benchmarks & Leaderboards Put things like:
- Collection3 items • Updated
-
openai/openai_humaneval
Viewer • Updated • 164 • 300k • 402 -
bigcode/humanevalpack
Viewer • Updated • 984 • 11.8k • 92 -
Human & GPT-4 Evaluation of LLMs Leaderboard
👩78 -
fingertap/GPQA-Diamond
Viewer • Updated • 198 • 6.61k • 15 -
hendrydong/gpqa_diamond
Viewer • Updated • 198 • 3.81k • 10 -
Idavidrein/gpqa
Benchmark • Updated • 1.25k • 127k • 522 -
Arena Leaderboard
🏆4.99kView the LMArena leaderboard in full‑screen
-
lmsys/chatbot_arena_conversations
Viewer • Updated • 33k • 3.25k • 480 -
lmarena-ai/search-arena-24k
Viewer • Updated • 24.1k • 462 • 40 -
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Paper • 2408.04682 • Published • 18 -
Tool Sandbox
🔁Play a looping sandbox demo video
-
tarsur385/toolsandbox-data
Updated • 70 -
Open LLM Leaderboard
🏆14.1kTrack, rank and evaluate open LLMs and chatbots
-
open-llm-leaderboard/results
Preview • Updated • 9.66k • 20 -
BigCodeBench Evaluator
🥇21Evaluate code samples using specified parameters
-
bigcode/bigcodebench-hard
Viewer • Updated • 740 • 8.74k • 3 -
BigCodeBench Leaderboard
🥇233Explore code-generation model leaderboards and task details
-
bigcode/bigcodebench-results
Viewer • Updated • 202 • 842 • 2 -
bigcode/bigcodebench
Viewer • Updated • 5.7k • 92.6k • 87 -
bubbleresearch/bigcodebench-plus
Viewer • Updated • 1.14k • 190 -
bigcode/bigcodebench-instruct-perf
Updated • 566 -
nuprl/MultiPL-E
Viewer • Updated • 12.7k • 50.4k • 71 -
bigcode/MultiPL-E-completions
Viewer • Updated • 20.3k • 1.86k • 8 -
tokhey/qwen2.5-1.5b-mbpp-reasoning-sft
Text Generation • Updated • 22 -
mradermacher/Llama-3.1-Nemotron-8B-mbpp-reasoning-GGUF
8B • Updated • 334 -
cruxeval-org/cruxeval
Viewer • Updated • 800 • 15.3k • 21 -
xhwl/cruxeval-x
Viewer • Updated • 13.5k • 206 • 4 -
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
Paper • 2401.03065 • Published • 11 -
CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution
Paper • 2408.13001 • Published -
evalplus/mbppplus
Viewer • Updated • 378 • 68.8k • 19 -
evalplus/humanevalplus
Viewer • Updated • 164 • 34.3k • 23 -
evalplus/evalperf
Viewer • Updated • 120 • 4.24k • 3 -
claudios/evalplus__mbppplus
Viewer • Updated • 378 • 18 -
google-research-datasets/mbpp
Viewer • Updated • 1.4k • 496k • 238 -
chrisjcundy/mbpp-synthetic-v3
Viewer • Updated • 2.66k • 3.53k -
ellamind/mbpp-multilingual
Viewer • Updated • 498 • 1.55k -
allenai/multilingual_mbpp
Viewer • Updated • 16.6k • 2.84k • 2 -
TIGER-Lab/MMLU-Pro
Benchmark • Updated • 12.1k • 234k • 512 -
MMLU-Pro Leaderboard
🥇256More advanced and challenging multi-task evaluation
-
cais/hle
Benchmark • Updated • 2.5k • 39.9k • 951 -
HLE Leaderboard for Agents with Tools
🥇8Humanity's Last Exam Leaderboard for LLM Agents with Tools
-
skylenage-ai/HLE-Verified
Viewer • Updated • 2.5k • 46.8k • 19 -
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
Paper • 2603.24621 • Published • 2 -
ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
Paper • 2505.11831 • Published • 1 -
ARC-AGI Without Pretraining
Paper • 2512.06104 • Published • 1 -
MathArena/aime_2025
Viewer • Updated • 30 • 28.7k • 17 -
MathArena/aime_2025_II
Viewer • Updated • 15 • 494 -
opencompass/AIME2025
Viewer • Updated • 30 • 17.6k • 56 -
yentinglin/aime_2025
Viewer • Updated • 60 • 33.5k • 12 -
openai/gsm8k
Benchmark • Updated • 17.6k • 1.24M • 1.6k -
ellamind/gsm8k-platinum-multilingual
Viewer • Updated • 5.34k • 1.87k • 1 -
gaia-benchmark/GAIA
Viewer • Updated • 932 • 10.2k • 786 -
GAIA Leaderboard
🦾624Submit and view GAIA model evaluation leaderboard
-
meta-agents-research-environments/gaia2
Viewer • Updated • 963 • 13.5k • 45 -
meta-agents-research-environments/gaia2_filesystem
Viewer • Updated • 9 • 23k • 1 -
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
Paper • 2510.02418 • Published • 2 -
WebArena: A Realistic Web Environment for Building Autonomous Agents
Paper • 2307.13854 • Published • 27 -
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Paper • 2404.07972 • Published • 52 -
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Paper • 2606.29537 • Published • 24 -
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
Paper • 2506.16042 • Published -
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
Paper • 2510.24563 • Published • 23 -
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26 -
eth-sri/agentbench
Viewer • Updated • 138 • 664 • 1 -
iFurySt/AgentBench
Viewer • Updated • 144 • 116 • 1 -
MMMU/MMMU_Pro
Benchmark • Updated • 5.19k • 22k • 66 -
MMMU/MMMU
Viewer • Updated • 11.6k • 75.2k • 333 -
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Paper • 2311.16502 • Published • 40 -
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Paper • 2409.02813 • Published • 34 -
AI4Math/MathVista
Viewer • Updated • 6.14k • 17.7k • 224 -
DocVQA: A Dataset for VQA on Document Images
Paper • 2007.00398 • Published • 2 -
miracl/miracl-corpus
Viewer • Updated • 77.2M • 5.02k • 54 -
nvidia/miracl-vision
Viewer • Updated • 695k • 952 • 13 -
MTEB Leaderboard
📊7.66kEmbedding Leaderboard
-
MTEB: Massive Text Embedding Benchmark
Paper • 2210.07316 • Published • 8 -
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
Paper • 2506.21182 • Published • 3 -
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
Paper • 2509.12340 • Published • 6 -
Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
Paper • 2607.04071 • Published • 1 -
MTEB-PT: A Text Embedding Benchmark for Brazilian Portuguese
Paper • 2607.04581 • Published -
MTEB-BR Leaderboard
🏆121Massive Text Embedding Benchmark for Brazilian Portuguese
-
MTEB Leaderboard
🥇1 -
Open ASR Leaderboard configuration for Lite-Whisper models
🎙 -
Leaderboard
📊1Display speech recognition leaderboard
-
mozilla-foundation/common_voice_17_0
Updated • 3.52k • 43 -
stanford-crfm/helm-scenarios
Viewer • Updated • 498 • 1.53k • 2 -
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Paper • 2404.04475 • Published -
lmsys/mt_bench_human_judgments
Viewer • Updated • 5.76k • 2.39k • 147 -
HuggingFaceH4/mt_bench_prompts
Viewer • Updated • 80 • 18.2k • 26 -
MT Bench
📊203Explore and compare AI model answers on benchmark questions
-
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Paper • 2406.04770 • Published • 28 -
allenai/WildBench
Viewer • Updated • 2.3k • 1.75k • 40 -
AI2 WildBench Leaderboard (V2)
🦁232Display LLM performance leaderboards with customizable views
-
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
Paper • 2506.11928 • Published • 25 -
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
Paper • 2606.20517 • Published • 61 -
livecodebench/code_generation
Viewer • Updated • 121 • 6.59k • 31 -
livecodebench/test_generation
Viewer • Updated • 442 • 734 • 8 -
Leaderboard
🐠42View the LiveCodeBench leaderboard rankings
-
livecodebench/code_generation_lite
Updated • 115k • 101 -
Code Generation Samples
🏢13Compare code generation models on coding problems
-
zai-org/LongBench
Updated • 60.4k • 190 -
zai-org/LongBench-v2
Viewer • Updated • 503 • 72.5k • 56 -
LongBench Pro Leaderboard
📊1Realistic and Comprehensive Bilingual Long-Context Benchmark
-
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Paper • 2308.14508 • Published • 2 -
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
Paper • 2412.15204 • Published • 39 -
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
Paper • 2601.02872 • Published -
ameyhengle/Multilingual-Needle-in-a-Haystack
Viewer • Updated • 67.9k • 136 • 3 -
dwzhu/needle_in_a_haystack_retrieval
Updated • 34 • 1 -
VenusChenyy/RULER_50
Preview • Updated • 682 • 1 -
lmarena-ai/VisionArena-Chat
Viewer • Updated • 199k • 4.49k • 15 -
lmarena-ai/vision-arena-bench-v0.1
Viewer • Updated • 500 • 1.63k • 3 -
lmarena-ai/VisionArena-Battle
Viewer • Updated • 29.8k • 198 • 10 -
google/simpleqa-verified
Viewer • Updated • 1k • 4.16k • 52 -
codelion/SimpleQA-Verified
Viewer • Updated • 1k • 1.05k • 2 -
Humanity's Last Exam
Paper • 2501.14249 • Published • 78 -
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
Paper • 2602.13964 • Published • 12 -
SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
Paper • 2507.05241 • Published • 5 -
BrowseComp-Plus
🔍46Fair and Disentangled Evaluation of Deep-Research Agents
-
Tevatron/browsecomp-plus
Viewer • Updated • 830 • 57.4k • 37 -
openai/BrowseCompLongContext
Viewer • Updated • 295 • 3.57k • 54 -
Halcyon-Zhang/BrowseComp-V3
Viewer • Updated • 300 • 5.75k • 6 -
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
Paper • 2508.06600 • Published • 44 -
BrowseComp-V^3: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents
Paper • 2602.12876 • Published • 14 -
BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
Paper • 2504.19314 • Published • 8 -
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Paper • 2411.04872 • Published • 6 -
LiveBench: A Challenging, Contamination-Free LLM Benchmark
Paper • 2406.19314 • Published • 23 -
livebench/reasoning
Viewer • Updated • 200 • 7.29k • 19 -
livebench/math
Viewer • Updated • 368 • 7.58k • 2 -
livebench/model_judgment
Viewer • Updated • 60.4k • 2.02k • 1 -
livebench/coding
Viewer • Updated • 128 • 6.45k • 9 -
LiveBench
🥇21