Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
All benchmarks

GPQA Diamond

GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. Each question is written by a subject-matter expert and designed so that even domain specialists need careful reasoning to identify the correct answer. We run the same fixed question set across provider endpoints to compare model capability, routing, and the practical cost of solving difficult scientific problems.

Last benchmark run Oct 2, 2026, 7:40 PM UTC

PaperGitHub

Which model leads the GPQA Diamond benchmark?

Google: Gemini 3.8 Flash leads GPQA Diamond at 95.6% as of Oct 2, 2026, 7:40 PM UTC. The same fixed GPQA Diamond question set and scoring are held fixed across production provider endpoints. The headline row uses default routing where available; otherwise, it uses a representative provider result.

Model comparisonCost efficiencyLeaderboardExample problemsWhy we run itWhat scores tell youMethodologyAPI accessDownload and citeFAQ

Model comparison

Most Accurate

Favicon for google
Google: Gemini 3.8 Flash

95.6%

Best Value

Favicon for typesafe
TypeSafe: Jev Router

$0.016/question

Fastest

Favicon for typesafe
TypeSafe: Jev Router

25s

Accuracy
Representative-run accuracy, best first.
Cost per question
Average cost per question, cheapest first.
Time per question
Average wall-clock time per question, fastest first.

Cost efficiency

Accuracy vs. cost (Pareto frontier)
One point per model, using default routing (not pinned to a provider) when available. The line is the Pareto frontier: no model beats these on both accuracy and cost.

Leaderboard

Top-level rows use default routing where available; click a row to expand provider-pinned results.

GPQA Diamond leaderboard: model accuracy, cost, time, and output tokens per question by provider
#ModelStd dev
1
Favicon for google
Google: Gemini 3.8 Flash
Pareto
95.6%--$0.0682.2m18k
2
Favicon for openai
OpenAI: GPT-6 Astra Pro
95.5%--$0.3375s4.53k
3
Favicon for sakana
Fugu Ultra
94.6%--$1.155.5m33.3k
4
Favicon for google
Google: Gemini 3.1 Pro Preview
94.5%±0.1pp$0.202.2m16.7k
5
Favicon for openai
OpenAI: GPT-6 Astra
94.4%--$0.1157s2.11k
6
Favicon for openai
OpenAI: GPT-6.1 Sol
Pareto
94.4%--$0.02152s1.98k
7
Favicon for typesafe
TypeSafe: Jev Router
Pareto
93.9%--$0.01625s2.22k
8
Favicon for openai
OpenAI: GPT-5.6 Sol Pro
93.9%±0.2pp$0.301.8m11.7k
9
Favicon for google
Google: Gemini 3.7 Flash
93.9%±0.4pp$0.02959s7.68k
10
Favicon for openai
OpenAI: GPT-5.5
93.7%±0.4pp$0.312.5m10.2k
11
Favicon for sakana
Fugu Ultra V2
93.3%--$0.503.2m13.8k
12
Favicon for anthropic
Anthropic: Claude Sonnet 5.5
92.9%--$0.02521s2.27k
13
Favicon for x-ai
SpaceXAI: Grok 4.6
92.8%±0.4pp$0.126.0m20.3k
14
Favicon for google
Google: Gemini 3.6 Flash
92.6%±0.4pp$0.06059s11.7k
15
Favicon for google
Google: Gemini 3.5 Flash
92.6%±0.4pp$0.1482s15.6k
16
Favicon for unbiased
Pareto
92.4%--$0.0341.9m4.36k
17
Favicon for openai
OpenAI: GPT-5.6 Sol
92.1%±0.5pp$0.06063s3.59k
18
Favicon for openai
OpenAI: GPT-6 Sol
91.9%--$0.02532s2.37k
19
Favicon for moonshotai
MoonshotAI: Kimi K3
91.9%±1.3pp$0.135.6m10.5k
20
Favicon for nvidia
NVIDIA: Switchyard
91.4%±3.0pp$0.01858s3k
21
Favicon for z-ai
GLM 5.3 Flash
Pareto
digitalocean
90.9%--$0.0053.9m10.7k
22
Favicon for qwen
Qwen: Qwen3.7 Max
AtlasCloud
90.9%--$0.225.0m29.4k
23
Favicon for anthropic
Anthropic: Claude Opus 5.5
90.6%--$0.05431s2.51k
24
Favicon for minimax
MiniMax: MiniMax M3
90.4%±2.0pp$0.0345.3m28.7k
25
Favicon for openai
OpenAI: GPT-5.4
90.3%±0.1pp$0.152.2m9.65k
26
Favicon for openai
OpenAI: GPT-5.6 Luna Pro
90.0%±0.6pp$0.0412.4m27.2k
27
Favicon for anthropic
Anthropic: Claude Fable 5.1
89.9%±0.7pp$0.1331s2.4k
28
Favicon for anthropic
Anthropic: Claude Opus 4.8
89.7%±1.3pp$0.201.6m7.87k
29
Favicon for deepseek
DeepSeek: DeepSeek V4 Pro 0813
89.6%±1.6pp$0.1210.4m36.4k
30
Favicon for anthropic
Anthropic: Claude Opus 4.7
89.4%±0.9pp$0.231.8m8.8k
31
Favicon for deepseek
DeepSeek: DeepSeek V4.1 Flash
89.1%±1.2pp$0.0274.6m26.3k
32
Favicon for tencent
Tencent: Hy4 preview
89.1%--$0.1312.3m50.2k
33
Favicon for xiaomi
Xiaomi: MiMo-V2.6-Pro
89.1%--$0.03213.1m37.8k
34
Favicon for amazon
Amazon: Nova Micro 1.0
89.0%±0.6pp$0.221.7m8.43k
35
Favicon for openai
OpenAI: GPT-5.6 Terra
88.8%±0.6pp$0.03450s3.31k
36
Favicon for sakana
Fugu Max
88.6%--$0.0763.8m12.3k
37
Favicon for qwen
Qwen: Qwen3.8 Flash
88.6%--$0.02111.2m45.5k
38
Favicon for google
Google: Gemini 3 Flash Preview
88.4%±0.1pp$0.134.1m43.8k
39
Favicon for openai
OpenAI: GPT-6 Luna Pro
88.4%--$0.00873s13.1k
40
Favicon for openrouterFavicon for openrouter
Auto Router
88.0%±2.6pp$0.0893.3m37.4k
41
Favicon for openai
OpenAI: GPT-5.2
OpenAI
87.9%--$0.132.3m9.43k
42
Favicon for qwen
Qwen: Qwen3.7 Plus
Alibaba
87.9%--$0.0449.8m34.2k
43
Favicon for openai
OpenAI: GPT-5.6 Luna
87.9%±0.6pp$0.00880s7.83k
44
Favicon for deepseek
DeepSeek: DeepSeek V4 Flash Vision Exp
87.8%±1.2pp$0.0223.8m25k
45
Favicon for anthropic
Anthropic: Claude Opus 5
87.4%±1.9pp$0.1363s4.98k
46
Favicon for openai
OpenAI: GPT-6 Luna
Pareto
87.4%--$0.00350s5.75k
47
Favicon for moonshotai
MoonshotAI: Kimi K2 Thinking
87.3%±0.7pp$0.1117.8m42.6k
48
Favicon for qwen
Qwen: Qwen3.8 2.4T A95B
86.8%±1.4pp$0.0823.2m13.2k
49
Favicon for thinkingmachines
Thinking Machines: Inkling Small
86.6%±0.8pp$0.0253.9m20.8k
50
Favicon for openai
OpenAI: GPT-5.1
86.4%±0.5pp$0.224.3m22k
51
Favicon for anthropic
Anthropic: Claude Opus 4.5
86.3%±0.9pp$0.846.9m33.6k
52
Favicon for minimax
MiniMax: MiniMax M2.1
86.2%±0.5pp$0.0405.9m24.2k
53
Favicon for deepseek
DeepSeek: DeepSeek V4 Pro 0423
85.9%±1.4pp$0.05412.0m32.9k
54
Favicon for deepseek
DeepSeek: DeepSeek V4 Flash 0731
85.9%±2.7pp$0.0088.3m24.6k
55
Favicon for anthropic
Anthropic: Claude Opus 4.6
85.8%±1.3pp$0.758.2m29.7k
56
Favicon for deepseek
DeepSeek: DeepSeek V4 Flash 0423
85.7%±1.2pp$0.00510.8m28.5k
57
Favicon for openai
OpenAI: GPT-5
85.7%±0.3pp$0.214.4m20.9k
58
Favicon for z-ai
Z.ai: GLM 5.3
85.7%±1.5pp$0.09411.2m28.7k
59
Favicon for z-ai
Z.ai: GLM 5.3 Flash
85.6%±1.8pp$0.01110.9m33.1k
60
Favicon for z-ai
Z.ai: GLM 5.2
85.3%±2.1pp$0.09110.5m41.4k
61
Favicon for anthropic
Anthropic: Claude Fable 5
85.0%±4.2pp$0.1744s3.25k
62
Favicon for inclusionai
inclusionAI: Ling 3.0 Flash VL
Pareto
84.8%--$0.00188s17.5k
63
Favicon for google
Google: Gemini 3.5 Flash Lite
84.1%±0.9pp$0.03751s14.6k
64
Favicon for qwen
Qwen: Qwen3.5 397B A17B
83.8%±2.3pp$0.0689.6m22k
65
Favicon for minimax
MiniMax: MiniMax M2.7
83.8%±1.8pp$0.04810.9m37.6k
66
Favicon for moonshotai
MoonshotAI: Kimi K2.6
83.7%±3.1pp$0.1816.8m53.4k
67
Favicon for openai
OpenAI: GPT-5.4 Mini
83.5%±0.6pp$0.06284s13.6k
68
Favicon for moonshotai
MoonshotAI: Kimi K2.5
83.5%±3.9pp$0.1119.7m43.2k
69
Favicon for thinkingmachines
Thinking Machines: Inkling
83.4%±1.1pp$0.0954.9m23.4k
70
Favicon for qwen
Qwen: Qwen3.5-122B-A10B
83.2%±3.4pp$0.0564.7m23.9k
71
Favicon for anthropic
Anthropic: Claude Sonnet 5
83.0%±4.2pp$0.172.7m16.3k
72
Favicon for minimax
MiniMax: MiniMax M2.5
82.9%±3.0pp$0.0258.0m23.6k
73
Favicon for qwen
Qwen: Qwen3.8 27B
82.5%±1.4pp$0.0354.0m13.8k
74
Favicon for z-ai
GLM 5.3
digitalocean
82.3%--$0.0262.7m9.04k
75
Favicon for anthropic
Anthropic: Claude Sonnet 4.5
82.2%±1.4pp$0.314.8m20.8k
76
Favicon for deepseek
DeepSeek: DeepSeek V3.2 Exp
82.2%±0.5pp$0.00810.3m19.2k
77
Favicon for xiaomi
Xiaomi: MiMo-V2.5-Pro
82.0%±6.1pp$0.02810.3m28.9k
78
Favicon for google
Google: Gemma 4 31B
81.9%±2.3pp$0.0068.4m17.3k
79
Favicon for google
Google: Gemini 3.1 Flash Lite
81.6%±2.1pp$0.04183s27.5k
80
Favicon for google
Google: Gemini 2.5 Pro
81.3%±2.2pp$0.263.8m25.5k
81
Favicon for nvidia
NVIDIA: Nemotron 3 Ultra
81.2%--$0.145.1m39.8k
82
Favicon for z-ai
Z.ai: GLM 4.7
81.1%±3.3pp$0.07615.9m39.4k
83
Favicon for z-ai
Z.ai: GLM 5.1
80.9%±2.7pp$0.1914.3m51.9k
84
Favicon for openai
OpenAI: GPT-5 Mini
80.8%±1.8pp$0.0443.6m22.1k
85
Favicon for meta
Meta: Muse Glimmer 30B
80.5%±1.4pp$0.0234.2m17.5k
86
Favicon for qwen
Qwen: Qwen3.6 27B
80.5%±4.9pp$0.08413.2m31.8k
87
Favicon for openai
OpenAI: GPT-5.3 Chat
80.5%--$0.02424s1.63k
88
Favicon for deepseek
DeepSeek: DeepSeek V3.2
80.4%±2.2pp$0.0078.4m18.2k
89
Favicon for openai
OpenAI: GPT-5.2 Chat
OpenAI
79.8%--$0.03633s2.48k
90
Favicon for qwen
Qwen: Qwen3.5-35B-A3B
79.7%±4.5pp$0.0276.0m25.5k
91
Favicon for moonshotai
MoonshotAI: Kimi K2.7 Code
79.3%±18.3pp$0.0918.3m24.6k
92
Favicon for qwen
Qwen: Qwen3.6 35B A3B
79.2%±4.0pp$0.0437.7m42.4k
93
Favicon for deepseek
DeepSeek: R1 0528
78.7%±1.5pp$0.05314.6m23.8k
94
Favicon for qwen
Qwen: Qwen3 235B A22B Thinking 2507
78.3%±2.1pp$0.05611.1m24.5k
95
Favicon for openai
OpenAI: GPT-5.4 Nano
77.7%±0.3pp$0.01368s10.3k
96
Favicon for deepseek
DeepSeek: DeepSeek V3.1 Terminus
77.6%±2.0pp$0.0167.3m15.9k
97
Favicon for qwen
Qwen: Qwen3.5-9B
77.2%±2.6pp$0.00611.8m37.1k
98
Favicon for anthropic
Anthropic: Claude Sonnet 4
76.7%±1.9pp$0.273.0m17.6k
99
Favicon for stepfun
StepFun: Step 3.7 Flash
76.6%--$0.0807.1m69.8k
100
Favicon for xiaomi
Xiaomi: MiMo-V2.5
76.0%±5.2pp$0.01112.2m32.3k
101
Favicon for deepseek
DeepSeek: DeepSeek V3.1
75.8%±3.5pp$0.0167.8m16.1k
102
Favicon for mistralai
Mistral: Mistral Small 4
75.8%--$0.0152.9m24k
103
Favicon for z-ai
Z.ai: GLM 4.6
75.3%±4.6pp$0.05316.3m27.4k
104
Favicon for xiaomi
Xiaomi: MiMo-V2.6-Flash
75.0%--$0.00730.4m23.9k
105
Favicon for inclusionai
inclusionAI: Ling 3.0 Flash
74.9%±1.9pp$0.0032.8m31k
106
Favicon for qwen
Qwen: Qwen3 Coder Next
74.7%±0.9pp$0.0151.8m14.4k
107
Favicon for z-ai
Z.ai: GLM 5
74.5%±7.6pp$0.1217.2m54k
108
Favicon for moonshotai
MoonshotAI: Kimi K2 0905
74.5%±1.5pp$0.0181.7m6.28k
109
Favicon for google
Google: Gemma 4 26B A4B
73.5%±5.6pp$0.01614.1m41.2k
110
Favicon for qwen
Qwen: Qwen3 235B A22B Instruct 2507
73.3%±1.5pp$0.0064.0m7.98k
111
Favicon for anthropic
Anthropic: Claude Haiku 4.5
72.7%±1.0pp$0.194.0m37.9k
112
Favicon for google
Google: Gemini 2.5 Flash
72.6%±0.4pp$0.0602.1m23.9k
113
Favicon for openai
OpenAI: gpt-oss-120b
72.3%±2.9pp$0.0083.9m16.6k
114
Favicon for qwen
Qwen: Qwen3 Next 80B A3B Instruct
70.8%±1.1pp$0.0111.7m10.5k
115
Favicon for openai
OpenAI: GPT-5 Nano
70.5%±0.4pp$0.0133.9m33.3k
116
Favicon for qwen
Qwen: Qwen3 VL 235B A22B Instruct
69.6%±1.9pp$0.0094.7m6.39k
117
Favicon for nvidia
NVIDIA: Nemotron 3.5 Lightning
68.6%±2.3pp$0.0106.0m46.1k
118
Favicon for google
Gemma 4 26B A4B IT (free)
Darkbloom
68.5%----6.4m6.99k
119
Favicon for z-ai
Z.ai: GLM 4.6V
66.8%--$0.0145.4m15.3k
120
Favicon for meta-llama
Meta: Llama 4 Maverick
66.0%±1.1pp$0.00380s2.91k
121
Favicon for ibm-granite
IBM: Granite 4.2 8B
65.7%--$0.01528.6m64.4k
122
Favicon for qwen
Qwen: Qwen3 VL 30B A3B Instruct
64.9%±1.8pp$0.0073.9m11.2k
123
Favicon for openai
OpenAI: GPT-4.1 Mini
64.9%±1.1pp$0.00429s2.35k
124
Favicon for openai
OpenAI: GPT-4.1
64.6%±1.4pp$0.01514s1.73k
125
Favicon for qwen
Qwen: Qwen3 30B A3B Instruct 2507
64.3%±1.1pp$0.0022.1m8.81k
126
Favicon for openai
OpenAI: gpt-oss-20b
64.1%±1.9pp$0.01112.0m55.1k
127
Favicon for z-ai
Z.ai: GLM 4.5 Air
63.2%±4.0pp$0.03013.6m36.4k
128
Favicon for nvidia
NVIDIA: Nemotron 3 Nano 30B A3B
63.0%±2.2pp$0.01710.6m83.5k
129
Favicon for deepseek
DeepSeek: DeepSeek V3
62.1%±1.2pp$0.00265s2.17k
130
Favicon for qwen
Qwen: Qwen3 30B A3B
62.1%±1.0pp$0.0086.2m15.8k
131
Favicon for anthropic
Anthropic: Claude Sonnet 4.6
61.4%±12.4pp$0.6210.5m40.9k
132
Favicon for qwen
Qwen: Qwen3 32B
60.9%±1.0pp$0.0084.5m17.5k
133
Favicon for qwen
Qwen: Qwen3 Coder 480B A35B
60.3%±0.7pp$0.00356s2.12k
134
Favicon for qwen
Qwen: Qwen3 14B
59.2%±0.3pp$0.0038.7m11.7k
135
Favicon for google
Google: Gemini 2.5 Flash Lite
54.9%±0.6pp$0.0172.0m42.9k
136
Favicon for z-ai
Z.ai: GLM 4.7 Flash
54.7%±4.5pp$0.01611.8m40.9k
137
Favicon for qwen
Qwen: Qwen3 Coder 30B A3B Instruct
52.0%±0.2pp$0.0041.5m2.72k
138
Favicon for openai
OpenAI: GPT-4o (2024-08-06)
52.0%±0.3pp$0.01816s1.56k
139
Favicon for qwen
Qwen: Qwen3 VL 8B Instruct
51.3%±0.0pp$0.0072.9m13.6k
140
Favicon for openai
OpenAI: GPT-4o (2024-05-13)
50.7%--$0.02411s1.32k
141
Favicon for openai
OpenAI: GPT-4o
50.6%±0.8pp$0.01313s1.1k
142
Favicon for inclusionai
inclusionAI: Ling 3.0 Flash Fin
50.5%--$0.01011.7m64.8k
143
Favicon for openai
OpenAI: GPT-4.1 Nano
Pareto
50.0%±1.5pp$0.0008420s1.9k
144
Favicon for deepseek
DeepSeek: DeepSeek V3 0324
48.3%±12.7pp$0.0041.8m3.48k
145
Favicon for meta-llama
Meta: Llama 3.3 70B Instruct
47.3%±1.1pp$0.00165s2.28k
146
Favicon for qwen
Qwen2.5 72B Instruct
46.0%±1.3pp$0.0011.6m1.85k
147
Favicon for qwen
Qwen: Qwen2.5 VL 72B Instruct
44.4%±1.5pp$0.00259s1.59k
148
Favicon for openai
OpenAI: GPT-4o-mini
43.2%±0.8pp$0.00121s1.48k
149
Favicon for tencent
Tencent: Hy-MT2-30B-A3B
Pareto
Reka
38.9%--$0.000239s711
150
Favicon for mistralai
Mistral: Mistral Nemo
Pareto
33.0%±0.8pp$0.0001212s504
151
Favicon for qwen
Qwen: Qwen2.5 7B Instruct
32.8%--$0.0005336s2.04k
152
Favicon for meta-llama
Meta: Llama 3.1 8B Instruct
28.7%±1.8pp$0.00164s19k
153
Favicon for sao10k
Sao10K: Llama 3 8B Lunaris
Pareto
27.3%±0.6pp$0.00006310s586
154
Favicon for meta-llama
Meta: Llama 3.2 3B Instruct
11.6%--$0.0001530s1.84k
155
Favicon for google
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
10.3%±0.0pp$0.01822s11.8k

Example problems

GPQA uses four-choice questions that require more than recalling a definition. These representative examples show the format and the range of scientific domains without reproducing items from the benchmark's protected question pool.

Biology

A researcher observes that a membrane protein is synthesized on ribosomes attached to the rough endoplasmic reticulum. Which destination is most consistent with this protein entering the secretory pathway?

  1. A.The cytosol, where it remains soluble
  2. B.The nucleus, after import through a nuclear pore
  3. C.A membrane of the endomembrane system or the cell surface
  4. D.The mitochondrial matrix through a TOM/TIM complex

Answer: C

Ribosomes on the rough ER synthesize proteins destined for secretion or insertion into the endomembrane system, including the plasma membrane.

Physics

A spacecraft is far from other bodies and fires its engine in the direction opposite to its velocity. Ignoring mass loss during the brief burn, what happens immediately to its speed?

  1. A.It increases because the exhaust carries away backward momentum
  2. B.It decreases because the thrust points opposite to its velocity
  3. C.It remains unchanged because thrust only changes direction
  4. D.It becomes zero because the spacecraft is in free space

Answer: B

An impulse opposite the velocity vector reduces the spacecraft’s momentum and therefore its speed during the burn.

Chemistry

Why does adding a small amount of a common ion generally reduce the solubility of a sparingly soluble ionic solid in water?

  1. A.The common ion increases the solid’s lattice energy
  2. B.The common ion shifts the dissolution equilibrium toward the solid
  3. C.The common ion converts every dissolved ion into a neutral molecule
  4. D.The common ion removes solvent molecules from the solution

Answer: B

The added ion raises the concentration of a dissolution product, so Le Chatelier’s principle shifts the equilibrium toward the undissolved solid.

Why we run this benchmark

GPQA is a broad graduate-level reasoning test across biology, physics, and chemistry, so it gives us a cheap, high-floor signal that a deployment is healthy. A model that normally clears these questions but suddenly drops usually points to something broken in the endpoint or routing rather than the questions themselves.

Because we run the same fixed question set across provider endpoints, a large accuracy gap between providers serving the same model is a quick way to catch a misconfigured or degraded endpoint. The cost and latency columns show what that reasoning quality costs to serve.

What the scores can and can't tell you

GPQA is a narrow, high-difficulty evaluation, not a complete measure of general intelligence or usefulness. A score reflects performance on expert-written multiple-choice science questions and should be considered alongside coding, instruction-following, factuality, and other evaluations.

Scores can be sensitive to sampling settings, answer-position handling, and the number of repeated runs. Small differences may not be meaningful when models have similar sample counts, so the leaderboard includes run variability and cost context rather than presenting accuracy alone.

The benchmark is publicly described, and some questions may eventually appear in training data. We avoid reproducing the private question pool here, but no public benchmark can guarantee that every future evaluation item is uncontaminated.

Methodology

Scores aggregate successful runs from the newest 90 days, weighted by question count, with a minimum sample threshold per model-provider pair. A model's headline score uses default routing when available; otherwise it falls back to the median provider. Cost, time, and output-token figures are per-question averages from the same runs. Best value is the cheapest Pareto-optimal model within five percentage points of the top score.

GPQA Diamond is described in the original paper. See the docs for routing details, or browse all models to try one.

API access

These scores are available through OpenRouter's public benchmarks API, so you can retrieve the same model-level results programmatically.

GET https://openrouter.ai/api/v1/benchmarks?source=openrouter
Authorization: Bearer <API key>

Use task_type=intelligence to filter to gpqa_diamond. Each item represents one model and includes accuracy, accuracy_stddev, avg_cost_per_task, total_tasks, and last_run_timestamp. See the benchmarks API docs.

Download and cite

A snapshot of this leaderboard is published in the benchmark-leaderboard open dataset (filter on the benchmark column), regenerated daily with no API key required. The current snapshot was generated on Oct 2, 2026, 1:46 AM UTC. It is licensed under CC BY 4.0, so you can reuse and republish it with attribution to OpenRouter.

  • Download CSV (text/csv)
  • Download JSON (application/json)
  • Snapshot manifest (SHA-256 checksums of the files above)
  • Latest manifest (mutable alias, updated as snapshots are published)

Column definitions, methodology and redaction rules are in the dataset README.

Cite this leaderboard

OpenRouter (2026). GPQA Diamond results on OpenRouter, 2026-10-02 snapshot. https://openrouter.ai/benchmarks/gpqa-diamond. Licensed under CC BY 4.0.

Frequently asked questions

GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. Each question is written by a subject-matter expert and designed so that even domain specialists need careful reasoning to identify the correct answer.

Every model answers the same fixed question set through real provider endpoints, and results are aggregated across repeated runs. A model’s headline score is a single representative result rather than its best-performing provider, and the cost, time, and output-token figures are per-question averages from those same runs.

Every run goes to a real provider endpoint, so provider behavior is part of the measurement. A large accuracy gap between providers serving the same model usually points to a misconfigured or degraded endpoint rather than to the questions.

Yes. The public benchmarks API returns the same model-level results from GET https://openrouter.ai/api/v1/benchmarks?source=openrouter with an API key. Filter with task_type=intelligence to reach gpqa_diamond.

It is a narrow, high-difficulty evaluation rather than a measure of general usefulness. Scores are sensitive to sampling settings, answer-position handling, and the number of repeated runs, so small differences between models with similar sample counts may not be meaningful.