AI Model Benchmark

vendor reported
Models
87
Rows
311
Last updated
2026-09-03

AI model benchmark union table

Column filters

87 / 87 model columns

Row filters

311 / 311 benchmark rows

Benchmark
Canonical namevariant · original source
OpenAI
Anthropic
Google
6 Astra
2026-09-03
5.6 Sol
2026-07-09
5.6 Terra
2026-07-09
5.6 Luna
2026-07-09
5.5 Pro
2026-04-23
5.5
2026-04-23
5.4 nano
2026-03-17
5.4 mini
2026-03-17
5.4 Pro
2026-03-05
5.4
2026-03-05
5.3 Spark
2026-02-12
5.3 Codex
2026-02-05
5.2 Codex
2025-12-18
5.2 Pro
2025-12-11
5.2 Think
2025-12-11
5.1 C-Max
2025-11-19
5.1 Codex
2025-11-13
5.1 Think
2025-11-12
5 Codex
2025-09-15
5 Pro
2025-08-07
5 nano
2025-08-07
5 mini
2025-08-07
5
2025-08-07
codex-1
2025-05-16
o4-mini
2025-04-16
o3
2025-04-16
4.1 nano
2025-04-14
4.1 mini
2025-04-14
4.1
2025-04-14
4.5
2025-02-27
o3-mini
2025-01-31
o1
2024-12-05
o1-mini
2024-09-12
o1-preview
2024-09-12
4o mini
2024-07-18
4o
2024-05-13
4 Turbo
2023-11-06
4
2023-03-14
3.5
2022-11-30
Instruct 175B
2022-01-27
Codex 12B
2021-07-07
GPT-3
2020-05-28
GPT-2
2019-11-05
GPT
2018-06-11
Fable 5.1
2026-09-01
Mythos 5.1
2026-09-01
Opus 5
2026-07-24
Sonnet 5
2026-06-30
Mythos 5
2026-06-09
Fable 5
2026-06-09
Opus 4.8
2026-05-28
Opus 4.7
2026-04-16
Mythos Prev.
2026-04-07
Sonnet 4.6
2026-02-17
Opus 4.6
2026-02-05
Opus 4.5
2025-11-24
Haiku 4.5
2025-10-15
Sonnet 4.5
2025-09-29
Opus 4.1
2025-08-05
Opus 4
2025-05-22
Sonnet 4
2025-05-22
Sonnet 3.7
2025-02-24
Haiku 3.5
2024-10-22
Sonnet 3.5 v2
2024-10-22
Sonnet 3.5
2024-06-20
Haiku 3
2024-03-13
Opus 3
2024-03-04
Sonnet 3
2024-03-04
Claude 2.1
2023-11-21
Instant 1.2
2023-08-09
Claude 2
2023-07-11
Claude 1.3
2023-05-11
Instant 1.1
2023-05-11
3.8 Flash
2026-09-02
3.7 Flash
2026-08-13
3.6 Flash
2026-07-21
3.5 Flash
2026-05-19
3.1 Pro
2026-02-19
3.1 Deep Think
2026-02-12
3 Flash
2025-12-17
3 Deep Think
2025-12-04
3 Pro
2025-11-18
2.5 Deep Think
2025-08-01
2.5 Pro
2025-06-17
2.0 Flash
2025-02-05
1.5 Pro 002
2024-09-24
1.0 Ultra
2023-12-06
Official AI model benchmark results by model and benchmark
Professional work 22

(release-page harness · version not disclosed)

64.4%adaptive · max55.3%effort n/d
FinanceAgentpartial3

(v1.1)

60%xhigh61.5%xhigh56%xhigh

(v2)

56.31%adaptive · max53.9%adaptive · max51.5%adaptive · max61.4%effort n/d59%effort n/d57.9%effort n/d43%effort n/d42.6%effort n/d
GDPvalpublic7

(ties allowed · wins or ties)

82.3%xhigh84.9%xhigh82%xhigh83%xhigh70.9%xhigh74.1%xhigh70.9%xhigh
GDPvalpublic2

(ties allowed · clear wins)

60%xhigh49.8%xhigh
GDPvalpublic2

(no ties)

67.6%xhigh61%xhigh
GDPval-AApublic8

(Artificial Analysis Elo)

1,932effort n/d1,890adaptive · max1,753adaptive · max1,633adaptive · max1,656effort n/d1,314adaptive · high1,204effort n/d1,195adaptive · high
Investment Banking Modeling Tasksinternal7

(Internal)

88.6%xhigh88.5%xhigh83.6%xhigh87.3%xhigh71.7%xhigh68.4%xhigh59.1%xhigh
OfficeQApartial2

(Full)

68.1%xhigh86.3%adaptive · max
OfficeQApartial3

(Pro)

54.1%xhigh57.9%effort n/d80.6%adaptive · max
59.3%best effort*52.7%effort n/d50.4%effort n/d50.3%effort n/d26.3%effort n/d24.2%effort n/d
GDPval-AA v2public12
1,747.8effort n/d1,593effort n/d1,591.8effort n/d1,853adaptive · max1,853shared · adaptive · max1,861adaptive · max1,618adaptive · max1,545effort n/d1,482effort n/d1,422effort n/d1,349effort n/d965effort n/d
Management Consulting Tasksinternal3

(Internal)

43.2%effort n/d37.2%effort n/d35.4%effort n/d
Big Finance Benchinternal3
53%effort n/d51%effort n/d36%effort n/d
58.9max55effort n/d51.2effort n/d
61.2best effort*60.9best effort*
AA-Briefcasepartial3
1,694adaptive · max1,694shared · adaptive · max1,720adaptive · max

(Full Public Set)

8.9%adaptive · max16.9%adaptive · max10%effort n/d8.8%effort n/d

(Harvey's Held-Out Set)

5.8%adaptive · max13.3%effort n/d
33.5%adaptive · high18.4%adaptive · high

(August 2026 snapshot)

56effort n/d52effort n/d

(Artificial Analysis leaderboard)

90.7%effort n/d85.1%effort n/d
Coding 36
Codeforcespartial5

(simulated contestant Elo)

2,036high1,673effort n/d1,650effort n/d1,258effort n/d808effort n/d
HumanEvalpublic21

(0-shot · pass@1)

99.3%effort n/d88.4%effort n/d97.6%effort n/d92.4%effort n/d92.4%effort n/d87.2%effort n/d91%effort n/d88.2%effort n/d67%effort n/d48.1%effort n/d28.8%effort n/d88.1%effort n/d93.7%effort n/d75.9%effort n/d84.9%effort n/d73%effort n/d58.7%effort n/d71.2%effort n/d56%effort n/d52.8%effort n/d74.4%no thinking · effort n/d

(full 500 · pass@1)

77.9%effort n/d73.7%effort n/d76.3%high74.5%high38%best effort*33.2%effort n/d96%adaptive · max85.2%adaptive · max95.5%adaptive · max95%adaptive · max88.6%adaptive · max87.6%adaptive · max93.9%adaptive · max79.6%adaptive · max80.8%adaptive · max80.9%no thinking · high73.3%adaptive · max77.2%high74.5%high72.5%effort n/d72.7%effort n/d40.6%effort n/d49%effort n/d33.4%effort n/d7.2%effort n/d22.2%effort n/d80.6%adaptive · high78%effort n/d76.2%effort n/d59.6%effort n/d21.4%no thinking · effort n/d22.3%no thinking · effort n/d

(full 500 · parallel test-time compute)

82%adaptive · max79.4%max80.2%max67.2%effort n/d34.2%no thinking · effort n/d34.2%no thinking · effort n/d

(489-task subset · single attempt)

63.7%high

(489-task subset · parallel test-time compute)

70.3%max

(OpenAI 477-task subset)

54.7%high71%high74.9%high72.1%medium68.1%high69.1%high23.6%high54.6%high

(diff editing format)

48.4%high71.6%high88%high58.2%high79.6%high6.2%high31.6%high52.9%high82.2%effort n/d21.3%no thinking · effort n/d16.9%no thinking · effort n/d

(v1 · Claude Code harness)

43.3%effort n/d43.2%effort n/d35.5%effort n/d
82.7%xhigh46.3%xhigh60%xhigh75.1%xhigh58.4%effort n/d77.3%xhigh64%xhigh58.1%xhigh52.8%high69.4%no thinking · effort n/d82%adaptive · max59.1%no thinking · max65.4%adaptive · max59.3%high41%effort n/d50%effort n/d68.5%adaptive · high47.6%effort n/d56.9%adaptive · high32.6%effort n/d
SWE-Lancerpublic5

(IC Diamond · percent earned)

81.4%xhigh76%xhigh74.6%xhigh69.7%xhigh32.6%best effort*
SWE-Lancerpublic2

(IC SWE)

79.9%effort n/d66.3%effort n/d
Expert-SWEinternal1

(Internal)

73.1%xhigh
80max77.4effort n/d74.6effort n/d
64.6%effort n/d63.4%effort n/d62.7%effort n/d58.6%xhigh52.4%xhigh54.4%xhigh57.7%xhigh56.8%xhigh56.4%xhigh55.6%xhigh50.8%xhigh81.2%adaptive · max81.2%shared · adaptive · max79.2%adaptive · max63.2%adaptive · max80.3%adaptive · max80%adaptive · max69.2%adaptive · xhigh64.3%adaptive · max77.8%adaptive · max53.4%adaptive · max52%no thinking · high58.7%effort n/d55.1%effort n/d54.2%adaptive · high49.6%effort n/d43.3%adaptive · high
89.1%adaptive · max89.1%shared · adaptive · max89.5%adaptive · max78.3%adaptive · max92.2%adaptive · max84.4%adaptive · max80.5%adaptive · max87.3%adaptive · max75.9%adaptive · max77.83%adaptive · max76.2%no thinking · high
54.7%adaptive · max54.7%shared · adaptive · max59.4%adaptive · max28.1%adaptive · max54.9%adaptive · max38.4%adaptive · max34.5%adaptive · max59%adaptive · max
DeepSWE v1.1public10
74.1%best effort*72.7%effort n/d69.6%effort n/d67.2%effort n/d68.8%adaptive · max73.7%adaptive · high65.3%adaptive · high48.6%adaptive · high37%adaptive · medium12%best effort*
88.8%effort n/d87.4%effort n/d84.7%effort n/d80.4%adaptive · xhigh88%adaptive · high84.3%adaptive · high74.6%adaptive · high66.1%effort n/d89.4%effort n/d85.8%effort n/d78%effort n/d76.2%effort n/d73.8%effort n/d58%effort n/d

(multi-agent · 4 agents)

91.9%ultra
57.7%best effort*37.3%best effort*55.8%adaptive · max60.9%adaptive · max19.1%best effort*11.2%best effort*

(Main)

53.3%best effort*47.5%best effort*53.4%adaptive · medium43.6%effort n/d34.4%effort n/d

(Extended · score)

64.5%best effort*60.6%best effort*
38.8%adaptive · max
FrontierCodepartial1

(Diamond subset · score · mean@5)

29.3%adaptive · xhigh
43.3%adaptive · max
73.4%adaptive · max
Database Migration Tasksinternal2

(Internal)

63.9%best effort*42.7%best effort*
67best effort*65.1best effort*

(2025-01-01–2025-05-01 · pass@1)

87.6%adaptive · ultra74.2%effort n/d29.1%no thinking · effort n/d29.7%no thinking · effort n/d

(Elo)

2,887adaptive · high2,316effort n/d2,439effort n/d1,775effort n/d
SciCodepublic2
59%adaptive · high56%adaptive · high
Code Arenapublic2

(Web development · Elo)

1,588effort n/d1,538effort n/d
14.9%best effort*5.4%best effort*
Natural2Codepartial1

(0-shot · chat preamble)

74.9%no thinking · effort n/d
Codeforcespublic1

(no tools · Elo)

3,455adaptive · ultra
Science & health 25
Frontier Science Researchpartial2
36.7%xhigh33%xhigh
GeneBenchpartial2
33.2%xhigh25%xhigh
BixBenchpublic1
80.5%xhigh
64.6%best effort*22.4%best effort*52.6%adaptive · max52.6%shared · adaptive · max

(lower-cost setting · effort not named)

61.1%effort n/d
GeneBench Propartial4
37.8%best effort*28.7%effort n/d23.3%effort n/d10.8%effort n/d
LifeSciBenchpartial4
60.3%best effort*59.9%effort n/d56%effort n/d51.2%effort n/d
MedChemBenchinternal4

(Internal)

49.3%best effort*48.3%effort n/d35%effort n/d30.4%effort n/d
60.5%effort n/d57.7%effort n/d55.7%effort n/d62.1%adaptive · max62.1%shared · adaptive · max59.8%adaptive · max57.8%adaptive · max66%adaptive · max

(length-adjusted · unclipped · GPT-5.4 grader)

63.4%best effort*60.5%best effort*

(Human Solvable)

90.3%effort n/d88.8%effort n/d87.1%effort n/d80.6%effort n/d

(Human Difficult)

44.1%effort n/d56.5%effort n/d43.5%effort n/d41.2%effort n/d

(Verified)

77.6%effort n/d
61.9%effort n/d
ProteinGympublic1

(Hard · rank correlation)

49.3%effort n/d
Protein Designinternal1

(Sequence Generation)

46%effort n/d
Protein Designinternal1

(Library Ranking)

49.3%effort n/d
Organic Chemistryinternal1

(v2)

69.2%effort n/d
Protocolsinternal1

(Troubleshooting)

70.2%effort n/d
Protocolspartial1

(Understanding · no external network)

77.2%effort n/d
LABBench2public3

(macro-average across 11 subtasks)

86.2%effort n/d82.1%effort n/d76.1%effort n/d
PhysicsFinalsinternal1

(0-shot)

41%no thinking · effort n/d

(theory)

87.7%adaptive · ultra

(condensed matter theory)

50.5%adaptive · ultra

(theory)

82.8%adaptive · ultra
Computer use 19

(Python tool)

86.3%xhigh64.2%xhigh87.9%adaptive · max87.6%adaptive · max
BrowseComppublic2

(single-agent · Anthropic 2026 compaction harness)

84.3%adaptive · max79.8%adaptive · max

(Anthropic May 2026 zoom fix · 128K/turn)

85%adaptive · max85%effort n/d83.4%adaptive · max82.8%adaptive · max
62.6%effort n/d50.2%effort n/d45.6%effort n/d70.6%adaptive · max47.9%effort n/d33.8%effort n/d

(v2026.08.08 · offline set · partial score)

72.6%best effort*65.7%best effort*

(Aug 2026 tasks · partial)

77.9%adaptive · max77.9%shared · adaptive · max

(Aug 2026 tasks · strict)

41.7%adaptive · max41.7%shared · adaptive · max
78.7%xhigh39%xhigh72.1%xhigh75%xhigh64.7%xhigh38.2%xhigh81.2%adaptive · max78%adaptive · max79.6%adaptive · max72.5%adaptive · max72.7%adaptive · max66.3%high50.7%effort n/d61.4%effort n/d83%effort n/d78.4%effort n/d76.2%effort n/d65.1%effort n/d
BrowseComppublic15

(single-agent)

91.5%best effort*90.4%effort n/d87.5%effort n/d83.3%effort n/d90.1%xhigh84.4%xhigh89.3%xhigh82.7%xhigh77.9%xhigh65.8%xhigh50.8%xhigh90.8%adaptive · max84.7%adaptive · max88%adaptive · max79.3%adaptive · max
BrowseComppublic4

(multi-agent)

92.2%ultra86.6%adaptive · max93.3%adaptive · max88.5%adaptive · ultra

(no tools)

92.7%best effort*76.9%best effort*82.3%adaptive · max79.5%adaptive · max69.1%effort n/d72.7%effort n/d11.4%effort n/d
BenchCADpublic3
70.6%effort n/d62.3%effort n/d63.1%effort n/d
BenchCADpublic3

(python tool)

83.4%effort n/d78.2%effort n/d73.9%effort n/d
BenchCADpublic2

(with tools · geometric overlap)

95.9%best effort*83.3%best effort*

(1 − OMR-NED)

0.84best effort*0.19best effort*
Design Tasksinternal2

(Internal)

50%best effort*47.4%best effort*
Data Science Tasksinternal2

(Internal)

40.9%best effort*30.5%best effort*
BrowseComppublic2

(Deep Research · Search + Python + Browse)

85.9%adaptive · high59.2%adaptive · high

(pre-2026-08-08 patch · partial score · batched tool calls)

59%effort n/d50.6%effort n/d
Cybersecurity 17
Capture-the-Flag Challengespartial6
96.7%effort n/d91.8%effort n/d85.2%effort n/d88.1%xhigh77.6%xhigh67.4%xhigh
SEC-Bench Propartial3
71.2%effort n/d57.7%effort n/d48.9%effort n/d
SEC-Bench Propartial1

(multi-agent · 16 agents)

74.3%ultra
ExploitBenchpartial3
73.5%effort n/d52.9%effort n/d33.2%effort n/d
ExploitGympartial3

(6-hour cap)

33.7%effort n/d23.2%effort n/d12.4%effort n/d

(mean capability flags · plain arm)

11.8effort n/d

(mean capability flags · AutoNudge arm)

12.61effort n/d

(Full ACEs · count out of 410 runs)

222effort n/d
OSS-Fuzzinternal1

(top-score targets · count)

17effort n/d
OSS-Fuzzinternal1

(vulnerability identification · score > 0)

78.7%effort n/d
Firefox 147 Exploitationinternal1

(full working exploits · 250 trials)

98%effort n/d
ExploitBenchpartial2

(without production safeguards · OpenAI Sep 2026)

100%best effort*78.5%best effort*
ExploitGympartial2

(without production safeguards · no 6-hour limit)

42.4%best effort*30.3%best effort*
ExploitBenchinternal2

(Internal Port · June–August 2026)

39%best effort*5.5%best effort*
SRE-Benchpublic2

(single attempt)

88%best effort*55.9%best effort*
SRE-Benchpublic2

(within 4 attempts)

99.2%best effort*68.7%best effort*
SEC-Bench Propartial2

(OpenAI Sep 2026 settings)

85.4%best effort*79.1%best effort*
AI R&D 6
Internal Research Debugging Evaluationinternal3
68.3%effort n/d67.8%effort n/d50.8%effort n/d
KernelGen 1Ppartial3
61.1%effort n/d49.2%effort n/d22.4%effort n/d
NanoGPTpartial3
9.69%effort n/d14.5%effort n/d1.66%effort n/d
PostTrainBench Litepartial3
50.3%effort n/d51.5%effort n/d29.6%effort n/d
RSI Indexpartial3
57.9%effort n/d56.3%effort n/d41.9%effort n/d
MLE-benchpublic3

(canonical Partial 30 · average position score)

63.9%effort n/d49.7%effort n/d42.6%effort n/d
Multimodal 66
MMMUpublic10

(validation set)

80.7%high73.2%effort n/d77.8%effort n/d77.1%effort n/d75%high70.4%effort n/d68.3%effort n/d59.4%effort n/d53.1%effort n/d59.4%no thinking · effort n/d
MMMUpublic13

(standard)

75.6%high81.6%high84.2%high81.6%high82.9%high55.4%high72.7%high74.8%high74.4%effort n/d59.4%effort n/d82%effort n/d69.3%no thinking · effort n/d67.7%no thinking · effort n/d
MathVistapublic9

(testmini)

56.2%effort n/d73.1%effort n/d72.2%effort n/d70.7%effort n/d67.7%effort n/d46.4%effort n/d50.5%effort n/d47.9%effort n/d53%no thinking · effort n/d
AI2Dpublic6

(test)

95.3%effort n/d94.7%effort n/d86.7%effort n/d88.1%effort n/d88.7%effort n/d79.5%no thinking · effort n/d
ChartQApublic6

(test · relaxed accuracy)

90.8%effort n/d90.8%effort n/d81.7%effort n/d80.8%effort n/d81.1%effort n/d80.8%no thinking · effort n/d
ChartQAPropublic2

(full test set · no tools · Chain-of-Thought prompting)

69.4%adaptive · max67.6%adaptive · max
ChartQAPropublic2

(full test set · Python tool · Chain-of-Thought prompting)

72.3%adaptive · max69.8%adaptive · max
DocVQApublic6

(test · ANLS)

94.2%effort n/d95.2%effort n/d88.8%effort n/d89.3%effort n/d89.5%effort n/d90.9%no thinking · effort n/d

(no tools)

82.1%xhigh67%xhigh40.5%effort n/d56.8%effort n/d56.7%effort n/d82.1%adaptive · max86.1%adaptive · max86.2%effort n/d84.5%effort n/d85.2%effort n/d84.2%effort n/d83.3%effort n/d80.3%effort n/d81.4%effort n/d69.6%effort n/d

(Python tool)

88.7%xhigh80.3%xhigh62.7%high75.5%high81.1%high72%high78.6%high91%adaptive · max93.2%adaptive · max
MMMU-Propublic5

(average across standard and vision sets)

62.6%high74.1%high78.4%high73.4%high76.4%high
Video-MMMUpublic5

(max 256 frames)

66.8%high82.5%high84.6%high79.4%high83.3%high
Video-MMMUpublic4

(no tools)

85.9%xhigh82.9%xhigh86.9%effort n/d87.6%effort n/d
ERQApublic5
50.1%high62.9%high65.7%high56.5%high64%high

(overall edit distance · no tools · lower is better)

0.24effort n/d0.13effort n/d0.11effort n/d0.12effort n/d0.12effort n/d0.15effort n/d
MMMU-Propublic16

(no tools)

83%effort n/d80.7%effort n/d78.4%effort n/d81.2%xhigh66.1%xhigh76.6%xhigh81.2%xhigh79.5%xhigh74.5%adaptive · max73.9%adaptive · max83.6%effort n/d80.5%adaptive · high81.5%adaptive · ultra81.2%effort n/d81%effort n/d68%effort n/d
MMMU-Propublic11

(with tools)

84.6%effort n/d82%effort n/d79.5%effort n/d83.2%xhigh69.5%xhigh78%xhigh82.1%xhigh80.4%xhigh79%xhigh75.6%adaptive · max77.3%adaptive · max
gdp.pdfpublic7
30.7%effort n/d24.7%effort n/d22.7%effort n/d29.8%adaptive · max35%effort n/d34%effort n/d22%effort n/d

(Andon Labs standard harness · weighted composite)

38.6%adaptive · max33.6%effort n/d26.5%effort n/d0%effort n/d
Vibe-Evalpublic3

(Reka subset · Gemini judge)

67.2%effort n/d55.4%no thinking · effort n/d55.9%no thinking · effort n/d
ZeroBenchpublic3
4.5%effort n/d1.25%no thinking · effort n/d1%no thinking · effort n/d
BetterChartQApartial3
72.4%effort n/d57.8%no thinking · effort n/d65.8%no thinking · effort n/d

(Google Search + code execution)

88.7%effort n/d89.4%effort n/d84.9%effort n/d83.2%effort n/d
LVBenchpublic3

(static · 1024 frames · no tools)

87.1%effort n/d85.4%effort n/d84.2%effort n/d
LVBenchpublic1

(agentic · 1024 frames)

87.8%effort n/d
MMMUpublic1

(validation · Maj1@32)

62.4%no thinking · effort n/d
TextVQApublic1

(validation · pixel only · no external OCR)

82.3%no thinking · effort n/d

(test · pixel only · no external OCR)

80.3%no thinking · effort n/d
VQAv2public1

(test-dev · pixel only · no external OCR)

77.8%no thinking · effort n/d
FLEURSpublic3

(53 languages · WER · lower is better)

6.66%effort n/d9.04%no thinking · effort n/d7.14%no thinking · effort n/d
CoVoST 2public3

(21 languages to English · BLEU)

38.48effort n/d36.35no thinking · effort n/d37.53no thinking · effort n/d

(test · 0-shot · 1 FPS · ≤1024 frames)

66.7%effort n/d56.4%no thinking · effort n/d57.3%no thinking · effort n/d
EgoTempopublic3

(test · 0-shot · 1 FPS · ≤256 frames)

44.3%effort n/d39.3%no thinking · effort n/d36.3%no thinking · effort n/d

(test · 0-shot · 1 FPS · ≤256 frames)

78.4%effort n/d68.8%no thinking · effort n/d69.4%no thinking · effort n/d

(validation · 4-shot · R1@0.5 · ≤256 frames)

75%effort n/d63.9%no thinking · effort n/d68.7%no thinking · effort n/d
Video-MMMUpublic3

(test · 0-shot · ≤256 frames)

83.6%effort n/d68.5%no thinking · effort n/d70.4%no thinking · effort n/d
Video-MMMUpublic1

(test · no tools · high media resolution)

83.6%effort n/d
1H-VideoQApartial3

(test · 0-shot · ≤7200 frames)

81%effort n/d67.5%no thinking · effort n/d72.2%no thinking · effort n/d
LVBenchpublic3

(test · 0-shot · audio + visual · ≤1024 frames)

78.7%effort n/d61.8%no thinking · effort n/d65.7%no thinking · effort n/d
Video-MMEpublic3

(long test subset · audio + visual · 0-shot · ≤1024 frames)

84.3%effort n/d72.8%no thinking · effort n/d73.2%no thinking · effort n/d
VATEXpublic3

(test · 4-shot · CIDEr · ≤64 frames)

71.3effort n/d56.9no thinking · effort n/d55.5no thinking · effort n/d
VATEX-ZHpublic3

(validation · 4-shot · CIDEr · ≤64 frames)

59.7effort n/d48.5no thinking · effort n/d52.2no thinking · effort n/d

(validation · 4-shot · CIDEr · ≤256 frames)

188.3effort n/d129no thinking · effort n/d170no thinking · effort n/d
Minervapublic3

(test · 0-shot · ≤1024 frames)

67.6%effort n/d52.4%no thinking · effort n/d52.8%no thinking · effort n/d
Neptunepublic3

(test · 0-shot · ≤1024 frames)

87.3%effort n/d83.1%no thinking · effort n/d82.7%no thinking · effort n/d
Video-MMEpublic3

(full test · audio + visual + subtitles · 0-shot · ≤1024 frames)

86.9%effort n/d78.8%no thinking · effort n/d79.8%no thinking · effort n/d
AI2Dpublic1

(test · transparent bounding box)

87.7%no thinking · effort n/d
ChemicalDiagramQAinternal1

(0-shot · internal)

55.2%no thinking · effort n/d
BetterChartQApartial1

(0-shot · Gemini 1.5 report setup)

47.9%no thinking · effort n/d
DocVQApublic1

(test · ANLS · Google Cloud OCR)

92.4%no thinking · effort n/d
DUDEpublic1

(test · 0-shot · ANLS)

44%no thinking · effort n/d
TAT-DQApublic1

(test · 0-shot)

13.2%no thinking · effort n/d
V* Benchpublic1

(test · 0-shot)

62.8%no thinking · effort n/d
BLINKpublic1

(validation · 0-shot)

51.7%no thinking · effort n/d

(test · 0-shot)

64.7%no thinking · effort n/d
VATEXpublic1

(test · 4-shot · CIDEr · Gemini 1.5 report setup)

62.7no thinking · effort n/d
VATEX-ZHpublic1

(validation · 4-shot · CIDEr · Gemini 1.5 report setup)

50.8no thinking · effort n/d

(validation · 4-shot · CIDEr · Gemini 1.5 report setup)

135.4no thinking · effort n/d

(test · 0-shot · Gemini 1.5 report setup)

52.2%no thinking · effort n/d
EgoSchemapublic1

(test · 0-shot · Gemini 1.5 report setup)

61.5%no thinking · effort n/d
OpenEQApublic1

(validation · 0-shot · LLM evaluation score)

61.3no thinking · effort n/d
YouTube ASRinternal1

(English · WER · lower is better)

4.7%no thinking · effort n/d
YouTube ASRinternal1

(52 languages · WER · lower is better)

21%no thinking · effort n/d

(English · WER · lower is better)

4.4%no thinking · effort n/d
FLEURSpublic1

(55 languages · WER · lower is better)

6%no thinking · effort n/d
CoVoST 2public1

(20 languages to English · BLEU)

41no thinking · effort n/d
Academic reasoning 76
GLUEpublic1

(test-set average)

72.8%effort n/d
LAMBADApublic1

(zero-shot · accuracy)

63.24%effort n/d
LAMBADApublic1

(few-shot · accuracy)

86.4%effort n/d
Human Preference vs GPT-3 175Bpartial1

(API prompt distribution · 175B PPO-ptx)

85%effort n/d
MMLUpublic9

(5-shot)

86.4%effort n/d70%effort n/d77.6%effort n/d88.7%effort n/d88.7%effort n/d75.2%effort n/d86.8%effort n/d79%effort n/d83.7%no thinking · effort n/d
MMLUpublic9

(5-shot chain-of-thought)

80.9%effort n/d90.5%effort n/d90.4%effort n/d76.7%effort n/d88.2%effort n/d81.5%effort n/d78.5%effort n/d77%effort n/d73.4%effort n/d
MMLUpublic9

(simple-evals · 0-shot chain-of-thought)

90.3%effort n/d93.3%effort n/d86.9%effort n/d91.8%effort n/d85.2%effort n/d90.8%effort n/d82%effort n/d87.2%effort n/d86.7%effort n/d
GSM8Kpublic7

(0-shot chain-of-thought)

88.9%effort n/d95%effort n/d92.3%effort n/d86.7%effort n/d88%effort n/d85.2%effort n/d80.9%effort n/d
GSM8Kpublic2

(5-shot chain-of-thought)

92%effort n/d57.1%effort n/d
MATHpublic9

(simple-evals · 0-shot chain-of-thought)

98.2%effort n/d98.1%effort n/d97.9%effort n/d96.4%effort n/d90%effort n/d85.5%effort n/d70.2%effort n/d76.6%effort n/d73.4%effort n/d

(simple-evals · 0-shot chain-of-thought)

81.3%effort n/d83.4%effort n/d77.2%effort n/d75.7%effort n/d60%effort n/d73.3%effort n/d40.2%effort n/d49.9%effort n/d49.3%effort n/d
GPQApublic1

(science · GPT-4.5 release harness)

71.4%effort n/d
AIME 2024public10

(pass@1)

29.4%effort n/d49.6%effort n/d48.1%effort n/d36.7%effort n/d74.4%effort n/d44.6%effort n/d9.3%effort n/d61.3%high5.3%effort n/d16%effort n/d
AIME 2024public6

(consensus@64)

83.3%effort n/d56.7%effort n/d13.4%effort n/d80%max10.1%effort n/d27.6%effort n/d
MMMLUpublic21

(average over 14 non-English languages)

89.6%xhigh89.5%xhigh66.9%effort n/d78.5%effort n/d87.3%effort n/d85.1%effort n/d91.5%adaptive · max92.7%adaptive · max89.3%adaptive · max91.1%adaptive · max90.8%high83%effort n/d89.1%effort n/d89.5%effort n/d88.8%high86.5%high86.1%high92.6%adaptive · high91.8%effort n/d91.8%effort n/d89.5%effort n/d
TriviaQApublic3

(5-shot)

87.5%effort n/d86.7%effort n/d78.9%effort n/d

(5-shot)

91%effort n/d90%effort n/d85.7%effort n/d
RACE-Hpublic3

(5-shot)

88.3%effort n/d88.8%effort n/d85.5%effort n/d
MATHpublic6

(0-shot chain-of-thought)

69.2%effort n/d78.3%effort n/d71.1%effort n/d38.9%effort n/d60.1%effort n/d43.1%effort n/d
MGSMpublic7

(0-shot chain-of-thought)

87%effort n/d85.6%effort n/d92.5%effort n/d91.6%effort n/d75.1%effort n/d90.7%effort n/d83.5%effort n/d
DROPpublic6

(3-shot · F1)

83.1 F1effort n/d88.3 F1effort n/d87.1 F1effort n/d78.4 F1effort n/d83.1 F1effort n/d78.9 F1effort n/d

(3-shot chain-of-thought)

86.6%effort n/d93.2%effort n/d93.1%effort n/d73.7%effort n/d86.8%effort n/d82.9%effort n/d
MMLU-Propublic2

(0-shot chain-of-thought)

65%effort n/d78%effort n/d
IFEvalpublic6
74.5%effort n/d84.1%effort n/d87.4%effort n/d93.2%high85.9%effort n/d90.2%effort n/d

(no extended thinking)

68%effort n/d

(64K extended thinking · pass@1)

78.2%high

(64K extended thinking · parallel test-time compute)

84.8%max
MATH 500public1
96.2%high
AIME 2025public17

(no tools)

100%xhigh100%xhigh94%xhigh85.2%high91.1%high94.6%high92.7%high88.9%high80.7%effort n/d87%effort n/d78%effort n/d95.2%effort n/d95%effort n/d99.2%adaptive · ultra88%effort n/d29.7%no thinking · effort n/d17.5%no thinking · effort n/d
AIME 2025public2

(Python tool)

96.3%effort n/d100%effort n/d
USAMO 2026partial1
97.6%adaptive · max
FrontierMathpartial5

(Python tool)

9.6%high22.1%high26.3%high15.4%high15.8%high
HMMT 2025public5

(no tools)

75.6%high87.8%high93.3%high85%high81.7%high
HMMTpublic3

(February 2025 · no tools)

100%xhigh99.4%xhigh96.3%xhigh

(o3-mini grader)

54.9%high62.3%high69.6%high57.5%high60.4%high31.1%effort n/d42.2%effort n/d46.2%effort n/d
OpenAI API Instruction Followinginternal8

(hard · internal)

56.1%high65.8%high64%high44.7%high47.4%high31.6%effort n/d45.1%effort n/d49.1%effort n/d
COLLIEpublic8
96.9%high98.5%high99%high96.1%high98.4%high42.5%effort n/d54.6%effort n/d65.8%effort n/d
15%effort n/d35.8%effort n/d38.3%effort n/d
Multi-IFpublic3
57.2%effort n/d67%effort n/d70.8%effort n/d

(hallucination rate · no tools · lower is better)

1%high0.7%high1%high3%high5.2%high

(hallucination rate · no tools · lower is better)

2.8%high1.3%high1.2%high8.9%high6.8%high
FActScorepublic5

(hallucination rate · no tools · lower is better)

7.3%high3.5%high2.8%high38.7%high23.5%high
GPQA Diamondpublic43
96%best effort*94.6%effort n/d92.9%effort n/d92.3%effort n/d93.6%xhigh82.8%xhigh88%xhigh94.4%xhigh92.8%xhigh93.2%xhigh92.4%xhigh88.1%high88.4%max71.2%high82.3%high85.7%high50.3%effort n/d65%effort n/d66.3%effort n/d93.6%adaptive · max94.2%adaptive · max94.5%adaptive · max89.9%adaptive · max91.3%adaptive · max87%high73%effort n/d83.4%high80.9%effort n/d79.6%high75.4%high41.6%effort n/d65%effort n/d59.4%effort n/d33.3%effort n/d50.4%effort n/d40.4%effort n/d94.3%adaptive · high90.4%effort n/d93.8%adaptive · ultra91.9%effort n/d86.4%effort n/d65.2%no thinking · effort n/d58.1%no thinking · effort n/d
ArXivMathpublic1

(June 2026 · no tools)

91.33%max
ArXivMathpublic1

(June 2026 · with tools)

93.88%max

(v2)

89%effort n/d84.9%effort n/d78.6%effort n/d52.4%xhigh51.7%xhigh50%xhigh47.6%xhigh40.3%xhigh31%xhigh

(v2)

97.6%best effort*83%effort n/d68.3%effort n/d58.5%effort n/d39.6%xhigh35.4%xhigh38%xhigh27.1%xhigh31.3%xhigh14.6%xhigh12.5%xhigh

(lower-cost setting · effort not named)

94.9%effort n/d

(no tools)

43.1%xhigh41.4%xhigh24.3%xhigh28.2%xhigh42.7%xhigh39.8%xhigh36.6%xhigh34.5%xhigh25.7%xhigh8.7%high16.7%high24.8%high14.7%high20.2%high60.9%adaptive · max60.9%shared · adaptive · max56.3%adaptive · max43.2%adaptive · max59%adaptive · max49.8%adaptive · max46.9%adaptive · max56.8%adaptive · max33.2%adaptive · max40.2%effort n/d44.4%adaptive · high48.4%adaptive · ultra33.7%effort n/d41%adaptive · ultra37.5%effort n/d34.8%adaptive · ultra21.6%effort n/d5.1%no thinking · effort n/d4.6%no thinking · effort n/d

(with tools)

57.2%best effort*57.2%xhigh52.2%xhigh37.7%xhigh41.5%xhigh58.7%xhigh52.1%xhigh50%xhigh45.5%xhigh42.7%xhigh65%adaptive · max65%shared · adaptive · max64.7%adaptive · max57.4%adaptive · max64.5%adaptive · max57.9%adaptive · max54.7%adaptive · max64.7%adaptive · max49%adaptive · max
SimpleQApublic3
54%effort n/d29.9%no thinking · effort n/d24.9%no thinking · effort n/d
87.8%effort n/d84.6%no thinking · effort n/d80%no thinking · effort n/d

(Lite)

89.2%effort n/d83.4%no thinking · effort n/d80.8%no thinking · effort n/d
ECLeKTicpublic3
46.8%effort n/d33.6%no thinking · effort n/d27%no thinking · effort n/d
HiddenMath-Hardinternal3
80.5%effort n/d53.7%no thinking · effort n/d44.3%no thinking · effort n/d
IMO 2025public1

(pass@1 · model grade)

60.7%adaptive · ultra

(Search (blocklist) + Code)

51.4%adaptive · high43.5%effort n/d45.8%effort n/d
AIME 2025public2

(code execution)

99.7%effort n/d100%effort n/d

(overall suite score)

61.9%effort n/d70.5%effort n/d63.4%effort n/d
68.7%effort n/d72.1%effort n/d54.5%effort n/d

(100+ languages and cultures)

92.8%effort n/d93.4%effort n/d91.5%effort n/d

(full 1,811-item verified set)

54.9%effort n/d53.6%effort n/d51.2%effort n/d
GPQApublic1

(4-shot)

35.7%no thinking · effort n/d
MATHpublic1

(4-shot · Minerva prompt)

53.2%no thinking · effort n/d
HiddenMathinternal1

(0-shot)

11.2%no thinking · effort n/d
Functional MATHpartial1

(December snapshot · 0-shot)

55.8%no thinking · effort n/d

(4-shot)

30%no thinking · effort n/d
GSM8Kpublic1

(11-shot)

88.9%no thinking · effort n/d

(3-shot)

83.6%no thinking · effort n/d
DROPpublic1

(variable-shot · F1)

82.4 F1no thinking · effort n/d
HellaSwagpublic1

(10-shot)

87.8%no thinking · effort n/d
WMT23public1

(1-shot · sentence-level translation · BLEURT)

74.4no thinking · effort n/d
MGSMpublic1

(8-shot)

79%no thinking · effort n/d
MMLUpublic1

(5-shot chain-of-thought · maj@32)

90%no thinking · effort n/d

(Search + code execution)

53.4%adaptive · ultra

(theory · model score)

81.5%adaptive · ultra
Tool use 13
τ-benchpublic6

(Retail)

22.6%effort n/d55.8%effort n/d68%effort n/d82.4%effort n/d81.4%high80.5%high
τ-benchpublic6

(Airline)

14%effort n/d36%effort n/d49.4%effort n/d56%effort n/d59.6%high60%high
τ²-benchpublic14

(Retail)

82%xhigh77.9%xhigh62.3%high78.3%high81.1%high70.5%high80.2%high91.7%adaptive · max91.9%adaptive · max88.9%high83.2%effort n/d86.2%effort n/d90.8%adaptive · high85.3%adaptive · high
τ²-benchpublic8

(Airline)

41%high60%high62.6%high60.2%high64.8%high67.9%high63.6%effort n/d70%effort n/d
τ²-benchpublic18

(Telecom)

98%xhigh92.5%xhigh93.4%xhigh98.9%xhigh98.7%xhigh95.6%xhigh35.5%high74.1%high96.7%high40.5%high58.2%high97.9%adaptive · max99.3%adaptive · max98.2%high83%effort n/d98%effort n/d99.3%adaptive · high98%adaptive · high
MCP-Atlaspublic15

(pass rate)

75.3%xhigh56.1%xhigh57.7%xhigh67.2%xhigh60.6%xhigh44.5%xhigh77.3%adaptive · max61.3%adaptive · max59.5%adaptive · max62.3%high83.6%effort n/d78.2%adaptive · high62%effort n/d54.1%effort n/d8.8%effort n/d
MCP-Atlaspublic3

(April 2026 Scale config · 100-tool budget · updated judge)

82.2%adaptive · max79.1%adaptive · max76.8%adaptive · max
5.7%effort n/d49.3%effort n/d65.5%effort n/d
41.4%best effort*18.1%effort n/d15.2%effort n/d14.9%effort n/d31.4%adaptive · max31.4%shared · adaptive · max26%adaptive · max13.5%adaptive · max17.4%max15.5%adaptive · max9.9%adaptive · max
Toolathlonpublic13
58%effort n/d53.1%effort n/d53.4%effort n/d55.6%xhigh35.5%xhigh42.9%xhigh54.6%xhigh46.3%xhigh36.1%xhigh56.5%effort n/d49.4%effort n/d36.4%effort n/d10.5%effort n/d
τ²-benchpublic3

(Retail + fixed Airline + Telecom average)

90.2%effort n/d90.7%effort n/d77.8%effort n/d

(mean ending net worth)

3,635effort n/d5,478effort n/d574effort n/d

(private set)

30.4%effort n/d17%effort n/d
Long context 27
Long-document factuality errorsinternal1

(reduction vs Claude 2)

30%effort n/d
QuALITYpublic3

(5-shot)

83.2%effort n/d84.1%effort n/d80.5%effort n/d
GraphWalks Parentspartial3

(256K · accuracy)

99.96%adaptive · max99.3%adaptive · max93.6%adaptive · max
GraphWalks Parentspartial2

(1M · accuracy)

83.3%adaptive · max56.6%adaptive · max
GraphWalks BFSpartial13

(<128K)

73.4%xhigh76.3%xhigh93%xhigh94%xhigh76.8%xhigh64%high73.4%high78.3%high62.3%high77.3%high25%effort n/d61.7%effort n/d61.7%effort n/d
GraphWalks BFSpartial4

(>128K)

21.4%xhigh2.9%effort n/d15%effort n/d19%effort n/d
GraphWalks Parentspartial13

(<128K)

50.8%xhigh71.5%xhigh89.8%xhigh89%xhigh71.5%xhigh43.8%high64.3%high73.3%high51.1%high72.9%high9.4%effort n/d60.5%effort n/d58%effort n/d
GraphWalks Parentspartial4

(>128K)

32.4%xhigh5.6%effort n/d11%effort n/d25%effort n/d
OpenAI MRCRpartial8

(2-needle · 128K)

43.2%high84.3%high95.2%high56.4%high55%high36.6%effort n/d47.2%effort n/d57.2%effort n/d
OpenAI MRCRpartial3

(2-needle · 256K)

34.9%high58.8%high86.8%high
OpenAI MRCRpartial3

(2-needle · 1M)

12%effort n/d33.3%effort n/d46.3%effort n/d
OpenAI MRCR v2partial3

(8-needle · 4K–8K)

97.3%xhigh98.2%xhigh65.3%xhigh
OpenAI MRCR v2partial3

(8-needle · 8K–16K)

91.4%xhigh89.3%xhigh47.8%xhigh
OpenAI MRCR v2partial3

(8-needle · 16K–32K)

97.2%xhigh95.3%xhigh44%xhigh
OpenAI MRCR v2partial3

(8-needle · 32K–64K)

90.5%xhigh92%xhigh37.8%xhigh
OpenAI MRCR v2partial5

(8-needle · 64K–128K)

44.2%xhigh47.7%xhigh86%xhigh85.6%xhigh36%xhigh
OpenAI MRCR v2partial5

(8-needle · 128K–256K)

33.1%xhigh33.6%xhigh79.3%xhigh77%xhigh29.6%xhigh

(128K)

92%xhigh90%xhigh80.4%high89.4%high90%high80%high88.3%high

(256K)

89.8%xhigh89.5%xhigh68.4%high86%high88.8%high
OpenAI MRCR v2partial5

(8-needle · 256K-512K)

100%best effort*91.5%effort n/d89.6%effort n/d41.3%effort n/d57.5%xhigh
OpenAI MRCR v2partial5

(8-needle · 512K-1M)

96.3%best effort*73.8%effort n/d72.5%effort n/d41.3%effort n/d36.6%xhigh
GraphWalks BFSpartial8

(256K · F1)

90.7 F1effort n/d76.9 F1effort n/d81.3 F1effort n/d73.7 F1xhigh91.1 F1adaptive · max85.9 F1adaptive · max76.9 F1adaptive · max80 F1adaptive · max
GraphWalks BFSpartial6

(1M · F1)

77.1 F1effort n/d71.2 F1effort n/d51.2 F1effort n/d45.4 F1xhigh68.1 F1adaptive · max40.3 F1adaptive · max
LOFTpublic3

(hard retrieval · ≤128K)

87%effort n/d58%no thinking · effort n/d75.9%no thinking · effort n/d
LOFTpublic3

(hard retrieval · 1M)

69.8%effort n/d7.6%no thinking · effort n/d47.1%no thinking · effort n/d

(8-needle · ≤128K cumulative average)

97%effort n/d91.8%effort n/d77.3%effort n/d84.9%adaptive · high67.2%effort n/d77%effort n/d58%effort n/d19%no thinking · effort n/d26.2%no thinking · effort n/d

(8-needle · 1M pointwise)

54%effort n/d26.6%effort n/d26.3%adaptive · high22.1%effort n/d26.3%effort n/d16.4%effort n/d5.3%no thinking · effort n/d12.1%no thinking · effort n/d
Abstract reasoning 4
ARC-AGI-1public10
98.5%best effort*94.5%xhigh93.7%xhigh90.5%xhigh86.2%xhigh72.8%xhigh97.5%adaptive · max97.5%shared · adaptive · max97.5%adaptive · max92%adaptive · max
ARC-AGI-2public19
95%best effort*83.3%xhigh73.3%xhigh54.2%xhigh52.9%xhigh17.6%xhigh90%adaptive · max90%shared · adaptive · max90.4%adaptive · max75.83%adaptive · max58.3%adaptive · max68.8%adaptive · max37.6%high72.1%effort n/d77.1%adaptive · high84.6%adaptive · ultra33.6%effort n/d31.1%effort n/d4.9%effort n/d
ARC-AGI-3public5
99.9%best effort*7.78%effort n/d0.8%effort n/d0.18%effort n/d30.2%adaptive · high
ARC-AGI-2public1

(code execution · ARC Prize Verified)

45.1%adaptive · ultra

Effort: max/high/xhigh are disclosed settings · ultra is multi-agent · best effort* means the source reports the maximum score across effort settings · shared marks one score reported for configurations with identical model weights · n/d means not disclosed. Scores are vendor-reported snapshots; no automatic winner highlighting.

Comparability

同一 benchmark 也可能因 task revision、harness、工具、上下文或评分方式而不可直接比较。因此不自动标冠军;hover 每个分数可查看完整条件。

Sources

OpenAI · 31 official sources
Anthropic · 21 official sources
Google · 13 official sources