Google: Gemini 3.1 Pro Preview

by google

2,268 claims submitted by 172 reviewers

Monitored by HumanJudge · Endpoint registered, 86 traces logged
Maintained by HumanJudge Admin
Enrolled in: C-pop Challenge: GPT-5.2 , Spanish Culture Challenge: GPT-5.2 , Japanese Culture Challenge: GPT-5.2 , Mexican Culture Challenge: GPT-5.2 , AP Biology Challenge: GPT-5.2 , Spanish Cinema Challenge: GPT-5.2 , Korean Culture Challenge: GPT-5.2 , Taiwanese Culture Challenge: GPT-5.2 , Brazilian Culture Challenge: GPT-5.2 , AI in Healthcare | Stanford I4UI 2026 , Argentine Culture Challenge: GPT-5.2 , Japanese Language Challenge: GPT-5.2 , Arabic Music Challenge: GPT-5.2 , Gulf Culture Challenge: GPT-5.2 , Latin Music Challenge: GPT-5.2 , Arabic Language Challenge: GPT-5.2 , Chinese Cinema Challenge: GPT-5.2 , Mexican Cinema Challenge: GPT-5.2 , Korean Language Challenge: GPT-5.2 , AP Government Challenge: GPT-5.2 , Chinese Language Challenge: GPT-5.2 , Korean Cinema Challenge: GPT-5.2 , Chinese Culture Challenge: GPT-5.2 , J-drama Challenge: GPT-5.2 , AP English Language Challenge: GPT-5.2 , C-drama Challenge: GPT-5.2 , 日本文化のヒーロー | Japanese Culture Hero , Levantine Culture Challenge: GPT-5.2 , Egyptian Culture Challenge: GPT-5.2 , Arab Cinema Challenge: GPT-5.2 , AP English Literature Challenge: GPT-5.2 , AP US History Challenge: GPT-5.2 , K-pop Challenge: GPT-5.2 , AP Calculus AB Challenge: GPT-5.2 , Spanish Language Challenge: GPT-5.2 , K-drama Challenge: GPT-5.2 , Spanish Music Challenge: GPT-5.2 , Humans Evaluation Benchmark for AI Marketing and Content Generation
Performance
Humans Evaluation Benchmark for AI Marketing and Content Generation 94%
1811 votes 107 flags 148 reviewers
AP Calculus AB Challenge: GPT-5.2 98%
118 votes 2 flags 18 reviewers
日本文化のヒーロー | Japanese Culture Hero 94%
98 votes 6 flags 19 reviewers
AI in Healthcare | Stanford I4UI 2026 94%
64 votes 4 flags 13 reviewers
AP Biology Challenge: GPT-5.2 94%
53 votes 3 flags 53 reviewers
AP English Language Challenge: GPT-5.2 98%
41 votes 1 flags 15 reviewers
AP English Literature Challenge: GPT-5.2 100%
27 votes 0 flags 9 reviewers
AP US History Challenge: GPT-5.2 96%
27 votes 1 flags 9 reviewers
K-pop Challenge: GPT-5.2 100%
19 votes 0 flags 3 reviewers
AP Government Challenge: GPT-5.2 100%
10 votes 0 flags 10 reviewers
Independent Claims
flag Does AI know AP Calculus AB? 9/6/2026

Flag — the core Riemann-sum explanation is correct, but it contains a materially misleading AP exam claim: it says the A...

— Sareena Bilal

flag Does AI know AP Calculus AB? 9/6/2026

Flag — the core FTC concepts are correct, but the response contains major false claims about the AP exam, especially tha...

— Sareena Bilal

flag AP Biology Challenge: GPT-5.2 9/6/2026

the response contains several significant factual overstatements and inaccuracies. For example, biological methods are n...

— Sareena Bilal

flag AI in Healthcare | Stanford I4UI 2026 8/13/2026

Missed to emphasize how there's no substitute for himan and professional advice

— Ekaterina Yael Lechtchiner

pass AI Marketing & Content Generation 7/20/2026

"The AI response satisfies all constraints of the prompt flawlessly. It delivers an engaging, well-structured 15-second ...

— Thuy Hang Vo

This evaluation was conducted independently. Google: Gemini 3.1 Pro Preview did not participate in or pay for this evaluation. All verdicts come from double-blind evaluation — reviewers did not know which AI produced each response.

We help people define what trustworthy AI looks like — publicly, transparently, together. Support this mission