xAI: Grok 4.1 Fast

by x-ai

2,200 claims submitted by 165 reviewers

Monitored by HumanJudge · Endpoint registered, 76 traces logged
Maintained by HumanJudge Admin
Enrolled in: Chinese Cinema Challenge: GPT-5.2 , Spanish Language Challenge: GPT-5.2 , Korean Cinema Challenge: GPT-5.2 , AP English Language Challenge: GPT-5.2 , Chinese Culture Challenge: GPT-5.2 , Spanish Music Challenge: GPT-5.2 , AP English Literature Challenge: GPT-5.2 , Arabic Language Challenge: GPT-5.2 , Mexican Culture Challenge: GPT-5.2 , C-drama Challenge: GPT-5.2 , Egyptian Culture Challenge: GPT-5.2 , Mexican Cinema Challenge: GPT-5.2 , Spanish Cinema Challenge: GPT-5.2 , Arab Cinema Challenge: GPT-5.2 , Humans Evaluation Benchmark for AI Marketing and Content Generation , AP Biology Challenge: GPT-5.2 , K-drama Challenge: GPT-5.2 , AP Calculus AB Challenge: GPT-5.2 , Korean Culture Challenge: GPT-5.2 , C-pop Challenge: GPT-5.2 , Spanish Culture Challenge: GPT-5.2 , AP US History Challenge: GPT-5.2 , Levantine Culture Challenge: GPT-5.2 , 日本文化のヒーロー | Japanese Culture Hero , Taiwanese Culture Challenge: GPT-5.2 , Japanese Culture Challenge: GPT-5.2 , Gulf Culture Challenge: GPT-5.2 , Brazilian Culture Challenge: GPT-5.2 , Chinese Language Challenge: GPT-5.2 , AP Government Challenge: GPT-5.2 , K-pop Challenge: GPT-5.2 , Arabic Music Challenge: GPT-5.2 , Latin Music Challenge: GPT-5.2 , Japanese Language Challenge: GPT-5.2 , Korean Language Challenge: GPT-5.2 , J-drama Challenge: GPT-5.2 , Argentine Culture Challenge: GPT-5.2
Performance
Humans Evaluation Benchmark for AI Marketing and Content Generation 90%
1808 votes 173 flags 152 reviewers
AP Calculus AB Challenge: GPT-5.2 99%
119 votes 1 flags 19 reviewers
日本文化のヒーロー | Japanese Culture Hero 94%
100 votes 6 flags 20 reviewers
AP Biology Challenge: GPT-5.2 88%
50 votes 6 flags 50 reviewers
AP English Language Challenge: GPT-5.2 100%
41 votes 0 flags 15 reviewers
AP US History Challenge: GPT-5.2 96%
28 votes 1 flags 10 reviewers
AP English Literature Challenge: GPT-5.2 100%
27 votes 0 flags 9 reviewers
K-pop Challenge: GPT-5.2 100%
18 votes 0 flags 2 reviewers
AP Government Challenge: GPT-5.2 100%
9 votes 0 flags 9 reviewers
Independent Claims
flag Does AI know AP Calculus AB? 9/6/2026

Flag — the core FTC explanation is correct, but it contains significant AP Calculus AB inaccuracies. It incorrectly says...

— Sareena Bilal

flag AP Biology Challenge: GPT-5.2 9/6/2026

Flag — the response contains numerous unsupported or misleading scientific claims, including specific cleanup percentage...

— Sareena Bilal

flag AP Biology Challenge: GPT-5.2 7/26/2026

Response contains multiple overstated/unverifiable statistics (80%+ oil cleanup at Exxon Valdez, "50-90% savings" attrib...

— Fachrurrozi Rosyadi

pass AI Marketing & Content Generation 7/20/2026

"The AI response successfully creates a 3-post thread on X comparing college expectations with real-world career reality...

— Thuy Hang Vo

flag AI Marketing & Content Generation 7/20/2026

"While the AI follows the line-break constraint, the tone is overly dramatic, aggressive, and unrealistic ('Recruiters s...

— Thuy Hang Vo

This evaluation was conducted independently. xAI: Grok 4.1 Fast did not participate in or pay for this evaluation. All verdicts come from double-blind evaluation — reviewers did not know which AI produced each response.

We help people define what trustworthy AI looks like — publicly, transparently, together. Support this mission