Home / Blog / Students prefer Gemini over ChatGPT and Claude for AI…
Tech News

Students prefer Gemini over ChatGPT and Claude for AI essays in blind

Points: 1 # Comments: 0 Students prefer Gemini over ChatGPT and Claude for AI essays in blind Coverage based on HN Claude/Codex/Fable reporting.

By Dillip Chowdary • Aug 26, 2026 • Source: HN Claude/Codex/Fable

Students prefer Gemini over ChatGPT and Claude for AI essays in blind

What happened

Students prefer Gemini over ChatGPT and Claude for AI essays in blind tests

College-essay season has its first data-backed model ranking. StudyArena, an AI study platform, collected 6,851 eligible blind student votes across ChatGPT, Claude, and Gemini — and as of August 2026, students chose Gemini-family writing answers more often than the competition. The platform, built by co-founders Pasha Rayan and Pennie Li, shows a clear preference gap when the model name is hidden until after the student has already judged the response.

This piece breaks down how the test was run, what the choice rates actually mean for students writing personal statements and argumentative essays, and why the ranking reshuffles depending on which stage of writing you are working on. It is relevant to any student currently choosing a default AI tool, and to educators and product builders trying to understand which models real users prefer when branding is removed.

How it works

StudyArena published an analysis of production data drawn from August 2026, covering 6,851 blind writing and essay comparisons run through the platform's arena interface. Students saw responses before they saw the model name and then picked the answer they preferred. Across those decisive writing ballots, Gemini finished first with a 39.6% blind writing choice rate. Claude came second at 31.8%. ChatGPT and other OpenAI models reached 29.2%. The analysis was written by Pasha Rayan, edited by Giovana Rabello, and reviewed by Pennie Li. Internal and admin activity was excluded from the count, as were ballots ineligible for the public leaderboard.

The models tested are the current-generation variants, not older releases. The arena includes GPT-5.6 Sol from OpenAI, Claude Opus 5 from Anthropic, and Gemini 3.1 Pro from Google. The methodology groups model variants by provider family and defines choice rate as the share of decisive ballots where a provider's answer was selected, meaning three-way splits and ties were filtered out before the percentages were calculated.

Students prefer Gemini over ChatGPT and Claude for AI essays in blind
Illustration · Pexels

StudyArena's arena interface sends the same prompt to multiple models simultaneously, strips identifying labels from the responses, and presents them to the student side by side. The student picks a preferred answer. Only after the selection is recorded does the platform reveal which model wrote which response. This design removes the brand-preference effect, which is the tendency for people to rate a response higher once they know it comes from the model they already favor. The 6,851 eligible votes reported here cover writing and essay tasks specifically, pulled from aggregate production data rather than a controlled lab study.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

Why it matters

Each provider's choice rate in the writing category reflects only decisive ballots, those in which one answer clearly won rather than the student declining to choose. Gemini's 41.7% lead in the writing feedback subcategory and Claude's 43.2% lead in assignment planning were derived the same way. GPT's 39.3% lead in research-related tasks came from a separate subset of ballots filtered for research prompts. StudyArena states it will update the recommendation when the results move, framing this as a live ranking rather than a one-time benchmark.

The 10-point spread between Gemini at 39.6% and ChatGPT at 29.2% is large enough to be operationally useful, not just statistically interesting. Students making a decision about which subscription to pay for, or which tab to keep open during an application cycle, now have a preference signal grounded in actual student behavior rather than benchmarks designed by the model providers themselves. The finding also runs against the assumption, common in developer circles, that more powerful reasoning models produce better prose. StudyArena's data shows the opposite: responses from models at a low reasoning setting won 40.7% of decisive writing comparisons, while responses at a high reasoning setting won only 29.5%.

Who is affected

The subcategory breakdown adds a second layer of utility. Claude leading assignment planning at 43.2% and ChatGPT leading research framing at 39.3% means the overall Gemini win is not a clean sweep. If a student is mapping counterarguments before writing a draft, the data suggests Claude is a better first stop at that stage. The ranking tells builders that no single model dominates every writing sub-task, and that task routing — sending different parts of a job to different models — is validated by real student preference data.

The most directly affected group is undergraduate and graduate students currently in or approaching an application cycle for college, graduate school, or scholarship programs. For those students, a blind preference signal covering 6,851 votes gives them a concrete default: start with Gemini when editing or improving a draft, then check Claude for structural feedback and ChatGPT for research framing. Students paying for a subscription to only one provider have data that points toward Gemini as the broadest single choice for writing tasks, though the subcategory results show meaningful tradeoffs.

Educators, advisors, and writing center staff who recommend AI tools to students are a second affected group. The blind methodology matters here because it reduces the risk that a recommendation is driven by brand familiarity. Builders working on edtech products that route prompts to AI providers now have evidence that students notice writing quality differences large enough to produce a 10-point gap between the top and bottom finishers, which sets a practical floor on the quality delta that a routing decision needs to clear in order to matter to end users.

What to watch next

The most important thing to verify is whether Gemini's 39.6% choice rate holds as the platform accumulates more ballots beyond the current 6,851. StudyArena says the recommendation will be updated when results move, so the percentages published in August 2026 should be treated as a snapshot rather than a settled ranking. Claude Opus 5 and GPT-5.6 Sol are recent releases, and user preferences across novel model versions often shift in the weeks after launch as students develop familiarity with each model's defaults. The current scores reflect preference during early exposure to this generation of models.

A second variable to watch is the reasoning-setting result. The finding that low-effort settings won 40.7% of decisive writing comparisons while high-effort settings won only 29.5% suggests that model providers tuning their default outputs for essay tasks may have more to gain from reducing reasoning verbosity than from increasing it. If any of the three providers adjusts defaults or introduces a writing-optimized mode before the next application season, the subcategory rankings — particularly Gemini's 41.7% writing-feedback lead — are the most likely figures to shift first.

Developer Action Items

  • Verify the claim on the official OpenAI / Anthropic / Claude page (or HN Claude/Codex/Fable), not from this recap alone.
  • Name the surface that moved — API, policy, model, hardware, or commercial terms — before you Slack the thread.
  • Assign one owner a day to read the primary material and decide: this-sprint, this-quarter, or noise.
  • Do not change production on day-one coverage. Watch the vendor changelog and one independent write-up first.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →