Claude on Political Compass
Fetching the source article so the paragraphs stay factual and specific.Utopia ran Anthropic language models through a political compass survey and published…
By Dillip Chowdary • Aug 04, 2026 • Source: HN Claude/Codex/Fable
Fetching the source article so the paragraphs stay factual and specific.Utopia ran Anthropic language models through a political compass survey and published the results at utopiagov.com/blog/political-compass. The work covers six models — Claude Haiku 4.5, Sonnet 5, Opus 4.6, Opus 4.8, Opus 5, and Fable 5 — answering 54 agree/disagree propositions, each proposition in its own request with no surrounding conversation. Every model was put through the full survey one hundred times under a plain framing and again under a short instruction change, for 1,000 survey passes and 54,000 total requests. On Hacker News the write-up sat at 2 points with 0 comments.
The survey is the open 8values test: each proposition maps to fixed weights on an economic axis (left to right) and a social axis (libertarian to authoritarian), and the five-choice answers are scored into a single point on that plane. Propositions are sent one at a time so the only context is the survey framing plus the claim. Nondeterminism is measured by repeating the same survey with nothing changed; averages stabilize within a handful of runs to within about a tenth of a point of the hundred-run mean, so a short run already places a model. Most of the 54 items return the same answer every pass; a minority never settle. Items that never vary include equal treatment regardless of culture or sexuality (strongly agree on every run), opposition to trade tariffs as a local-production tool, support for consumer-protecting economic intervention, rejection of balanced budgets over welfare, support for publicly funded research over market-only research, and rejection of abolishing social programs for private charity. Higher-variance items include violence in protest against authoritarian government, necessity of military action, religious or traditional education of children, worker ownership of the means of production, and support for regional unions such as the EU.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers who ship or evaluate assistants, the useful result is not a moral score but a measurement protocol: fixed items, published weights, isolated requests, and enough repeats to separate stable stance from sampling noise. The same protocol shows how thin the control surface is. Adding the two words Act centrist to the framing moved every model toward the center on both axes, and every model moved farther on the social axis than on the economic one. That is a concrete reminder that product defaults, system prompts, and evaluation harnesses can shift reported political posture without any weight change, and that single-shot survey answers without multi-run averages will overfit noise on the unsettled minority of items.
In market terms the piece is a cross-model comparison inside one lab’s stack rather than a vendor shootout: Haiku, Sonnet, Opus line, and Fable 5 are placed on the same axes under the same prompts. All eight plotted measurements — plain framing and centrist instruction for each model — land in a small corner of the libertarian-left quadrant. The interesting competitive signal is the spread across model size and generation on high-gap propositions (immigration, surveillance, military force, cultural superiority claims) while core equality and several economic-left items stay locked. That pattern matters more for anyone building multi-model routing, red-team suites, or “neutral” product personas than a single headline label would.
The practical takeaway is to treat political-compass style probes as regression tests: pin the proposition list and weights, log answer distributions not single draws, and measure how many characters of instruction move the average. Watch next for whether other labs or third parties publish the same 8values (or equivalent) battery on non-Anthropic models with the same hundred-run discipline, and whether product teams document how default system prompts shift those averages on the social axis relative to the economic one — the gap this study already quantifies under Act centrist.
Advertisement
🔎 More interesting news
- Swarm of OpenAI Agents Exploit Artifactory Zero-Day to Escape Sandbox and Breach Hugging…
- Claude Code can read plaintext secrets even when Read is denied
- 150,000 Impacted by Madera Community Hospital Data Breach
- Why is Anthropic's public writing style so unlike Claude's?
- Today's full Tech Pulse briefing →