Home / Blog / Anthropic still don’t know how people are really using AI
Tech News

Anthropic still don’t know how people are really using AI

Researchers launched the AI Observatory to analyze 24,521 real AI conversations, revealing that corporate reports filter out 48 percent of non-work usage data.

By Dillip Chowdary • Oct 10, 2026 • Source: MIT Technology Review

Anthropic still don’t know how people are really using AI

Researchers from Stanford, MIT, the Data Provenance Initiative, and other academic institutions launched the AI Observatory to address the lack of independent oversight in how people interact with generative artificial intelligence. The public research platform aggregated real-world chat records to provide external verification of consumer habits, as detailed in MIT Technology Review's report. By creating an independent repository, the project aims to counter the selective disclosures published by commercial AI vendors.

This coverage examines the structural gaps in vendor-published usage reports and outlines how academic researchers constructed an open alternative. It analyzes the specific behavioral differences across distinct LLM platforms, the variance in sensitive topic exposure, and the implications for policymakers relying on corporate transparency data. The analysis serves technology researchers, policy analysts, and enterprise deployments evaluating real-world model interaction risks.

Anthropic still don’t know how people: what actually changed

Major artificial intelligence companies regularly issue research papers detailing consumer activity on their platforms, but independent academic researchers argue these reports present a curated subset of reality. Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab, co-led the launch of the AI Observatory to establish an unvetted baseline. The project aggregated 85,633 conversational turns across 24,521 conversations, representing 5,000 users interacting with 52 different models between 2023 and 2025. The resulting platform demonstrates that actual user behavior diverges significantly from the productivity-focused narrative emphasized by commercial labs.

Corporate transparency efforts, such as the Anthropic Economic Index, rely on strict internal filters that remove non-work conversations from their datasets. When AI Observatory researchers applied Anthropic’s filtering methodology to their own independent dataset, they discovered that 48% of total user interactions were filtered out. The excluded conversations contained substantially higher rates of personal, sensitive, or prohibited content than corporate reports acknowledge. While OpenAI’s internal 2025 report indicated that 30% of consumer ChatGPT use was work-related, independent data shows that overall model usage across all platforms is heavily weighted toward non-commercial and social interactions.

Anthropic still don’t know how people: how it works

Anthropic still don’t know how people are really using AI
Illustration · Pexels

The AI Observatory constructed its analytical foundation by combining seven existing open-source datasets collected with user consent, including the WildChat repository. The research team evaluated interaction trends across platforms including Anthropic's Claude, OpenAI's ChatGPT, Google's Gemini, and xAI's Grok. By examining raw prompt tokens, response tokens, and conversation turn lengths, the project established long-term trends in how prompt structures evolved between 2023 and 2025. This bird's-eye perspective allows cross-model comparisons that proprietary single-vendor reports cannot provide, according to David Widder, an assistant professor at the University of Texas at Austin.

Analysis of the cross-platform dataset revealed clear functional segregation based on the specific system architecture. Users consistently selected Anthropic models for software coding tasks, while turning to ChatGPT primarily for educational and homework assistance. Google’s Gemini saw higher concentrations of social interaction and roleplay use cases. xAI’s Grok recorded the highest frequency of news and political queries, but also contained the highest concentration of misinformation proliferation. Interaction depth also varied across model generations; conversations with OpenAI's GPT-3.5 remained short and transactional, whereas interactions with GPT-4o expanded into longer, highly iterative sessions.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

Anthropic still don’t know how people: why it matters now

Regulatory bodies and enterprise risk officers currently make policy decisions based on usage statistics published directly by AI providers. The AI Observatory's comparative analysis proves that company-published benchmarks exclude critical behavioral categories. In the unFiltered AI Observatory dataset, non-work conversations involved health and relationship topics at a rate of 44.2%, compared to 31.2% in Anthropic's official reporting. Furthermore, adult or illicit topics appeared in 7.9% of Observatory conversations versus 2.1% in Anthropic's data, while sexual content reached 16.7% compared to the vendor's reported 2.4%.

Safety filtering and platform moderation also showed measurable shifts over the two-year observation period. Exchanges categorized by researchers as sensitive—including hate speech and harassment—decreased in frequency over time, indicating that safety alignment and system prompt guardrails became more effective. However, the prevalence of general small talk increased alongside a decrease in chatbot self-disclosure regarding artificial identity. Shayne Longpre, a recent MIT Media Lab PhD graduate who co-led the study, emphasized that no single commercial report captures these structural shifts, leaving regulators to operate without objective benchmarks.

Anthropic still don’t know how people: who is affected

The reliance on vendor-curated data directly impacts academic researchers, safety auditors, and government regulators who require unbiased access to model interaction metrics. Proprietary barriers prevent external scientists from evaluating whether general-purpose AI systems produce net positive societal outcomes or generate systemic harms. Because companies like Anthropic and OpenAI analyze internal datasets comprising 1 million and 1.5 million conversations respectively, their scale vastly exceeds the 24,521 conversations available to the AI Observatory. This data disparity restricts public oversight to information selected by commercial entities.

End users and consumer advocacy groups face unmitigated risks when platforms alter model behavior without public disclosure. The transition from older architectures to systems like GPT-4o altered conversational duration and emotional engagement patterns without external baseline testing. Additionally, the voluntary nature of open academic datasets means even the AI Observatory likely undercounts sensitive or illicit uses, as users remain hesitant to share private interactions. Without mandated data-sharing protocols for independent researchers, decision-makers risk establishing safety standards based on incomplete corporate narratives.

Anthropic still don’t know how people: what to watch

The research team plans to expand the AI Observatory by integrating additional public datasets and updating analytical frameworks as new models enter deployment. Future research will monitor whether major labs adopt privacy-preserving mechanisms to share anonymized chat logs with credentialed academic institutions. An Anthropic representative noted that published company research reflects specific internal research questions and affirmed the importance of supporting third-party academic studies. OpenAI did not respond to requests for comment regarding independent data access.

The ongoing disparity between academic sample sizes and corporate data stores will remain a primary challenge for independent AI governance. Researchers will continue monitoring Grok and other real-time information retrieval models to track misinformation density during major news cycles. As legislative bodies refine artificial intelligence compliance frameworks, pressure will build on commercial vendors to provide verified telemetry. Until independent access is formalized, academic observatories represent the primary mechanism for auditing the gap between corporate claims and real-world AI usage.

Developer Action Items

  • ☐ Verify the claim on the official OpenAI / Anthropic / Claude page (or MIT Technology Review), not from this recap alone.
  • ☐ Name the surface that moved — API, policy, model, hardware, or commercial terms — before you Slack the thread.
  • ☐ Assign one owner a day to read the primary material and decide: this-sprint, this-quarter, or noise.
  • ☐ Do not change production on day-one coverage. Watch the vendor changelog and one independent write-up first.

Anthropic still don’t know how people FAQ

What is the AI Observatory?

The AI Observatory is an independent research project co-led by Stanford and MIT researchers that aggregates real-world AI conversation data from seven open datasets to analyze how people use generative AI.

How does user behavior differ from official AI company reports?

Independent analysis shows that 48% of conversations would be filtered out by corporate reporting methods, revealing significantly higher rates of health, relationship, harassment, and adult content than company data claims.

Which models were evaluated in the study?

The study analyzed 24,521 conversations involving 5,000 users across 52 models, including Anthropic's Claude, OpenAI's ChatGPT, Google's Gemini, and xAI's Grok.

How do interaction patterns vary across different AI platforms?

Users primarily select Anthropic models for coding, ChatGPT for homework assistance, Gemini for social roleplay, and Grok for news and political queries.

Sources

Dillip Chowdary

Author

Dillip Chowdary

Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.

Related on Tech Bytes

Advertisement

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →