OpenAI announced on June 18 that GPT-5.5 Instant — the default model for free ChatGPT users — now outperforms doctor-written health responses on accuracy, communication clarity, and completeness in the company's own evaluations. The rate of health responses flagged for factuality errors has fallen 71% over two months, OpenAI said, based on monitoring of live production traffic. What the announcement did not lead with: OpenAI designed the benchmark, organized the physician review panel, and controlled what test conversations were evaluated.
What OpenAI actually measured
HealthBench, the framework OpenAI used to assess this upgrade, is an open-source evaluation it built and maintains. It draws on 5,000 multilingual health conversations and more than 48,000 physician-written rubrics that define what a good response looks like. For the doctor-vs-AI comparison specifically, OpenAI asked a panel of physicians to write responses to representative health questions, then had a separate group of doctors score both the AI-written and human-written answers. Across 3,500 reviewed interactions, the evaluators rated GPT-5.5 Instant higher on every criterion, with the model scoring up to 89.9% on instruction-following.
The physician network supporting that evaluation numbers more than 260 doctors spanning 60 countries, 49 languages, and 26 specialties, who have collectively reviewed over 700,000 model responses. Their role is not clinical care but assessment: they score AI outputs and help OpenAI define what quality health communication requires.
The independent evidence tells a different story
The framing of 'beats doctors' rests entirely on that self-administered test. A 2026 Mass General Brigham study, conducted without OpenAI's involvement, found that AI chatbots got initial diagnoses right under one-fifth of the time when given limited patient information — the kind of partial scenario typical of a real ChatGPT conversation. OpenAI's own evaluation compared against generic physician-written answers, not specialist consultations, in-person diagnosis, or cases where a doctor had access to a full patient history.
The 71% reduction in flagged factuality errors is drawn from production traffic monitoring, which suggests real-world improvement at scale but is not the same as clinical accuracy. OpenAI did not disclose whether the physician evaluators knew they were being compared against an AI.
Where GPT-5.5 Instant actually ranks
Even on OpenAI's own HealthBench Professional leaderboard — the harder, clinician-focused variant — GPT-5.5 Instant is not the leader. Anthropic's Claude Opus 4.8 currently tops the board at 55.8%. GPT-5.5 Instant scores 38.4%. That gap is notable in a story framed around the model matching 'frontier' performance.
What the update does deliver is a genuine access shift. Capabilities that previously required a paid ChatGPT subscription are now available on the free tier. Given that more than 230 million people use ChatGPT weekly for health-related questions — from decoding lab results to preparing for appointments to navigating insurance — the practical reach of this improvement is significant even if the headline comparison is carefully constructed.
Why the market timing matters
The announcement comes as ChatGPT's share of global AI chatbot web traffic has slipped to roughly 54.7% — down from over 76% in early 2024 — with Google Gemini holding about 27.4% and Anthropic's Claude at around 8%. Health is one of the highest-stakes verticals for retaining free users who might otherwise migrate. OpenAI also offers ChatGPT for Clinicians and a broader OpenAI for Healthcare product line, making this free-tier upgrade partly a pipeline move for those commercial offerings.
The upgrade reflects a real technical advance, documented across multiple independent evaluators even if they were organized by OpenAI. But the language of 'beating doctors' is doing more rhetorical work than the underlying evidence supports.
Key facts
- 71% reduction in health responses flagged for factuality errors over two months (OpenAI's own monitoring)
- 3,500 physician-reviewed response pairs used in the doctor-vs-AI comparison
- 260+ doctors, 60 countries, 49 languages, 26 specialties in the evaluation network
- 700,000+ model responses reviewed by that physician network
- GPT-5.5 Instant scores 38.4% on HealthBench Professional; Claude Opus 4.8 leads at 55.8%
- 230 million+ weekly ChatGPT users ask health-related questions
- 2026 Mass General Brigham study: AI chatbots correct on initial diagnoses under 20% of the time
What to watch: Whether a peer-reviewed journal or an independent health AI auditor can replicate the 71% error reduction figure — and whether the HealthBench Professional gap between GPT-5.5 Instant and Claude Opus 4.8 narrows in the next model cycle.
Source: the_decoder








