General-purpose large language models outperform specialized clinical AI tools on medical benchmarks
Krithik Vishwanath et al.
ATTENZIONE ! UNO STUDIO RANDOMIZZATO CON 12 MEDICI SU NATURE CONFRONTA LE PIATTAFORME GENERALI CON QUELLE MEDICHE SPECIALIZZATE (OPENEVIDENCE, UPTODATE). AND THE WINNER IS….
______________________________________________________________________

KEY POINTS FROM IA
Uno studio di giugno 2026 pubblicato su Nature Medicine ha rilevato che i modelli di intelligenza artificiale generalisti, tra cui Gemini 3.1 Pro, GPT-5.2 e Claude hanno superato le piattaforme cliniche specializzate (OpenEvidence, UpToDate) nella conoscenza medica e nell’allineamento con i medici. Lo studio ha impiegato una valutazione alla cieca e randomizzata da parte di 12 medici, utilizzando esami abilitativi e quesiti clinici reali, dimostrando che la scala dei modelli di base attualmente supera la sintonizzazione su domini ristretti in medicina.
_______________________________________________________________________
Abstract
Specialized clinical artificial intelligence (AI) tools are entering medical practice despite scarce independent evaluation. We quantitatively evaluate two clinical AI tools, OpenEvidence and UpToDate Expert AI, built on large language models (LLMs) against three frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. Our evaluation has three stages: (1) 500 MedQA questions testing medical knowledge, (2) 500 HealthBench items measuring alignment with clinicians and (3) the real clinical queries (RCQ) benchmark, built from 100 de-identified queries from physicians to a general-purpose language model in a live clinical environment. For the RCQ benchmark, 12 US clinicians performed randomized, blinded review of model outputs, producing 1,800 model–question annotations. Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ. These findings highlight the need for independent, real-world evaluation of AI tools before they enter clinical settings.
Main
Specialized clinical artificial intelligence (AI) tools are entering medical practice at scale1,2. These proprietary large language model (LLM)-based tools promise superior clinical performance to general-purpose frontier LLMs as a result of domain-specific training or retrieval-augmented generation (RAG)3. Yet, their architectures, base models and training pipelines are not public. Clinicians and health systems must therefore assess their value and safety without independent evidence. Conversely, large training corpora and extensive alignment of frontier LLMs may enable them to challenge clinical AI tools without domain-specific modification. We test this hypothesis by comparing clinical AI tools (OpenEvidence1 and UpToDate Expert AI2) to leading general-purpose LLMs (OpenAI GPT-5.2, Google Gemini 3.1 Pro Preview and Anthropic Claude Opus 4.6). Later, we include auto-enabled Google Search AI Overview as a real-world control frequently encountered by physicians.
Our evaluation (Fig. 1) has three stages: (1) 500 US Medical Licensing Examination-style MedQA4 questions assessing medical knowledge, (2) 500 HealthBench5 items evaluating agreement with expert clinicians and (3) 100 real clinical queries (RCQ) drawn from physician LLM queries during live clinical deployment. The RCQ stage underwent randomized, blinded review by 12 US clinicians, producing 1,800 model–question annotations. The combined analysis spans multiple-choice reasoning, expert clinical judgment and everyday clinician use.
Fig. 1: Clinical LLM evaluation pipeline.
Schematic overview of the comparative analysis between frontier, clinical-specific and search-embedded AI models. The framework integrates automated scoring (MedQA and HealthBench) with high-fidelity blinded and randomized clinician reviews (RCQ) to assess model performance across accuracy, safety and reliability metrics. N, no; USMLE, US Medical Licensing Examination; Y, yes. Blindfold/blinded icon from Tailwind Labs under an MIT license (©Tailwind Labs); other icons from React Icons under a CC BY 4.0 (brain, stethoscope, book, bar chart, doctor, justice scale, chat, shield) or Apache 2.0 (browser, API badge).

Fig. 2: Comparative evaluation of AI systems across benchmark performance and real-world clinical use.




