Radiologists brought in to evaluate synthetic chest X-rays produced by a new AI model performed at near-chance accuracy — essentially flipping a coin. That result, reported in a paper posted to arXiv on June 17, 2026, marks the clearest evidence yet that AI-generated medical images have crossed a clinical realism threshold that researchers have been chasing for years.
The model, called RadiT XL, was built by a team led by Fabio De Sousa Ribeiro, Emma A. M. Stanley, and Charles Jones at Imperial College London's BioMedIA group. It uses a rectified flow transformer architecture — the same family of techniques behind recent high-fidelity image generators in consumer AI — scaled to 1.3 billion parameters and trained from scratch on 1.2 million radiographs drawn from seven datasets, processed using 1.6 trillion tokens.
Why the data problem matters
Medical AI has long faced a paradox: the models most needed in clinical settings — tools that diagnose pneumonia, detect lesions, or flag rare pathologies — require enormous, diverse training sets that are legally and logistically difficult to assemble. Patient privacy rules restrict data sharing. Rare conditions mean some hospitals see only a handful of cases per year. And existing public datasets skew heavily toward specific demographics, which means AI trained on them can underperform on patients who do not match the training distribution.
The consequences are real. A 2025 study showed that radiographic AI models trained on skewed datasets produce a measurable 'underdiagnosis fairness gap' across sex, age, and race — a gap that synthetic data training pipelines can reduce by roughly 19 percent when done correctly.
RadiT XL attacks this bottleneck directly. The model supports controllable generation across 12 pathologies and multiple demographic subgroups and acquisition views, meaning researchers can, in principle, dial up synthetic images of rare conditions in underrepresented patient populations without requiring new clinical data collection.
What the evaluation actually showed
The paper reports that clinical experts evaluated synthetic and real images in a 'visual Turing test' format, achieving near-chance accuracy with low Cohen's kappa scores — a statistical measure of inter-rater agreement that approaches zero when raters are essentially guessing. The paper does not disclose exact accuracy percentages or Frechet Inception Distance scores, an omission that limits direct comparison with competing models.

That gap matters competitively. RoentGen-v2, a text-to-image diffusion model developed by Stanford researchers and disclosed in a 2025 preprint, generated a dataset of 565,000 demographically balanced chest radiographs and demonstrated that synthetic pretraining improved downstream classifier accuracy by 6.5 percent across five institutions. RadiT XL's claim to the highest realism is plausible given its scale, but without published FID benchmarks the comparison remains informal.
Specialization at scale
What makes RadiT XL structurally different from prior work is the combination of size and domain focus. General-purpose image generators from major AI labs can produce photorealistic images across any subject, but they cannot reliably condition output on specific pathological features — a pneumothorax on the left versus right side, a nodule at a specific anatomical location. RadiT XL is trained exclusively on radiographs and conditions generation on detailed radiologist-guided metadata, which gives it a precision that general models cannot yet replicate.

This is the same logic that has driven the development of specialist language models for clinical text: broad scale helps, but domain alignment at the data level closes the last gap between plausible-looking output and clinically faithful output.
The reproducibility caveat
The paper does not release model weights or the CXR7-1M training dataset. It also does not report compute costs or training time beyond the headline parameter and token counts. In a field where reproducibility is already strained by data access restrictions, a closed model release limits the community's ability to validate the clinical realism claims independently or build on them quickly.

The authors are transparent about this: the heterogeneous training data required careful harmonization, and releasing the underlying radiographs would require navigating the same patient privacy constraints the model is designed to help circumvent. Whether that justifies a fully closed release is a question the field will need to answer as synthetic medical data becomes more commercially valuable.
What to watch
If RadiT XL weights are released — even under a research license — expect a rapid wave of downstream studies testing whether synthetic pretraining at this scale transfers performance gains to rare-disease classifiers and underrepresented patient subgroups. The FDA has not yet approved purely synthetic data for model validation; that regulatory threshold is the next meaningful unlock for the technology.
Source: arxiv_cv









