Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

AI Generates Chest X-Rays Clinicians Cannot Tell Apart From Real Ones
AI ResearchScore: 85

AI Generates Chest X-Rays Clinicians Cannot Tell Apart From Real Ones

RadiT XL, a 1.3B-parameter rectified flow transformer trained on 1.2 million chest radiographs, produces synthetic images that clinical experts cannot reliably distinguish from real ones — a milestone that could break the data bottleneck limiting medical AI fairness and generalization.

·Jun 19, 2026·5 min read··126 views·AI-Generated·Report error
Share:
Source: arxiv.orgvia arxiv_cvWidely Reported
Can a 1.3B-parameter generative model produce chest X-rays that fool clinical experts?

A 1.3B-parameter rectified flow transformer (RadiT XL) trained on 1.2M chest radiographs generates synthetic images clinical experts cannot distinguish from real ones, achieving near-chance accuracy on real-vs-synthetic tests.

TL;DR

A 1.3-billion-parameter model trained on 1.2 million radiographs fools clinical experts, pointing toward a solution to the chronic shortage of annotated medical imaging data.

Radiologists brought in to evaluate synthetic chest X-rays produced by a new AI model performed at near-chance accuracy — essentially flipping a coin. That result, reported in a paper posted to arXiv on June 17, 2026, marks the clearest evidence yet that AI-generated medical images have crossed a clinical realism threshold that researchers have been chasing for years.

The model, called RadiT XL, was built by a team led by Fabio De Sousa Ribeiro, Emma A. M. Stanley, and Charles Jones at Imperial College London's BioMedIA group. It uses a rectified flow transformer architecture — the same family of techniques behind recent high-fidelity image generators in consumer AI — scaled to 1.3 billion parameters and trained from scratch on 1.2 million radiographs drawn from seven datasets, processed using 1.6 trillion tokens.

Why the data problem matters

Medical AI has long faced a paradox: the models most needed in clinical settings — tools that diagnose pneumonia, detect lesions, or flag rare pathologies — require enormous, diverse training sets that are legally and logistically difficult to assemble. Patient privacy rules restrict data sharing. Rare conditions mean some hospitals see only a handful of cases per year. And existing public datasets skew heavily toward specific demographics, which means AI trained on them can underperform on patients who do not match the training distribution.

The consequences are real. A 2025 study showed that radiographic AI models trained on skewed datasets produce a measurable 'underdiagnosis fairness gap' across sex, age, and race — a gap that synthetic data training pipelines can reduce by roughly 19 percent when done correctly.

RadiT XL attacks this bottleneck directly. The model supports controllable generation across 12 pathologies and multiple demographic subgroups and acquisition views, meaning researchers can, in principle, dial up synthetic images of rare conditions in underrepresented patient populations without requiring new clinical data collection.

What the evaluation actually showed

The paper reports that clinical experts evaluated synthetic and real images in a 'visual Turing test' format, achieving near-chance accuracy with low Cohen's kappa scores — a statistical measure of inter-rater agreement that approaches zero when raters are essentially guessing. The paper does not disclose exact accuracy percentages or Frechet Inception Distance scores, an omission that limits direct comparison with competing models.

Figure 2:Clinical experts’ performance on the real-vs-synthetic task across 2 presentations.Near-chance accuracy and

That gap matters competitively. RoentGen-v2, a text-to-image diffusion model developed by Stanford researchers and disclosed in a 2025 preprint, generated a dataset of 565,000 demographically balanced chest radiographs and demonstrated that synthetic pretraining improved downstream classifier accuracy by 6.5 percent across five institutions. RadiT XL's claim to the highest realism is plausible given its scale, but without published FID benchmarks the comparison remains informal.

Specialization at scale

What makes RadiT XL structurally different from prior work is the combination of size and domain focus. General-purpose image generators from major AI labs can produce photorealistic images across any subject, but they cannot reliably condition output on specific pathological features — a pneumothorax on the left versus right side, a nodule at a specific anatomical location. RadiT XL is trained exclusively on radiographs and conditions generation on detailed radiologist-guided metadata, which gives it a precision that general models cannot yet replicate.

Figure 12: Difference in Rad-DINOAP{}^{\text{AP}} ROCAUC between the 5K images sampled from our models, and their edited

This is the same logic that has driven the development of specialist language models for clinical text: broad scale helps, but domain alignment at the data level closes the last gap between plausible-looking output and clinically faithful output.

The reproducibility caveat

The paper does not release model weights or the CXR7-1M training dataset. It also does not report compute costs or training time beyond the headline parameter and token counts. In a field where reproducibility is already strained by data access restrictions, a closed model release limits the community's ability to validate the clinical realism claims independently or build on them quickly.

Figure 8: Rectified flow transformer architectures.(a) Latent-space rectified flow models operate on Rad-VAE latent tok

The authors are transparent about this: the heterogeneous training data required careful harmonization, and releasing the underlying radiographs would require navigating the same patient privacy constraints the model is designed to help circumvent. Whether that justifies a fully closed release is a question the field will need to answer as synthetic medical data becomes more commercially valuable.

What to watch

If RadiT XL weights are released — even under a research license — expect a rapid wave of downstream studies testing whether synthetic pretraining at this scale transfers performance gains to rare-disease classifiers and underrepresented patient subgroups. The FDA has not yet approved purely synthetic data for model validation; that regulatory threshold is the next meaningful unlock for the technology.


Source: arxiv_cv


Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from multiple verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The paper represents a notable scaling effort in medical image generation, but several critical details are absent. The lack of quantitative fidelity metrics (FID, IS, or any distributional distance) is a significant omission for a computer vision paper. The claim of 'indistinguishable' relies solely on a small-scale human evaluation with undisclosed sample sizes and expert demographics. The paper also does not compare against existing generative models for chest radiographs (e.g., diffusion models like Med-DDPM) on standard benchmarks, making the 'state of the art' claim difficult to verify. The architectural choice of rectified flow transformers over standard diffusion or GANs is interesting but underexplored in the paper — no ablation studies justify this choice. The Rad-DINO perceptual loss is novel but not compared against other perceptual losses (LPIPS, PSNR). The paper's strength lies in its dataset curation and scale, but the evaluation methodology is thin for a 'foundation model' claim. The lack of model or data release reduces the paper's immediate impact on the field, though it sets a benchmark for future work in specialist medical generative models.
Compare side-by-side
Fabio De Sousa Ribeiro vs Emma A. M. Stanley
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all