Alibaba DAMO Academy’s RADAR is a generalist AI model for abdominal CT analysis, detecting 146 findings across 18 anatomical structures. Research shows strong diagnostic performance and improved radiologist sensitivity and reading speed, highlighting AI’s role as a clinical decision-support tool.

Abdominal CT scans are among the most complex imaging studies in medicine. A single scan can cover the liver, pancreas, kidneys, stomach, intestines, and several other anatomical structures. Radiologists have to look for a wide range of possible findings, often while working under considerable time pressure. Misses can occur, fatigue matters, and the challenge becomes greater in healthcare systems where specialist radiologists are in short supply.

Alibaba’s DAMO Academy has now introduced a model that could significantly influence this field. Called RADAR, the vision-language system is designed to analyse contrast-enhanced abdominal CT scans and identify 146 imaging findings across 18 anatomical structures. The research, published in Science, reports expert-level performance across large real-world datasets and a multicentre reader study involving radiologists. Importantly, the researchers have also made the model code and pretrained weights publicly available for research and non-commercial use.

This is not another narrowly focused medical AI system designed to detect a single disease. RADAR represents a broader attempt to build a generalist artificial intelligence system for abdominal medical imaging.

How RADAR Was Trained Without Massive Manual Annotation

Most medical AI projects still depend heavily on expert annotation. Radiologists and other specialists may spend months or years marking lesions, organs, abnormalities, and other clinical features before a model can be properly trained.

RADAR took a different approach. The researchers assembled 424,911 contrast-enhanced abdominal CT examinations, producing approximately 1.5 million image-text pairs and more than 15 million anatomy-specific image-text pairs. Instead of asking radiologists to create entirely new labels or manually annotate every abnormality, the researchers aligned CT images with the clinical reports that had already been written for those examinations.

The model therefore learned by matching visual patterns in CT scans with the language radiologists used in their reports. This approach allowed the researchers to train the system across a much wider range of clinical findings without requiring additional manual annotation for the core training process.

That choice matters. Expert annotation remains one of the largest bottlenecks in medical AI development. If models can reliably learn from clinical reports that hospitals already generate, researchers may be able to build broader medical imaging systems much more efficiently. Similar approaches could eventually be explored in areas such as chest CT, brain imaging, MRI, and other diagnostic modalities.

RADAR’s Performance Across Real-World CT Data

RADAR was evaluated on an internal real-world test set containing nearly 40,000 examinations. Across 146 imaging findings, the model achieved an average area under the receiver operating characteristic curve, or AUC, of 0.913.

An AUC of 1.0 represents perfect discrimination, while an AUC of 0.5 represents performance roughly equivalent to random discrimination. Importantly, the researchers did not limit testing to data from the environment where the model was developed. RADAR was also evaluated across eight external medical centres, where reported AUC values ranged from approximately 0.874 to 0.912.

The researchers further tested the model on pathology-confirmed cases involving four common cancers:

  • liver cancer
  • pancreatic cancer
  • gastric cancer
  • colorectal cancer

Across these cancer cohorts, reported AUC values ranged from approximately 0.891 to 0.984. RADAR was also evaluated on acute abdominal conditions that were not specifically included in its original training objective. In that setting, the model achieved an AUC of approximately 0.904.

A separate cross-population evaluation reported an AUC of approximately 0.883 without additional fine-tuning. These results are important because medical AI systems often perform well on carefully selected internal datasets but struggle when deployed across different hospitals, scanners, populations, and clinical environments. RADAR’s external validation suggests that the model can retain substantial diagnostic performance outside its original development setting, although further prospective clinical validation will still be necessary.

How RADAR Compared With Radiologists

One of the most closely watched parts of the study was its multicentre reader evaluation.

A total of 26 radiologists from 14 medical centres participated. In the study, RADAR’s diagnostic performance exceeded that of 23 of the 26 participating radiologists on the evaluated cases. However, the more clinically important question was not simply whether the AI system could outperform individual readers.

The researchers also tested what happened when radiologists used RADAR as an assistant. Radiologists working with the system showed an approximately 10% improvement in diagnostic sensitivity, meaning more abnormalities were detected.

The study also reported a reduction in evaluation time of more than 30%. Less-experienced readers appeared to benefit particularly from AI assistance, suggesting that systems such as RADAR could potentially reduce some of the performance gap between junior and experienced radiologists.

This does not mean that RADAR replaces radiologists. The more realistic near-term application is as a clinical decision-support system that helps radiologists identify findings, prioritise attention, and reduce the cognitive burden associated with reviewing complex abdominal CT studies. That could become particularly relevant in high-volume hospitals and healthcare systems facing shortages of specialist radiologists.

Why RADAR’s Public Release Matters

Another important aspect of the project is accessibility. Many commercial medical AI systems operate as proprietary platforms, often requiring hospitals to purchase licences or use externally hosted infrastructure.

RADAR takes a different approach. The project’s software code has been released on GitHub under the Apache 2.0 licence, while pretrained model weights are publicly available through Hugging Face.

However, there is an important distinction. The pretrained weights are distributed under a CC BY-NC-SA 4.0 licence, which permits research and other non-commercial uses but places restrictions on commercial deployment. Therefore, RADAR should not be interpreted as an unrestricted commercial open-source product.

Still, making the model architecture, implementation, and pretrained weights available gives universities, hospitals, researchers, and medical AI teams an opportunity to evaluate the technology without having to build a comparable system from the beginning. Because the inference code and model weights are publicly available, institutions can also explore running and validating the system on their own infrastructure rather than relying exclusively on a proprietary hosted service. Actual hardware requirements, however, will depend on the implementation and deployment environment.

The Limits of RADAR and Medical Imaging AI

No single research paper solves radiology.  RADAR was developed primarily using contrast-enhanced abdominal CT examinations. Its performance on non-contrast CT studies, other anatomical regions, or entirely different imaging modalities remains to be established.

Real-world clinical deployment may also expose challenges involving scanner differences, image quality, unusual patient populations, rare diseases, and workflow variations that are difficult to capture completely in retrospective studies.

Clinical use would require appropriate regulatory review, local validation, workflow integration, data governance, and clinical oversight, with requirements varying across countries and healthcare systems. The system also relies on upstream image-processing and anatomical localisation steps. Errors introduced earlier in the pipeline could potentially influence downstream predictions.

Most importantly, the published results demonstrate diagnostic model performance and benefits within a controlled radiologist reader study. They do not yet prove that deploying RADAR in routine hospitals will directly improve patient outcomes, reduce mortality, shorten hospital stays, or improve treatment decisions.

Those questions require prospective clinical studies. The researchers themselves position the work as a foundation rather than an endpoint. The same training strategy could potentially be extended to other imaging modalities and anatomical regions, which may eventually move medical imaging AI away from hundreds of isolated disease-specific systems toward broader diagnostic platforms.

What RADAR Signals for the Future of Medical AI

Medical AI has spent years producing systems capable of impressive results under relatively narrow conditions. RADAR is significant because it advances several areas at once. It was trained using hundreds of thousands of real-world CT examinations.

It learned from existing clinical reports rather than depending entirely on newly created expert annotations. It covers 146 different imaging findings across 18 anatomical structures.It demonstrated substantial performance across external medical centres and different patient populations.

And, perhaps most importantly, it improved the performance of radiologists when used as an assistive system. The larger story is therefore not simply about an AI model outperforming individual doctors in a study.

It is about a potential shift in how medical imaging AI systems are built. Learning from the clinical reports that doctors already produce may offer a more scalable path than relying exclusively on expensive pixel-level annotation. At the same time, broader generalist models could reduce the need to deploy separate AI tools for every disease or organ.

The strongest implication from the RADAR study may ultimately be the most practical one. Radiologists are unlikely to disappear. Medical imaging volumes continue to grow, while specialist availability remains uneven across hospitals and countries. The more immediate opportunity is therefore to use AI systems to help radiologists review large numbers of examinations more efficiently and reduce the risk of overlooked findings.

A system that improves sensitivity while reducing reading time could be valuable even if final clinical responsibility remains with the physician. That human-AI model may prove far more realistic than the idea of replacing specialists entirely.

RADAR offers an early example of what such a system could look like.The real test will come as independent research groups and medical institutions evaluate the model on their own patient populations, identify where it performs well, and document where it fails.

That process has only begun. For now, RADAR represents an important step toward generalist medical imaging AI — systems designed not merely to detect one disease, but to assist specialists across a much broader diagnostic landscape.


Discover more from Poniak Times

Subscribe to get the latest posts sent to your email.