Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

Cardiac MRI In, Chamber Overlays LGE Flags and On-Scanner Report Out in About 90 Seconds

Hang sequence labels, chamber and scar overlays, LGE and disease flags, then a radiologist-style report written back as a DICOM series on the scanner GPU in about 90 seconds, with no external network call. That is the clinic ask Demirel, Horst, and colleagues take on in arXiv:2609.23950, posted around late September 2026. The paper is ORION-CMR: On-scanner Reporting with Integrated Foundation Model for End-to-End Cardiac MRI Analysis and Interpretation. The stack is a shared self-supervised ViT backbone plus a locally hosted Qwen2.5-14B via Ollama that only sees structured findings, so the language model does not invent measurements or scars on its own.

CMR still sits behind fragmented post-processing tools and scarce reader coverage. A recent US analysis they cite reports fewer than 1,000 physicians interpreting CMR, with tens of millions of people living far from a service site. ORION-CMR tries to collapse sequence triage, ventricular function, LGE detection, multiclass disease labels, and report drafting into one scanner-native pass on a Philips 1.5T Ambition X with an NVIDIA RTX A6000.

What hangs on the viewer

Upstream is a multi-sequence CMR study: localizers, cine 2/3/4-chamber and short-axis (SAX), LGE (PSIR, IR, MOCO), mapping, perfusion, and flow. The first head classifies the series into eleven sequence types so the right downstream head fires without a technologist hand-picking series. Cine-SAX segmentation returns LV, myocardium, and RV masks that feed volumes, ejection fractions, and indexed LV mass. LGE and disease heads emit binary infarct presence and a four-way label among normal, congenital heart disease, dilated cardiomyopathy, and myocardial infarction. Those numbers and labels land in a deterministic structured study object. The local LLM turns that object into a clinical impression, exported as a new DICOM series on the console.

For a DICOM viewer or cardiac AI rail, the hangable pieces are chamber overlays on cine-SAX, scar-related LGE flags, disease tags next to the study, and a draft report series you can open beside the original images. Nothing in that path leaves the scanner network.

How it works in plain words

Pretraining used a DINO-style teacher-student setup on 12,896,733 CMR slices from 9,258 multi-vendor Mayo Clinic studies (GE and Siemens, 2017-2025). Two ViT-Base patch sizes were tried; ViT-B8 beat ViT-B16 on the benchmarks they care about and became the shared encoder. Downstream, the encoder stays fixed. Lightweight heads do the work: an MLP for sequence classification from concatenated late-block embeddings, a small convolutional decoder for cine-SAX and LGE-SAX segmentation, and calibrated linear SVMs on pooled CLS tokens for LGE and disease labels.

Ventricular volumes and mass are checked against Healthy Hearts Consortium 2024 age-, sex-, and ethnicity-specific reference ranges before they enter the structured findings. LGE classifier outputs and disease predictions join the same object. Qwen2.5-14B-Instruct, running locally through Ollama, writes the report only from that protected schema. Image analysis and language generation stay separated on purpose: the LLM never sees raw pixels when it drafts the impression.

On the Ambition X console the pipeline runs as series arrive. Sequence classification takes about 11.2 ± 3.9 s, cine-SAX segmentation about 51.6 ± 15.5 s, LGE and disease classification about 9.7 ± 3.3 s, and report generation about 11.2 ± 3.2 s. End to end is about 90 seconds per subject in their timing.

What the numbers say

Clinical hold-out is 68 subjects (25 normal, 7 congenital heart disease, 15 dilated cardiomyopathy, 21 myocardial infarction) drawn from a larger 426-subject multi-vendor in-house set. Normal-versus-abnormal AUC is 0.960 (0.906-0.998) versus 0.761 for a ResNet-18 baseline. Multiclass one-vs-rest AUCs are 0.960 (normal), 0.848 (congenital), 0.845 (dilated cardiomyopathy), and 0.877 (myocardial infarction); the paper summarizes multiclass disease classification at AUC 0.88. Clinical LGE classification AUC is 0.826 (0.722-0.918) versus 0.706 for ResNet-18.

Three expert readers scored generated impressions against the original clinical reports on a three-level scale. Concordant share is 81.4% ± 1.7%, intermediate 13.7% ± 3.7%, discordant 4.9% ± 2.2%. Pairwise observed agreement sat between 76.5% and 88.2%, with Cohen’s κ from 0.41 to 0.61. On public sets, ORION-CMR reaches EMIDEC LGE classification 0.930 and EMIDEC scar Dice 0.799 / myocardium Dice 0.915, ahead of the prior published SoTA they cite. ACDC disease overall is 0.820 versus 0.700 for CMR-FM. Sequence classification accuracy on the clinical cohort is 0.992 overall.

Where it fails and what not to trust

The clinical hold-out is single-center and small, especially congenital heart disease (n=7). Ventricular measurements were checked against values written in the original reports, not against fresh expert contours. Discordant cases in Fig. 3 include a missed acute LAD infarct with transmural scar where the generated impression claimed normal size and function with no LGE, a bicuspid aortic valve case misframed as LGE-positive dilated cardiomyopathy, and an over-called scar when the original report had none. Intermediate cases often disagree on LGE presence while the disease flag still hints at abnormality.

Public ACDC segmentation sits close to published SoTA but does not beat it on every structure; myocardium Dice trails the cited SoTA. In-house cine-SAX contours were unavailable, so the ACDC-trained decoder was applied without fine-tuning. Ablations of individual pipeline pieces and deeper LLM prompt or schema studies are left for later work. Treat the ~90 s timing as hardware-specific to their Ambition X plus A6000 setup. This is a research deployment description, not a cleared medical device.

For a viewer or clinic AI shop

Wire multi-sequence CMR in, sequence tags plus LV/MYO/RV overlays plus LGE and disease labels out, then a draft report series written locally from structured findings. Prefer keeping the LLM behind a verified measurement and classifier schema rather than letting it free-read pixels. Surface concordance risk next to the draft: green concordant examples look clinically usable; red discordant misses (missed infarct, wrong scar) need a hard human gate before anything reaches the chart.

If you already ship a DICOM viewer, the natural hang points are chamber overlays on cine-SAX, an LGE positive/negative badge, the multiclass disease tag, and an imported report series with provenance that names ORION-CMR outputs versus radiologist edits. Rebuild timing and thresholds on your own vendor mix. Audit report concordance with your readers using the same three-level scale before you advertise hand-off.

Rebuild from arXiv:2609.23950. PDF: https://arxiv.org/pdf/2609.23950.

Sources

  • Demirel, O. B., Horst, K. K., Perazzolo, A., Bruno, E., Kaya, K., Ouyang, R., Ahmed, E., Smink, J., Waddle, S. L., Kallumpurath, Z., Chao, T. C., Wang, D., Langer, S. G., Kline, T. L., Korfiatis, P., Browne, J., Isgum, I., Leiner, T. ORION-CMR: On-scanner Reporting with Integrated Foundation Model for End-to-End Cardiac MRI Analysis and Interpretation. arXiv:2609.23950, 2026. https://arxiv.org/abs/2609.23950. PDF: https://arxiv.org/pdf/2609.23950.
  • Public benchmarks: ACDC (Bernard et al., IEEE TMI 2018); EMIDEC (Lalande et al., Data 2020). Prior CMR foundation model comparison: CMR-FM (Jacob et al., JCMR 2025).
  • Reference ranges: Healthy Hearts Consortium 2024 (Raisi-Estabragh et al., JACC Cardiovasc Imaging). Local LLM stack: Qwen2.5-14B-Instruct via Ollama on-scanner.

We build custom medical imaging platforms — advanced DICOM viewers, AI segmentation, and the clinical systems around them.

Get in Touch

Copyright © 2026 PYCAD. All Rights Reserved.