Hang a whole fetal ultrasound study into a DICOM viewer or prenatal AI queue and the shop question is not whether you have a four-chamber still. It is whether the case gets a study-level CHD versus healthy screen when nobody pre-picked the cardiac frames, plus a short stack of the frames that drove the call for a reader to review. Mohamed Azzam, Ruobing Liu, Esther C. Ugwueke, Ziyang Xu, Shibiao Wan, Alex Foy, Abraham Zabih, Jason Christensen, Neil Hamill, Ling Li, and Jieqiong Wang (University of Nebraska Medical Center) take that hang seriously in arXiv:2609.31376, posted 25 September 2026. They call the setting whole-study cardiac-gated multiple instance learning (MIL) on their multi-source FUSE cohort: a frozen MAE frame encoder, a disease-aware cardiac-frame gate, and a transformer MIL aggregator trained from study-level labels alone. On the internal FUSE test set of 177 studies the gated model reaches AUC 0.985 with specificity 0.990, ahead of reproduced NATMED (0.861 / 0.600) and FetalCLIP (0.867 / 0.710). On external CARDIUM, label-free CORAL adaptation lifts the same model from AUC 0.513 to 0.944.
What hangs on the viewer
Upstream is a full B-mode fetal ultrasound study: hundreds of frames (median 298 in FUSE), no guarantee that standard cardiac planes were saved, and plenty of abdominal and non-cardiac anatomy in the bag. Downstream hangables for a viewer or clinic AI shop include: a binary case-level CHD versus healthy flag; an optional critical versus non-critical severity grade among predicted CHD cases; and the highest-scoring frames ranked by last-layer class-token attention so a clinician can open the evidence instead of hunting the sweep.
The shop claim is not another plane classifier bolted onto curated stills. Prior pipelines often assume a clinician or a standard-plane detector already isolated the cardiac frames. Plane detectors trained mostly on normal anatomy can drop the abnormal cardiac frames that matter most for CHD. This paper removes that assumption and scores the whole study.
How it works in plain words
Stage 1 pre-trains a ViT-L/16 masked autoencoder from scratch on unlabeled fetal ultrasound (369,846 masked B-mode frames from 1,650 fetuses), then freezes it. Each frame becomes a 1024-d embedding from the final class token. An MLP cardiac-frame identifier, trained with focal loss on the view-annotated subset, gates which frames enter the bag. The gate is meant to stay usable on abnormal hearts rather than only on textbook standard planes.
Stage 2 is transformer MIL on the gated bag. A permutation-invariant aggregator with self-attention produces one subject-level prediction from study-level labels only. No frame-level CHD labels are required. The classification token’s attention mass over frames ranks the supporting evidence returned for review. A hierarchical severity head first detects CHD, then separates critical CHD (needs neonatal catheter or surgical intervention) from non-critical CHD among predicted positives.
Cross-site transfer to CARDIUM uses Deep CORAL: a label-free second-order feature alignment that pairs FUSE source bags with unlabeled CARDIUM bags. Adaptation is restricted to identified cardiac frames. Without the gate, the same MIL setup sits near chance on CARDIUM even before adaptation.
What the numbers say
FUSE mixes an in-house multi-institution Nebraska cohort with public fetal ultrasound sets (FETAL PLANES DB, MFUSPAC, RFCHD): 869 subjects total, split 692 train / 177 internal test (105 healthy, 72 CHD). CARDIUM is an independent external set of 790 subjects with about 2.9% CHD prevalence and curated cardiac frames rather than full sweeps.
Table III (internal detection): proposed cardiac-gated MIL AUC 0.985, sensitivity 0.972, specificity 0.990; NATMED 0.861 / 0.906 / 0.600; FetalCLIP 0.867 / 0.828 / 0.710. Table IV (CARDIUM AUC): gated MIL 0.513 without adaptation, 0.944 after CORAL; NATMED 0.506 to 0.557; FetalCLIP 0.571 to 0.616. Ablation Table V: whole-study MIL alone is already 0.958 internal / 0.18 CARDIUM; adding cardiac gating reaches 0.985 / 0.513; gating plus CORAL reaches 0.985 / 0.944.
Hierarchical severity on the internal test set (Fig. 6): versus a flat three-class head, the two-stage head raises non-critical sensitivity from 0.188 to 0.688 and critical sensitivity from 0.846 to 0.923 at specificity 0.968, with CHD-subclass balanced accuracy 0.805 and AUC 0.841. Fig. 7 shows the five highest-attention frames for two healthy and two CHD subjects; top-ranked frames are cardiac views at diagnostic quality. Fig. 8: attention mass on cardiac frames exceeds their share of the bag across the 177 test studies.
The cardiac-frame identifier itself beats NATMED’s view classifier and SonoNet on view-annotated internal studies (cardiac-view AUC 0.925 on healthy, 0.838 on CHD).
Where it fails and what not to trust
This is research software on retrospective, de-identified imaging under an IRB protocol. It is not a cleared medical device. CARDIUM transfer only recovers after label-free adaptation, and without cardiac gating the external AUC collapses. CHD prevalence differs sharply between FUSE training (~41%) and CARDIUM (~3%), so operating points need re-fitting per site. Severity labels come from confirmed postnatal diagnosis by clinical collaborators; frame-level CHD annotations are not used and mostly do not exist. The cardiac gate’s recall drops on CHD frames versus healthy ones (0.584 vs 0.757), so abnormal anatomy remains harder. Public code is not linked in the preprint we read. Validate on your scanners, sweep lengths, and local prevalence before you hang a case-level flag next to the study in a clinical viewer.
For a viewer or clinic AI shop
Wire a whole fetal ultrasound study in; hang a case-level CHD versus healthy screen, optional critical versus non-critical severity, and the top attention-ranked frames out for review. Prefer this pattern when your queue already receives uncurated studies rather than curated four-chamber stills, when plane classifiers discard the abnormal cardiac frames you care about, and when you can run a frozen encoder plus a light MIL head (they note one MIL run takes minutes on a single GPU once features are extracted). Keep a QC path that shows the gated bag size, the attention-ranked stills beside the study timeline, and a site-specific threshold after any CORAL-style adaptation. Start from the preprint. Rebuild from arXiv:2609.31376. PDF: https://arxiv.org/pdf/2609.31376.
Sources
- Azzam, M., Liu, R., Ugwueke, E.C., Xu, Z., Wan, S., Foy, A., Zabih, A., Christensen, J., Hamill, N., Li, L., Wang, J. Towards Whole-Study Screening for Congenital Heart Disease in Fetal Ultrasound Using Multiple Instance Learning. arXiv:2609.31376, 2026. https://arxiv.org/abs/2609.31376. PDF: https://arxiv.org/pdf/2609.31376.
- Internal FUSE test (177): gated MIL AUC 0.985 / specificity 0.990 vs NATMED 0.861 / 0.600 and FetalCLIP 0.867 / 0.710 (Table III).
- External CARDIUM: gated MIL AUC 0.513 without adaptation, 0.944 after label-free CORAL (Table IV); ablation Table V isolates gating and CORAL.
- Fig. 7: top-5 attention frames for HC and CHD subjects. Fig. 8: attention mass by cardiac vs non-cardiac view group.