Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

Undersampled Cardiac MRI In, Sharper Reconstruction Out From a Frozen Pretrained Vision Encoder With About Half the Trainable Weights

Two panels of cardiac MRI slices with absolute error maps below, each showing the ground truth, a UNETR trained from scratch, and a reconstruction from a pretrained DINOv2 encoder, at 4x acceleration in-domain and 10x acceleration across datasets (Hashmi et al., arXiv:2610.08109, Fig. 3)

Faster cardiac MRI means measuring less of k-space, and the missing data has to be filled in by a reconstruction model. Most of those models are trained from scratch on a particular scanner and protocol, and they tend to degrade when either changes. A group at Dublin City University and University College Dublin (Anam Hashmi, Mayug Maniparambil, Julia Dietlmeier, Kathleen Curran, and Noel O’Connor) posted arXiv:2610.08109 on 6 October 2026 asking whether a vision model pretrained on large collections of ordinary images is a better starting point than a transformer trained on MRI alone. In their tests it is.

What goes in and what comes out

The input is a single 2D cardiac MRI slice reconstructed from undersampled multi-coil k-space. The authors take an inverse FFT of each coil, combine the coils with root-sum-of-squares, normalise the magnitude image to the range 0 to 1, and center-crop or pad it to 224 by 224 pixels. The output is a cleaned 224 by 224 image meant to match the fully sampled reference. They test 4x, 8x, and 10x acceleration.

The featured image is Fig. 3 of the paper, with every tile taken from the paper without retouching. Each panel shows the ground truth, a UNETR trained from scratch, and a reconstruction from a pretrained encoder, with absolute error maps underneath. The left panel is in-domain at 4x acceleration on CMRxRecon2024. The right panel trains on CMRxRecon2024, tests on CMRxRecon2023, and uses 10x acceleration.

How it works

The baseline is a 2D UNETR: a six-layer vision transformer encoder with patch size 16, which gives 196 patches of 768 dimensions, plus a U-Net-style decoder with skip connections. The authors keep the decoder and swap the encoder for a ViT-B from CLIP, BiomedCLIP, or DINOv2. The grayscale slice is copied to three channels to fit the pretrained input. Features from all transformer layers are fused with a layer norm and a two-layer MLP to form the bottleneck, and intermediate layers feed the skip connections. DINOv2 uses patch size 14, so its 16 by 16 token grid is bilinearly interpolated to fit the decoder.

They compare three setups: the scratch UNETR, the frozen pretrained encoder (only the fusion layers and decoder train), and the pretrained encoder with LoRA adapters in the attention projections, with ranks 8, 64, and 128 tried. Training is SSIM loss with AdamW for up to 300 epochs, batch size 8, on one RTX 4090.

Data and evaluation

CMRxRecon2023 has 120 fully sampled training cases with cine, T1 mapping, and T2 mapping in several views, split at the patient level 70/10/20 into 2,940, 420, and 852 samples. CMRxRecon2024 has 200 healthy volunteers on a 3T Siemens Vida scanner, with cine, aorta, T1 and T2 mapping, and tagging, split the same way into 12,537, 1,803, and 3,597 samples. Quality is SSIM, PSNR, and NMSE.

What the numbers say

Every frozen encoder beat the scratch UNETR at every acceleration on both datasets, with about 49 percent fewer trainable parameters (33.7M against 65.7M). At 4x on CMRxRecon2023, frozen DINOv2 reaches SSIM 0.9073 and 32.21 dB, against 0.8904 and 31.14 dB for UNETR. At 10x on CMRxRecon2024 the pair is 0.8514 and 29.16 dB against 0.8343 and 28.56 dB. The paper says the gains grow with acceleration. In Table 1 the SSIM gap for DINOv2 on CMRxRecon2023 is 0.0169 at 4x and 0.0148 at 10x, so I did not see that growth.

BiomedCLIP, the encoder trained on biomedical image-text pairs, did not come out on top. DINOv2 was best overall, CLIP was close, and BiomedCLIP trailed CLIP in the cross-dataset test (SSIM 0.8360 against 0.8510, with UNETR at 0.8293).

At 10x acceleration with 5 percent of the training data, frozen DINOv2 gets SSIM 0.8348 on CMRxRecon2023 and 0.8175 on CMRxRecon2024, against 0.7900 and 0.7803 for UNETR. LoRA does not help there: it lowers DINOv2 to 0.8242 and 0.8094. With more data it pays off on CMRxRecon2024, where LoRA DINOv2 reaches SSIM 0.8550 and 29.58 dB with 75 percent of the data, against 0.8503 and 29.32 dB frozen. On CMRxRecon2023 LoRA usually costs a little SSIM and gains PSNR.

When trained on CMRxRecon2024 and tested on CMRxRecon2023 at 10x, LoRA DINOv2 reaches SSIM 0.8634 and 29.60 dB, against 0.8293 and 27.99 dB for UNETR. The reverse direction is harder for every model: UNETR falls to 0.7596 and 25.89 dB, frozen DINOv2 gets 0.8016 and 27.21 dB.

Speed

The paper does not report inference time or memory. It reports only training on one RTX 4090 with mixed precision. The trainable parameter count is lower than the baseline’s, but the frozen encoder still has to run on every slice, so I would not assume it is faster to run. We would measure it before saying anything about scanner-side use.

What the pictures show

The authors say that in Fig. 3 the foundation models give sharper reconstructions and lower errors than UNETR, that DINOv2 comes closest to the ground truth, and that UNETR tends to oversmooth and lose fine detail. At the size the figure is printed, the images look close to me, and the difference is easier to see in the error maps: the UNETR maps are brighter around the heart and along the body outline, and the DINOv2 maps are darker. I would not read a diagnostic difference from these two slices.

Where it falls short

The authors state the main limit themselves: they evaluate with average SSIM and PSNR and no radiological assessment, so the gains should not be read as clinical superiority. Beyond that, the model works on 2D magnitude slices after coil combination, and I did not see a data-consistency step with the measured k-space, so the output is not guaranteed to agree with what the scanner measured. Everything is center-cropped or padded to 224 by 224, which a real pipeline has to handle with the original matrix size and pixel spacing. CMRxRecon2024 is healthy volunteers on one scanner model, and the paper does not describe disease cases. The tables show single numbers with no standard deviations that I could find, and an SSIM gap of about 0.015 comes without a significance test. The cross-dataset test moves between two datasets from the same challenge series, not between vendors or field strengths.

Code, weights, and data

The paper links github.com/Hashmi360/BeyondScratch-CMR, and the README says the work was accepted at ECCVW 2026. When I checked on 10 October 2026, the repository held a README and nothing else: two commits from 14 August 2026, no code, no weights, no license. The encoders themselves come from public CLIP, BiomedCLIP, and DINOv2 ViT-B checkpoints, and the paper describes the design in enough detail that a team could rebuild it, but I would not expect a rebuild to match the tables without the authors’ code. The data comes from the CMRxRecon challenges, and I did not check their access terms.

How we would wire it into a viewer

We would not put this on a clinical read path. For a reconstruction research view, we would store the network output as a derived series flagged as AI-reconstructed, next to the vendor reconstruction, and add a difference toggle. On real accelerated scans there is no fully sampled reference, so that map would show disagreement with the vendor image and not error, and the label would say so.

Before using it on our own data we would test four things: the model on our scanner’s undersampled data against the vendor reconstruction at the accelerations we run, disease cases with small structures such as scar and thrombus where oversmoothing would hurt, how the 224 by 224 crop interacts with our matrix sizes and orientation, and left ventricular volumes and ejection fraction measured on its output against the fully sampled images.

Sources

We build custom medical imaging platforms — advanced DICOM viewers, AI segmentation, and the clinical systems around them.

Get in Touch

Copyright © 2026 PYCAD. All Rights Reserved.