Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

One End-Diastole Cardiac MRI Frame In, a Full 30-Frame Beating-Heart Cine Out Without ECG Gating by Warping the Real Pixels With a Phase-Conditioned Flow Model

Cardiac MRI cine frames at end-diastole, mid-systole, end-systole and mid-diastole for a hypertrophic cardiomyopathy and a normal example: real cine on top, a blurry VAE-decoder result in the middle, a sharp deformation-decoder result below, and predicted displacement heat maps at the bottom (Wang et al., arXiv:2610.09397, Fig. 4)

Give PhaseFlow one end-diastolic cardiac MRI frame and it generates a 30-frame cine of the heartbeat with no ECG gating. It predicts a motion field and warps the real pixels, which keeps the image sharp, and it takes its timing from a mean phase curve for the patient’s pathology group. Tested on ACDC with 30 test patients and one training run, the ejection fraction error is 17.79 points, above the paper’s own clinical threshold of about 10, and I found no public code. Wang et al., arXiv:2610.09397.

Handheld Endoscope or Microscope Video In, One Wide-Field Tissue Mosaic Out at About Ten Frames per Second From a Modality-Tuned Optical Flow Model

Before and after mosaics of fetoscopy, dermoscopy and light-sheet microscopy video: the top row, stitched with an untuned optical flow model, shows grey patches and broken pieces; the bottom row, from the fine-tuned FloVMos model, is continuous (Liu et al., arXiv:2610.04258, Figs. 3 and 5)

Feed FloVMos video from a handheld endoscope, dermoscope or confocal microscope and it stitches the frames into one growing wide-field mosaic, at about ten frames per second on 1 megapixel frames in the authors’ research code. The trick is fine-tuning an optical flow network per device on synthetic motion, so no hand-labelled flow is needed. Several real test sets are one or two videos, accuracy is a feature distance check, a GPU is required, and the repository linked in the paper returned a 404 when I checked. Liu et al., arXiv:2610.04258.

Undersampled Cardiac MRI In, Sharper Reconstruction Out From a Frozen Pretrained Vision Encoder With About Half the Trainable Weights

Two panels of cardiac MRI slices with absolute error maps below, each showing the ground truth, a UNETR trained from scratch, and a reconstruction from a pretrained DINOv2 encoder, at 4x acceleration in-domain and 10x acceleration across datasets (Hashmi et al., arXiv:2610.08109, Fig. 3)

Hand an undersampled cardiac MRI slice to a UNETR-style network whose encoder is a frozen pretrained CLIP, BiomedCLIP or DINOv2 model, and the reconstruction beats the same network trained from scratch, with about 49 percent fewer trainable parameters. The gap holds with 5 percent of the training data and across two challenge datasets. It is a 2D study on public CMRxRecon data scored with SSIM and PSNR only, with no reader study, and the linked GitHub repository held only a README when I checked. Hashmi et al., arXiv:2610.08109.

Eight X-ray Projections In, a Generated 3D Lung CT Volume Out With No Training but About 22 Minutes of GPU Time

Three rows of lung CT crops, a sagittal slice and two zooms, comparing the ground-truth CT, X2CT, the diffusion prior alone, DAPS, and PhyDiCT reconstructed from eight X-ray projections (Dai et al., arXiv:2610.09253, Fig. 2)

Give it eight simulated X-ray projections around the chest and it steers a frozen, text-conditioned lung CT diffusion model until its own rendered X-rays match, returning a 3D volume with no paired training. It takes about 22 minutes per volume on an RTX 6000, is built for lung CT only, and was tested on X-rays simulated from CT-RATE scans, so the output is a plausible estimate, not a measurement. Apache 2.0 code. Dai et al., arXiv:2610.09253.