Hang a carotid ultrasound frame into a viewer and the grain is not optional. Speckle rides along with the echo, and clean, speckle-free targets are not sitting in your PACS waiting to supervise a denoiser. Classical filters and many blind-spot networks either leave residual grain or smear the fine walls you still need for reading and for downstream segmentation. Xuesong Li, Yingtai Xu, Zhongliang Jiang, Nassir Navab, and Yuan Bi take that mismatch seriously in arXiv:2609.32844, posted 26 September 2026. They call the method Mask2Restore: Self-Supervised Ultrasound Despeckling via Inpainting. The shop-facing idea is plain. Treat despeckling as block-wise contextual inpainting on a single noisy image, add cross-resolution context regularization (CRCR), and emit a despeckled frame that keeps fine anatomy without needing a clean target at train time. This is a paper walkthrough for practitioners, not a product claim.
Training uses a standard U-Net backbone. Block-wise masking samples anchors at ratio r=0.01 with N×N blocks of size 7, fills masked patches from a uniform intensity draw, and at inference drops masking and multi-resolution branches so cost matches a single forward pass.
What hangs on the viewer
Upstream is a noisy envelope ultrasound frame. The paper’s main clinical hang is in vivo carotid US (523 images at 512×512, no clean reference) plus a simulated envelope set with references for metrics. Downstream is a despeckled frame meant to stay structurally faithful enough for viewing and for later analysis. For a DICOM or US viewer rail, the hangable pieces are: a noisy-versus-despeckled scrubber with a fixed ROI zoom (as in their Figure 3 insets), a model card that names block-wise masked inpainting plus CRCR instead of a generic “AI denoise” toggle, and a path that does not demand paired clean US you will never acquire.
The practical claim for a clinic AI shop is not a new probe. Mask2Restore hands you a self-supervised stack that trains on single noisy images, then at inference applies the trained network directly to the original frame. You can still overlay the despeckled output against the input for QC.
How it works in plain words
Blind-spot networks hide individual pixels and ask the model to predict them from neighbors. That works when noise is roughly pixel-independent. Ultrasound speckle is not. Speckle grains cover multiple neighboring pixels because of the point-spread function and coherent scattering, so a masked center pixel can still be guessed from the local grain, and the network learns to copy speckles rather than remove them.
Mask2Restore changes the spatial unit. It samples compact N×N blocks, masks those whole regions, fills them with random intensities, and trains the network to inpaint the masked area from broader anatomical context. That block-wise objective interrupts local speckle replication. Because the visible context can still carry correlated grain, the authors add CRCR: the same masked image is also processed at half and quarter resolution, predictions are upsampled back, and an L1 consistency term pulls the full-resolution output toward those broader-context branches. Reconstruction loss stays on the full-resolution masked regions only (α=1, β=0.1 in their joint loss).
In shop language: hide contiguous speckled patches, force inpainting from surrounding tissue layout, then regularize with multi-resolution predictions so residual grain bias is less likely to win. Inference is a single forward pass on the untouched noisy frame.
What the numbers say
On the simulated set with clean structural references (Table I), Mask2Restore reports PSNR 20.36 and SSIM 0.810. Those sit above Speckle2Self (19.27 / 0.730) and the listed blind-spot and classical baselines on those reference metrics. On an unseen fine-structure simulated test (Figure 4), they report SSIM 0.8241 while red-arrowed small structures stay sharper than Speckle2Self, which tends to over-smooth.
On in vivo carotid data there is no clean target, so the paper leans on qualitative Figure 3 strips and no-reference metrics read jointly with visuals. Their carotid NIQE is 6.54; EPI and gCNR sit in a mixed band versus Speckle2Self, and the authors themselves warn that EPI and gCNR can reward over-smoothing. Downstream, after retraining despecklers on CAMUS, left-atrium segmentation with identical SAM 2 prompts reaches Dice 0.8207 for Ours versus 0.7664 on noisy input (Table IV). Treat panels and metrics as paper evidence, not as a promise on your local vendor protocol.
Where it fails and what not to trust
This is research software on simulated envelopes, a public-style carotid cohort, and a CAMUS segmentation probe. It is not a cleared medical device. Rebuild and validate on your own acquisition depth, frequency, and probe mix before you hang despeckled frames next to reportable reads. Informal clinical feedback in the paper notes that some collaborators still preferred a lighter touch closer to the original speckled look, because sonographers train on speckled images every day.
Do not treat SSIM 0.810 or Dice 0.8207 as a guarantee on a different scanner stack. Do not skip a local holdout with visual review of vessel wall, plaque edges, and small hypoechoic pockets against the noisy input. And do not hide that no-reference scores alone can favor blurry outputs; the paper’s own Figure 3 and Figure 4 are the check you want next to any number.
For a viewer or clinic AI shop
Wire speckled ultrasound in; hang an anatomy-keeping despeckled frame out. Prefer Mask2Restore when you want self-supervised US despeckling without clean targets: block-wise masked inpainting to break multi-pixel grain copy, CRCR for broader-context consistency, U-Net forward pass at inference. Keep a QC scrubber so support can compare despeckled slices to the noisy input and to any downstream mask. Log dataset (carotid vs cardiac vs other), masking hyperparameters (r, N), and the exact checkpoint.
Validate on at least one internal cohort before you claim despeckle hang in production. Start from the preprint. Rebuild from arXiv:2609.32844. PDF: https://arxiv.org/pdf/2609.32844.
Sources
- Li, X., Xu, Y., Jiang, Z., Navab, N., Bi, Y. Mask2Restore: Self-Supervised Ultrasound Despeckling via Inpainting. arXiv:2609.32844, 2026. https://arxiv.org/abs/2609.32844. PDF: https://arxiv.org/pdf/2609.32844.
- Table I simulated (reference): Ours PSNR 20.36, SSIM 0.810; Speckle2Self 19.27 / 0.730.
- Fig. 4 unseen fine-structure SSIM 0.8241 for Ours; Fig. 3 carotid qualitative noisy vs Ours vs baselines.
- Table IV CAMUS left atrium after despeckling (SAM 2): Ours Dice 0.8207, IoU 0.7002; noisy baseline Dice 0.7664.
- Carotid US 523 images 512×512; simulated 538 with references; masking r=0.01, N=7; CRCR scales {1, 0.5, 0.25}.