Anyone who has scrolled through a postoperative pelvic CT with a hip or fixation implant knows the problem. Bright streaks and dark bands fan out from the metal and cover the exact bone and soft tissue the surgeon asked about. Scanner vendors ship metal artifact reduction (MAR), but the strongest methods need raw projection data at reconstruction time. Once a study lands in PACS, a research archive, or a multi-site dataset, the raw projections are usually gone and only the reconstructed DICOM series is left.
Amritesh Banerjee, Abdul Basit, Renil Renji Joseph, Nouhaila Innan, and Muhammad Shafique (UMass Amherst, NYU Abu Dhabi, IIT Delhi Abu Dhabi) target that situation in arXiv:2610.01512, posted 1 October 2026 and accepted at BHI 2026. VoxelSynth3D is a training-free, deterministic 3D correction that works on reconstructed CT in Hounsfield units. It leaves the detected implant voxels untouched, only edits tissue inside an implant-adaptive support region, and exposes every mask and weight it uses. The authors also build Synthetic CLINIC-Metal, a paired evaluation set made by inserting simulated implants into metal-free pelvic CT from CTPelvic1K, so they can score corrections against a known clean scan.
What goes in and what comes out on the viewer
The input is one reconstructed CT volume in HU plus its voxel spacing. The method needs no sinogram, vendor reconstruction kernel, or trained weights. The output is a corrected volume on the same grid, so a viewer can hang it as a derived series next to the original and let the reader flip between them.
Three properties make it easy to reason about in a clinical viewer. Every voxel the method classifies as metal comes back with its original value. Voxels outside the correction support come back unchanged, which keeps distant anatomy identical to the source series. The intermediate products (metal mask, support region, tissue prior, blend weights) are all inspectable volumes, so a developer could render them as overlays when a reader asks why a region changed.
How it works in plain words
The pipeline has four fixed stages, shown in the paper’s Fig. 1 on a pelvic CT slice.
First, it finds the metal. On real scans the seed is an adaptive HU threshold (the 99.5th percentile of voxels above 400 HU, clipped to 1000 to 1500 HU), cleaned with 3D connected-component rules. From the physical metal volume it computes an effective implant radius and scales the correction zone to it. A wider zone (8 to 22 mm) catches dark bands and streaks, and a narrower one (4 to 10 mm) limits where edge recovery is allowed. A small screw and a large hip stem therefore get different footprints.
The second stage estimates what the tissue should look like. A normalized 3D Gaussian average over valid soft tissue (between minus 300 and 500 HU, metal excluded) gives a smooth prior that is not dragged up by the 3071 HU metal values. The smoothing scale is set in millimeters, so anisotropic spacing is handled.
Next, the prior is blended in only where a voxel looks corrupted. A deviation gate compares each voxel with expected tissue statistics, and the blend weight fades with distance from the implant and drops near strong real edges, based on a Sobel gradient map.
The last stage puts some sharpness back. Smoothing softens bone and tissue boundaries, so a guided local linear model reinjects high-frequency structure from the original input, but only for high-gradient voxels inside the support. The final step pastes the original metal voxels back.
Every parameter is fixed. The authors picked the edge refinement settings from a 27-setting validation sweep and chose a conservative point that sits 0.10 HU from the numerical minimum on validation.
What the numbers say
Synthetic CLINIC-Metal uses patient-level splits (124 development, 42 validation, 40 test) with one to three simulated ellipsoidal implants per case and a projection-based artifact simulation. The primary score is RMSE in a fixed tissue region around the implant with the metal itself excluded, because metal is only 0.218% of that region yet would otherwise account for a large share of any error reduction.
On the 40 held-out cases, with the exact inserted-metal mask supplied, VoxelSynth3D lowered tissue RMSE from 801.48 HU (uncorrected input) to 786.18 HU. The paired gain of 15.30 HU is small in absolute terms, but it appeared in all 40 cases and beat a 3D Gaussian smoother by 13.58 HU. Public checkpoints of DICDNet, InDuDoNet, and InDuDoNet+, applied without retraining, did worse than the input on this RMSE score, although InDuDoNet had the best SSIM. The authors label those learned rows as transfer tests and say they do not represent in-domain retrained performance.
Runtime is about 53 seconds per volume on CPU in their implementation, with metric bookkeeping included. A lighter variant without edge refinement (VoxelSynth3D-A) runs in about 33 seconds with nearly the same RMSE.
Where it fails and what to watch
The paper measures its own limits, and three of them matter for a viewer team.
Edges suffer right next to the metal. In the 2 to 5 mm band, clean-edge F1 fell from 0.548 for the input to 0.352 after correction, so many true boundaries closest to the implant were lost. Beyond 5 mm, edge agreement was better than the input. The paper frames the result as artifact suppression with a localized structural cost, and says it does not show anatomical recovery. That band is exactly where a radiologist looks for loosening or a fracture line at the bone and implant interface.
The headline numbers assume a perfect metal mask. With the automatic HU threshold that a real deployment would use, localization precision was 0.285 because dense bone also crosses the threshold. Tissue RMSE inside the scored region improved further, but 1.04% of voxels outside it changed by more than 50 HU, compared with 0.082% under the exact mask. A wider correction zone can look better on the target metric while altering normal tissue, and the authors list accurate localization as an open deployment problem.
Evidence on real scans is qualitative. The 75 real implant volumes in the CLINIC-Metal subset have no clean reference, so Fig. 7 shows visual comparisons only. On the external AAPM CT-MAR challenge images (2D slices replicated into a five-slice stack), only 12 of 29 images improved and the change was not significant (p = 0.381). In the Fig. 7 real-case panels the implant itself also looks less bright in the VoxelSynth3D column than in the input, so check how the metal pass-through behaves on your own series before trusting that region. A spacing-aware artifact model kept the broad-region gain on average but exposed weaker behavior near the metal. All synthetic data is pelvis, so other anatomies and implant shapes are untested, and no blinded radiologist reading has been done yet.
What this means for a DICOM viewer or clinic integration
Because VoxelSynth3D needs only the reconstructed series, it can run as a post-processing step on studies pulled from PACS, with no vendor integration. The output should be stored as a separate derived series with its own SeriesDescription and a clear “processed” flag, never written over the original, because the paper itself shows it removes true edges within a few millimeters of the implant.
The design also suits audit. Since the method is deterministic and the support and weight maps are explicit, a viewer can show the reader exactly which voxels changed and by how much, for example as a semi-transparent difference overlay. That is easier to defend in a quality review than an opaque network output, and the same input always gives the same output.
Two practical points come before wiring it in. Plan a better metal segmentation step than a plain HU threshold, or at least restrict the threshold to a user-confirmed region around the implant. Treat the corrected series as a reading aid for soft tissue a few centimeters from the hardware, and send readers back to the original series and vendor MAR reconstructions for the bone and implant interface itself.
Code and data availability
The preprint does not give a public code repository or a download link for Synthetic CLINIC-Metal. The source volumes come from the public CTPelvic1K pelvic CT collection, and the paper documents the default parameters (Table II) and the artifact simulation settings in enough detail to attempt a reimplementation. Check the arXiv page for later updates from the authors.
Sources
- Banerjee, A., Basit, A., Joseph, R. R., Innan, N., Shafique, M. VoxelSynth3D: Interpretable Volumetric Image-Domain Metal Artifact Reduction with a Paired Synthetic CLINIC-Metal Benchmark. arXiv:2610.01512, 2026. Accepted at BHI 2026. https://arxiv.org/abs/2610.01512. PDF: https://arxiv.org/pdf/2610.01512.
- Liu, P., et al. Deep learning to segment pelvic bones: Large-scale CT datasets and baseline models (CTPelvic1K). arXiv:2012.08721, 2020. https://arxiv.org/abs/2012.08721.
- Featured image: Fig. 7 of arXiv:2610.01512, real CLINIC-Metal case M0052, uncorrected input versus VoxelSynth3D output (window minus 900 to 1300 HU).