Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

One Click on One CT Slice In, a Full 3D Organ Mask Out That Is Carried Slice by Slice From the Neighbours and Still Takes Class Prompts, Reference Images and Corrections in the Same Model

Pancreas CT slices after one click: ground truth, three baseline propagation models and UniPro at five slice positions, with liver, kidney and brain tumor success and failure examples on the right (Guo et al., arXiv:2610.06938, Fig. 5 and Fig. 6)

A segmentation tool inside a DICOM viewer usually goes wrong in one of two ways. The automatic model gets the organ roughly right and a reader has to fix many slices by hand, or the click-based tool needs a click on every slice. A team at Rutgers University, Stanford University, the University of Texas at Arlington and New York University (Bangwei Guo, Yunhe Gao, Meng Ye, Yang Zhou, Difei Gu, Guoning Zhang, Leon Axel and Dimitris Metaxas) posted arXiv:2610.06938 on 3 October 2026 with UniPro. You click once on one slice, or hand it a class name or a reference example, and it carries the mask through the rest of the volume. One 2D model handles all of those entry points. The code repository exists, but its README still says “To be released”.

What goes in and what comes out

The input is a CT or MR volume plus one starting signal on one slice. That signal can be a click, a class prior (“pancreas”), or a reference image with its mask from another case. The output is a 3D mask for the target. The same model also works on plain 2D images such as chest X-ray, dermoscopy, ultrasound and endoscopy images, where there is no volume to propagate through.

The featured image is built from Fig. 5 and Fig. 6 of the paper, with every tile taken as published and only cropped and resized. On the left, the same single click on a pancreas CT is given to three other propagation models and to UniPro. On the right are the authors’ selected success and failure cases on liver, kidney and brain tumors.

How it works

UniPro is a slice-based network with a shared image encoder and a transformer decoder whose feed-forward layers are a mixture of experts. The four ways of starting a segmentation differ only in what is fed to that decoder: class priors, reference image and mask pairs, user clicks, or neighbouring-slice predictions. Each is turned into dense features and sparse queries for the same decoder.

The central step treats propagation as in-context segmentation inside one volume. To predict slice q, the model gets a reference window of five image and mask pairs: the seed slice, which is always kept as an anchor, plus the most recent predictions. It then runs forward and backward from the seed slice, which in the paper’s tests is the median axial slice of the target. Keeping the seed in the window the whole way is meant to stop the mask drifting as it moves away from the click.

Training has two stages. Stage 1 trains semantic, in-context and interactive segmentation together for 100 epochs on 8 Quadro RTX 8000 GPUs. Stage 2 runs 10 epochs of propagation training with two extra losses: a backward pass that has to return to the seed mask, and a 3D consistency loss on the stacked slice predictions. Stage 2 batches are mixed with ordinary 2D batches so the model keeps its 2D skills. The training data is nine public datasets, and four more (BTCV, ACDC, UW-SC and BUS) plus BraTS brain tumors were held out for testing.

What the numbers say

The main test starts every method from one positive click on the seed slice and scores the whole volume. In that setting UniPro is best on 7 of 8 datasets. On abdominal CT (AbdomenCT-1K) it scores 88.62 against 86.29 for MedSAM2, the strongest baseline, and on M&Ms cardiac MRI it scores 87.64 against 82.97. MedSAM2 wins on CHAOS (90.64 against 89.70). On BraTS brain tumors, which the model never saw in training, UniPro reaches 54.89 against 40.74 for MedSAM2, and 51.25 for the 3D interactive model nnInteractive. SAM2 without medical training is far behind on most datasets, with 45.00 on AbdomenCT-1K.

The paper also simulates an editing session. Starting from an empty mask, the simulated user clicks until the 3D overlap reaches a target. On ACDC, editing slice by slice takes 27.7 clicks and 8.5 edited slices, and re-propagating from each corrected slice takes 25.2 clicks and 3.3 edited slices. On BraTS the numbers are 92.0 clicks and 62.4 slices against 57.0 clicks and 12.1 slices. These are simulated clicks, not readers.

In the pancreas example in the featured image, the 3D Dice (volume overlap) is 0.834 for UniPro against 0.665 for MedSAM2, 0.541 for PAM, 0.503 for iSegFormer, 0.445 for nnInteractive and 0.410 for SAM2.

What the pictures show

In the pancreas rows, the baselines fail in different ways. MedSAM2 covers only part of the organ, PAM’s mask balloons at the far slice (z=48) into the tissue around the pancreas, and nnInteractive loses the pancreas on slices far from the click. UniPro stays on the organ across the whole range from z=22 to z=48, though its mask at z=43 includes a small extra island next to the main blob. This is one case picked by the authors. All baselines except SAM2 were retrained on the same data.

The tumor panels show where it breaks. In the kidney and liver successes the mask follows the main tumor across a range of slices. The two failure rows are small, separate lesions: the model keeps part of the seed region and loses the lesions that sit apart from it, and the brain case shown as a failure in the paper misses the fragmented parts of a heterogeneous tumor. Even the success rows have local errors, including missed small foci and false extensions, as the authors say in their appendix.

Where it falls short

The overlap scores are computed over the slices that contain the target, so they are higher than a full-volume score. The paper reports the difference: it ranges from 0.17 points on ACDC to 3.90 on KiTS. The confidence intervals are wide on the hard datasets. LiTS has a mean of 66.95 with a standard deviation of 28.60, and paired tests against MedSAM2 are significant on five of the eight datasets at the unadjusted 0.05 level, and not on CHAOS, LiTS or KiTS. Every number comes from a single training run per method.

In plain semantic mode, where there is no user input, UniPro trails the fully supervised 3D nnU-Net on most datasets. Propagation training also costs some 2D skill: on ISIC skin lesions the Dice drops from 87.54 to 82.40 when stage 2 uses propagation batches only, and to 86.11 with the 2D batches mixed in, still 1.43 points under the stage-1 model. Inference is slower than a 3D nnU-Net on AbdomenCT-1K: 28.0 seconds per volume against 9.3 on one Quadro RTX 8000, with 1.93 GB of peak GPU memory against 3.60. The model does not take slice spacing or physical distance as input, and the authors say large gaps between slices weaken the match between neighbouring slices. A single click can also miss disconnected regions. No clinician study is reported, and the intestine is listed as untested.

Code, weights, and data

The paper links github.com/bangwayne/UniPro. On 11 October 2026 that repository had two commits from 3 October, a README that says “To be released …”, and no code or weights. A GitHub repository search for UniPro with segmentation and propagation returned nothing else, and a Hugging Face search showed only unrelated UniProt protein models. The datasets are public: the paper trains on nine of them and tests on BTCV, ACDC, UW-SC, BUS and BraTS. The baselines it retrains are nnInteractive, PAM, MedSAM2 and iSegFormer, plus SAM2 as released.

How we would wire it into a viewer

We would keep the clinician’s seed slice as the anchor and store it as its own object: the click or prompt, the seed mask, and the model version. The propagated masks go into a DICOM SEG object marked as machine-generated, with the seed slice flagged. When a reader fixes a slice, the viewer re-runs propagation from the corrected slices as new anchors, which matches how the paper’s editing simulation works. We would show the slices far from the seed with a visible warning rather than flat full-confidence overlays, since that is where the baselines and UniPro both fail.

Before trying it on our own studies we would test four things: thick-slice and uneven-spacing CT, which the paper does not model, multi-lesion cases where the second lesion is far from the first click, how long a reader needs to fix a bad propagation compared with a fresh start, and whether 28 seconds per volume fits a reading workflow. The code is not out, so for now this is a paper to track. For earlier work on one-slice-to-volume editing on this site, see our post on scribbling on one slice and propagating adapted volume masks.

Sources

  • Guo, B., Gao, Y., Ye, M., Zhou, Y., Gu, D., Zhang, G., Axel, L., Metaxas, D. N. UniPro: Unified Multi-Mode Medical Image Segmentation from 2D Images to 3D Volumes via Propagation. arXiv:2610.06938, 2026. https://arxiv.org/abs/2610.06938. PDF: https://arxiv.org/pdf/2610.06938.
  • Project repository named in the paper (README only, checked 11 October 2026): https://github.com/bangwayne/UniPro.
  • Featured image: Fig. 5 (pancreas, single-click propagation) and Fig. 6 (success and failure cases) of arXiv:2610.06938, as published.

We build custom medical imaging platforms — advanced DICOM viewers, AI segmentation, and the clinical systems around them.

Get in Touch

Copyright © 2026 PYCAD. All Rights Reserved.