Open a 3D CT or MR volume in a viewer and the expensive part is rarely the hang. It is painting a usable mask on a branching vessel tree, a tangled NF1 tumor burden, or some other long-tailed structure that no task-specific U-Net was trained for. Point and box interactive models promise a way out: a few clicks, then a full-volume overlay you can correct. Gong, Su, Zhang, Han, Sun, Li, and Yu from Deepwise Healthcare (with Yu also at The University of Hong Kong) take that claim seriously in arXiv:2609.25743, posted 22 September 2026. They call the model SAMI3D-DW: Interactive Segmentation of Any 3D Medical Images. The shop-facing idea is plain. Drop one point or a 2D box on a CT/MR volume, optionally add a few corrective clicks, and get a full-volume mask across a very wide taxonomy of anatomy and pathology. This is a paper walkthrough for practitioners, not a product claim. Architecture and weights are not released in the report.
Training used Deepwise proprietary data: 849 in-house datasets screened down to about 115k 3D volumes, trained from scratch without pretrained weights, then fine-tuned on harder lesion subsets. Evaluation is the part the clinic AI shop can actually compare.
What hangs on the viewer
Upstream is a CT or MR volume. Downstream is a full-volume mask from spatial prompts. Two interaction modes are reported. Point-only uses one to five cumulative foreground/background points. Box-initialized starts with one 2D bounding box on the axial slice with the largest target area, then up to five corrective points. For a DICOM viewer rail, the hangable pieces are: a prompt tool that accepts points and a 2D box, a previous-mask feedback loop so each correction edits the volume rather than one slice, and a model card that names simulated-interaction Dice on a category-balanced taxonomy instead of a single public dataset leaderboard.
The practical claim for a clinic AI shop is not a new scanner. SAMI3D-DW is aimed at targets that task-specific networks miss, including intracranial vessel trees on CT and MR angiography, with a few clicks. In a preliminary in-house NF1 comparison, assisted tumor annotation took minutes per case and about one-fifteenth of the time required for full manual annotation. Treat that as an early workflow hint, not a timed multi-site study.
How it works in plain words
The report keeps the network architecture private. What it discloses is the interaction contract and the evaluation design. Prompts are spatial only: points, or a 2D box plus corrective points. The evaluation simulator places prompts from reference masks using an error-driven rule (largest connected error component, then an interior point via distance transform), so scores describe a deterministic annotator rather than real-user timing variability.
Labels across 219 source datasets are mapped through a medical taxonomy: dataset label to specific target to reporting category. The scored benchmark has 4,326 cases, 13,826 instances, 405 specific targets, and 107 active categories (63 anatomical, 43 pathological, plus a small other bucket). Scoring uses category-macro Dice: average valid object Dice inside each category, then weight categories equally. That keeps rare vascular trees and long-tailed pathology visible instead of letting liver and kidney dominate the mean.
In shop language: one spatial prompt language for many CT/MR targets, measured with equal category weight so the hard tree-like and neoplastic cases still move the score.
What the numbers say
On point-only interaction (Table 3), SAMI3D-DW leads the reported pack. With one point it reaches category-macro Dice 0.5764 versus 0.5315 for nnInteractive, the strongest baseline. With five total points the scores are 0.7771 versus 0.7494. MedSAM2, SAM-Med3D, SegVol, and VISTA3D sit lower on the same axes (first-point band roughly 0.28 to 0.37).
With box initialization (Table 4), SAMI3D-DW scores 0.7130 versus 0.6530 for nnInteractive at the initial box, then 0.8002 versus 0.7868 after five corrective clicks. The report notes it is the only evaluated box-compatible model to clear 0.80 under that protocol. Gains are largest before correction and narrow as clicks accumulate.
Category-level tables matter for viewer shops. Head vascular trees reach 0.8228 versus 0.4455 after box plus five corrections; neck vascular trees reach 0.9576 versus 0.8805. Whole-body NF1-related tumor burden (381 objects from 31 cases) scores 0.6441 at the box and 0.7351 after five corrections, versus 0.5512 and 0.7077 for nnInteractive. Some categories favor the baseline (jaw bones and some CNS demyelinating lesions), and several vascular wins rest on small object counts. Treat panels and metrics as paper evidence, not as a promise on your local protocol.
Where it fails and what not to trust
This is a Deepwise technical report on proprietary training data. Architecture, code, and weights are not public in the preprint. The benchmark is CT/MR only; PET and 3D ultrasound are future work. Training and test cases are disjoint, but some source datasets contribute to both splits, so transfer to unseen institutions is untested. Simulated prompts do not capture human click noise or real annotation time. The NF1 timing note is preliminary and task-specific.
Do not treat 0.8002 category-macro Dice as a guarantee on your vessel protocol or NF1 cohort. Do not skip human review of predicted masks. Do not hide that VISTA3D and other baselines have different valid-score coverage and, in VISTA3D’s case, can use target-type cues absent from the spatial-only protocols. And do not ship this as a cleared medical device; it is research software until you rebuild, validate, and clear it under your own quality system.
For a viewer or clinic AI shop
Wire few clicks (points or a 2D box) on a 3D CT/MR volume in; hang complex anatomy and pathology overlays out. Prefer this class of interactive 3D model when the target set is long-tailed and task-specific networks leave vascular trees, NF1 burden, or odd neoplastic shapes for the reader to paint by hand. Keep a correction loop, log prompt type and click count, and keep a QC view that compares the mask to the volume across MPR.
Validate on at least one internal CT and one MR cohort before you claim interactive hang in production. Start from the preprint. Rebuild from arXiv:2609.25743. PDF: https://arxiv.org/pdf/2609.25743. Public code was not linked in the report at the time of writing.
Sources
- Gong, P., Su, S., Zhang, F., Han, X., Sun, H., Li, Y., Yu, Y. SAMI3D-DW: Interactive Segmentation of Any 3D Medical Images. Deepwise technical report DW-AILAB-TR-2026-001; arXiv:2609.25743, 2026. https://arxiv.org/abs/2609.25743. PDF: https://arxiv.org/pdf/2609.25743.
- Table 3 point-only category-macro Dice: SAMI3D-DW 0.5764 / 0.7398 / 0.7771 at 1/3/5 points; nnInteractive 0.5315 / 0.6917 / 0.7494.
- Table 4 BBox-initialized: SAMI3D-DW 0.7130 → 0.8002 after five corrections; nnInteractive 0.6530 → 0.7868.
- Benchmark: 4,326 cases, 219 source datasets, 107 categories, 13,826 instances; CT 9,289 / MR 4,537 instances in modality split.
- Head vascular tree BBox+5: 0.8228 vs 0.4455; NF1 tumor burden BBox+5: 0.7351 vs 0.7077; NF1 assisted annotation ~1/15 manual time (preliminary).
- Training: ~115k curated 3D volumes from 849 Deepwise datasets; from-scratch, no pretrained weights; architecture not disclosed.