Hang a portal-venous liver CT in the viewer. Drop a loose point on a hypodense met, draw a rough box, or leave the prompt empty. A tiny trainable adapter on a frozen Segment Anything Model returns colorectal liver-metastasis overlays you can hang for response assessment, surgical planning, or follow-up. Ramtin Mojtahedi, Mohammad Hamghalam, Jacob J. Peoples, Natalie Gangai, Mithat Gonen, Yun Shin Chun, HyunSeon Christine Kang, Richard K. G. Do, and Amber L. Simpson (Queen’s, MSK, Alberta/Amii, MD Anderson, and collaborators) posted arXiv:2609.11703 on 10 September 2026. The paper is Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed Tomography. Two adapters do the work: SiGA for the clinical hook, DiSECT when you care about a near-zero trainable footprint.
Manual CRLM contours burn time and drift between readers. Full fine-tuning of a vision-transformer SAM backbone is often the wrong trade for a clinic stack with limited GPU budget. This work keeps SAM ViT-B frozen and trains only spectral residual adapters in the image encoder and mask decoder, starting from public Medical Adapter Zoo / Med-SA weights.
What goes in, what comes out
Input is a contrast-enhanced axial liver CT slice (portal venous phase in this study), optionally with a point, a few points, or a bounding box. Output is a tumor mask overlay on that slice. The authors train and select under a single-point regime, then report held-out numbers under no-prompt inference on tumor-positive 2D slices. Click once when you care about a lesion, or run automatic overlay on already-cropped liver slices.
The cohort is 446 portal-venous CECT volumes with CRLM from Memorial Sloan Kettering, MD Anderson, and TCIA public data (355 train / 91 test). For the SAM path they crop the liver, keep tumor-positive axial slices, window to [-150, 250] HU, and resize to 1024×1024. Masks were auto-generated then expert-verified. A 3D nnU-Net (Residual Encoder U-Net ensemble of five) sits as the fully trained CNN reference on the same test cases.
How SiGA and DiSECT adapt SAM without a full retrain
Both adapters start from an SVD of each frozen linear weight and keep only the leading singular directions. The residual update lives in that spectral subspace instead of learning a free low-rank matrix from scratch.
DiSECT multiplies those directions by a trainable global gate vector. It is extremely light: 0.14 million trainable parameters, about 0.14% of the model, and the lowest FLOPs in their adapter table. SiGA adds a second gate predicted from the current input features through a small MLP, then combines that instance gate with the global gate. Metastases vary in size, shape, and contrast, so the adapter routes spectral directions differently per case.
They also compare LoRA, QLoRA, and a convolutional adapter (CAD). Under single-point training, SiGA leads with 0.77 Dice and 35.39 mm HD95. On the 91-case no-prompt test set, SiGA reaches 0.76 Dice, in the same ballpark as the 3D nnU-Net tumor Dice of 0.758, while still keeping the promptable SAM interface. DiSECT stays at about 0.70 Dice in both regimes; that is the price of the ultra-small trainable set.
Prompt regimes that matter in a viewer
The paper evaluates five prompt styles: one foreground point near the lesion center, three points (center plus two refinements), boxes loosened to IoU near 0.50 or 0.75 against the ground-truth mask, and no prompt at all. Single-point was the best validation setting and needs almost no annotation effort, so it became the training reference. No-prompt is the test they emphasize for high-throughput workflows.
Point and box prompts are the interactive overlay tools radiologists and surgeons already know. No-prompt is the batch rail on liver-cropped, tumor-positive slices. Do not confuse that with open-set detection on full uncropped volumes with tumor-negative slices; the paper does not claim that.
Where it still fails
Everything is 2D on axial slices, so through-plane context is gone. The study stays on portal-venous CRLM; multi-phase CT and broader protocol shifts are future work. Cross-site generalization is only partly covered by the multi-institutional mix. Boundary error is still large: SiGA’s no-prompt HD95 is 46.76 mm, and DiSECT’s is higher. Calibration and uncertainty are not the focus. Adapter ranks and gating capacity were not exhaustively swept. If you need true 3D volumetric consistency or automatic detection on whole-abdomen studies, this preprint is not that product yet.
For a viewer shop
Treat SiGA as the accuracy-first spectral adapter when you want point, box, or empty-prompt CRLM overlays on frozen SAM with a reusable backbone. Reach for DiSECT when the trainable budget has to stay tiny and a lower Dice is acceptable. Fail closed on non-portal-venous phases you never saw, on slices without a liver crop, on full-volume detection without a separate detector, and when boundary millimeter error matters more than overlap for your surgical margin workflow.
Rebuild from arXiv:2609.11703. As of 11 September 2026 the abstract and PDF respond. No public code repository is linked in the preprint skim. The preprint is the source of record until a camera-ready version exists.
Sources
- Mojtahedi, R., Hamghalam, M., Peoples, J. J., Gangai, N., Gonen, M., Chun, Y. S., Kang, H. C., Do, R. K. G., Simpson, A. L. Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed Tomography. arXiv:2609.11703, posted 10 September 2026. https://arxiv.org/abs/2609.11703. PDF: https://arxiv.org/pdf/2609.11703.
- Kirillov, A., et al. Segment Anything. ICCV 2023.
- Ma, J., et al. Segment Anything in Medical Images. Nat. Commun. 2024 (MedSAM).
- Wu, J., et al. Medical SAM Adapter (Med-SA) / Medical Adapter Zoo (cited initialization).
- Isensee, F., et al. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods 2021.