Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

Liver CT In, Organ Tumor Vessel Overlays Physical Volumes and Resection Plan Out

Hang a liver CT into a viewer and the shop question is not “does an LLM know liver anatomy.” It is whether organ, tumor, and vessel overlays, physical volumes, diameters, and tumor-vessel distances stay tied to the same DICOM or NIfTI case, and whether a proposed resection sequence can be checked before anyone trusts it. Binghong Qian, Xuanhe Liu, Yifan Xing, Wenjie Deng, Jian Wu, and Haochao Ying at Zhejiang University take that workflow gap seriously in arXiv:2609.37648, posted 29 September 2026. They call the system VoxelSage: Tool-Augmented 3D CT Analysis and Simulator-Shielded Sequential Resection Planning for Liver Tumors. The shop-facing idea is plain. Dual-port design: Port A orchestrates language and skill selection; Port B does the image math on CT volumes and masks and returns structured results. Physical measurements never leave Port B, so the language model cannot invent millilitres or millimetres. Eight built-in skills cover quantification, slice evidence, interactive mask edit, 3D reconstruction, and confirmation-gated resection sequence planning. Public code: https://github.com/ZJUMAI/VoxelSage. This is a paper walkthrough for practitioners, not a product claim.

Default segmentation backend is VISTA3D (TotalSegmentator available as an alternative). Evaluation cases for the online workflow use public colorectal liver metastasis CT from TCIA. Simulator planning numbers below are system-level, not clinical efficacy.

What hangs on the viewer

Upstream is a DICOM series or NIfTI liver CT. After upload, Port A assigns a case UUID and Port B standardizes orientation (RAS+), retains the affine, runs segmentation for liver, hepatic tumor, portal vein, and hepatic vein, then post-processes masks onto the CT grid. Downstream hangables for a DICOM or clinic AI shop include: axial slices with organ (cyan), tumor (red), and vessel (blue) overlays as in their Figure 4 for CRLM-CT-1036; per-lesion measurement cards (volumes, diameters, tumor-vessel distances) next to the image rather than baked into the overlay; an interactive segmentation editor with explicit save before other skills see the edit; a Three.js 3D scene with liver, tumors, and vessels; and editable bicubic Bézier resection surfaces on a 4×4 control grid that only enter sequence planning after the user saves them.

The practical claim is not a new scanner. VoxelSage hands you a case-centered loop where every quantitative claim stays linked to masks, geometry, and execution records, and where the LLM is an orchestrator, not a calculator.

How it works in plain words

Port A (agent orchestration) interprets the natural-language request, builds context from case history and cached evidence, and talks to Qwen. When evidence is missing, it selects skills from Port B’s /api/skills/list manifests and calls /api/skills/run. The agent loop is capped at six rounds. Cache keys are case UUID + skill name + input parameters, so equivalent requests reuse Port B results instead of recomputing.

Port B owns deterministic computation. Four measurement skills compute from masks and the affine: liver_analysis (integrated volumes, diameters, distances, report), tumor_diameter, tumor_vessel_distance, and vessel_volume. Volume uses the absolute determinant of the affine’s linear part so voxel volume is in mm³ even on anisotropic grids. Diameters transform lesion coordinates through the affine before measuring maximum 3D extent. Tumor-vessel proximity uses a Euclidean distance transform with per-axis spacing. slice_selection ranks axial slices and returns top-K overlays. segmentation_modification edits in a working session; other skills keep reading the original mask until save (with a .bak). three_d_reconstruction builds the shared 3D scene and can overlay candidate Bézier resection surfaces.

For sequence planning, the system proposes candidate resection surfaces from tumor-vessel geometry. The user edits and saves a surface. A frozen behaviour-cloned neural ranker orders candidate targets on a 2D grid derived from that surface, and a simulator-based “lazy exact” shield checks them against predefined constraints (including an episode blood budget) before accepting a target. Unconfirmed surfaces cannot enter later computation.

What the numbers say

On 256 unseen frozen simulator scenes, the shielded learned controller (C4) reduced mean simulated time from 34.274 to 33.388 min (−0.886 min, −2.59%) and mean simulated blood loss from 300.847 to 183.852 mL (−116.995 mL, −38.89%) versus a deterministic baseline. The paper states these results demonstrate system integration and simulator-level performance, not clinical efficacy or safety. C4’s median controller wall time was also 54.7% lower than an exact-search comparator (C3) in their latency check (21.21 vs 46.85 s median across 64 scene medians).

Separate Port A experiments with reference masks (not automatic segmentation) on AbdomenAtlas / DeepTumorVQA-style tasks (Table 4): lesion existence 44/44, total lesion volume 27/27, max 3D diameter 25/25, lesion counting 30/34 (88.2%), liver volume 22/24 (91.7%), multi-case larger-liver identification 59/63 (93.7%), no invented measurements on missing cases 137/137, and cache reuse on 130/137 repeats with median response time 8.16 → 2.56 s. Treat these as engineering agreement under their experimental thresholds, not as cleared clinical accuracy.

Where it fails and what not to trust

This is research software on public CRLM CT, reference-mask Port A probes, and a frozen 2D planar simulator for planning. It is not a cleared medical device. The image-analysis path uses liver, tumor, and vessel masks only; evaluation on colorectal liver metastases alone does not establish generalization to primary liver cancer or other origins. The planner does not model liver function, segmental perfusion, deformable anatomy, instrument mechanics, or postoperative outcomes. Future liver remnant volume and function sit outside the current geometric planner.

Do not treat the −38.89% simulated blood-loss delta as a patient-level prediction. Do not skip local validation of automatic segmentation before you trust Port B volumes on your own scanner mix. And do not let an LLM free-text millilitres when Port B already owns the affine math; that separation is the point of the dual-port design.

For a viewer or clinic AI shop

Wire a liver CT volume in; hang organ, tumor, and vessel overlays, physical volumes and distances, and a simulator-checked resection sequence out. Prefer VoxelSage when you want case-linked measurement and planning with the language model kept off the calculator: dual-port orchestration, eight inspectable skills, explicit save gates on mask edits and resection surfaces, public GitHub. Keep a QC path so support can compare overlays to the source CT and re-run Port B after a saved mask edit. Log case UUID, skill name, parameters, and whether a result was cached or freshly computed.

Validate on at least one internal cohort before you claim hang in production. Start from the preprint. Rebuild from arXiv:2609.37648. PDF: https://arxiv.org/pdf/2609.37648. Code: https://github.com/ZJUMAI/VoxelSage.

Sources

  • Qian, B., Liu, X., Xing, Y., Deng, W., Wu, J., Ying, H. VoxelSage: Tool-Augmented 3D CT Analysis and Simulator-Shielded Sequential Resection Planning for Liver Tumors. arXiv:2609.37648, 2026. https://arxiv.org/abs/2609.37648. PDF: https://arxiv.org/pdf/2609.37648. Code: https://github.com/ZJUMAI/VoxelSage.
  • Simulator (256 unseen scenes): mean simulated time 34.274 → 33.388 min (−2.59%); mean simulated blood loss 300.847 → 183.852 mL (−38.89%) vs deterministic baseline. Paper caveat: system/simulator-level, not clinical efficacy.
  • Table 4 Port A (reference masks): lesion existence 44/44; lesion count 30/34; liver volume 22/24; total lesion volume 27/27; max 3D diameter 25/25; multi-case 59/63; evidence faithfulness 137/137; cache 130/137, median 8.16 → 2.56 s.
  • Eight skills include liver_analysis, tumor_diameter, tumor_vessel_distance, vessel_volume, slice_selection, segmentation_modification, three_d_reconstruction, and resection sequence planning. Default seg: VISTA3D. Illustrative cases: CRLM-CT-1012, CRLM-CT-1036 (TCIA Colorectal-Liver-Metastases).

We build custom medical imaging platforms — advanced DICOM viewers, AI segmentation, and the clinical systems around them.

Get in Touch

Copyright © 2026 PYCAD. All Rights Reserved.