Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

Abdominal CT Report Sentence In, 3D Finding Mask Out, Learned From the Arrows and Measurement Lines Already in PACS

Most abdominal CT reports already point at the image. A radiologist writes “series 2 image 41, a 1.3 x 1.2 cm hypoattenuating lesion” and, while reading, drops two measurement lines on that slice in PACS. The sentence lives in the RIS and the lines live in a DICOM grayscale softcopy presentation state (GSPS) object, and nothing in a typical viewer connects the two. A referring clinician who opens the study later still has to scroll to find the lesion the report describes.

Sam Church, Danyal Maqbool, Joshua D. Warner, Andrew Voter, Junjie Hu, Meghan G. Lubner, and Tyler J. Bradshaw (University of Wisconsin-Madison, Computer Sciences and Radiology) use those everyday marks as training data in arXiv:2610.04095, posted 2 October 2026. Their pipeline matches each arrow or measurement line to the report sentence it belongs to, grows the 2D mark into a 3D mask, and trains a model called LocusCT that takes a free-text finding phrase plus an abdominal CT volume and returns a 3D mask of that finding. They also built LocusBench, a radiologist-checked test set of 500 abdominal findings.

What goes in and what comes out

At inference, LocusCT needs two things: an abdominal CT volume and one sentence describing one positive finding, written the way radiologists dictate (“The appendix is mildly dilated up to 1.0 cm, with surrounding mesenteric fat stranding”). It returns a voxel mask on the CT grid. The volume is resampled to 1.5 x 1.5 x 3.0 mm and fed as three channels, windowed for soft tissue, lung, and bone.

The model reuses the VoxTell architecture, a 3D nnU-Net backbone where text features from a frozen language encoder are multiplied into the image features at several decoder levels. The authors trained it from scratch on 4 H200 GPUs. The user gives no click, box, or seed point, and the sentence alone tells the model where to look.

How the training data gets built

Most of the work went into the data. The team pulled every abdominal CT from their institution over ten years, 91K exams split into oncology (26K), emergency department (31K), and everything else (34K).

An open-weight LLM (Qwen3.8-27B), run inside the hospital network, pulls out report sentences that describe exactly one non-negative finding and extracts the clues needed for matching: series and slice numbers, measurement values, and slice positions. Then two matching rules link sentences to marks. A measurement line matches a sentence that reports the same size, allowing rounding when only one line in the exam could round to that value. Two crossing lines at roughly 90 degrees count as one bidimensional measurement and match “1.2 x 3.4 cm” style sentences. An arrow matches when the sentence cites a slice and exactly one arrow sits on that slice.

A 2D mark gives a location but no boundary, so the pipeline hands it to SAM2CT, the group’s earlier promptable model built on SAM2 that accepts arrows and lines as prompts. It segments the annotated slice and propagates the mask up and down through the volume. The yield was 105K phrase, mask, and volume triplets from 59K exams and 37K patients, with no new annotation work. Fig. 2 of the paper, used as this post’s featured image, shows three of these conversions: a liver lesion measured on slice 41, a kidney lesion marked with an arrow on slice 83, and a 14 cm hernia measured on slice 97.

Two board-certified radiologists checked raw pipeline output for the LocusBench cases. Masks averaged 4.27 out of 5 on the oncology set and 4.32 on the ED set, and 11.7% and 16.9% of cases needed segmentation corrections before approval.

How well it finds things

The paper scores grounding with Dice and a hit rate, where a hit means the predicted mask overlaps the reference with Dice of at least 0.1. Existing text-promptable 3D models did poorly here because they were trained on organ names and category labels, not report language. The best of them, VoxTell used zero-shot, hit 18.8% of oncology findings and 20.9% of ED findings. LocusCT hit 72.5% and 77.3%.

More weak data kept helping. Training on 10% of the triplets gave a 45.8% oncology hit rate, 30% gave 55.0%, and the full set gave 72.5%. A smaller, cleaner subset restricted to findings that resemble SAM2CT’s own training data did worse than the larger, noisier full set.

The Dice distribution is bimodal. Median Dice was 0.58 (oncology) and 0.59 (ED), but cases cluster either at a good overlap or near zero. When LocusCT finds the right structure, it usually segments it well. Most errors come from looking in the wrong place or returning nothing.

On Merlin, a public abdominal CT dataset with reports and no masks, a radiologist rated 78 predictions on a five-point scale. Mean score was 3.65, and 80.8% were at least useful for localization. Of the 11 worst cases, 8 had no predicted mask at all and 3 landed in the wrong location.

Where it fails

Diffuse findings are the weak spot. Diverticulitis had a median Dice of 0.07 on LocusBench-ED, and abscess was also near the bottom, because their margins blend into surrounding inflammation. Cysts, aneurysms, and masses scored best. On Merlin, kidney stones did poorly despite doing well in-house, though each Merlin category had only six cases.

The pipeline inherits radiologist habits. Arrow matching only works where radiologists dictate slice numbers, which is common at this institution but not everywhere. Lines are often drawn next to a finding, or drawn on it and dragged aside so the lesion stays visible, and that shift can push SAM2CT onto the wrong structure. Arrows show a location but no extent, so SAM2CT sometimes segments a larger neighboring structure. Ellipses and text marks were left out entirely, and only 33 to 40% of all GSPS annotations per cohort became training masks. Findings that radiologists rarely mark will be underrepresented.

Everything comes from one institution, so scanners, protocols, and reporting style are site-specific, and Merlin is the only external check so far. Each phrase in training and testing describes a single positive finding, so sentences with several findings or negations (“no free fluid”) are untested. The authors state that LocusCT is a research model and not intended for clinical use.

What this would look like in a viewer

For a DICOM viewer team, the first feature I would build is a clickable report. A reader or referring physician clicks a finding sentence in the report pane, the viewer jumps to the slice, and a 3D outline appears on the axial, coronal, and sagittal views. The same mask gives you a volume or longest diameter to show next to the dictated value, and a starting region when the next follow-up study needs comparing.

The failure pattern is friendly to a viewer. Because LocusCT tends to return an empty mask rather than a confident wrong one, the viewer can show “not localized” and fall back to the plain report. I would still store outputs as a separate DICOM SEG or overlay series, tagged as machine-generated, and never write them into the original study.

A second use is for anyone running a PACS archive. If your radiologists measure and mark findings, years of GSPS objects plus reports already form a weakly labeled grounding dataset. The paper’s matching rules and LLM prompts are written out in the appendix, and the prompts run on a local open-weight model, so no report text has to leave the hospital.

Code and data availability

Training and evaluation code is public at github.com/samdchurch/LocusCT. As of 6 October 2026 the README lists the model weights and LocusBench as coming soon, and the paper says LocusBench (CT volumes, radiologist-verified masks, and report phrases), the weights, and the code will be released upon acceptance. The 105K-triplet training set will not be released because of institutional data governance. Merlin, used for the external test, is publicly available.

Sources

We build custom medical imaging platforms — advanced DICOM viewers, AI segmentation, and the clinical systems around them.

Get in Touch

Copyright © 2026 PYCAD. All Rights Reserved.