When I add a 3D view to a DICOM viewer, the first control people reach for is the transfer function, the curve that decides which Hounsfield range shows up as bone, which turns into a translucent vessel, and which disappears. Cinematic path tracing makes those renders look like an anatomy atlas, but each frame is expensive. Gaussian splat proxies fix the speed problem by fitting a cloud of small 3D Gaussians (roughly 80,000 to 280,000 per scan in this paper) to a set of rendered images and rasterizing them in real time. But the fitted proxy bakes in the one preset it was trained on. Ask it to show the aorta in a different color or hide the ribs and you are back to fitting a new proxy.
Zhongpai Gao, Benjamin Planche, Meng Zheng, Anwesa Choudhuri, Terrence Chen, and Ziyan Wu at United Imaging Intelligence (Boston) go after that limitation in arXiv:2610.02382, posted 1 October 2026. Their method, FactorSplat, is a per-scan Gaussian proxy that takes region-specific intensity-to-color and opacity curves as an input at render time. One checkpoint per scan handles new presets, including combinations and edits it never saw in training, without refitting.
What goes in and what comes out on the viewer
Upstream, the method needs three things: the CT or MR volume, a segmentation that splits it into regions (heart chambers, vessels, bone and so on), and a bank of authored transfer functions, one intensity-to-RGBA curve per region per preset. For CT the curve is defined on calibrated HU. For MR it uses the renderer’s normalized intensity. The reference images come from cinematic volume path tracing of the original volume and masks.
Downstream you get a compact 3D Gaussian representation that a viewer can rasterize interactively. The user can change a region’s color, change its opacity, or switch a region off, and the proxy updates its appearance without new training. Geometry stays fixed across presets, so moving the camera after an edit reuses the same primitives. In the paper’s Fig. 1 (the featured image at the top of this post), the same vascular CT proxy highlights selected tissue in one edit and recolors tissue in another. The highlight edit is one of the held-out out-of-distribution edits in the paper’s Fig. 3. The baseline it is compared against drifts the edited vessels toward green.
How it works in plain words
Every Gaussian carries a small fingerprint of the voxels around it. At initialization the method samples a 3 x 3 x 3 neighborhood around each primitive and records which region each sample belongs to and which intensity bin it falls in, weighted by how much of the Gaussian covers that voxel. That record stays attached to the primitive for the whole training run, and children created by splitting inherit it.
When a user edits a preset, FactorSplat computes the difference between the new curves and a base preset and reads that difference through each Gaussian’s fingerprint. Two branches then act on the primitive’s base color and opacity. The first is a direct lookup that applies the authored change, so if you make bone redder, Gaussians sitting on bone get redder. The second is a small learned network shared across the scene, combined with low-rank per-Gaussian factors (rank 8 in the paper), that corrects for how a cinematic renderer responds to that change once lighting, shading, and occlusion are involved. Only the view-independent color term and opacity change. Directional appearance and geometry are shared across all presets, and at the base preset both branches output exactly zero, so the base render is preserved.
Hiding a structure is handled by a separate rule instead of learning. If at least half of a Gaussian’s recorded material belongs to regions the user switched off, its opacity is set to zero. Training uses a pruning rule that keeps a Gaussian if it is visible under any training preset. Without that rule, structures that are transparent in the base preset (and that a later preset should reveal) would be thrown away early.
The authors also run region-aware VEG, the closest transfer-function-agnostic Gaussian method (Dyken et al., 2026), as the main baseline. They extended the official VEG code with one fixed region per Gaussian and fixed a numerical bug in its specular term that otherwise stopped training.
What the numbers say
The evaluation covers seven scans: five CT (heart, vascular, intestine, lower body, knee joint) and two MR (nose, hand), with 3 to 17 region curves per scan. Each scan has 41 conditions. Twenty-four presets are used for training, and the held-out ones split into validation (4), interpolation within the training range (4), unseen compositions of operations seen separately (4), and five out-of-distribution edits (grayscale, opacity sharpening, isolating a target, suppressing occluders, and emphasizing a target with translucent context). Ground truth is path-traced at 1600 x 1600 and every method is scored on the same 40 held-out cameras.
Against region-aware VEG, FactorSplat had higher full-frame PSNR in 24 of 28 scan and split comparisons. Mean gains were 1.45, 1.52, 1.10, and 1.41 dB on validation, interpolation, composition, and out-of-distribution edits. The authors also score edit fidelity in the pixels the preset was supposed to change, and FactorSplat’s error there was lower in 19 of 28 comparisons, with lower means on every split for both CT and MR.
On speed, the cached renderer averaged 524 frames per second at 1600 x 1600 against 144 for VEG, and a preset switch cost 1.17 ms on average. Training took 9.8 minutes per scan versus 43.8 for VEG. Fitting a separate fixed-preset proxy for each of the 17 evaluated presets would have cost an estimated 88.1 minutes per scan. All timings are on NVIDIA B200 GPUs.
In the ablations, removing the direct lookup cost 2.57 and 2.65 dB on composition for heart and vascular. Replacing the curve input with a nearest-preset embedding stayed competitive on edits close to the training bank but lost 1.02 and 2.84 dB on out-of-distribution edits.
Where it falls short
New kinds of edits remain the weak spot. Proxies fitted directly to one target preset kept a 5.2 dB mean advantage on the out-of-distribution edits. VEG had lower changed-region error than FactorSplat throughout the heart and knee joint scans, and it scored higher PSNR on the heart out-of-distribution edits and several knee joint splits. If your users push sliders far outside the presets you trained on, expect visible drift.
Conditioning is local. A Gaussian only reacts to edits on the materials it covers, so when an edit to one structure changes the lighting on its neighbors, those neighbors do not respond. The authors also note that the attached voxel fingerprints can fall out of alignment as Gaussians move during training, and that hiding a region removes whole Gaussians instead of clipping cleanly at the voxel boundary. Near a hidden structure you may see a ragged edge.
Each scan still needs its own fitted proxy, which means a path-traced training bank (24 presets with 48 views each) and about ten minutes on a B200 per case. Feed-forward construction, building on the same group’s Render-FM work, is listed as future work. Frame rates on the workstation or laptop GPU a radiologist or surgeon uses were not reported.
The paper’s ethics statement sets the clinical scope. FactorSplat is evaluated against rendered CT and MR images, it is not a validated diagnostic tool, and approximation errors or transfer-function edits can change what anatomy is visible. They say clinical use would need task-specific validation, expert oversight, and access to the original scan.
How I would wire it into a DICOM viewer
I would treat FactorSplat as a fast presentation layer that sits next to the real volume renderer, with the true render and the source slices one click away. The pipeline would run server-side after a study and its segmentation are in place. Path-trace the training bank from the original series and masks, fit the proxy, then store it as a sidecar artifact keyed to the source SeriesInstanceUID and the segmentation it was built from. If either one changes, the proxy is stale and gets rebuilt. We have written before about choosing between server-side and in-browser rendering, and this design fits a split path well. The heavy fitting happens on the server and the light rasterization happens in the browser.
In the viewer, the per-region curves become per-structure color, opacity, and visibility controls. Since a preset switch costs about a millisecond on the paper’s hardware, sliders can feel live. I would constrain those sliders to the range the training bank covered, because interpolation and composition held up far better than out-of-distribution edits, and show a clear “approximate render” label on the 3D panel. Any measurement, and any question about whether a structure is present, goes back to the original slices or a true volume render.
The uses I have in mind are patient explanation, tumor boards, surgical planning review, and case sharing, where people want to recolor or hide structures many times on one case and a slow path tracer gets in the way. For primary reading I would not use it.
Code and data availability
As of 5 October 2026 there is no public code. The project page links only the arXiv paper and a PDF copy, and its GitHub repository holds the page itself. The repository README says the paper is under review at ICLR 2027. The paper does not name or link the seven source scans or the authored preset banks, so outside teams cannot rerun the benchmark yet. The reproducibility statement mentions a repository with experiment scripts and preset manifests, but gives no link to it. The method section and appendix give enough detail (descriptor sampling, loss, rank, pruning rule, VEG adaptation) to attempt a reimplementation.
Sources
- Gao, Z., Planche, B., Zheng, M., Choudhuri, A., Chen, T., Wu, Z. FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering. arXiv:2610.02382, 2026. https://arxiv.org/abs/2610.02382. PDF: https://arxiv.org/pdf/2610.02382.
- Project page: https://gaozhongpai.github.io/FactorSplat/.
- Dyken, L., Sewell, A., Usher, W., Debardeleben, N., Petruzza, S., Kumar, S. Volume Encoding Gaussians: Transfer function-agnostic 3D Gaussians for volume rendering. IEEE TVCG 32(6), 2026.
- Featured image: Fig. 1 of arXiv:2610.02382, vascular CT, region-aware VEG baseline (left) versus FactorSplat (right) on two transfer-function edits (top half: highlight selected tissue, bottom half: change tissue colors).