A cone-beam CT in the room can give you a handful of 2D projections, and sometimes that is all you have. Turning a few projections into a full volume is ill-posed. Classical sparse-view methods want many more views (the paper cites one that needs 30), and learned methods want large collections of paired X-ray and CT cases, which are hard to get. A group at Boston University (Weicheng Dai, Shantanu Ghosh, and Kayhan Batmanghelich) posted arXiv:2610.09253 on 7 October 2026 with a different approach. PhyDiCT takes a diffusion model that already generates lung CT and steers it until X-rays rendered from its volume match the X-rays you have. The prior is used without any training or fine-tuning.
What goes in and what comes out
The input is a set of X-ray projections taken at known angles, plus an optional piece of text, a radiology report. The default in the paper is eight projections spread evenly around the patient, 448 by 448 pixels each, with the source 1020 mm from the receiver plane. The output is a 3D lung CT volume at 256 by 256 by 256 voxels, generated by the diffusion model under the guidance of your projections.
The featured image is Fig. 2 of the paper, with every cell taken from the paper without retouching. The left column is the ground-truth CT. Then come X2CT (a fully trained network), the diffusion prior alone with no X-ray guidance, DAPS (another training-free method on the same prior), and PhyDiCT. The top row is a sagittal slice, the middle row zooms into the boxed region that the paper uses to show atelectasis, and the bottom row zooms into lung tissue with fissure lines.
How it works
The prior is MedSyn, a text-conditioned 3D CT diffusion model from the same lab, used frozen. It generates a 64 cubed volume first and a second branch upsamples it to 256 cubed. At every denoising step the model predicts a clean volume. PhyDiCT then renders that predicted volume as X-rays with a differentiable renderer based on the Beer-Lambert law, compares the renders to your projections with a mean-squared error plus a perceptual loss (LPIPS), and takes two gradient steps on the volume. The loop alternates this likelihood step with a normal denoising step, in the style of split Gibbs sampling.
The authors point to two choices. The guidance acts on the predicted clean volume, while DAPS and PnP-DM, the other training-free methods they compare against, guide the noisy sample, and the authors credit this for better steering. The LPIPS term is there to cover the gap between the idealised renderer and the X-rays you feed in.
A last step, which the authors call extra compute refinement, adds noise back to the finished volume (timestep fraction 0.8) and denoises it again with no X-ray guidance. That pulls the output back onto what the prior treats as a realistic CT and removes small artefacts. It costs roughly 20 seconds.
Data and evaluation
There is no public dataset of paired X-rays and CT, so the authors simulated the X-rays. They took 400 CT volumes with radiology reports from CT-RATE and rendered projections with two tools: DiffVox, which is more consistent with the renderer inside PhyDiCT, and DeepDRR, which adds noise and beam-hardening effects and is the more realistic test. The text prompt is the radiology report of the CT scan, used as a stand-in for an X-ray report.
What the numbers say
On the DiffVox projections, PhyDiCT reaches an SSIM of 0.4268 against 0.3969 for X2CT, the best baseline, and 0.2466 for the prior with no X-ray guidance. The paper reports this as a 7.5 percent SSIM gain over the trained baselines. The absolute values are low, so the volume is still far from the reference.
The trained network still wins on some metrics. X2CT has higher semantic scores from CLIP (16.16 against 13.75 for text-to-image, 89.04 against 84.66 for image-to-image). The authors put this down to X2CT being trained on CT-RATE while the prior was trained on other data.
The more realistic DeepDRR projections hurt every method on the perceptual metrics. PhyDiCT stays ahead on SSIM there, 0.3534 against 0.3131 for Xray2CTPA and 0.2619 for X2CT, whose image-to-image CLIP score falls to 71.98, which the authors blame on GAN collapse under domain shift.
More views help a little and cost a lot. Going from 8 to 32 views moves SSIM (without the refinement step) from 0.3983 to 0.4073 while run time grows from 1,342 s to 4,709 s. Two views take 453 s and score 0.3340.
Speed
The paper’s timing for eight views on a single NVIDIA RTX 6000 is 1,342 seconds for the main run plus about 20 seconds for refinement, so around 22 minutes per volume. That is a batch job. Nobody waits for it at the console.
What the pictures show
In the paper’s qualitative figure, the prior alone exaggerates disease, which the authors describe as exaggeration artefacts, since nothing ties it to the X-rays. They report that the other baselines fail to reconstruct the atelectasis region and that PhyDiCT recovers it without the overemphasis, and that it reproduces the fissure lines the others miss. Without the refinement step the output shows minor artefacts, and with it they are suppressed. In the zoomed rows the PhyDiCT crop looks softer than the ground truth to me, so I would not read fine detail from it.
Where it falls short
Every test in the paper uses X-rays simulated from CT scans. No real X-ray paired with a CT appears, so nothing here shows how it behaves with real scatter, patient motion, or the geometry of a real machine. The reports used as text prompts come from the CT scans themselves, which an X-ray-only workflow would not have. Without the report, SSIM drops to 0.3747, below X2CT in that setting.
The method is built for lung CT, and the authors say other organs need a matching pretrained prior. The output comes from a generative prior that, left alone, exaggerates disease, so treat it as a plausible volume consistent with the projections and not as a measurement. I did not see a reader study or any clinical validation.
Code, weights, and data
The repository is github.com/batmanlab/PhyDiCT under the Apache 2.0 license. The README walks through the steps: download the MedSyn weights separately, create the conda environment, prepare the X-ray projections, extract a text feature from a report sentence, then run eval_low_PhyDiCT.py for the 64 cubed result and eval_super_res.py for 256 cubed. It includes eight example X-ray images, marked as reference only. The README also warns you to visualise the projections and the 3D features in case a transpose or flip is needed, which matters when you connect scanner orientation conventions to a model. I did not run it, and I did not find a documented route from DICOM input. The CT data is CT-RATE, under its own terms.
How we would wire it into a viewer
We would run it as a batch service that creates a derived series. A job takes the projections and geometry from the acquisition, runs for roughly 20 to 25 minutes on the paper’s GPU, and writes back a CT series flagged as derived and AI-generated in the DICOM header and the series description, with a reference to the source projections. In the viewer, the generated volume would sit beside the best conventional reconstruction we have, with a visible label, and it would stay out of any measurement or report.
Before it goes near a clinic we would test it on our own projections: the angle and orientation conventions against the paper’s renderer, how it handles real noise and scatter, and how often the generated volume adds or removes a finding that the projections do not support.
Sources
- Dai, W., Ghosh, S., Batmanghelich, K. PhyDiCT: Plug-and-Play CT Reconstruction from Sparse X-Rays via Differentiable Rendering and Strong Priors. arXiv:2610.09253, 2026. https://arxiv.org/abs/2610.09253. PDF: https://arxiv.org/pdf/2610.09253.
- Code and example X-rays: https://github.com/batmanlab/PhyDiCT.
- Xu, Y., et al. MedSyn: Text-guided Anatomy-aware Synthesis of High-Fidelity 3D CT Images. IEEE Transactions on Medical Imaging, 2024. https://doi.org/10.1109/TMI.2024.3415032.
- Hamamci, I. E., et al. Developing generalist foundation models from a multimodal dataset for 3D computed tomography (CT-RATE). arXiv:2403.17834, 2024. https://arxiv.org/abs/2403.17834.
- Ying, X., et al. X2CT-GAN: Reconstructing CT from biplanar X-rays with generative adversarial networks. CVPR 2019. https://doi.org/10.1109/CVPR.2019.01087.
- Featured image: Fig. 2 of arXiv:2610.09253, all five columns and all three rows as published.