Hang a weak ultrasound lesion overlay from a frozen segmentor and the shop question is not whether mean Dice looks fine on the test split. It is whether the bad cases, gross miss, incomplete lesion, and false-positive blobs, can be repaired without retraining the base model. Ziliang Wang, XuJiang Tang, Lu Yuting, Weixin Xu, Yongqiang Zhao, Ying Fu, and Kehua Guo (Central South University and collaborators) take that failure tail seriously in arXiv:2609.18256, posted 16 September 2026. They call the method Failure-Aware Progressive Repair (FAPR). The shop-facing idea is plain. Keep U-KAN or U-Net frozen. Treat the current mask as an evolving error state. Run ordered specialists for gross-miss recovery, false-negative expansion, then false-positive suppression, with conditional routing and a lightweight rollback when repair would hurt. Mean DSC rises by 1.52% across BUSI, BUSIS, and TN3K. On the very-hard subsets of BUSI and TN3K the average gain reaches 13.77%. Code and weights are promised upon acceptance; this is a paper walkthrough for practitioners, not a product claim.
What hangs on the viewer
Upstream is a 2D ultrasound frame with a lesion mask from a frozen Stage-1 segmentor (U-KAN in the main tables; U-Net transfer in Table 5). Downstream hangables for a DICOM or ultrasound AI shop include: the Stage-1 overlay M0 next to ground truth; intermediate repaired states M1 (gross-miss), M2 (FN expansion), and M3 (FP suppression) with TP green / FP red / FN blue overlays as in their Figure 3 on two BUSI cases; and a final accepted mask after a risk-aware rollback controller that can restore M0 when M3 would worsen Dice. In Figure 3 row (a), a miss-activated case moves from DSC 0.234 at M0 to 0.927 at M3. Row (b) bypasses gross-miss repair when support is already present, then still cleans FP contamination.
The practical claim is not a new ultrasound network. FAPR hangs a progressive repair chain on top of whatever base segmentor you already ship, and concentrates lift on the hard tail instead of polishing easy cases.
How it works in plain words
Stage 1 is frozen. It emits the initial mask M0 and a multi-scale feature pyramid F. FAPR never updates those weights. Three non-exchangeable specialists then run in order. A strict-miss router decides whether M0 lacks reliable lesion support; if so, the gross-miss expert proposes Q1 and a pixel gate blends it into M1. A tail router and FN expert expand under-segmented regions into M2. A second tail router and FP expert suppress unsupported foreground into M3. Sample-level gates decide whether a stage fires; pixel-level gates decide where pixels change. Later stages see the updated mask, so recovering a miss changes the support available for FN completion, which then changes what FP suppression should remove.
Because severe failures are rare in training, FAPR builds a training-only failure replay set from hard Stage-1 predictions (DSC < 0.7) and, with probability 0.55, replaces the Stage-1 prior with controlled corruptions (gross miss, under-segmentation, over-segmentation, boundary jitter). After the full chain, a lightweight risk-aware controller looks at ten mask-change descriptors (base foreground ratio, centroid, largest-component ratio, entropy, plus add/remove balance and related deltas) and either accepts M3 or rolls back to M0, capped at a 5% rollback rate on validation. No ground truth is used at inference.
What the numbers say
With the same frozen U-KAN, FAPR reaches 83.78 DSC on BUSI, 92.62 on BUSIS, and 83.72 on TN3K (Table 1, mean over three seeds). The headline mean gain versus Stage 1 across the three public ultrasound lesion benchmarks is +1.52% DSC. Failure-conditioned Table 2 is the shop number: very-hard BUSI cases (< 0.5 DSC at M0, n=9) rise from 17.04 to 35.78 (+18.74 points); very-hard TN3K (n=53) from 25.97 to 34.76 (+8.79); average very-hard gain 13.77%. Easy cases barely move (+0.93 BUSI, −0.35 TN3K). Hard subsets (< 0.7) also improve (BUSI 42.62 → 53.69; TN3K 42.85 → 49.39).
Ablations (Table 3) show removing any specialist drops very-hard BUSI DSC by 5.38–14.20 points; removing failure replay or controlled corruption drops it by 16.27–16.83. Easy DSC stays inside 89.77–90.11. Versus matched static correctors (SegRefiner, CausalBridgeNet) on the same frozen backbone, FAPR wins on the low-accuracy tail. Single-image FP32 at 256×256 uses 7.4% fewer parameters than their CausalBridgeNet reproduction and cuts Conv/Linear FLOPs, latency, and peak memory by 81.3%, 59.3%, and 73.9%. Table 5 transfers the repair onto a frozen U-Net Stage-1: BUSI DSC 80.91 → 82.86 (very-hard 16.38 → 30.37); TN3K 80.48 → 82.55 (very-hard 21.83 → 37.83).
Where it fails and what not to trust
This is research software on public 2D breast and thyroid ultrasound (BUSI, BUSIS, TN3K). It is not a cleared medical device. Gains concentrate on sparse hard tails; on nearly saturated BUSIS the method stays within 0.01 DSC of SegRefiner and should not be sold as a universal uplift. The rollback controller is calibrated offline with a 5% rate cap; it does not guarantee patient-level safety. Code and trained weights are not public yet (promised upon acceptance), so you cannot reproduce the hang from a GitHub clone today. Do not treat +18.74 very-hard BUSI points as a guarantee on your scanner mix, frequency, or lesion protocol. Validate the full M0→M3 chain, including rollback decisions, on an internal cohort before you trust overlays in a viewer.
For a viewer or clinic AI shop
Wire a frozen ultrasound lesion segmentor in; hang progressively repaired overlays out: gross-miss recovery, FN expansion, FP suppression, then optional rollback to the Stage-1 mask. Prefer FAPR when mean Dice already looks acceptable but a few clinically costly misses and FP blobs still escape QC, and when you refuse to retrain the base model. Keep a QC path that shows M0 next to M3 with TP/FP/FN color, logs which stages fired, and records rollback yes/no. Start from the preprint. Rebuild from arXiv:2609.18256. PDF: https://arxiv.org/pdf/2609.18256. Watch the authors’ acceptance note for code release.
Sources
- Wang, Z., Tang, X., Lu, Y., Xu, W., Zhao, Y., Fu, Y., Guo, K. Evolving Error States: Failure-Aware Progressive Repair for Ultrasound Lesion Segmentation. arXiv:2609.18256, 2026. https://arxiv.org/abs/2609.18256. PDF: https://arxiv.org/pdf/2609.18256.
- Mean DSC +1.52% across BUSI, BUSIS, TN3K; very-hard BUSI/TN3K average gain 13.77% (BUSI +18.74, TN3K +8.79). Table 1 FAPR: BUSI 83.78, BUSIS 92.62, TN3K 83.72 DSC.
- Figure 3 BUSI qualitative: row (a) M0 DSC 0.234 → M3 0.927; row (b) miss bypassed then FP cleaned to M3 DSC 0.766. Colors: TP green, FP red, FN blue.
- Rollback: ten mask-change descriptors; max 5% rollback rate on validation. Code/weights: to be released upon acceptance.