Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

Handheld Endoscope or Microscope Video In, One Wide-Field Tissue Mosaic Out at About Ten Frames per Second From a Modality-Tuned Optical Flow Model

Before and after mosaics of fetoscopy, dermoscopy and light-sheet microscopy video: the top row, stitched with an untuned optical flow model, shows grey patches and broken pieces; the bottom row, from the fine-tuned FloVMos model, is continuous (Liu et al., arXiv:2610.04258, Figs. 3 and 5)

A handheld probe that touches the tissue sees only a small window at a time. Dermatology confocal microscopes, fetoscopes and endoscopes all give you high resolution or a wide field of view, not both. The usual workaround is to record a video while the operator sweeps the probe and stitch the frames into one big picture afterwards. A group at Northeastern University and Memorial Sloan Kettering Cancer Center (Jinyang Liu, Sandesh Ghimire, Chaman Singh, Jennifer Dy, Milind Rajadhyaksha, Dana Brooks, Octavia Camps, and Kivanc Kose) posted arXiv:2610.04258 on 3 October 2026 with FloVMos, a stitching pipeline that is meant to work across imaging devices. The contribution is a recipe for retraining an existing optical flow network on synthetic motion for each device.

What goes in and what comes out

The input is a video from a small field of view device, with the probe moving freehand or loosely guided. The output is one wide field mosaic that grows as frames arrive, so the operator can see where they have already been. The authors tested seven kinds of imaging: reflectance confocal microscopy of skin (RCM), open-top light-sheet microscopy (OTLS), fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy (SSM), and endoscopy.

The featured image is built from two figures of the paper, with every tile taken as published and only cropped to the mosaic and resized. The top row is Fig. 5, which the caption says shows the untuned optical flow model on the same video sequences as Fig. 3. The bottom row is the matching mosaic from Fig. 3 made with the fine-tuned model. Fetoscopy, dermoscopy and light-sheet microscopy are shown.

How it works

There are four steps: read a new frame, register it, warp it, and blend it into the mosaic. Instead of chaining each frame to the previous one, which lets small errors pile up until the mosaic breaks, FloVMos tracks where the last stitched frame sits in the mosaic, crops that region, and estimates dense optical flow between the crop and the incoming frame. Because the flow is dense, it can follow tissue that moves under the probe as well as the camera move.

The flow network is FlowNet2, fine-tuned per device. RAFT, GMFlow and FlowFormer were also tried. Warping uses bilinear sampling. Stitching uses a graph-cut seam through low-detail parts of the overlap, so pixel values come from one frame or the other rather than a blend, which avoids blur.

The authors also added a guard for jumps from motion blur or depth changes. They measure how smooth the flow field is, using its second derivatives, and flag a frame when the value passes 0.1. The pipeline then buffers about a second of frames, picks the one with the smoothest flow, and if it does not overlap the mosaic enough, starts a new mosaic.

Training without labels

Hand-labelled optical flow does not exist for these devices, so the authors fabricate it. When a large mosaic is available, as with RCM from a stage-mounted microscope and OTLS, they slide a virtual camera across it along a random path. When only videos exist, they build a 2 by 2 grid out of one flipped frame, or out of four frames taken far apart in the video, and treat that as a small mosaic. Either way each step uses a known flow, made from a global homography (translation, rotation, shear, with smoothed random speeds) plus a smooth non-rigid warp. The non-rigid level is chosen per device, 3 for dermoscopy and 10 for RCM.

Training uses 256 by 256 images, Adam at 1e-4 for 100 epochs, on a single RTX 2080 Ti. Most modalities used a few hundred synthetic frames (250 videos and 500 frames for endoscopy, for example). RCM used 11,000 videos and 40,000 frames because real mosaics were available.

What the numbers say

The authors score accuracy without a ground truth mosaic by matching SIFT features between the source frames and the final mosaic and checking whether distances between matched points are preserved. In all seven modalities the average difference was under 5 pixels per match, under 0.5 percent of the field of view.

The comparison against other stitchers was run on 800 simulated RCM videos, against Parallax, APAP and AVM. FloVMos had higher SSIM and lower LPIPS than the best of them, Parallax, with p below 0.0001. On speed, the paper says Parallax fell from 0.2 to 0.05 frames per second as videos got longer, and AVM often collapsed on 50-frame videos. FloVMos handles 1 megapixel frames at about ten frames per second in the research code.

The fine-tuning effect is clearest in the optical flow table. On simulated RCM, off-the-shelf FlowNet2 has a mean squared flow error of 0.940 and the fine-tuned one 0.046. The other three networks show the same pattern: RAFT 0.921 to 0.102, FlowFormer 2.503 to 0.063, GMFlow 1.259 to 0.129. The speed column lists 32.8 frames per second for FlowNet2, 35.1 for GMFlow, 14.5 for FlowFormer and 10.8 for RAFT.

What the pictures show

In the featured image the untuned model leaves visible damage. The fetoscopy mosaic has grey patches where frames landed wrong, the dermoscopy strip has a misplaced piece on the left, and the light-sheet mosaic is broken into islands, while the fine-tuned versions are continuous. The Fig. 5 caption explains why: devices with large black or white featureless regions fail most, and even the RGB-like ones, endoscopy, dermoscopy and laparoscopy, show wrong alignment in the last frames or artifacts. I could not check that each pair is the same video, because the paper’s captions say so but the mosaics have different shapes and scales. Treat it as the authors’ qualitative comparison.

Where it falls short

The authors say the reliance on simulated training data is a major limitation. They only tested close-proximity devices, and the simulation may miss distortions from devices farther from tissue. Real clinical data and clinician review are listed as future work, so nothing here measures diagnostic value. Several real test sets are tiny: one endoscopy video with 350 frames, one laparoscopy video with 428, two fetoscopy videos, two SSM videos. For everything except the simulated RCM and OTLS videos, accuracy is a feature-distance sanity check, with no ground truth mosaic to compare against. Real-time speed needs a GPU, which most clinical imaging systems do not have. Consecutive frames need at least 50 percent overlap, and sequences with strong vertical camera motion were excluded. The result is a 2D mosaic.

Code, weights, and data

The paper says data generation, mosaicking code and the fine-tuned flow models are public at github.com/JinyangMarkLiu/OpticalFlowBasedMosaicking. When I checked on 10 October 2026, that URL returned a 404, and a GitHub search for the repository name and for FloVMos found nothing. It may be private or not yet pushed, so I could not review code, weights or a license. Base FlowNet2 weights come from NVIDIA’s public flownet2-pytorch repository. The endoscopy, laparoscopy and fetoscopy videos come from the EndoVis 2024 challenge, the Hamlyn Centre, and the FetReg dataset. The RCM and dermoscopy data are from MSKCC, and the OTLS and SSM data were provided by partner labs, and the paper does not say they are released.

How we would wire it into a viewer

A mosaic from this kind of pipeline should not enter a viewer as if the scanner acquired it. We would store it as a derived image labelled as stitched, keep the per-frame transforms with it so a click on the mosaic opens the source frame, and show the first-frame marker and any points where the pipeline restarted a new mosaic. Distances measured across a warped mosaic would carry a warning, because the paper checks distance preservation only through matched features.

Before using it with our own data we would test five things: whether the fine-tune on our device’s frames holds on videos from other operators, the minimum frame overlap and frame rate at our device speed, the GPU we have at the point of care, how often the jump guard restarts a mosaic on real motion, and how the mosaic compares with a slower offline stitch.

Sources

We build custom medical imaging platforms — advanced DICOM viewers, AI segmentation, and the clinical systems around them.

Get in Touch

Copyright © 2026 PYCAD. All Rights Reserved.