Most brain MRI archives can be searched by patient name, date, study description, and maybe a diagnosis code, but not by what the brain looks like. If a neuroradiologist wants to see five prior cases with hippocampi shaped like the one on screen, or a research coordinator wants every scan in the archive that resembles a given patient, the answer today is a lot of manual scrolling or nothing at all.
Félix Nieto-del-Amor, Jingru Fu, J.-Sebastian Muehlboeck, Eric Westman, Daniel Ferreira, and Rodrigo Moreno (KTH Royal Institute of Technology, Karolinska Institutet, and the Martinos Center at MGH and Harvard) describe a system for that in arXiv:2610.06502, posted 5 October 2026. NeuroCBIR takes a 3D T1-weighted brain MRI, turns it into a 32-number fingerprint, and ranks a reference set of more than 26,000 real scans by similarity. The search can run on the whole brain or on any one of 103 cortical and subcortical regions. The code, the trained models, and the precomputed embeddings are public under Apache 2.0.
What goes in and what comes out
The input is one 3D T1w volume. It is skull-stripped, bias-corrected, resampled to 1 mm isotropic with the FreeSurfer pipeline, affinely registered to Talairach space, and cropped to 160 x 176 x 208 voxels. For region queries, a FreeSurfer or FastSurfer segmentation (aseg plus aparc) supplies 103 structures, each cropped with a fixed bounding box and resampled to 64 x 64 x 64.
The output is a ranked list of the closest scans in the reference set, with their cosine similarity scores. Every result is a real, curated scan from a research cohort. The featured image is Fig. 5 of the paper. The blue frame is the query, green frames are other scans of the same person, and red frames are scans of different people that NeuroCBIR ranked as anatomically close. In row A, a subject with seven scans, the five top hits all belong to the same person. In row B, a subject with three scans, the two other scans come first and the rest of the list is lookalikes from other people. That is the intended order: the same person first, then the closest other brains.
How it works
The first training stage is a variational autoencoder trained with an adversarial critic, using MONAI’s AutoencoderKL. It compresses the whole-brain volume to a latent tensor of shape 8 x 20 x 22 x 26 (8 x 8 x 8 x 8 for a region). Retrieval uses only the encoder’s latent means. Training the VAE on whole brains took about 63 GB of GPU memory at a batch size of one, so the contrastive stage runs on stored latents instead of raw volumes, which allows batches of 128.
The second stage is a small convolutional encoder plus a projection head that maps those latents to a 32-dimensional, unit-length embedding. It is trained with a loss the authors call Multi-Positive Ranking Contrastive Loss. For each anchor scan there are three kinds of neighbours. Hard positives are other scans of the same subject. Soft negatives are different subjects that already look close in the VAE latent space. Hard negatives are different subjects that look far away. The loss pulls hard positives closest, keeps soft negatives at a middle distance, and pushes hard negatives furthest. No diagnosis or age labels are used in training, only subject identity.
At inference the frozen encoder and projector (the Q2E module) embed the query, and cosine similarity against the stored embeddings gives the ranking. A whole-brain volume ends up as 32 floats.
The ablation in Table 3 shows why both negative terms are there. Without the soft-negative term, validation mAP@5 drops from 98.6% to 79.0%. Without the hard-negative term, retrieval accuracy stays high but the agreement between embedding similarity and MS-SSIM image similarity collapses from -0.71 to -0.01, so the ranking of other people’s scans stops tracking how similar they look.
Data
Five public T1w cohorts: ADNI (20,367 scans from 2,389 subjects), OASIS3 (2,642 scans, 1,303 subjects), AIBL (1,276 scans, 685 subjects), MIRIAD (706 scans, 69 subjects), and SLIM (1,015 scans, 571 healthy young adults around 21 years old). Scanners are Siemens, GE, and Philips at 1.5T and 3T. ADNI and OASIS3 were split by subject into roughly 80% training, 10% validation, and 10% test. AIBL, MIRIAD, and SLIM were held out entirely as external test sets.
Apart from SLIM, the population is older adults, cognitively normal or with mild cognitive impairment or Alzheimer’s disease. These cohorts were collected to study ageing and dementia, not tumours, strokes, or post-surgical brains.
How well it finds the same patient
The main benchmark is re-identification. A query counts as a hit when a returned scan belongs to the same subject. On the whole-brain test split, mAP@5 is 98.4%, the top result is correct 97.2% of the time, and at least one correct scan appears in the top five for 99.8% of queries. OASIS3, with one scan per session, is the hardest cohort at 92.5%.
On the same test set, BrainIAC embeddings reach mAP@5 of 72.4%, FOMO2JOMO 80.3%, and a MONAI ResNet-50 10.8%, against 98.3% for NeuroCBIR, which does it with 32 features against 768 for BrainIAC and 256 for FOMO2JOMO. Where differences between scanner vendors or between 1.5T and 3T were statistically detectable, the effect sizes were negligible.
Region search is uneven. Large, well-defined structures behave like the whole brain: the left hippocampus reaches 98.3% mAP@5 on the test split and the cerebral white matter 99.5% or more. Small or homogeneous structures do worse: left amygdala 75.5%, central corpus callosum 62.1%, left pallidum 61.1%, and left accumbens area 40.8%.
What else the embedding carries
The authors also checked whether the neighbours are useful beyond re-identification, without any task training. For brain age, they averaged the ages of the 500 most similar cognitively normal scans, with weights to correct for uneven age sampling. Whole-brain mean absolute error was 4.3 years with a Pearson r of 0.54. In the left hippocampus analysis, queries from Alzheimer’s patients pulled back cognitively normal brains about six years older than the patient.
For diagnosis, a weighted vote over the top 500 neighbours gave a three-class (normal, MCI, Alzheimer’s) balanced accuracy of 56.3%, against 33.3% chance, and 76.7% for normal versus Alzheimer’s. I would treat these numbers as a check that the neighbours carry clinical signal. They are well short of what a diagnostic tool needs.
Speed and hardware
All timings in the paper are on a CPU, an Intel Xeon Gold 6226R. Embedding a whole-brain scan takes 66.4 seconds on one core and 18.7 seconds on four, with peak memory around 6.8 GB. A single region takes 3.6 seconds on one core and 1.2 seconds on four, peaking at 1.4 GB. The top-5 search over the stored embeddings takes under 0.01 seconds. The timing table covers loading the model, running it, and the search. Skull stripping, registration, and segmentation come on top.
Where it falls short
It only handles T1w. FLAIR, T2, diffusion, or contrast-enhanced series would need retraining. Skull stripping is required, partly so the embedding does not learn the face instead of the brain, and the authors acknowledge it is extra time and complexity in a clinical workflow. They also say it is unclear how the model behaves on data preprocessed with a substantially different pipeline.
The benchmark defines “relevant” as “same subject”. That is a clean, label-free target. A radiologist asking for similar cases usually wants similar disease. The zero-shot age and diagnosis numbers suggest the neighbours carry clinical signal, but nobody has measured whether clinicians find the retrieved cases useful. The authors describe the study as technical validation and say clinical evaluation is still needed.
The reference population is mostly older research participants. A query with a glioma, a large infarct, or a resection cavity will still return its nearest neighbours, and those neighbours will come from a population without those findings.
Privacy needs thought too. The property that finds a patient’s prior scan also means a 32-number vector can link a de-identified research scan back to other scans of that person. The repository notes that the shared features cannot identify anyone without authorised access to the source datasets. Anyone building on this should treat the embeddings as identifiable data.
Code, weights, and data
The code is at github.com/minnelab/NeuroCBIR under Apache 2.0. Release v1.1.0 ships the whole-brain and region VAE and contrastive checkpoints plus the precomputed embedding tables (a zip of about 690 MB), and the repository supports Docker, Snakemake with Apptainer, and a pip-installable Python package. The embeddings are keyed to dataset image identifiers, so viewing the actual neighbour images requires access to ADNI, OASIS3, AIBL, MIRIAD, or SLIM under their own terms.
How we would wire it into a viewer
The cost is all on ingest, so I would run it there. When a 3D T1w series lands, convert it to NIfTI, run the preprocessing and the Q2E embedding as a background job, and store one 32-float vector for the whole brain plus one per region in the study database next to the series UID. At 32 floats per embedding, the whole-brain vectors for an entire hospital archive fit in memory, and the search itself is already under 0.01 seconds in the paper.
In the viewer, the feature is a “similar cases” panel: a dropdown for whole brain or a named region, a thumbnail strip of the top matches with their similarity scores and, where the user is allowed to see them, diagnosis and age. The reference set for a clinic should be its own archive rather than the research cohorts, which the authors support, since new scans can be added without retraining.
The second feature is cheaper and probably more useful on day one. It is archive hygiene. The authors point out that NeuroCBIR can flag scans that look like the same person but carry different subject IDs, which they say is not uncommon in ADNI. In a clinical archive the same check would catch merged or split patient records before they reach a radiologist. Before shipping either, we would test it on our own scanners and preprocessing, and keep it away from anything outside the T1w, adult, no-gross-lesion envelope it was trained on.
Sources
- Nieto-del-Amor, F., Fu, J., Muehlboeck, J.-S., Westman, E., Ferreira, D., Moreno, R. NeuroCBIR: A Fast and Accurate Image Retrieval System for Whole-Brain and Region-Specific MRI. arXiv:2610.06502, 2026. https://arxiv.org/abs/2610.06502. PDF: https://arxiv.org/pdf/2610.06502.
- Code, checkpoints, and embeddings: https://github.com/minnelab/NeuroCBIR.
- Featured image: Fig. 5 of arXiv:2610.06502, rows A and B, coronal and axial T1w views of two query scans and their top-ranked retrievals. Frame colours are the paper’s.