Medical image datasets are named public (or application-gated) collections of scans a team can train or test on: TCIA, MIMIC-CXR, OpenNeuro, MURA, and a short list of siblings. This page is that list. It is not a ranked “top resources 2025.” It is not a PYCAD dataset product.
If you meant the tool that draws the labels → image annotation tools. If you meant how you label (box vs mask) → medical image annotation. If you meant one chest-X-ray reporting benchmark → PadChest-GR. If you meant synthetic images → medigan.
PYCAD builds custom web DICOM viewers and the annotation → model → deploy path. It does not host a dataset, and it does not rank TCIA. The datasets & annotation hub is an empty landing (HTTP 200, no body). This URL is the list. /sample-dicom-dataset-library/ 404s.
Eight collections, unranked
Counts below are the numbers the hosts publish. They move. Treat them as order-of-magnitude, not a leaderboard. Access is the real filter: some are a click, some are a CITI course, some are a research application.
| Collection | What it actually is | Access | Catch |
|---|---|---|---|
| TCIA | NCI cancer-imaging archive. DICOM collections, often with clinical / genomic / pathology sidecars. NBIA retriever for bulk pull | Most collections open; some restricted | Oncology-weighted. Quality and labels vary by collection. Not one dataset — a library of them |
| OpenNeuro | Human brain imaging in BIDS: MRI, fMRI, EEG, and siblings | Public download | Neuro, not a CT abdomen dump. BIDS is the point |
| Stanford AIMI | A catalogue of Stanford-shared medical-imaging sets (CheXpert, MURA, and later releases live here or are linked) | Per-dataset agreement | A hub, not one tarball. Read the license on the set you want |
| UK Biobank | Population cohort. Imaging is a subset (brain, cardiac, abdominal MRI, DXA) on an application | Approved researcher access, fees | Not a public dump. Do not list it as “free CT” |
| MIMIC-CXR | PhysioNet chest radiographs plus the free-text report. CheXpert-style structured labels shipped beside it | Credentialing (human-subjects training) | DICOM. One hospital. Labels from text, not a pixel mask |
| MedPix | NLM teaching file: cases with history, findings, a diagnosis | Open | A teaching file, not a training dump. Case text, not a segmentation |
| NIH ChestX-ray14 | Frontal chest X-rays with 14 NLP-derived findings. PNG, not DICOM | Open | Noisy labels. Widely used as a benchmark; do not treat the tags as truth |
| MURA | Upper-extremity musculoskeletal radiographs, study-level normal / abnormal | Research agreement | Study-level label. A study is several images; do not score per PNG against that label |
The old page listed TCIA twice and put MONAI in as item 12. MONAI is a PyTorch framework, not a dataset. Dropped from the table. Use it to load a set you already have a license for.
Finders, not collections
When the organ you need is not in the eight:
- grand-challenge.org — challenge host. The data lives with the challenge. Read the license; a leaderboard is not a redistribution right.
- NCI Imaging Data Commons — cloud-hosted cancer imaging (DICOM, often TCIA-overlapping) you query instead of downloading a truck.
- re3data — a registry of repositories. It does not host pixels.
A public set is not a site hold-out. Scanner, protocol, and population shift. De-id is not a research license to ship the pixels to a third-party annotator. Annotation as a job is the other two URLs above.
What this page is not
- Not a ranking. Dropped “12 best 2025,” Outrank screenshots,
/portfolio, zemith, and a dump of LinkedIn short-links. - Not a PYCAD dataset, annotator, or CRM.
- Not the empty
/datasets-annotation/hub and not the 404 sample-library page.
If the missing piece is a viewer or a model on a set you already have the right to use, that is the imaging piece. Case studies.