PatchCLIP enables region specific contrastive health record and image joint training with patch embedding loss
Scientific Reports / 2026 / journal
Bhat S, Mansoor A, Georgescu B, et al.
Scientific summary
PatchCLIP extends CLIP-style medical image-text training with patch-level alignment to improve region-specific evidence localization.
Abstract
PatchCLIP adds patch-level contrastive learning so image-text models can localize medical findings, not only classify whole images.
Why it matters
Patch-level localization strengthens the evidentiary value of medical image-text models by connecting global predictions to local image regions.
Contribution
The paper introduces a patch embedding loss for region-specific contrastive image-report training.
Method overview
Patch embeddings are contrasted with text embeddings, generating patch-wise prediction maps during inference in addition to global classification scores.
Key findings
- PubMed reports state-of-the-art performance across eight chest X-ray abnormality detection tasks.
- Patch prediction maps reduced false positives at comparable sensitivity compared with saliency maps.
- The paper includes public code from Siemens Healthineers.
Citation
Bhat S, Mansoor A, Georgescu B, et al.. (2026). PatchCLIP enables region specific contrastive health record and image joint training with patch embedding loss. Scientific Reports.
TODO: Add BibTeX.