Back to publications
Vision-Language ModelsContrastive LearningMedical AI

PatchCLIP enables region specific contrastive health record and image joint training with patch embedding loss

Scientific Reports / 2026 / journal

Bhat S, Mansoor A, Georgescu B, et al.

Patch-based output and evaluation diagram for PatchCLIP
Local visual summary of PatchCLIP patch-level contrastive learning. Nature source

Scientific summary

PatchCLIP extends CLIP-style medical image-text training with patch-level alignment to improve region-specific evidence localization.

Abstract

PatchCLIP adds patch-level contrastive learning so image-text models can localize medical findings, not only classify whole images.

Why it matters

Patch-level localization strengthens the evidentiary value of medical image-text models by connecting global predictions to local image regions.

Contribution

The paper introduces a patch embedding loss for region-specific contrastive image-report training.

Method overview

Patch embeddings are contrasted with text embeddings, generating patch-wise prediction maps during inference in addition to global classification scores.

Key findings

  • PubMed reports state-of-the-art performance across eight chest X-ray abnormality detection tasks.
  • Patch prediction maps reduced false positives at comparable sensitivity compared with saliency maps.
  • The paper includes public code from Siemens Healthineers.

Citation

Bhat S, Mansoor A, Georgescu B, et al.. (2026). PatchCLIP enables region specific contrastive health record and image joint training with patch embedding loss. Scientific Reports.

TODO: Add BibTeX.

Related publications