Back to publications
MammographyVision-Language ModelsReport Generation

MammoBLIP: End-to-End Mammography Report Generation with Vision-Language Models and Public Multi-Institutional Datasets

IEEE Medical Imaging Conference (MIC), Yokohama, Japan / 2025 / conference

Bhandary Panambur A, Wind S, Bayer S, Maier A

Presentation on MammoBLIP at IEEE Medical Imaging Conference 2025
Presentation photo from IEEE MIC 2025. The paper PDF contains the MammoBLIP pipeline overview. FAU PDF

Scientific summary

MammoBLIP investigates end-to-end mammography report generation by aligning breast images with standardized radiology-style language.

Abstract

MammoBLIP is an end-to-end mammography report generation framework built on MedBLIP-style vision-language modeling and a curated multi-source public dataset.

Why it matters

Structured report generation provides a research path toward consistent multimodal breast imaging assistants and reduced documentation burden.

Contribution

The work curates 81,076 mammography images from five public sources and trains a report-generation pipeline for structured mammography text.

Method overview

Images are encoded with a frozen EVA-CLIP Vision Transformer, aligned with text embeddings through a lightweight transformer, and used to condition BioMedLM report generation.

Key findings

  • The FAU-hosted manuscript reports overall BLEU 65.36, ROUGE-1 0.75, BERT-F1 0.88, and SBERT similarity 0.91.
  • Training used standardized data from VinDR-Mammo, RSNA, CMMD, InBreast, and KAU.
  • Only the transformer and projection heads were trained while the major backbone models remained frozen.

Citation

Bhandary Panambur A, Wind S, Bayer S, Maier A. (2025). MammoBLIP: End-to-End Mammography Report Generation with Vision-Language Models and Public Multi-Institutional Datasets. IEEE Medical Imaging Conference (MIC), Yokohama, Japan.

TODO: Add BibTeX.

Related publications