MammoBLIP: End-to-End Mammography Report Generation with Vision-Language Models and Public Multi-Institutional Datasets
IEEE Medical Imaging Conference (MIC), Yokohama, Japan / 2025 / conference
Bhandary Panambur A, Wind S, Bayer S, Maier A
Scientific summary
MammoBLIP investigates end-to-end mammography report generation by aligning breast images with standardized radiology-style language.
Abstract
MammoBLIP is an end-to-end mammography report generation framework built on MedBLIP-style vision-language modeling and a curated multi-source public dataset.
Why it matters
Structured report generation provides a research path toward consistent multimodal breast imaging assistants and reduced documentation burden.
Contribution
The work curates 81,076 mammography images from five public sources and trains a report-generation pipeline for structured mammography text.
Method overview
Images are encoded with a frozen EVA-CLIP Vision Transformer, aligned with text embeddings through a lightweight transformer, and used to condition BioMedLM report generation.
Key findings
- The FAU-hosted manuscript reports overall BLEU 65.36, ROUGE-1 0.75, BERT-F1 0.88, and SBERT similarity 0.91.
- Training used standardized data from VinDR-Mammo, RSNA, CMMD, InBreast, and KAU.
- Only the transformer and projection heads were trained while the major backbone models remained frozen.
Citation
Bhandary Panambur A, Wind S, Bayer S, Maier A. (2025). MammoBLIP: End-to-End Mammography Report Generation with Vision-Language Models and Public Multi-Institutional Datasets. IEEE Medical Imaging Conference (MIC), Yokohama, Japan.
TODO: Add BibTeX.