
Abstract
This multimodal vision-language transformer classifies breast cancer from a single thermal image by fusing thermal features with structured clinical metadata encoded as text prompts. Cross-attention blocks combine patient-specific cues with image tokens, enabling robust predictions across different viewpoints. Experiments on the DMR-IR and Breast Thermography datasets show strong accuracy and sensitivity, outperforming visual-only and prior multimodal baselines.
ColCACI






