RP39 - 3D Multi-Modal Vision Language Models (MMVLMs)

By combining computer vision models’ perceptual abilities with LLMs’ generative capabilities, multi-modal vision language models (MMVLMs) have claimed improved performance in various multimodal tasks [1]. This synergy has captured significant attention from researchers, who recognise the vast potential of MMVLMs in applications such as image-text retrieval, report generation, and visual question answering. Despite their success, existing MMVLMs are often fine-tuned on publicly available 2D data and evaluated on tasks that do not fully capture the complexities of real-world medical settings. The widespread use of 3D medical images, such as computed tomography (CT) and magnetic resonance imaging (MRI), poses a significant challenge for these models [2]. As a result, existing MMVLMs often struggle to analyse and interpret 3D medical images effectively. In response to this challenge, our project builds upon RP21 and aims to develop MMVLMs with 3D capabilities that can effectively integrate image and text data from medical imaging modalities. We will design and implement novel architectures that capture spatial information in 3D medical images by leveraging state-of-the-art techniques such as self-supervised learning and transfer learning. To further enhance the performance of our proposed models, we will incorporate medical knowledge of spatial anatomic relationships and metastases spread routes. This domain-specific knowledge will be encoded using medical ontologies (e.g., RadLex) and expert annotations [3], enabling our models to understand better the complex relationships between anatomical structures and disease progression. For example, melanoma, but also other cancers such as colorectal cancer, have a high risk of causing metastasis in the liver. This critical piece of information can inform diagnosis and treatment. A key aspect of our proposed models is the ability to establish links between language representations and image segmentations. By grounding the results of language in the images, we can enable confirmability and trustworthiness of generated descriptions. For evaluation, we will conduct extensive experiments using real-world data sets.


[1] J. N. Acosta, G. J. Falcone, P. Rajpurkar, and E. J. Topol, “Multimodal biomedical AI,” Nat. Med., vol. 28, no. 9, pp. 1773–1784, Sep. 2022, doi: 10.1038/s41591-022-01981-2.

[2] J. Egger et al., “Medical deep learning—A systematic meta-review,” Comput. Methods Programs Biomed., vol. 221, p. 106874, Jun. 2022, doi: 10.1016/j.cmpb.2022.106874.

[3] Z. Marinov, P. F. Jäger, J. Egger, J. Kleesiek, and R. Stiefelhagen, “Deep Interactive Segmentation of Medical Images: A Systematic Review and Taxonomy,” IEEE Trans. Pattern Anal. Mach. Intell., pp. 1–20, 2024, doi: 10.1109/TPAMI.2024.3452629.

Next