RP31 - Preparation for Molecular Tumour Board

For the decision about the further treatment of patients in molecular tumour boards and any other treatment decisions and diagnoses that are or will be based on information at the level of genomic variants, it is currently established practice to consider posterior latent variable distributions, technical observations like sequence reads, estimated impacts on gene products and prior knowledge (e.g. known pathogenicity). Displaying this information can overwhelm or lead to inefficient interaction with the displaying interface. While the 2nd cohort project (RP19), we are currently investigating a novel graph-based approach to better estimate and visualise the impact on gene products that takes haplotypes of subclones into account, our goal in the next funding period is to develop adaptive interfaces to all of the above-listed levels of information while taking the assessor’s background and the context into account. This also entails introducing the ability of the assessor to interact with the interface in natural language, e.g. in a conversational mode. For example, asking questions like “Why is this variant considered to be a germline variant?”, “Why is this variant considered pathogenic?”, “Why is there no variant found in TP53?” or “Why do you think this variant is a sequencing artefact?”. For this purpose, we will develop an approach that automatically encodes all variant-associated information as a knowledge graph. A particular challenge will be encoding uncertainty information from our underlying statistical variant calling model [1], the impact predictions developed in RP19, and haplotyping information inferred in RP06. We strive to integrate those variables along with their posterior distributions into the knowledge representation and develop context and background-specific solutions to communicate and reason over the underlying uncertainty in both natural language and graphical ways. We plan to use rule-based methods, visualisation, and LLMs. In particular, we aim to combine in-context learning [2] with knowledge-graph-to-text approaches [3] and guided output generation [4].


[1] J. Köster, L. J. Dijkstra, T. Marschall, and A. Schönhuth, “Varlociraptor: enhancing sensitivity and controlling false discovery rate in somatic indel discovery,” Genome Biol., vol. 21, no. 1, p. 98, Apr. 2020, doi: 10.1186/s13059-020-01993-6.

[2] Q. Dong et al., “A Survey on In-context Learning,” Jun. 18, 2024, arXiv: arXiv:2301.00234. doi: 10.48550/arXiv.2301.00234.

[3] F. Zhao, H. Zou, and C. Yan, “Structure-aware Knowledge Graph-to-text Generation with Planning Selection and Similarity Distinction,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds., Singapore: Association for Computational Linguistics, Dec. 2023, pp. 8693–8703. doi: 10.18653/v1/2023.emnlp-main.537.

[4] B. T. Willard and R. Louf, “Efficient Guided Generation for Large Language Models,” arXiv.org. Accessed: Sep. 09, 2024. [Online]. Available: https://arxiv.org/abs/2307.09702v4

Next
Previous