Publications
Filtered · 2
A scalable multimodal framework for unbiased risk biomarker discovery across multiple cancer types
Constantin Petrescu · Lisa Schmunk · Jack Monahan · Abbas Salami · Tom Stubbs, PhD
Background: Most existing cancer risk models are built on single modalities and hand-selected features. Systematic, unbiased integration of germline genetics, plasma proteomics, and deep clinical phenotyping holds promise for revealing novel risk biomarkers across diverse cancer types. Methods: We developed a multi-modal biomarker discovery engine that can be used for discovering risk, diagnostic, prognostic, predictive and monitoring biomarkers. Currently the framework handles: • Germline genetics and polygenic risk scores • High-dimensional plasma proteomics (Olink) • Longitudinal primary-care records, hospital episodes, laboratory results, lifestyle questionnaires, and cancer registry linkages Key design features include modular cohort handling, automated data preprocessing, and machine-learning models (including: gradient boosting and neural networks). Application: The platform is currently deployed on the UK Biobank (n = 502,505 participants; >46,000 incident cancers across 22 cancer types) with active model training and biomarker discovery in progress. The architecture is cohort-agnostic and ready for direct application to emerging large-scale resources including Our Future Health and the All of Us Research Program. Poster presentation: We will demonstrate the platform’s configurability through examples of cancer-risk modelling in the UK Biobank, showcasing: (i) comparative performance of individual modalities versus multimodal ensembles, (ii) cancer-specific patterns of modality contribution, and (iii) the effect of time-window filtering on separating true predictive signals from prevalent disease effects. Conclusions: By eliminating bias in feature engineering and supporting seamless integration of diverse health data streams, this scalable framework provides a robust foundation for data-driven discovery of multimodal cancer risk biomarkers, paving the way for next-generation precision prevention strategies.
Cancer Research 86 (7_Supplement): Abstract 1116 · 2026 · doi.org/10.1158/1538-7445.AM2026-1116
PresentationOn the epigenetic ageing clock in humans(opens in a new tab)
Dani Martin-Herranz, PhD
Epigenetic clocks are mathematical models that predict the biological age of an organism using DNA methylation data, and which have emerged in the last few years as the most accurate biomarkers of the ageing process. However, little is known about the molecular mechanisms that control the rate of such clocks. In this thesis I focus on the study of the epigenetic ageing clock in humans. First, I review and benchmark statistical and computational tools required for the analysis of DNA methylation data in the context of human ageing. Next, I validate the performance of the Horvath epigenetic clock, the most widely used multi-tissue epigenetic clock in humans, in a control blood dataset and test its behaviour in patients with a variety of developmental disorders, which harbour mutations in proteins of the epigenetic machinery. I demonstrate that loss-of-function mutations in the H3K36 methyltransferase NSD1, which cause Sotos syndrome, substantially accelerate epigenetic ageing. Furthermore, I show that the normal ageing process and Sotos syndrome share methylation changes and the genomic context in which they happen. These results suggest that the H3K36 methylation machinery is a key component of the epigenetic ... (continues)
2019 · repository.cam.ac.uk
Thesis
29 publications · 22 peer-reviewed · 4 preprint · 1 thesis · 1 presentation · 1 other