This presentation explores how population-scale data can inform the prediction, prevention and survival of cancer. It highlights the use of machine learning to screen thousands of potential risk factors, identify previously overlooked predictors and complement conventional epidemiological approaches. It also discusses gene–environment interactions, and how inherited susceptibility may combine with modifiable exposures such as lifestyle, nutrition and environmental factors to influence cancer risk. Together, these approaches demonstrate how large biobanks, machine learning and genetic epidemiology can help move cancer research towards earlier detection and support strategies for cancer prevention and improved survival