The search for mutations linked to cancer seems straightforward enough: pinpoint changes in the DNA compared with normal cells. But defining what is normal can differ from person to person, with each individual potentially carrying unique but otherwise benign genetic variants.
To account for this individuality, the computational process of flagging mutations in sequencing data, known as somatic variant calling, works best when tumour sequences can be compared with matched normal samples from patient blood or non-tumour tissue. This is crucial for measuring tumour mutation burden (TMB)—the total number of genetic errors within a given amount of DNA—which can help inform treatment decisions.
However, matched controls are not easy to come by, as routine clinical testing and biobanks often collect or preserve only tumour tissue. “Without a ‘normal’ DNA sample to act as a baseline, finding a true cancer mutation can be like trying to spot a needle in a haystack,” said Anders Skanderup, Associate Director at the A*STAR Genome Institute of Singapore (A*STAR GIS).
To distinguish true mutations from noise, Skanderup and colleagues developed VarNet-T, a deep learning framework for identifying cancer-associated variants in tumour sequencing data without requiring a matched normal sample.
The team trained VarNet-T using data from over 300 pairs of tumour and normal tissue samples. These tumour-normal pairs were used to help label the data, much like an answer key, but only the tumour sequences were fed into the model. “Real tumour samples contain complex biological ‘noise’ and unique error patterns that the model must experience firsthand to perform accurately in actual clinical scenarios,” said Kiran Krishnamachari, A*STAR GIS Scientist and first author of the study.
This use of real clinical samples proved advantageous: VarNet-T outperformed other tumour-only models trained on cell line or computer-simulated data, producing TMB estimates closest to those obtained using both tumour and matched normal samples.
“By more accurately calculating this mutation burden, our method can act as a sharper lens that helps doctors correctly identify which patients are the best candidates for life-saving immunotherapy treatments,” said Skanderup.
Moreover, the A*STAR team’s model achieved 88 percent accuracy in classifying TMB-high tumours, when tested on 1000 samples spanning 10 solid cancer types, compared with 55 percent for the next-best tumour-only method. VarNet-T also achieved a misclassification rate of five percent, around three times lower than that of the next-best model.
“We are refining VarNet-T to handle even more complex genetic anomalies and adapting it to work seamlessly with different types of samples,” Skanderup said, adding that testing across more diverse, real-world patient populations will be important for further development.
The researchers have since filed a patent for the approach and are exploring commercialisation opportunities to bring VarNet-T closer to clinical use.
The A*STAR-affiliated researchers contributing to this research are from the A*STAR Genome Institute of Singapore (A*STAR GIS).