Combining machine learning and iterative experiments to keep pace with emerging viral variants of concern
Keywords:
Computer and Data Science, Antibodies, Viral evolution, Epistasis, SARS CoV 2, Neural networks, Machine learning, Virus testing, Binding analysisAbstract
Modeling and predicting viral mutations before they emerge plays a crucial role in pandemic preparedness, enabling the early identification of emerging variants of concern (VOCs) and guiding timely updates to vaccines, diagnostic tests, and therapeutic strategies. However, existing machine learning models and large-scale experiments lose their predictive power as viral variants evolve further from the original strains in sequence space. Here, we present a scalable framework that integrates random forest and neural network machine learning models with targeted high-throughput experimentation to anticipate and evaluate emerging SARS-CoV-2 receptor-binding domain (RBD) variants. Using public datasets, we trained predictive models for binding to human Angiotensin-converting enzyme 2 (ACE2), RBD expression, and antibody escape, and refined these models through iterative integration of experimental data focused on over 200 variants derived from wild-type (WT) and Omicron strains. Through an indirect transfer learning approach, our machine learning models achieved high accuracy having correlation coefficients of up to 0.79 for antibody binding. The models were also generalizable across diverse antibody types including heavy-chain-only antibodies (HCAbs) by encoding complementarity-determining regions (CDRs) as input features. This dynamic approach enables rapid assessment of emerging variants, facilities prioritization of the therapeutic strategies, and supports a proactive, data-driven response to evolving viral threats. Author summary: The COVID-19 pandemic highlighted the threat posed by rapidly evolving viruses. SARS-CoV-2, the virus that causes COVID-19, continues to mutate in ways that can reduce the effectiveness of neutralizing antibodies from vaccines or prior infection. Predicting which viral variants might escape immune detection and identifying antibodies that can still work against them remains a major challenge. In this study, we developed a machine learning framework that helps forecast how mutations in the virus’s spike protein affect its ability to bind to human cells and avoid antibody recognition. We trained our models using large public datasets and then improved them with targeted lab experiments. Instead of treating each dataset separately, we used predictions from public data as features for training new models - a method known as indirect transfer learning. This allowed us to identify and validate antibodies with strong binding to emerging variants. Our approach supports faster, data-driven responses to viral evolution and can be applied to future outbreaks.
Original publication: PLOS Computational Biology (2026-06-17). Source. Source DOI: 10.1371/journal.pcbi.1014394.
Downloads
Published
Issue
Section
License
This article is dedicated to the public domain. You may copy, redistribute, adapt and use it for any lawful purpose without copyright attribution requirements.