Computational Biology at Khoury College of Computer Sciences
Uncovering the secrets of life through the power of computation
Computational biology bridges the gap between computer science and life sciences. Khoury College researchers explore both fundamental and applied questions in all aspects of studies of living systems. Their interdisciplinary approach brings together multiple computer science disciplines, including machine learning, computer vision, statistics, data science, and data visualization, as well as connections to programming languages, mathematics, physics, and network sciences.
Computational biology uses the power of large and complex datasets generated by modern biotechnologies to make complex biological processes legible and to provide insight for life science researchers, clinicians, policymakers, and educators.
Using biological data to explore fundamental questions
Computational biology fuels discoveries across a wide range of fields, including studies of genomes, proteomes, metabolomes, diagnosis and prognosis of disease, discovery of molecular mechanisms of disease, vaccines, and therapeutics, and advances toward personalized medicine. Computational biology is also providing insights into agriculture, climate science, and ecosystem analysis.
Sample research areas
- Causal inference in biomolecular systems
- Computational mass spectrometry-based proteomics
- Macromolecular structure and function
- Network biology and medicine
- Studies of molecular mechanisms of disease
- Variant and genome interpretation
- Systems biology
- Genomics
- Biomedical imaging

Current project highlights
This project aims to support the development, maintenance, and outreach efforts of the Bioregistry, a tool designed to facilitate data integration and standardization in biomedical research. We are developing a semi-automated curation framework to keep pace with the rapid introduction of new semantic spaces, harmonizing inconsistent registries, and developing new ways to access the Bioregistry through different formats, programming languages, and user interfaces. In parallel, we are expanding our training and outreach efforts to engage academic and industry groups, major academic publishers, and funding agencies to promote wider adoption of Bioregistry.
Funded by the Chan Zuckerberg Initiative (CZI) under award 2023-329850. (2023-)
Northeastern’s Barnett Institute for Chemical and Biological Analysis sponsors the national May Institute in computation and statistics for mass spectrometry and proteomics.
This research explores what molecular, clinical, and genetic factors increase the risk of adverse pregnancy outcomes. Using large data sets from pregnant women and the power of machine learning, this research has the potential to make a direct impact on maternal health.
Research publications
A sampling of research papers from the last one to two years, primarily drawn from area conferences and intended to be illustrative; see individual faculty websites (in bios below) for robust publications lists.
A generative deep learning approach to de novo antibiotic design (Cell, 2025)
Authors: Aarti Krishnan, Melis N. Anahtar, Jacqueline A. Valeri, Wengong Jin et al
Researchers developed a generative artificial intelligence framework for designing de novo antibiotics through two approaches: a fragment-based method to comprehensively screen >107 chemical fragments in silico against Neisseria gonorrhoeae or Staphylococcus aureus, subsequently expanding promising fragments, and an unconstrained de novo compound generation, each using genetic algorithms and variational autoencoders.
Eliater: A Python package for estimating outcomes of perturbations in biomolecular networks (Bioinformatics, 2024)
Authors: Sara Mohammad-Taheri, Pruthvi Prakash Navada, Charles Tapley Hoyt, Jeremy Zucker, Karen Sachs, Benjamin M. Gyori, Olga Vitek
This research introduces and showcases Eliater, a Python package for estimating the effect of perturbation of an upstream molecule on a downstream molecule in a biomolecular network.
Beyond protein lists: AI-assisted interpretation of proteomic investigations in the context of evolving scientific knowledge (Nature Methods, 2024)
Authors: Benjamin M. Gyori, Olga Vitek
Mass spectrometry-based proteomics provides broad and quantitative detection of the proteome, but its results are mostly presented as protein lists. Artificial intelligence approaches will exploit prior knowledge from literature and harmonize fragmented datasets to enable mechanistic and functional interpretation of proteomics experiments.
An MSstats workflow for detecting differentially abundant proteins in large-scale data-independent acquisition mass spectrometry experiments with FragPipe processing (Nature Protocols, 2024)
Authors: Devon Kohler, Mateusz Staniak, Fengchao Yu, Alexey I. Nesvizhskii, Olga Vitek
Khoury College is leading work on a new open-source software tool, MSstats, designed for researchers working with quantitative mass spectrometry-based proteomics, a technology used to measure protein levels in samples.
Improving transparency of computational tools for variant effect prediction (Nature Genetics, 2024)
Authors: Rachel Karchin, Pedrag Radivojac, Anne O’Donnell-Luria, Marc S. Greenblatt, Michael Y. Tolstorukov, Dmitriy Sonkin
Efforts to integrate computational tools for variant effect prediction into the process of clinical decision-making are in progress. However, for such efforts to succeed and help to provide more informed clinical decisions, it is necessary to enhance transparency and address the current limitations of computational predictors.