Close-up of a DNA Strand

In FY2026, the Center for Precision Medicine and Data Sciences advanced a tightly integrated portfolio spanning AI-enabled drug discovery, computational cardiology, protein design, clinical informatics, and precision-medicine software. Eight contributing researchers and trainees produced peer-reviewed science, publicly available tools, competitive grant activity, and a growing network of national and international partnerships.

Signature Accomplishments

  • 11 articles published, accepted, or in press — in journals including Nucleic Acids Research, Nature Reviews Methods Primer, The Journal of Precision Medicine: Health and Disease,  eLife, Journal of General Physiology, Pharmacological Reviews, and the Journal of Physiology.
  • Two precision-medicine platforms launched publicly — CATVariant (Nucleic Acids Research) for integrated variant interpretation and BoltzOmics (iScience) for AI-based drug-binding prediction.
  • Digital twins of human excitable cells published at eLife — a machine-learning framework for personalized cardiac electrophysiology.
  • Sustained NIH training investment, including two T32 fellowships (pre- and post-doctoral) and an active federal grant pipeline.
  • Collaborations across five countries — New Zealand, Spain, the Netherlands, Australia, and India — plus a digital-health book developed with the World Health Organization and the Government of India.
  • 10+ open tools, models, and datasets released, alongside the GPU/computing backbone powering it all.
  • 13+ talks, seminars, and posters and 7+ students mentored across graduate, undergraduate, and pre-med training.
  • Built and maintained the Center’s GPU/server computing backbone and released open research software with 80+ publication-ready analysis workflows.

Bottom line: FY2026 demonstrates a productive, well-networked Center converting computational and AI methods into published science, deployed tools, funded training, and translational impact — a strong foundation for grant renewals and continued institutional investment.

CATVariant graphical abstract

Every person carries small differences in their DNA. Some of these differences, often called variants, do very little. Others can change the amino-acid sequence of a protein, which may affect how that protein is built, folded, moved inside the cell, or how well it does its job. Because proteins carry out many of the body’s essential functions, even a small change can sometimes have important biological or medical consequences.

The challenge is figuring out which changes matter. Researchers often need to weigh many different clues, including whether a variant is rare or common in human populations, whether it has been reported in people with disease, whether it falls in an important or highly conserved part of the protein, whether it may change the protein’s shape or nearby interactions, what experiments have measured, and what the scientific literature says. Each type of evidence has strengths and caveats, and no single source is usually enough on its own. In practice, this often means moving between many separate databases and analysis tools, then manually piecing together a fragmented trail of evidence to decide whether a variant is likely to be harmless, disruptive, or still uncertain.

CATVariant was created to make that process easier and more informative. The platform uses automated data mining to retrieve and organize variant-related evidence from genetic variant databases, protein resources, population datasets, experimental assay collections, disease and pharmacology knowledge bases, and the scientific literature. It then goes further by mapping variants onto the protein sequence and available protein models, comparing them with known functional regions and nearby reported changes, and analyzing broader patterns such as mutation-sensitive regions, structural clusters, and residue connections across the protein. The result is an interactive report that helps users move from a broad protein-level view to detailed review of individual variants without manually stitching the evidence together across multiple resources.

CATVariant is especially useful when direct laboratory or clinical evidence is limited, which is true for many variants. The platform brings together a broad set of computational predictors, with 12 directly surfaced predictor or effect-estimation inputs, and interprets them alongside the rest of the evidence rather than in isolation. These models draw on different kinds of biological signal, including evolutionary conservation, protein sequence patterns, biochemical context, protein shape, and RNA splicing. Because the models capture different signals, CATVariant lets users see where the computational evidence agrees, where it conflicts, and how those predictions line up with structural, population, experimental, and literature evidence.

In short, CATVariant is designed to help researchers turn scattered clues into testable ideas about how a genetic change might affect protein function. The platform is open access and free to use.