Researchers at the University of California, Berkeley released GPN-Star, a genomic language model that encodes species trees and whole-genome alignments to predict which DNA variants are likely functional or disease-related.

Why phylogeny matters

Earlier genomic transformers treated DNA like text, requiring huge compute while sometimes losing to classical evolutionary scores. GPN-Star embeds alignment structure directly, training separate checkpoints on vertebrate, mammalian, and primate timescales.

In Nature on September 9, the team reported state-of-the-art results on coding and noncoding variant tasks, including fine-mapped genome-wide association hits tied to schizophrenia and other complex traits.

Clinical and research uses

Berkeley News emphasized efficiency: GPN-Star reaches strong performance with roughly 200 million parameters on human references, a fraction of the largest DNA models. Labs studying rare disease can prioritize variants faster before running costly functional assays.

The Song Lab published genome-wide prediction tracks on GitHub, inviting biologists to overlay annotations on their cohorts. Models for mouse, chicken, fly, worm, and Arabidopsis show the architecture travels beyond humans.

Timescale trade-offs

Deeper evolutionary training helped on conserved protein mutations, while primate-focused training better predicted common complex-trait variants in noncoding DNA. That split guides clinicians on which checkpoint to consult.

Heritability analyses showed unusually strong enrichment signals, suggesting the model captures regulatory constraints previous scores missed.

Ethics and access

Because weights are public, hospitals must still validate predictions in diverse populations. Ancestry gaps in alignment data could skew scores if applied blindly in the clinic.

Funding from the National Institutes of Health underpinned the work, keeping academic access open while commercial genomics firms evaluate licensing.

GPN-Star will not replace functional experiments, but it gives geneticists a sharper map of where to look when a patient’s exome looks inconclusive.

Collaboration plans

UC San Francisco clinicians partnering through the Bakar Computational Biomedicine Initiative said they would pilot GPN-Star scores in undiagnosed disease clinics this fall, comparing predictions with existing CADD and phyloP tracks.

Cloud vendors hosting alignment files reported increased downloads of vertebrate whole-genome alignments as labs seek to reproduce training data.

Lead author Yun S. Song told Berkeley News that publishing genome-wide scores should accelerate hypothesis generation for biologists who lack machine learning staff. Teams can overlay GPN-Star tracks on single-cell atlases to see whether regulatory variants coincide with cell-type-specific expression.

Competing labs praised the open release but cautioned that pathogenicity thresholds still require calibration on diverse cohorts, especially for variants rare in European reference panels.

Investors in genomics startups said GPN-Star’s efficiency could lower the cost of population-scale sequencing interpretation, a pitch that matters as biobanks add millions of new samples.

Teaching hospitals said they will add GPN-Star tracks to resident training on variant interpretation, giving junior clinicians a shared vocabulary with bioinformatics teams.