Training Data

What UniProt Coverage Means for Your Protein Family (And When It Runs Out)

by Proteinvue

Abstract phylogenetic tree visualization showing protein family distribution across organisms

Every generative protein design model is, at its core, a learned compression of the evolutionary record. What the model knows about which sequences are viable is encoded in the sequences it was trained on. This means the quality of your design run is bounded above by how well your protein family is represented in the training data, which is primarily the UniProt sequence database and its reviewed subset, Swiss-Prot.

Most practitioners understand this in the abstract. Fewer have a practical handle on how to assess training data coverage before starting a design campaign, and what to do when coverage is thin.

How Coverage Translates to Model Behavior

Protein language models and co-evolutionary models both learn statistical regularities from aligned sequences. Those regularities capture two things: which amino acids are tolerated at each position (positional preferences), and which positions co-vary with each other (co-evolutionary coupling). Deeper and more phylogenetically diverse coverage improves the quality of both signals.

When your protein family is well-covered in UniProt, the model can distinguish conserved positions from variable ones, identify which variable positions covary with functional properties, and generate candidate sequences that respect structural constraints without needing explicit structure input. The signals are redundant and mutually reinforcing.

When coverage is thin, positional preferences are based on fewer observations and more likely to reflect sampling bias rather than biology. Co-evolutionary signals become noisy because the statistical test underlying coupling detection requires enough sequence pairs at each position combination to distinguish genuine coupling from coincidence. Generative sampling from a low-coverage family produces sequences that look statistically reasonable but may systematically miss constraints that are simply not captured in the sparse data.

We are not saying that low-coverage families cannot be engineered computationally. We are saying that the confidence interval around any prediction is wider, and you should design your wet-lab validation to account for that.

The Coverage Assessment We Run Before a Campaign

Before we start generating candidates for a new protein family, we run a coverage characterization that answers four questions.

First: how many sequences are in UniProt for this family? For most common enzyme families and antibody frameworks, the answer is thousands to hundreds of thousands. For more specialized families, particularly viral proteins, extremophile-specific enzymes, or recently characterized subfamilies, counts can be in the hundreds or lower. There is no universal threshold below which you should stop, but below roughly 500 sequences we start applying additional scrutiny to every prediction.

Second: what is the phylogenetic breadth of the coverage? A family with 5,000 sequences all from closely related bacteria in one phylum is not better than a family with 800 sequences spread across bacteria, archaea, fungi, and plants. Phylogenetic diversity drives informative variation in the MSA. We use a rough check: how many distinct phyla are represented at better than 1% frequency? Families where one phylum dominates at over 80% should be treated as having coverage quality closer to their within-phylum count than their total count.

Third: what is the sequence identity distribution? The effective number of sequences in an MSA is not the raw count; it is corrected for redundancy. A standard approach clusters at 80% sequence identity and counts clusters rather than sequences. When Swiss-Prot reviewed sequences for a family cluster down to fewer than 50-80 non-redundant entries, the evolutionary signal available is comparable to a shallow MSA even if the raw count looks larger.

Fourth: is coverage uniform across the functional region you are designing in? Families with well-characterized catalytic domains often have patchy coverage of regulatory or linker regions because those regions are less conserved and less studied. If your design targets a loop or subdomain that is underrepresented in the alignments, the coverage deficit is worse than the family-level statistics suggest.

When Coverage Is Good: What You Can Expect

For well-covered families, generative design runs produce candidate lists where the top-ranked sequences have reasonably high probability of expressing correctly and showing measurable activity. In retrospective validations we have run on published deep mutational scanning datasets from well-covered families, rank-correlation coefficients between predicted and observed fitness typically fall in the range of 0.55 to 0.75 for single-point variants, depending on the family and the fitness assay used.

That is not perfect. It means roughly the top third of the ranked list has a good chance of being genuinely better than wild type on the target objective, the bottom third has a good chance of being neutral or worse, and the middle third is where the noise lives. A practitioner who understands this can design their wet-lab round accordingly: synthesize and screen the top 40-60 of a 120-variant ranked list, confirm the rank-correlation with measured data, and use the confirmed data to refine the next round.

When Coverage Is Thin: Practical Adjustments

Thin coverage does not make generative design impossible, but it changes how you should set up the campaign.

One adjustment is to supplement the natural sequence database with structural analogs from related families. If your target protein shares structural homology with a better-covered family, the structural alignment can be used to map positions in the well-covered family onto positions in your thin-covered family. This is not perfect, but it brings in additional co-evolutionary signal that would otherwise be absent.

A second adjustment is to run smaller, more exploratory initial rounds before committing to a refined design strategy. With thin coverage, the confidence in any ranked list is lower. Rather than synthesizing 60 of the top-ranked candidates from round 1, you might synthesize 20-30 with intentional structural diversity within the top-ranked group, use the assay results to confirm that the ranking is informative, and only then commit resources to a larger exploitative round.

A third adjustment applies specifically to new protein families where virtually no UniProt data exists. Here the right tool is often structure-guided design rather than sequence-statistics-driven design. If a high-quality predicted or experimental structure is available, residue-level structural analysis of the active site can guide mutations in the absence of evolutionary co-variation data. The prior is weaker, but at least it is grounded in physical chemistry rather than extrapolating from noise.

The Danger of Treating Coverage as a Yes/No Decision

The most common mistake is treating UniProt coverage as a binary qualifier. Either the family has "enough" coverage to proceed, or it does not. In practice, coverage is a continuous variable that shifts how you interpret every output from the design pipeline.

A well-covered family lets you trust ranked predictions with moderate confidence and design wet-lab rounds based on that confidence. A thin-covered family requires you to hold predictions more loosely, run more exploratory rounds, and update the model more aggressively when wet-lab data comes back. The science of the protein does not change; the epistemics of how much you trust the computational output does change, and that should drive how you structure the campaign from the start.

We include a coverage report as a standard part of the Proteinvue project setup. It takes a few hours to generate and it directly shapes the recommended campaign structure. Skipping it to get to the design step faster is one of the more consistent ways to set up a campaign for frustration in rounds two and three.

Interested in generative protein design?

Start with a free retrospective validation run on your sequence-activity data.

Request validation More articles