We were running 64-well plates on every CDR variant cycle. Proteinvue cut that to a 6-compound shortlist in the first program we ran it on. The retrospective validation gave us enough confidence to commit without a long pilot phase.
AI protein engineering
Design proteins before the first assay runs
Proteinvue ranks candidate sequences in silico so your wet lab tests the five most likely to work instead of fifty, turning protein engineering into a guided search.
Trial-and-error protein engineering is expensive by design.
When a protein engineer defines a fitness objective, the sequence space they need to explore is combinatorially vast. Even with directed evolution tools, teams routinely synthesize and assay dozens of variants before finding one that meets their criteria. Each plate that fails costs time, consumables, and grant-funded runway.
The sequences that would have worked were already there in the fitness landscape, distributed across a high-dimensional manifold that random mutagenesis walks through inefficiently. The field knows this. Most teams are still doing it the slow way because the alternative, a generative model that has actually learned the fitness landscape for their specific target, has been out of reach.
The bottleneck is not the biology, it is the lack of ranked signal before the assay.
Feed in a target. Get back ranked sequences.
Define your fitness objective, specify any known positive control sequences, and let the generative search explore the landscape. The model returns a ranked shortlist with predicted stability and binding scores for each candidate.
Scores represent predicted fitness relative to your input objective. Illustrative only.
How the scoring model works
The model uses masked language modeling trained on 200M+ UniProt sequences to learn the co-variation structure of protein families. It samples candidate sequences that score well on your fitness objective by iterative refinement of residue probabilities.
Unlike structure-prediction proxies, Proteinvue scores sequence fitness directly from learned sequence-function co-variation, which is more informative than folding geometry alone for most engineering objectives.
Two R&D contexts, one platform
Therapeutic protein design
Antibody CDR optimization, nanobody and VHH engineering, cytokine scaffold design for drug discovery teams running 96-well plate assays. We handle the sequence search so your team focuses on synthesis and assay.
Typical teams see 5 to 8x reduction in screening burden. Your results depend on your fitness landscape.
Learn moreIndustrial enzyme engineering
Thermostability-activity trade-off navigation, substrate specificity broadening or narrowing, soluble expression improvement in E. coli or yeast. For bioprocess development scientists at specialty chemical or contract manufacturing sites.
Typical teams see 5 to 8x reduction in screening burden. Your results depend on your fitness landscape.
Learn more"We do not use structure prediction as a proxy for fitness. We train directly on sequence-function co-variation."
Elena Marchetti, CEO and Co-Founder
The generative model approach
The model uses masked language modeling across 200M+ UniProt sequences to learn how residues co-vary across related protein families. Fitness is inferred from co-variation patterns, not from structure prediction intermediates.
This matters because many engineering objectives, such as binding affinity in a CDR loop or substrate specificity in an active site, are determined by fine sequence-level relationships that folding geometry does not fully capture.
Read our approachFrom the design loop
The honest handling of model limitations was what sold me. They ran a retrospective on our held-out data before we committed and told us upfront when coverage for our target family was thin. That kind of transparency is not common in this space.
Nanobody engineering is hard to iterate by hand because the scaffold is small and every residue matters. Having a ranked shortlist that accounts for VHH-specific co-variation patterns made our assay queue feel deliberate instead of exploratory.
Your next protein hit is already in sequence space.
Get early access while we expand capacity.
Or write to us directly: [email protected]