AI protein engineering

Design proteins before the first assay runs

Proteinvue ranks candidate sequences in silico so your wet lab tests the five most likely to work instead of fifty, turning protein engineering into a guided search.

Abstract protein ribbon structure diagram, alpha helices in deep forest green and beta sheets in pale sage, rendered on a near-black background
5-40x fewer sequences to screen before a viable hit, depending on fitness landscape
200M+ UniProt sequences in generative model training corpus
Hours from fitness objective to ranked candidate shortlist
3 modes antibody CDR, nanobody VHH, and industrial enzyme programs supported

Trial-and-error protein engineering is expensive by design.

When a protein engineer defines a fitness objective, the sequence space they need to explore is combinatorially vast. Even with directed evolution tools, teams routinely synthesize and assay dozens of variants before finding one that meets their criteria. Each plate that fails costs time, consumables, and grant-funded runway.

The sequences that would have worked were already there in the fitness landscape, distributed across a high-dimensional manifold that random mutagenesis walks through inefficiently. The field knows this. Most teams are still doing it the slow way because the alternative, a generative model that has actually learned the fitness landscape for their specific target, has been out of reach.

The bottleneck is not the biology, it is the lack of ranked signal before the assay.

Feed in a target. Get back ranked sequences.

Define your fitness objective, specify any known positive control sequences, and let the generative search explore the landscape. The model returns a ranked shortlist with predicted stability and binding scores for each candidate.

01 Define your fitness objective: target protein, activity of interest, and any known positive controls.
02 Generative search samples from the learned fitness landscape, producing a diverse candidate pool.
03 Receive a ranked shortlist with per-residue mutation impact and stability predictions.
Explore the platform

How the scoring model works

The model uses masked language modeling trained on 200M+ UniProt sequences to learn the co-variation structure of protein families. It samples candidate sequences that score well on your fitness objective by iterative refinement of residue probabilities.

Unlike structure-prediction proxies, Proteinvue scores sequence fitness directly from learned sequence-function co-variation, which is more informative than folding geometry alone for most engineering objectives.

Two R&D contexts, one platform

Therapeutic protein design

Antibody CDR optimization, nanobody and VHH engineering, cytokine scaffold design for drug discovery teams running 96-well plate assays. We handle the sequence search so your team focuses on synthesis and assay.

Typical teams see 5 to 8x reduction in screening burden. Your results depend on your fitness landscape.

Learn more

Industrial enzyme engineering

Thermostability-activity trade-off navigation, substrate specificity broadening or narrowing, soluble expression improvement in E. coli or yeast. For bioprocess development scientists at specialty chemical or contract manufacturing sites.

Typical teams see 5 to 8x reduction in screening burden. Your results depend on your fitness landscape.

Learn more
"We do not use structure prediction as a proxy for fitness. We train directly on sequence-function co-variation."

Elena Marchetti, CEO and Co-Founder

The generative model approach

The model uses masked language modeling across 200M+ UniProt sequences to learn how residues co-vary across related protein families. Fitness is inferred from co-variation patterns, not from structure prediction intermediates.

This matters because many engineering objectives, such as binding affinity in a CDR loop or substrate specificity in an active site, are determined by fine sequence-level relationships that folding geometry does not fully capture.

Read our approach

From the design loop

We were running 64-well plates on every CDR variant cycle. Proteinvue cut that to a 6-compound shortlist in the first program we ran it on. The retrospective validation gave us enough confidence to commit without a long pilot phase.

Protein Design Lead, an oncology biologics program

The honest handling of model limitations was what sold me. They ran a retrospective on our held-out data before we committed and told us upfront when coverage for our target family was thin. That kind of transparency is not common in this space.

Enzyme Science Director, specialty industrial biotech

Nanobody engineering is hard to iterate by hand because the scaffold is small and every residue matters. Having a ranked shortlist that accounts for VHH-specific co-variation patterns made our assay queue feel deliberate instead of exploratory.

Computational Biology Lead, a nanobody discovery program

Your next protein hit is already in sequence space.

Get early access while we expand capacity.