How it works

One input. A ranked shortlist of sequences worth testing.

Define your fitness objective, run the search, receive candidates with predicted residue-level stability and binding annotations.

Three steps: objective to ranked shortlist

Define your fitness objective

Specify your target protein, the activity of interest, and any known positive controls from prior assays. Prior sequence-activity data as a CSV and a target structure in PDB format are optional inputs that sharpen the search. No special bioinformatics tooling required on your side: plain FASTA is enough to start.

Generative search explores sequence space

The model samples from the learned fitness landscape for your protein family, using masked language modeling trained on 200M+ UniProt sequences. Iterative refinement of residue probabilities generates a diverse candidate pool, each scored against your objective. The search focuses on regions of sequence space where evolutionary co-variation indicates functional viability, not random mutagenesis walks.

Ranked shortlist with annotations

Receive a score-sorted candidate list in FASTA and CSV, with per-residue mutation impact annotations and predicted stability delta relative to your reference. Each candidate includes a model confidence note for your protein family's training coverage. You test the top five, not fifty.

Design objectives the platform handles

Improve binding affinity or reduce immunogenicity in CDR loops without full re-screening from scratch. Submit a lead antibody or nanobody sequence, define whether the objective is affinity maturation, developability improvement, or both, and receive ranked CDR variants informed by learned co-variation across immunoglobulin and VHH families.
Navigate the thermostability-activity trade-off by sampling sequence regions where the model has learned that thermostable variants in evolutionary relatives preserved or enhanced catalytic function. This does not guarantee simultaneous gains in both properties: training coverage for your enzyme family determines whether useful signal exists. The retrospective validation tells you before you commit to a full program run.
Generate candidates with simultaneous mutations at multiple positions, covering fitness landscape regions that single-residue walks cannot reach efficiently. Useful when your engineering objective requires epistatic combinations.
Receive a mutation matrix visualization showing predicted fitness impact per residue position. This tells your team which regions of the sequence are most constrained by co-variation, where substitutions are tolerated, and where the model has low confidence due to thin evolutionary coverage. Use it to guide your next synthesis round before the assay results are even back.
Abstract molecular surface map showing protein binding pocket in forest green tones with amber highlight region indicating optimized residues

Standard formats. No custom bioinformatics setup.

What you send us

  • Reference sequence (FASTA), required
  • Prior sequence-activity pairs (CSV), optional, sharpens the search
  • Target structure (PDB), optional
  • Known positive control sequences, optional

What we deliver back

  • Ranked candidate FASTA, score-sorted
  • Per-residue score CSV with mutation impact annotations
  • Mutation matrix visualization (PDF or SVG)
  • Model coverage note for your protein family, plus a direct conversation about what it means
Request access

Start with a free retrospective validation.

Send us a held-out sequence-activity dataset. We run the model and return a rank correlation report within 5 business days. No payment required.