Notes from the design loop.

Protein engineering methods, sequence co-variation intuitions, and what we have learned from watching wet-lab teams use generative search in real programs.

Abstract 3D fitness landscape topographic surface in forest green tones

How to Read a Fitness Landscape Before You Design a Single Variant

Before generating candidates, understanding the shape of the fitness landscape tells you whether single-residue walks will work or whether you need combinatorial diversity.

Abstract comparison of two protein scaffolds rendered as glowing ribbon structures

Nanobody vs. Antibody CDR Engineering: Where Generative Search Helps Most

The smaller scaffold of nanobodies makes sequence space coverage different in ways that affect how you should set up a generative design run.

Abstract dual-axis scatter plot concept visualizing the thermostability-activity trade-off in deep forest green and amber

Thermostability and Activity: Why You Rarely Get Both, and When You Can

The classic trade-off between thermal stability and catalytic activity has structural roots. Generative models can sometimes find sequences that thread the needle, but not always.

Abstract visualization of rank-correlation between predicted and observed protein activity scores

Retrospective Validation: What Rank Correlation on Your Own Data Actually Tells You

Before committing to any computational design tool, running a retrospective on held-out data is the only honest quality check. Here is how to interpret the result.

Abstract grid of colored tiles representing a sequence co-variation contact map

A Primer on Sequence Co-variation for Protein Engineers Who Are Not ML Researchers

Evolutionary co-variation in protein sequences encodes structural and functional constraints. Understanding it helps you interpret why a model ranks certain mutations favorably.

Abstract visualization of an enzyme active site pocket with substrate molecule docking

Broadening Substrate Specificity in Industrial Enzymes: A Case Study Framing

When a specialty chemical process needs an enzyme to accept a slightly different substrate, the design problem is narrower than therapeutic engineering but the stakes are equally high.

Abstract circular feedback loop diagram showing dry-lab computation and wet-lab validation cycles

Closing the Wet-Lab Dry-Lab Loop: How Fast Should the Feedback Cycle Be?

The latency between an assay result and the next computational design run determines how many iterations you can afford. We think about this as a scheduling problem.

Abstract phylogenetic tree visualization showing protein family distribution across organisms

What UniProt Coverage Means for Your Protein Family (And When It Runs Out)

A generative model trained on public sequence databases performs well when your target protein family has broad evolutionary representation. Here is how to check before you start.

Abstract overhead view of a protein engineering lab bench with pipettes and microplates in muted scientific tones

Why We Started Proteinvue

Elena and Marcus describe the specific moment in a protein engineering lab that made it obvious a computational shortlist tool had to exist.

Abstract multiple sequence alignment grid showing deep and shallow coverage zones

MSA Depth and Model Confidence: How Many Homologs Do You Actually Need?

Multiple sequence alignment depth is a proxy for how much evolutionary signal a model can draw on. Shallow MSAs degrade confidence in novel regions of sequence space.

Abstract target-and-arrow concept showing a fitness objective being defined before protein design begins

What Do We Mean by Fitness Objective, and How Do You Define Yours?

The single most important input to a generative protein design run is a clear definition of what you are optimizing for. Getting this wrong is more expensive than any model limitation.