Therapeutics

Nanobody vs. Antibody CDR Engineering: Where Generative Search Helps Most

by Proteinvue

Abstract comparison of two protein scaffolds rendered as glowing ribbon structures

When you move from antibody CDR engineering to nanobody CDR engineering, the sequence space changes in ways that are not simply a scaled-down version of the antibody problem. The smaller scaffold, the absence of the light chain, and the different role of CDR3 all create a distinct set of challenges and opportunities for generative design. Getting these differences wrong at the campaign setup stage leads to poorly calibrated runs and wasted screens.

This article is for computational biologists and protein engineers who already work with both modalities, not an introduction to what nanobodies are. We want to focus on the practical differences in how generative sequence search operates across the two scaffolds.

The Sequence Space Is Smaller but Not Simpler

A camelid single-domain antibody (VHH, or nanobody) has a total length of roughly 110 to 125 amino acids. A conventional IgG VH domain is similar in length, but the nanobody has a longer CDR3 on average, it lacks the VL pairing interface, and it has characteristic substitutions in the former VH-VL interface positions (usually Phe37 to Tyr, Gly44 to Glu, Leu45 to Arg, Trp47 to Gly in the IMGT numbering). These framework substitutions are what allow the VHH to be stable without a light chain partner.

For generative sequence design, this matters because the sequence constraints are distributed differently. In a conventional VH, a significant fraction of the framework positions are structurally constrained by the VL interface, and your CDR diversity is shaped in part by what the framework can accommodate in the context of a paired domain. In a VHH, all of the structural context comes from a single chain. CDR3 can be longer, more flexible, and it can form extended loops that reach into protein cavities, which is one reason nanobodies access epitopes that conventional antibodies cannot.

But longer CDR3 loops with higher intrinsic flexibility are harder for generative models to sample accurately, because there is less co-evolutionary signal about long, variable loops compared to shorter, conserved ones. A generative model trained on a broad antibody dataset will have good signal on CDR1 and CDR2 and weaker signal on CDR3 diversity, and this effect is amplified for VHH because the CDR3 length distribution is wider and the functional modes include structural elements like disulfide-bridged loops that are rare in human VH sequences.

Where the Models Perform Differently

Sequence language models pre-trained on the full protein universe (ESM-class models) treat a VHH and a conventional VH with similar evolutionary context, because both are immunoglobulin folds. The single-chain nature of VHH does not fundamentally change how these models encode structural plausibility.

Antibody-specific models (trained on paired or unpaired repertoire data) carry a different bias. If the training distribution is predominantly human IgG, the model has seen many VH sequences but relatively few VHH sequences, and the VHH-specific framework substitutions at the former interface positions may score as unusual. Running a human-antibody-trained model on a VHH campaign can produce scores that penalize sequences precisely because they are VHH-appropriate, not because they are structurally unstable.

We have seen this in practice when evaluating third-party models on VHH engineering tasks. Sequences that are clearly within the camelid VHH repertoire score lower than equivalent human VH sequences because the model learned the human distribution. The fix is either a VHH-specific fine-tune or, at minimum, a calibration step where you provide the model with a set of well-characterized VHH sequences and confirm that the scoring distribution aligns with known functional variants before you use it for a campaign.

CDR3 Length and Generative Diversity

In human antibodies, CDR3 length in the VH domain ranges roughly from 3 to 28 amino acids with the distribution peaking around 12 to 14 (IMGT). Camelid VHH CDR3 ranges from roughly 6 to 24 residues with a flatter distribution and many functional binders at lengths above 17 residues. Some published VHH sequences have CDR3 loops that form protruding beta-hairpins or disulfide-bridged stalks with a tip-loop architecture, enabling cavity and groove engagement that short CDR3 loops cannot access.

When you set up a generative design run for a VHH campaign, constrain CDR3 length sampling to the VHH-appropriate range for your target class. If you are targeting a surface epitope with a relatively shallow binding interface, CDR3 lengths of 14 to 17 residues will cover most of the useful space. If you are targeting a cavity or channel-like structure, you may want to weight the generation toward longer CDR3s, 18 to 22 residues, and ensure the model can produce disulfide-constrainable loops if your expression system can support them.

Do not simply port the CDR3 length distribution from your VH design experience into a VHH campaign. The functional space is not equivalent.

Framework Stability and Expression

One area where nanobody engineering is genuinely easier than conventional antibody engineering is single-chain stability. VHH scaffolds express well in bacterial systems (E. coli periplasm, yeast display) precisely because they do not require a paired chain for stability. The intrinsic Tm values of well-engineered VHH frameworks are typically in the range of 65 to 80 degrees Celsius, which is high for a single-domain binding protein of this size.

This means thermostability is usually not the primary optimization target in a VHH campaign the way it can be in enzyme engineering. The CDR design problem is predominantly about binding affinity, selectivity, and off-rate, with expression and stability as constraints rather than objectives. Generative runs for VHH can therefore focus selection on CDR sequence quality relative to the target structure rather than on framework stabilization.

That said, humanization is a real issue for therapeutic VHH. Camelid-specific framework residues are immunogenic in humans. The framework substitutions at the former VL interface positions need to be handled carefully: naively humanizing them back toward human VH consensus can destabilize the VHH because those positions are structural adapters for single-chain context, not just VL interface residues. Generative models that have seen VHH humanization datasets handle this more gracefully than those that have not.

Practical Setup Differences for Generative Runs

When we set up a VHH generative design campaign versus a VH campaign, the main differences in our approach are as follows.

For VHH, we assess the training representation of VHH sequences in whatever base model we are using before accepting its scores as calibrated. We run a held-out set of known VHH binders and non-binders through the model and check that the score separation is meaningful. If it is not, we either fine-tune or switch to a protein-universe model and rely less on sequence-level fitness scores, using structural models more heavily for the VHH-specific constraints.

We also allow longer CDR3 sampling windows for VHH than for VH, and we pay specific attention to whether the generated CDR3 sequences are consistent with the expected secondary structure propensity for the target length: extended loops versus hairpin-forming sequences have different amino acid biases (hairpin-forming sequences favor Gly, Asn at the turn, with hydrophobic residues on the adjacent strands) and the generative run should not suppress these features.

For conventional VH campaigns, the interplay with VL pairing matters more. Generated VH sequences need to be checked against VL compatibility, either by screening in paired format or by pre-screening against a known VL partner that will be used in the final construct. Generating in isolation for VH and then pairing later is a common source of expression and stability surprises.

The Shared Bottom Line

We are not arguing that one modality is inherently better for generative design than the other. Both benefit from computational pre-filtering. The differences are in how you configure the run: the CDR length distributions, the model calibration requirements, and what the primary fitness objectives are. Treating the two as interchangeable at the campaign setup stage leads to miscalibrated runs that either over-represent sequences that look like human IgG (bad for VHH) or under-sample the single-chain stability space (bad for VH development).

Know which scaffold you are on before you generate the first sequence. The downstream assay queue will be shorter for it.

Interested in generative protein design?

Start with a free retrospective validation run on your sequence-activity data.

Request validation More articles