Protein Engineering

Why Your Assay Queue Is Too Long (And It Is Not a Capacity Problem)

by Proteinvue

Abstract visualization of a long queue of microtiter plates in rows suggesting assay bottleneck

Every protein engineering team we talk to has the same complaint: the assay queue is too long. They want more plate readers, more HPLC time, more hands at the bench. The ask is always for throughput. But when we look closely at what is sitting in those queues, the diagnosis shifts. The bottleneck is not throughput. It is decision quality upstream.

Here is what we mean by that.

What Actually Fills a Queue

A typical screen in therapeutic antibody engineering might put 200 to 400 variants through a primary binding assay in a given sprint. Of those, maybe 15 to 30 advance to a secondary functional readout. Of those, a handful get expression titers measured. The funnel is steep, and every step costs real resources.

Now ask a harder question: how were those 200 to 400 variants selected in the first place? In most teams, the answer involves some combination of intuition from experienced engineers, a homology sweep around a known sequence, a handful of alanine scans, and maybe a random library if you are doing phage or yeast display. All of that is perfectly defensible. But it treats the assay as the primary information source. You are asking the assay to tell you which variants matter.

The cost of that posture is queue length. When you do not pre-filter computationally, you submit everything plausible and let the biology sort it out. That works. It has worked for decades. But it means your assay infrastructure is spending a meaningful fraction of its cycles confirming what was already probable, and discarding what was already improbable.

The Information Problem Framing

Reframe the assay queue as an information budget. You have a fixed number of assay slots per sprint, say 384 in a standard high-throughput format. Each slot consumes reagents, instrument time, and analyst attention. The question becomes: which 384 variants give you the most information about your fitness landscape?

If you submit 384 variants that cluster tightly in sequence space, you are learning about a small patch of that landscape. If 100 of those are effectively equivalent because they share the same few substitutions at non-critical positions, you are spending a third of your budget learning something you already suspected. That is not a capacity problem. That is a prioritization problem.

Computational pre-filtering changes what goes into the queue rather than how fast the queue moves. The variants you submit are those where the model has either high predicted fitness and moderate confidence (worth confirming) or moderate predicted fitness and high uncertainty (worth exploring to calibrate the model). Variants with low predicted fitness and high model confidence are not worth the slot, regardless of how fast your plate reader is.

Where Teams Get This Wrong

The most common failure mode is treating computational scores as pass/fail thresholds rather than as prioritization signals. A team will run a generative model, set a score cutoff, and submit everything above the line. That sounds like pre-filtering, but if the cutoff is permissive it is really just adding a noisy early step before the same crowded queue.

The second failure mode is running computation and wet lab on completely decoupled cadences. The computational run happens once at the start of a campaign. Assay data comes back over six to eight weeks. The model never sees that data until the campaign is finished. You have lost the feedback loop that would let you sharpen your selection criteria in round two.

A third issue is using computational predictions only for positive selection, when they are equally useful for negative selection. If your model assigns low fitness scores to a large swath of sequence space and you have good calibration data to back that up, you can confidently deprioritize that region. That is not a miss. That is a queue slot freed up for something more informative.

A Concrete Scenario

Consider an enzyme engineering campaign targeting a thermophilic beta-glucosidase for industrial saccharification. The starting point is a well-characterized wild-type from a Bacillus relative, with solid expression in E. coli and a known Tm of around 68 degrees Celsius. The engineering goal is to push Tm above 75 while preserving kcat at the target substrate.

The team generates a combinatorial library covering 12 positions in the barrel domain, estimated at roughly 4,000 double-mutant combinations. Without any pre-filtering, a primary thermostability screen alone would require 10 or more 384-well plates. With computational ranking using a structure-informed fitness model, the top 10 percent of predicted variants covers maybe 400 sequences. The first plate screen is now a realistic single-sprint effort, not a two-month commitment.

More importantly, if the first 400 return a good hit rate, round two can use those results to recalibrate the model and propose another 400 with improved confidence. The queue has not gotten shorter because the plate reader got faster. It has gotten shorter because the variants in it are better chosen.

We Are Not Saying Wet Lab Is the Problem

This is worth being direct about: the assay is not at fault, and more throughput is not categorically wrong. If your campaign is genuinely exploratory, if you are characterizing a poorly understood protein family with no prior sequence-activity data, then broad screening is exactly right and throughput matters a great deal. You need to cover ground. There is no model to pre-filter with.

But if you are in a campaign with any existing SAR data, any prior rounds of screening, any published homolog data in your family, then the queue you are running is not a capacity-constrained queue. It is an information-constrained queue. Buying another plate reader is the wrong answer.

How We Think About It at Proteinvue

When we work with a new campaign, the first question we ask is not "how many variants do you want to test?" It is "what do you already know?" That includes prior screening data (even from failed campaigns), published fitness data for homologs, structural data if available, and any MSA-depth assessment we can run on the target family. All of that shapes the quality of the initial computational rank list, and the quality of that rank list determines how many assay slots you actually need.

We have seen teams cut their first-round screen size by 40 to 60 percent while holding constant or improving the hit rate in downstream confirmation assays. That is not magic. It is using information you already had, processed in a way that the wet lab calendar does not naturally support.

The assay queue is a symptom. The root is upstream decision quality. Fix the prioritization, and the queue problem resolves without a single new piece of lab equipment.

Interested in generative protein design?

Start with a free retrospective validation run on your sequence-activity data.

Request validation More articles