Company

Why We Started Proteinvue

by Elena Marchetti

Abstract overhead view of a protein engineering lab bench with pipettes and microplates in muted scientific tones

This is not the polished origin story version. It is the actual version, which is less tidy and more useful if you are trying to understand what we built and why.

Before starting Proteinvue, I spent several years in structural biochemistry research, working on enzyme engineering for metabolic pathway applications. Marcus was doing ML research with a focus on sequence modeling. We had known each other since graduate school but had not worked together directly until a specific project threw us into the same problem from two different directions.

The Moment That Made the Problem Obvious

The project was a cytochrome P450 variant campaign. The goal was to shift regioselectivity toward a non-natural substrate for a synthetic biology application. The biology team had a library, a screening assay, and a plan to run error-prone PCR followed by three or four rounds of directed evolution. Standard workflow for this type of work.

Marcus had been working on a protein language model fine-tuned on cytochrome P450 families. I asked him, half expecting nothing, whether his model could produce a ranked list of single-point variants likely to improve activity on the new substrate. He ran an analysis overnight and sent me a spreadsheet with 80 variants ranked by predicted fitness, with positional confidence estimates attached.

We picked the top 40 and added them to the first round of screening alongside about 200 random library members. The correlation between Marcus's ranking and measured activity was not perfect, but the top-ranked variants were substantially enriched among the active hits compared to the unranked random set. It was not a miracle result, but it was clearly and reproducibly better than random sampling. The wet-lab team noticed immediately. They started asking Marcus for ranked lists before every subsequent round.

What struck us was not the performance of the model itself. It was the friction. Every time the wet lab wanted a ranked list, Marcus had to set up the analysis manually, handle the sequence alignment by hand, deal with the fact that the family's coverage in UniProt was patchy in the relevant subfamilies, and produce output in a format that could be loaded into the lab's plate-mapping software. The analytical capability existed; the infrastructure to use it routinely did not.

What We Thought Was Missing

The protein engineering community had tools for structure prediction, tools for sequence analysis, and increasingly capable language models. What it mostly lacked was a workflow layer that connected those tools to the cadence of a working protein engineering project: a team that runs assays on a schedule, wants a ranked candidate list ready before the next synthesis order goes out, and needs the computational output in a format the wet lab can directly act on.

We talked to people running similar projects. The pattern held. Teams that had someone with Marcus's background on staff were doing informal versions of what we had done. Teams without that person were running fully random library screens, which work eventually but waste enormous amounts of wet-lab time on low-probability experiments. The gap was not the science. It was the tooling and workflow.

This is a mundane observation but an important one. A computational method that produces correct predictions 65% of the time on a ranked list is already useful if the wet lab can access it in their existing workflow. The same method sitting on a researcher's local machine, requiring custom scripting to use, is practically useless for most teams even if the underlying accuracy is identical.

Why Baltimore, and Why Bootstrapped

We started Proteinvue in Baltimore in early 2024. Marcus had connections to the biotech ecosystem that had been building up around Johns Hopkins and the University of Maryland, and I had been based here for several years. The cost of living for a small founding team is significantly lower than in Cambridge or South San Francisco, and the research community around us is genuinely excellent for this type of work. We have not found any disadvantage to being here rather than in one of the coastal biotech hubs.

On bootstrapping: we did not raise outside money at the start and have not done so. This was a deliberate choice. The product we wanted to build required us to go deep on a technical problem that is not obviously "VC-backable" on a two-year horizon. Building a protein design workflow platform means building integration work that is unsexy but critical, getting the data format handoffs right, handling edge cases in protein family coverage. That work does not compress well under investor timeline pressure. We wanted to do it carefully.

We are not saying fundraising is wrong. We are saying we wanted to build the thing first, understand what we actually had, and then think about scale. That is the sequence that made sense for the specific problem we were solving.

What We Actually Built

Proteinvue is a platform that takes in sequence-activity data from prior wet-lab rounds, characterizes the protein family's evolutionary context, generates a ranked candidate list for the next screening round, and returns that list in a format that integrates with standard lab workflows. The output is not a prediction with a confidence score floating in a UI. It is a ranked sequence file with synthesis-ready formatting, positional confidence annotations, and a coverage report that tells the wet lab team how much weight to put on the computational ranking versus running a broader exploratory screen.

The retrospective validation run we offer at the start of a new project is designed to answer a specific question before a team commits to a multi-round campaign: does our model's ranking on your protein family's existing data correlate with the actual fitness measurements you have? That check matters because not every family has sufficient evolutionary coverage for our approach to be predictive. If the retrospective shows weak rank-correlation, we say so upfront rather than running a campaign that will disappoint.

What Comes Next

We have been working on the platform for about a year and are running it with early-stage collaborators on campaigns in both therapeutic and industrial enzyme contexts. The scientific problems are genuinely varied and continue to surface things we had not anticipated when we set up the initial architecture.

The goal from the beginning was not to build a general-purpose computational biology platform. It was to build the specific tool that would have made that cytochrome P450 campaign faster and less wasteful if it had existed two years earlier. We think we have a version of that tool now. Whether it holds up across a wider range of protein families and campaign types is what the next year will tell us.

Interested in generative protein design?

Start with a free retrospective validation run on your sequence-activity data.

Request validation More articles