ProGen2 Is Now Available on Vecura — Design and Score Proteins Without the Infrastructure Overhead
This update enables biotech researchers and protein engineers to generate novel protein sequences and evaluate their fitness through a guided workflow inside Vecura, without setting up complex GPU infrastructure or managing large-scale language model deployments.
What is ProGen2?
ProGen2 is a suite of autoregressive transformer-based protein language models developed by Salesforce Research, ranging from 151 million to 6.4 billion parameters. Trained on diverse protein sequence databases — including UniRef90, BFD30, BFD90, and the Observed Antibody Space (OAS) — the models treat amino acid sequences as token streams to enable both de novo sequence generation and likelihood-based scoring. A specialist variant, progen2-oas, is fine-tuned exclusively on antibody repertoires, making it particularly well-suited for antibody engineering tasks.
It helps users design novel protein candidates from scratch, continue partial sequences, and rank variants using zero-shot bidirectional log-likelihood scoring — all without labelled training data or structural inputs. It is especially useful for early-stage protein engineering, antibody design, and computational screening pipelines that need to prioritize candidates before committing to wet-lab synthesis.
What can users do with ProGen2 on Vecura?
With ProGen2 on Vecura, users can:
- Generate novel protein sequences from scratch or from a context-seeded prefix using configurable nucleus sampling, controlling the trade-off between sequence diversity and similarity to natural proteins.
- Score and rank candidate sequences using bidirectional log-likelihood as a zero-shot fitness proxy — no labelled data, alignments, or structural information required.
- Design antibody sequences using the
progen2-oasspecialist variant trained on the Observed Antibody Space repertoire. - Fine-tune generation parameters — temperature, top-p, sequence length, and number of samples — to explore sequence space systematically and balance novelty against naturalness.
What the output means
The output provides:
- Generated sequences: Novel amino acid strings produced by the model, ready to be passed to structure predictors or wet-lab screening pipelines.
- Bidirectional log-likelihood scores: Including sum and mean values for left-to-right, right-to-left, and averaged directions. Units are nats (sum) or nats per residue (mean). Higher (less negative) values indicate sequences that are more probable under the model and more consistent with natural protein distributions, serving as a calibrated proxy for functional fitness.
This output should be used to support scientific decision making. It does not replace experimental validation.
Why this matters
Protein engineering has traditionally relied on costly and time-consuming experimental screening to identify functional variants from vast sequence spaces. Language models like ProGen2 offer a powerful computational shortcut: by learning the statistical grammar of natural proteins from over a billion sequences drawn from genomic, metagenomic, and immune repertoire databases, these models can propose plausible candidates and flag unlikely ones before any physical experiment takes place. The ProGen2 research demonstrated that scaling autoregressive modelling into the multi-billion parameter regime yields measurably lower perplexity and higher experimental hit rates — a finding that underscores the importance of model scale in protein design.
The availability of ProGen2 on Vecura lowers the barrier to entry for research teams that lack dedicated ML infrastructure or access to high-VRAM GPU clusters (the 6.4B parameter model alone requires ~40 GB of VRAM). Scientists can now access state-of-the-art protein language models through a simple, guided interface, accelerating the design-build-test cycle and enabling rapid iteration on therapeutic proteins, industrial enzymes, and antibody candidates.
- Developed by: Salesforce Research
- Source: ProGen2 GitHub Repository · Pretrained Checkpoints
- Reference: Nijkamp, E., Ruffolo, J.A., Weinstein, E.L., Madani, N., & Bandyopadhyay, A. (2023). ProGen2: Exploring the boundaries of protein language models. Cell Systems. arXiv:2206.13517
Try ProGen2 on Vecura.
Open the model workspace and start evaluating it with your own inputs.