Vecura Biotech Insiders #08: A binder with a near-perfect interface score, on entirely the wrong face
For Biotech Insiders #08, Goh brought us a campaign that did not work, and a reasonable person reading the metrics at stage three would have called it finished. That is exactly why we asked to publish it

A binder with a near-perfect interface score, on entirely the wrong face
De novo miniprotein design against IL-33. Marcus Goh’s best candidate scored an interface pTM of 0.94 and an interface PAE of 3.26 angstrom. It also missed every one of the seventeen hotspot residues it was designed to engage.
For Biotech Insiders #08, Goh brought us a campaign that did not work, and a reasonable person reading the metrics at stage three would have called it finished. That is exactly why we asked to publish it.
Why IL-33
IL-33 is an alarmin, released from epithelial cells when tissue is damaged by infection, smoke or other environmental insult. It signals through the ST2 receptor and sits upstream of both type 2 and non-type 2 airway inflammation, which has made it one of the most closely watched targets in respiratory medicine.
It has also been an unforgiving one. AstraZeneca’s tozorakimab met its primary endpoint in the OBERON, TITANIA and MIRANDA Phase III COPD trials in early 2026, and it blocks both the reduced form of IL-33 signalling through ST2 and the oxidised form signalling through RAGE and EGFR. Sanofi and Regeneron’s itepekimab, which blocks primarily the reduced form, succeeded in AERIFY-1 and missed in AERIFY-2. Roche’s astegolimab, which binds the ST2 receptor rather than the ligand, was set aside after two negative trials.
Why a miniprotein
All three are antibodies given by injection every two to four weeks. A miniprotein is roughly a tenth the size, potentially amenable to inhaled delivery straight to the airway where IL-33 is released, and cheaper to produce. Whether that is achievable is open, and the first question is whether a binder can be designed against the right surface at all.
Goh ran that first step on Vecura, using PDB 4KC3, the IL-33 and ST2 co-crystal, as the structural starting point. The pipeline ran in four stages, each with a different model.
Stage 1. De novo generation with PXDesign
Goh ran PXDesign against IL-33 chain A with seventeen hotspot residues specified across two patches, A142 to A153 and A220 to A228, corresponding to the ST2-competing surface. Binder length was set to 130 residues, with ten samples over 400 steps under the extended preset and equal weighting between the AF2 and Protenix objectives.
The outcome was thin. Not one of the ten designs passed the strict success criteria on either engine, and only the top-ranked design reached the practical threshold.
| Rank | ipTM | ipAE (Å) | Binder RMSD (Å) | Grade |
| 1 | 0.570 | 18.72 | 5.86 | PROMISING |
| 2 | 0.220 | 22.19 | 0.31 | UNLIKELY |
| 3 | 0.210 | 23.37 | 3.58 | UNLIKELY |
| 4 | 0.120 | 25.61 | 1.26 | UNLIKELY |
| 5 to 10 | 0.100 to 0.170 | 26.19 to 26.73 | 5.80 to 19.04 | UNLIKELY |
Table 1. PXDesign output, ranked. Only rank 1 reaches the practical interface threshold.
Rank 2 is the more instructive row, and it is the one Goh points to first. It folds almost perfectly on its own, at monomer pLDDT 0.961, and AF2 re-predicts its backbone to within 0.31 angstrom of the design. It is a well-designed protein. Its interface pTM is 0.22. It sits beside IL-33 without engaging it, which is the characteristic failure mode of a design objective that has satisfied structural self-consistency without solving interface energy.
Rank 1 had the opposite profile. Reasonable interface confidence, and a binder RMSD of 5.86 angstrom, meaning the sequence did not recapitulate the backbone it was designed onto. That is a sequence design problem, and it has a standard fix.
Stage 2. Sequence redesign with ProteinMPNN, and a cysteine problem
Goh ran ProteinMPNN on the rank 1 backbone to find sequences that encode that geometry more reliably, using the all-atom model at temperature 0.1 for conservative, high-fidelity output. All 100 designs improved on the parent sequence score of 1.1462, with the best falling to 0.7985, which confirms the original sequence was suboptimal for its own backbone.
Then a target-specific problem appeared that no generic scoring function would have flagged. 98 of 100 designs reintroduced cysteine. The parent sequence contained none. IL-33 is redox-sensitive, and its biology depends on the distinction between the reduced and oxidised forms, which is precisely the distinction that separates tozorakimab from itepekimab clinically. A cysteine-containing binder risks forming spurious disulfides with IL-33 in the oxidising environment of airway surface liquid, scrambling both molecules. Only two designs came back cysteine-free.
Goh carried those two forward, with design 17 at score 0.8244 and design 20 at score 0.8519. Neither was the top-scoring sequence. They were selected on a constraint that came from the target’s biology rather than from any number the model produced.
Stage 3. Co-folding validation with Boltz-2
Goh refolded both designs in complex with IL-33 using Boltz-2, ten samples each, against a target of interface pTM above 0.6 and interface PAE below 15 angstrom. Eight of ten passed for each design. The best samples were not marginal.
| Model | Interface pTM | Interface PAE (Å) |
|---|---|---|
| PXDesign rank 1 parent | 0.570 | 18.72 |
| Design 17, best sample | 0.9445 | 3.26 |
| Design 20, best sample | 0.9002 | 4.35 |
Table 2. Best Boltz-2 sample from each design, against the PXDesign parent.
On these numbers the campaign looked finished. Sequence redesign followed by co-folding had moved a borderline candidate to a high-confidence complex, and it had done so reproducibly across two independent sequences.
FIGURE 1: Boltz-2 refolded complex of IL-33 (green) with design 17 (orange). Interface pTM 0.9445, interface PAE 3.26 angstrom.
FIGURE 2: Boltz-2 refolded complex of IL-33 (green) with design 20 (orange). Interface pTM 0.9002, interface PAE 4.35 angstrom.
Stage 4. The epitope check
The remaining question, as Goh framed it, was not whether the binder engaged IL-33 but where. Each Boltz-2 complex was superposed onto 4KC3 with US-align so the binder could be read in the crystal frame, and the contacted residues compared against the ST2 footprint.
The superposition itself was clean. TM-score 0.924 and 0.905, RMSD 1.52 and 1.67 angstrom, 133 of 134 IL-33 residues aligned in both cases. The IL-33 conformation in the predicted complexes is essentially that of the crystal, so the comparison is valid.
| Measure | Design 17 | Design 20 |
| IL-33 residues contacted | 40 | 38 |
| Overlap with ST2 footprint (46 residues) | 2 (4%) | 3 (7%) |
| Design hotspots engaged (17 specified) | 0 (0%) | 0 (0%) |
| ST2-blocking verdict | no | no |
Table 3. Epitope engagement against the ST2-blocking footprint.
Both binders dock on the wrong face of IL-33. Design 17 contacts forty IL-33 residues with an interface PAE of 3.26 angstrom, which is very high confidence, and not one of those residues is among the seventeen hotspots the design was directed at. It touches residues 129 and 130, peripheral to the ST2 interface. Design 20 behaves the same way, contacting 118, 119 and 130. Both would leave ST2 signalling intact.
A design can score at the top of the range on both and engage a surface with no functional relevance, and there is no threshold on either metric that would have caught it.
What Marcus Goh would change
Epitope verification belongs inside the loop, not after it. The check that discriminated here was a structural superposition against a reference complex, and it is cheap. Run in the design loop as a filter rather than at the end as a report, it would have redirected the campaign at stage 1 instead of stage 4.
Constraint satisfaction should be scored, not assumed. Seventeen hotspots were specified at the outset and the pipeline never checked whether any were engaged until the final step. Hotspot coverage is a number the optimiser could have been given.
Target biology outranks the generic metric. The cysteine filter came from IL-33’s redox chemistry, not from any score in the pipeline, and it eliminated 98 of 100 otherwise higher-ranked sequences. Whoever writes the objective has to know the target.
Caveats and scope
Goh is explicit about the limits of what this campaign shows:
-
Both surviving designs derive from the same PXDesign rank 1 backbone, so they are two sequences on one scaffold rather than two independent solutions. The consistency of the failure across them is suggestive but does not establish that the ST2 surface is intrinsically unfavourable for miniprotein binders.
-
Testing that would need several independent generation runs, ideally with a different preset, since nine of ten designs in this run were alpha-helical and the two-patch hotspot geometry may suit a beta-rich scaffold better.
-
This campaign was computational from end to end. No protein was expressed and no binding was measured, so nothing here demonstrates engagement with IL-33. Predicted complexes are hypotheses and interface confidence metrics are not evidence of affinity. The epitope analysis compares a predicted complex against a crystal structure and inherits the uncertainty of the prediction.
Run de novo binder design on your own target, with epitope verification in the loop.
Every model in this campaign, PXDesign, ProteinMPNN, Boltz-2 and US-align, is available in the Vecura model catalog at app.vecura.com
立即试用 Vecura。
带上您自己的输入,开始探索 Vecura 的能力。


