Vecura
PricingSolutions
Resources
Vecura

Product

  • Solutions
  • Pricing

Company

  • Contact us
  • Academic Research Program

Resources

  • Updates
  • News
  • Insights
  • Use Cases
  • AI4Life Bootcamp
  • Community

Legal

  • Privacy Policy
  • Terms of Service
  • Citation Guidelines
  • Trust Center

© 2026 NYB AI. All rights reserved.

All systems operational
Vecura
PricingSolutions
Resources
Back to insights

Vecura Biotech Insiders Spotlight #11: Grafting the Enzyme I active site onto a new scaffold

For Vecura Biotech Insiders Spotlight #11, Ethan Lee shares a nine step campaign logged exactly as it ran, including the gate that failed every single design and the comparison that only existed because three candidates went forward instead of one.

Oct 9, 2026

A motif scaffolding campaign set out to lift the active site of a bacterial enzyme off its native protein and mount it on a scaffold a quarter the size. Three designs were carried the whole way through, from backbone generation to a docking run against the partner protein. None of the three works. The part worth reading is the order they finished in: the design that cleared the fewest structural quality checks was the only one that could plausibly do the chemistry, and the design with the tightest catalytic geometry in the entire campaign never brought its catalytic residue within 10 Å of the target.

What the campaign was trying to build

Enzyme I of the bacterial phosphotransferase system does one job in this story. It carries a phosphoryl group on a specific histidine, His189, and hands that group to His15 on a much smaller partner protein called HPr. For the handoff to happen, the two histidines have to come close enough to pass the group directly, which in practice means a few ångström apart.

The question Ethan set out to answer was whether that active site can be taken off Enzyme I altogether and mounted on a smaller protein that still performs the handoff. Nine steps, five models: RFdiffusion to build a scaffold around the motif, ProteinMPNN to design a sequence for it, ColabFold to predict what that sequence would actually fold into, GROMACS to check the fold holds up in water, and HADDOCK3 to dock the result against HPr.

Most of the work is plumbing, and the plumbing is where campaigns break

Five models, four file formats, and three different residue numbering schemes in one campaign. His189 in the 1EZC reference structure becomes residue A48 once the motif is grafted into a 79 residue scaffold. HPr's His15 shows up in the HADDOCK3 output as residue 315, because chain B numbering in that file starts at 301. Both offsets are correct, and both are easy to get wrong, and one mistake anywhere in that chain turns a measured distance into a meaningless number that still looks like a result.

Vecura carries the mapping across tool boundaries, so a contact reported between His48 and His315 traces back to His189 and His15 without anyone doing the arithmetic by hand, and a flagged residue in a designed sequence reports where it sits in the original structure. Almost everything that follows is a story about numbers that needed careful reading. None of it is a story about numbers that were lined up wrong.

Steps 1 and 2. Building the scaffold, then designing a sequence for it

RFdiffusion ran in motif scaffolding mode with 1EZC as the motif source. His189 was confirmed as the phospho accepting residue, the active site shell was computed as the 31 residues sitting within 10 Å of it, and ten backbones were generated.

The number that invites the wrong conclusion at this step is motif RMSD. One backbone, backbone_06, came out worst of the ten on motif fidelity at 0.30 Å. But 0.30 Å is an excellent result in absolute terms, so if that is the worst of the set, all ten reproduce the motif essentially perfectly and ranking them on that column just sorts noise. The column that actually carried information was scaffold support around the catalytic residue, and there backbone_06 was the only design matching native density at the inner shell. It went forward at 79 residues.

ProteinMPNN then produced twenty sequences against that backbone with six residues pinned: Pro165, Thr168, His189, Thr190, Ile192 and Met193 in 1EZC numbering, which map to A24, A27, A48, A49, A51 and A52 in the new scaffold. All twenty sequences carried the correct residue at all six pinned positions.

Two more numbers invite wrong conclusions here. The first is a poly glycine reference that scored 2.46 against designs scoring between 0.93 and 1.00. That looks like a control, but it is not one, because scoring poly glycine with ProteinMPNN will always come back poor. The second is the ranking itself.

image.png

Figure 1. ProteinMPNN score against core-6 RMSD for all 20 designs. The horizontal line is the 1.0 Å gate on catalytic geometry.

Design_12 took first place on ProteinMPNN score and finished fourteenth of twenty on catalytic geometry, at 5.13 Å. Design_5, ranked dead last of twenty by the same scoring function, became the lead candidate. Design_12 and Design_15 sit 0.002 units apart on ProteinMPNN score and 4.5 Å apart on catalytic geometry. Across the full set, the score explains roughly six percent of the variance. That is not a defect in ProteinMPNN. The score measures how well a sequence fits a given backbone, not whether that sequence will fold to it. It is worth showing because the ranked table is the first thing a user sees, it looks authoritative, and it is what shaped the choice of which designs to carry forward.

The cysteine cluster

All twenty designs placed a cysteine at scaffold positions 40, 42 and 47. Those sit inside the grafted motif at offsets 18, 20 and 25, mapping back to 1EZC residues 181, 183 and 188, which is immediately next to the catalytic histidine. In the native structure those positions are Asp181, Ala183 and Ala188. Three cysteines clustered around an active site can form disulfides during expression and bury the catalytic center.

A design decision sits behind that. The motif spans 35 residues, but only six were held fixed, which left ProteinMPNN free to redesign 29 positions inside the exact region that was grafted to preserve geometry in the first place.

Step 3. Structure prediction, and four gates

ColabFold ran on all twenty sequences, five models each, rank 1 taken, with two alignments per design: self consistency RMSD against backbone_06 across the full 79 residue chain, and core-6 RMSD against the 1EZC motif reference at residues 163 to 197.

GateCriterionResult
G1Global pLDDT above 7015 of 20 pass
G1mMotif pLDDT above 807 of 20 pass
G2Self consistency RMSD below 2.0 Å0 of 20 pass
G3core-6 RMSD below 1.0 Å4 of 20 pass

Table 1. The four structural quality gates and how the twenty designs fared against each.

G2 failed completely, and it is reported here as a failure. Self consistency RMSD is the standard success criterion for this kind of pipeline, and zero of twenty is a failed design round, not a reason to go back and loosen the gate.

One caveat belongs with that column, stated plainly because the campaign log is the product. The reported self consistency values span 6.332 to 6.434 Å across eleven designs, a coefficient of variation of 0.50 percent, while core-6 RMSD across those same eleven designs varies by 37 percent. Two pairs of designs report values identical to three decimal places. A whole chain measure that stays that flat across structures whose local geometry differs that much is worth rechecking before anyone builds an interpretation on it. The values are reported as produced and the gate result stands. The explanation for them does not yet.

G3 is the gate that actually discriminated in this run. Four designs reproduced the catalytic six residue arrangement to better than 1.0 Å. It is measured after superposing the 35 residue motif region, so passing it means the catalytic residues hold the right geometry relative to each other. It does not mean the protein presents them correctly to a partner, and that distinction turned out to be the whole story.

Which three went forward, and why not just the best one

Rather than carrying the single strongest candidate, Ethan Lee picked three designs to span the range:

  • Design_5. Highest global and motif pLDDT, three of four gates passed, core-6 RMSD 0.837 Å. Ranked last of twenty by ProteinMPNN.

  • Design_15. The tightest catalytic geometry in the entire set at 0.609 Å, two of four gates passed.

  • Design_12. Best ProteinMPNN score of the twenty and the weakest structural candidate of the three. One of four gates passed, core-6 RMSD 5.129 Å, pLDDT 70.27, and the g-plus rotamer rather than g-minus.

That choice is the reason the rest of the campaign says anything at all.

image.png

Figure 2. ColabFold models of Design_5, Design_12 and Design_15.

Steps 4 to 6. Do the designs hold their shape in water

Two nanoseconds of explicit solvent molecular dynamics per design at 300 K, with backbone analysis across all 79 residues.

Design_5Design_12Design_15
Production RMSD1.63 ± 0.27 Å2.10 ± 0.15 Å1.58 ± 0.28 Å
Radius of gyration1.262 nm1.280 nm1.281 nm
Rg coefficient of variation0.75 %0.79 %0.86 %
Residues below 1.0 Å RMSFnot reported41 of 7960 of 79
His48 or active site loop RMSF0.53 Å1.25 Å1.61 Å loop mean
Trajectory plateaued at 2 nsno, slow driftyesno, +0.35 Å per ns

Table 2. Molecular dynamics summary across the three designs.

All three hold their fold. Radius of gyration stays constant to better than one percent in every run, nothing unfolds, and the compact core survives in all three. Two of the three had not reached a plateau at 2 ns, so these runs establish that the proteins do not fall apart, rather than that they have settled. The active site flexibility numbers are worth setting out plainly, because they span a threefold range and the campaign treated all three as acceptable. Design_5 is rigid at the catalytic histidine, Design_12 is moderately mobile, and Design_15 has a mobile active site loop averaging 1.61 Å with a maximum of 2.18 Å. A criterion that accepts 0.53 Å and 1.61 Å equally is not discriminating between the designs, which is why the numbers appear here without a verdict attached. One claim in particular has to be ruled out. It is tempting to read Design_5's 0.53 Å backbone RMSF as evidence that its histidine side chain is locked into the g-minus rotamer ColabFold predicted. It is not. Chi2 is a side chain dihedral and the analysis group here was backbone only, and a 2 ns trajectory cannot sample chi2 barrier crossings at any useful rate anyway. Settling the rotamer question needs 50 ns or more, or enhanced sampling on that dihedral. It stays open.

Steps 7 to 9. Docking against the partner protein

Each design was docked to HPr from 3EZE under HADDOCK3, same protocol, same surface restraints. No distance restraint was placed between the two active site histidines in any of the three runs, so nothing in the setup nudged the catalytic geometry toward existing.

Design_5Design_12Design_15
Clusters and poses5 and 333 and 293 and 29
Best HADDOCK score-82.49-118.93-78.73
Buried surface, best pose1140 Ų1795 Ų1104 Ų
Transfer-competent poses01 of 290 of 29
Closest His48 to His1518.17 Å3.90 Å10.06 Å

Table 3. Docking results against HPr.

image.png

Figure 3. Docking poses against HPr, all superposed on HPr. (A) Design_5, His48 to His15 18.17 Å. (B) Design_12, pose C1-P5, rank 8 of 29, 3.90 Å. (C) Design_15, closest of 29 poses, 10.06 Å. His48 and His15 shown as sticks in all panels.

Design_5 assembles a well scored, well packed complex that satisfies every surface restraint, and then parks His48 18 Å from its partner. His48 is genuinely at the interface and makes a 1.70 Å hydrogen bond, but to the backbone carbonyl of HPr Lys24. HPr's His15 makes no contact with the design at all. The complex forms, through the wrong face. Design_15 is worse. Across all 29 poses, not one brought the two histidines within even 8 Å, and the closest approach anywhere in the run was 10.06 Å. Its dominant binding mode packs His48 against the hydrophobic face of HPr, roughly 12 Å from His15. This is the design with the tightest catalytic geometry in the whole campaign. Design_12 produced the only transfer-competent pose of the three, at 3.90 Å nitrogen to nitrogen. It also scored 36 units better than Design_5 and 40 better than Design_15, and buried roughly 60 percent more surface area. Its own top scoring pose is not the transfer-competent one, so this is not a working enzyme either, but on every interface measure it is the strongest of the three by a wide margin.

What the three-way comparison says

Design_5Design_12Design_15
ProteinMPNN rank20 of 201 of 203 of 20
core-6 RMSD0.837 Å5.129 Å0.609 Å
Global pLDDT86.7070.2784.09
Rotamerg-minusg-plusg-minus
Structural gates passed3 of 41 of 42 of 4
Best HADDOCK score-82.49-118.93-78.73
Transfer-competent poses010

Table 4. Structural quality on top, functional outcome at the bottom, read down each column.

The design that passed the fewest structural gates produced the best interface and the only catalytically plausible pose. The design with the tightest catalytic geometry in the campaign could not bring its histidine within 10 Å of the target. Across these three candidates, the structural filters ranked them close to the reverse of the functional result.

Three designs is not a result, and what follows from it is a direction rather than a finding. There is also a plain alternative explanation that deserves testing. Design_12 has the lowest pLDDT of the three and buries far more surface, so it may simply be presenting more exposed area for a docking engine to work with, which would inflate both the score and the interface size without saying anything about catalysis. Its radius of gyration is only 0.02 nm larger than Design_5's, which does not obviously support that reading, but it has not been ruled out either.

What the comparison does establish is narrower and still useful: the gates used here, applied to these three designs, carried no predictive signal for the one functional test. A pipeline filtering on core-6 RMSD and rotamer state would have thrown out the only design that produced a transfer-competent pose.

What Ethan would run next

  • Native Enzyme I docked to HPr under the identical HADDOCK3 protocol. With three scores now spread across 40 units and nothing to calibrate them against, this is the most valuable single run available.

  • A full His48 to His15 scan across all 33 Design_5 poses, so that row of the comparison is drawn the same way as the other two.

  • A recheck of the self consistency RMSD calculation, for the reason given in step 3.

  • More designs through the same three-way comparison. Three is enough to notice the inversion and nowhere near enough to believe it.

  • A larger fixed residue set in ProteinMPNN, covering more of the 35 residue motif than six positions.

  • Fifty nanoseconds, or metadynamics on the chi2 dihedral, to settle the rotamer question the 2 ns runs could not.

On scale

Ten backbones, twenty sequences and three full characterizations is a tutorial. Production motif scaffolding campaigns generate thousands of backbones and tens of thousands of sequences, filter hard, and carry a few dozen into the lab. The filter is severe by design, because the pass rate is low.

This campaign is worth reading because it shows that filter operating on a sample small enough to inspect by hand, and because carrying three designs instead of one revealed something about the filter that a single candidate would have hidden completely.

Scope

  • This campaign is computational end to end. Nothing was expressed, purified or assayed, and no phosphoryl transfer has been measured in a biological system.

  • "Transfer-competent" here means a docking pose that places the two histidine nitrogens close enough for transfer to be geometrically plausible. It is a distance measurement, not evidence of activity.

  • Every gate result, every score and every distance quoted here came out of the run log. Nothing was reconstructed afterward and nothing was adjusted once the results were in.

Try Vecura now.

Bring your own inputs and start exploring what Vecura can do.

Try Vecura now

On this page

What the campaign was trying to buildMost of the work is plumbing, and the plumbing is where campaigns breakSteps 1 and 2. Building the scaffold, then designing a sequence for itThe cysteine clusterStep 3. Structure prediction, and four gatesWhich three went forward, and why not just the best oneSteps 4 to 6. Do the designs hold their shape in waterSteps 7 to 9. Docking against the partner proteinWhat the three-way comparison saysWhat Ethan would run nextOn scaleScope

Try Vecura now.

Try Vecura now

Related insights

Vecura Biotech Insiders Spotlight #10: The fragments never bridged. They circled back instead.

Sep 24, 2026

Vecura Biotech Insiders #09: It cleared every filter. Then the most confident synthesis step turned out to be impossible.

Sep 16, 2026

Harnessing Vecura’s AI Platform to Design FGFR4-Targeting Peptides for Hepatocellular Carcinoma Therapy

Sep 10, 2026

Vecura

Product

  • Solutions
  • Pricing

Company

  • Contact us
  • Academic Research Program

Resources

  • Updates
  • News
  • Insights
  • Use Cases
  • AI4Life Bootcamp
  • Community

Legal

  • Privacy Policy
  • Terms of Service
  • Citation Guidelines
  • Trust Center

© 2026 NYB AI. All rights reserved.

All systems operational