What was posted on 5 October?

On 5 October 2026 a DeepMind-led team posted “Designing enzymes for new-to-nature chemistry and non-natural substrates with AlphaProtein Novo” on bioRxiv. The server marks it as a preprint that has not been certified by peer review. Corresponding author Jue Wang is at Google DeepMind. Co-authors include Frances H. Arnold at Caltech, Demis Hassabis and Pushmeet Kohli, and further Caltech and University of Pittsburgh affiliations.

The same day, google-deepmind/alphaprotein-novo appeared on GitHub. We inspected the bioRxiv HTML, the PDF, the README and the license files on 7 October and did not find a matching DeepMind or Google newsroom post. The abstract’s claim that de novo design can, “for the first time,” outperform natural-sequence mining is the authors’ line on an unreviewed manuscript.

A preprint plus a code drop is still a research announcement, not a product launch. For how to keep those layers apart, see our note on reading AI company announcements.

How does AlphaProtein Novo work?

AP Novo scaffolds a catalytic motif—side chains plus ligand poses chosen for a hypothesized mechanism—inside a new backbone. A diffusion model co-generates coordinates and sequence. Optional LigandMPNN redesigns the sequence. Candidates are then filtered with AlphaFold 3, or a leaving-atom fine-tune the repository calls AF3-LA, using metrics the authors say track catalytic geometry across predicted states.

Filtering is the methodological core. Multi-seed scoring repeats AlphaFold 3 with different random seeds. Multi-resequence scoring asks whether several LigandMPNN sequences on the same backbone still pass those filters. Partial diffusion adds a little noise and re-denoises nearby sequences for the same test. The authors say combining those ensemble filters raised later-round hit rates.

They distinguish novel-scaffold designs, which start from a motif only, from recycled-scaffold designs that reuse a backbone already known to be active. Computational proposals still need wet-lab checks; that is the same split as in our earlier note on Claude’s enzyme-discovery work.

What did the piperidine and DEHP campaigns show?

The campaigns they put against natural-sequence mining are piperidine synthesis by intramolecular nitrene transfer and partial hydrolysis of the plasticiser DEHP. A screen of 188 natural and engineered heme proteins found nine above 5% C–H amination relative to wild-type TamPgb; none exceeded a 30:70 piperidine:pyrrolidine ratio. AP Novo then tested 86 novel-scaffold histidine-heme designs and 90 more aimed at piperidine. GDM_NT_121 reached up to 94:6 piperidine:pyrrolidine. The lead, GDM_NT_270, delivered 22 total turnovers, a 99:1 regioisomeric ratio and 94% enantiomeric excess—more piperidine turnover than TamPgb wild type, they say, through complementary regioselectivity rather than a higher overall rate.

On DEHP they monitored MEHP, the first hydrolysis product. Later rounds reached an 11% novel-scaffold hit rate and a 39% recycled-scaffold hit rate against a 2% of QHH21706.1 threshold. Purified GDM_DEHP_378 made 6% of QHH’s MEHP after four hours at room temperature, or 26% of Q7SIG1 (13% and 34% after 24 hours). GDM_DEHP_376 (191 residues) was 14-fold more active at 90 °C than at room temperature and twice as active in 75% acetonitrile; both natural controls were inactive under those conditions. Table 1 lists 3.3, 47 and 40 turnovers for the 23 °C / 4 h, 90 °C / 4 h and 75% acetonitrile / 24 h settings. The authors say full conversion to phthalic acid is still required for bioremediation.

Author-reported AlphaProtein Novo outcomes from Table 1 of the 5 October 2026 bioRxiv preprint. Hit-rate definitions differ by reaction and are not comparable across rows.
ReactionWhat they testedLead result they report
Piperidine nitrene transfer188 natural or engineered proteins vs 176 AP Novo designsGDM_NT_270: 22 TTN, 99:1 r.r., 94% e.e.
DEHP hydrolysis165 novel-scaffold and 184 recycled-scaffold designsGDM_DEHP_376: 47 TTN at 90 °C / 4 h; naturals inactive there
Kemp elimination1,536 novel-scaffold designsGDM_KE_1872: kcat/Km 44,940 ± 220 M−1 s−1 at pH 10
Serine esterase (4MU-Ac)1,569 novel-scaffold designsGDM_SE_2937: kcat/Km 360,000 ± 70,000 M−1 s−1 (NSE topology)
Carbene transfer91 novel-scaffold designsGDM_CT_103: >99:1 d.r., 96% e.e. S,S

What did the model reactions add?

While building the pipeline they also ran Kemp elimination, 4-methylumbelliferyl acetate hydrolysis and styrene cyclopropanation. On 1,536 novel-scaffold Kemp designs, hit rates (multiple-turnover kobs above 0.038 s−1) rose from 3% to 34% in the best facet. GDM_KE_1483 reached kcat/Km 14,600 ± 1,100 M−1 s−1 at pH 7.5—70-fold above earlier novel-scaffold reports, they say, and 10- to 100-fold below evolved designs. At pH 10, GDM_KE_1872 reached 44,940 ± 220 M−1 s−1.

Among 1,569 novel-scaffold serine esterases the best facet hit 44%. The overall lead, GDM_SE_2937 (kcat/Km 360,000 ± 70,000 M−1 s−1), has natural serine-esterase topology (TM-score above 0.45 and triad RMSD under 5 Å to a PDB hydrolase); six of their top seven designs did. The best non-NSE design, GDM_SE_2754, reached 4,900 ± 700 M−1 s−1 with substrate inhibition. On 91 carbene-transfer designs, 85% bound heme; GDM_CT_103 and GDM_CT_162 reached the stereoselectivities in the table in one round. Selected active-site knockouts cut activity, and tested designs kept activity after ten minutes at up to 80 °C.

How can researchers access the code and weights?

The repository LICENSE is Apache 2.0 and covers the pipeline code, not the trained generator. Weights (generator.bin.zst) sit on Google Cloud Storage under a 5 October terms file: non-commercial use by universities, nonprofits, research institutes, educational, journalism and government bodies; no commercial research; no training similar de novo models on the outputs; no publishing the parameters outside the licensee’s organisation. AlphaFold 3 and AF3-LA weights have their own terms. That split matches our explainer on open weights versus open source.

The paper’s “Model availability” paragraph says code and weights are available freely for non-commercial use. Designed sequences and predicted structures “will be made publicly available upon publication,” which has not happened. We have not run the pipeline. Neighbouring DeepMind biology releases this month include SynthID Bio protein watermarks and the AlphaGenome Atlas; those are separate systems.

What do the authors say is still missing?

The conclusions are blunt. Motifs still need reaction-specific expertise that current filters miss. Many samples must be generated before a design passes. Choosing an optimal backbone, then an optimal sequence, “remains far from solved.” Catalytic activities, they write, are still orders of magnitude below natural or evolved enzymes, and “our models do not meaningfully model the physics of catalysis.”

Author-affiliated entities have filed a US provisional patent on this class of de novo design. All authors other than Z.-Q. Li, Masy Domecillo, Ariane N. Mora, Julia C. Reisenbauer, Yu Zhang, Peng Liu and Frances H. Arnold “have commercial interests in the work described.” The bioRxiv record is all-rights-reserved. The next evidence is independent assays, a peer-reviewed version, and the promised sequence release—not a buy decision from the abstract.

Common questions

Can a company use AlphaProtein Novo to design commercial enzymes?

The repository code is Apache 2.0. The generator weights and their outputs are limited to non-commercial organisations and non-commercial use, and they bar training similar de novo design models on those outputs. Review both documents, plus the AlphaFold 3 terms, before any commercial plan. This is not legal advice.

Did the team produce a finished DEHP bioremediation enzyme?

No. They report first-step hydrolysis to MEHP and say higher activity and full degradation to phthalic acid are still required. The stability in heat and 75% acetonitrile is the property they contrast with two natural DEHPases.

Is the paper peer-reviewed?

No. bioRxiv posted version 1 on 5 October 2026 and labels it as not certified by peer review. Designed sequences are promised upon journal publication.

THE TAKEAWAY

What to remember

AlphaProtein Novo is a documented, author-measured de novo enzyme pipeline with Apache-licensed code and non-commercial weights. Use the wet-lab tables as their evidence, keep the “outperforms nature” line as an unreviewed claim, and wait for sequences and independent repeats before treating any design as a starting point you can order.

Sources & further reading

  1. Designing enzymes for new-to-nature chemistry and non-natural substrates with AlphaProtein Novo ↗
  2. google-deepmind/alphaprotein-novo ↗
  3. AlphaProtein Novo Generator Model Parameters Terms of Use ↗
  4. Apache License 2.0 (repository LICENSE) ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories