Announced 5 Oct 2026 · Sources checked
What was posted on 5 October?
On 5 October 2026 a DeepMind-led team posted “Designing enzymes for new-to-nature chemistry and non-natural substrates with AlphaProtein Novo” on bioRxiv. The server marks it as a preprint that has not been certified by peer review. Corresponding author Jue Wang is at Google DeepMind. Co-authors include Frances H. Arnold at Caltech, Demis Hassabis and Pushmeet Kohli, and further Caltech and University of Pittsburgh affiliations.
The same day, google-deepmind/alphaprotein-novo appeared on GitHub. We inspected the bioRxiv HTML, the PDF, the README and the license files on 7 October and did not find a matching DeepMind or Google newsroom post. The abstract’s claim that de novo design can, “for the first time,” outperform natural-sequence mining is the authors’ line on an unreviewed manuscript.
A preprint plus a code drop is still a research announcement, not a product launch. For how to keep those layers apart, see our note on reading AI company announcements.
How does AlphaProtein Novo work?
AP Novo scaffolds a catalytic motif—side chains plus ligand poses chosen for a hypothesized mechanism—inside a new backbone. A diffusion model co-generates coordinates and sequence. Optional LigandMPNN redesigns the sequence. Candidates are then filtered with AlphaFold 3, or a leaving-atom fine-tune the repository calls AF3-LA, using metrics the authors say track catalytic geometry across predicted states.
Filtering is the methodological core. Multi-seed scoring repeats AlphaFold 3 with different random seeds. Multi-resequence scoring asks whether several LigandMPNN sequences on the same backbone still pass those filters. Partial diffusion adds a little noise and re-denoises nearby sequences for the same test. The authors say combining those ensemble filters raised later-round hit rates.
They distinguish novel-scaffold designs, which start from a motif only, from recycled-scaffold designs that reuse a backbone already known to be active. Computational proposals still need wet-lab checks; that is the same split as in our earlier note on Claude’s enzyme-discovery work.
What did the piperidine and DEHP campaigns show?
The campaigns they put against natural-sequence mining are piperidine synthesis by intramolecular nitrene transfer and partial hydrolysis of the plasticiser DEHP. A screen of 188 natural and engineered heme proteins found nine above 5% C–H amination relative to wild-type TamPgb; none exceeded a 30:70 piperidine:pyrrolidine ratio. AP Novo then tested 86 novel-scaffold histidine-heme designs and 90 more aimed at piperidine. GDM_NT_121 reached up to 94:6 piperidine:pyrrolidine. The lead, GDM_NT_270, delivered 22 total turnovers, a 99:1 regioisomeric ratio and 94% enantiomeric excess—more piperidine turnover than TamPgb wild type, they say, through complementary regioselectivity rather than a higher overall rate.
On DEHP they monitored MEHP, the first hydrolysis product. Later rounds reached an 11% novel-scaffold hit rate and a 39% recycled-scaffold hit rate against a 2% of QHH21706.1 threshold. Purified GDM_DEHP_378 made 6% of QHH’s MEHP after four hours at room temperature, or 26% of Q7SIG1 (13% and 34% after 24 hours). GDM_DEHP_376 (191 residues) was 14-fold more active at 90 °C than at room temperature and twice as active in 75% acetonitrile; both natural controls were inactive under those conditions. Table 1 lists 3.3, 47 and 40 turnovers for the 23 °C / 4 h, 90 °C / 4 h and 75% acetonitrile / 24 h settings. The authors say full conversion to phthalic acid is still required for bioremediation.
| Reaction | What they tested | Lead result they report |
|---|---|---|
| Piperidine nitrene transfer | 188 natural or engineered proteins vs 176 AP Novo designs | GDM_NT_270: 22 TTN, 99:1 r.r., 94% e.e. |
| DEHP hydrolysis | 165 novel-scaffold and 184 recycled-scaffold designs | GDM_DEHP_376: 47 TTN at 90 °C / 4 h; naturals inactive there |
| Kemp elimination | 1,536 novel-scaffold designs | GDM_KE_1872: kcat/Km 44,940 ± 220 M−1 s−1 at pH 10 |
| Serine esterase (4MU-Ac) | 1,569 novel-scaffold designs | GDM_SE_2937: kcat/Km 360,000 ± 70,000 M−1 s−1 (NSE topology) |
| Carbene transfer | 91 novel-scaffold designs | GDM_CT_103: >99:1 d.r., 96% e.e. S,S |
What did the model reactions add?
While building the pipeline they also ran Kemp elimination, 4-methylumbelliferyl acetate hydrolysis and styrene cyclopropanation. On 1,536 novel-scaffold Kemp designs, hit rates (multiple-turnover kobs above 0.038 s−1) rose from 3% to 34% in the best facet. GDM_KE_1483 reached kcat/Km 14,600 ± 1,100 M−1 s−1 at pH 7.5—70-fold above earlier novel-scaffold reports, they say, and 10- to 100-fold below evolved designs. At pH 10, GDM_KE_1872 reached 44,940 ± 220 M−1 s−1.
Among 1,569 novel-scaffold serine esterases the best facet hit 44%. The overall lead, GDM_SE_2937 (kcat/Km 360,000 ± 70,000 M−1 s−1), has natural serine-esterase topology (TM-score above 0.45 and triad RMSD under 5 Å to a PDB hydrolase); six of their top seven designs did. The best non-NSE design, GDM_SE_2754, reached 4,900 ± 700 M−1 s−1 with substrate inhibition. On 91 carbene-transfer designs, 85% bound heme; GDM_CT_103 and GDM_CT_162 reached the stereoselectivities in the table in one round. Selected active-site knockouts cut activity, and tested designs kept activity after ten minutes at up to 80 °C.
How can researchers access the code and weights?
The repository LICENSE is Apache 2.0 and covers the pipeline code, not the trained generator. Weights (generator.bin.zst) sit on Google Cloud Storage under a 5 October terms file: non-commercial use by universities, nonprofits, research institutes, educational, journalism and government bodies; no commercial research; no training similar de novo models on the outputs; no publishing the parameters outside the licensee’s organisation. AlphaFold 3 and AF3-LA weights have their own terms. That split matches our explainer on open weights versus open source.
The paper’s “Model availability” paragraph says code and weights are available freely for non-commercial use. Designed sequences and predicted structures “will be made publicly available upon publication,” which has not happened. We have not run the pipeline. Neighbouring DeepMind biology releases this month include SynthID Bio protein watermarks and the AlphaGenome Atlas; those are separate systems.
Common questions
Can a company use AlphaProtein Novo to design commercial enzymes?
The repository code is Apache 2.0. The generator weights and their outputs are limited to non-commercial organisations and non-commercial use, and they bar training similar de novo design models on those outputs. Review both documents, plus the AlphaFold 3 terms, before any commercial plan. This is not legal advice.
Did the team produce a finished DEHP bioremediation enzyme?
No. They report first-step hydrolysis to MEHP and say higher activity and full degradation to phthalic acid are still required. The stability in heat and 75% acetonitrile is the property they contrast with two natural DEHPases.
Is the paper peer-reviewed?
No. bioRxiv posted version 1 on 5 October 2026 and labels it as not certified by peer review. Designed sequences are promised upon journal publication.
What to remember
AlphaProtein Novo is a documented, author-measured de novo enzyme pipeline with Apache-licensed code and non-commercial weights. Use the wet-lab tables as their evidence, keep the “outperforms nature” line as an unreviewed claim, and wait for sequences and independent repeats before treating any design as a starting point you can order.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





