What was announced on 8 October?

The dated consumer-facing object is the Illumina press release “Illumina releases SpliceAI2 to help advance rare disease research,” on PR Newswire at 09:15 ET on 8 October 2026, datelined San Diego. It says the model helps researchers find disease-relevant splice variants that otherwise may go unnoticed. Quotes from Rami Mehio and Kyle Farh in that release are Illumina’s.

The science note, published at 13:32:47 UTC the same day, is the longer lab write-up. It names Laurel Mastro, Lisa Eldridge, and Robert Yamulla as blog authors and lists Jaganathan, Jianbin Chen, Xin Liu, and Yan Zhang as joint first authors, with collaborators at Oxford, UCSF, and the New York Genome Center. GitHub’s API record for Illumina/SpliceAI2, retrieved 9 October, gives created_at 2026-10-07T18:47:33Z, pushed_at 2026-10-08T17:31:53Z, language Python, stargazers_count 10, and license Other.

This is a genomic annotation model, not a new sequencer. For another 8 October DNA-sequence annotator, see Hugging Face Bio’s Carbon-A. For a searchable catalog of predicted variant effects, see Google’s AlphaGenome Atlas.

What does the manuscript say the model does?

The PDF linked from the repository README as Manuscript is titled “A unified framework for quantitative splicing and transcript prediction.” The abstract we opened says SpliceAI2 is trained end-to-end on splice site, splice junction, and transcript usage across multiple mammalian species, and that it reconstructs complete transcripts directly from genomic sequence. The main text says the network processes approximately 200 kb of sequence, predicts position-wise splice-site usage and pairwise junction usage, and builds a splice graph whose paths are candidate transcripts.

Training counts in that paper and in the science note match: 314,745 RNA-seq samples across human and nine other mammalian species, and 330 long-read samples from human and mouse. The science note adds “more than 46 million observed splice junctions after filtering” and names the public ENCODE project as the long-read source. The paper says the model perfectly reconstructed the most abundant transcript for 82% of held-out genes, against 78% when trained without long-read data. Those percentages are the authors’.

The science note says SpliceAI2 also conditions on 147 splicing regulators, and that it was weaker at predicting how a given variant’s impact changes between tissues than at capturing the tissue’s splicing program. The paper’s tissue section uses 147 RNA-binding-protein levels and 48 tissues. We did not rerun those regressions.

Which numbers are Illumina’s?

On GTEx, the manuscript’s first variant-effect benchmark is 4,545 cryptic splice variants and 573,967 null variants created by rare GT or AG dinucleotides. The authors report auPRC 0.77 against 0.66 for original SpliceAI, and Spearman ρ 0.63 for usage against 0.47 for Pangolin. On GTEx sQTLs they report auROC 0.76 against 0.73 for AlphaGenome. An MPRA line in the same paragraph is Spearman ρ 0.59 against 0.55 for SpliceAI. The science note says Oxford collaborators ran the AlphaGenome rows.

The PR’s 34% GTEx quantification sentence is not a number we extracted from the PDF pages we opened. Keep it in the PR column. The science note, echoed by the paper’s 627,000-genome analysis, says the highest-scoring variants were depleted toward loss-of-function levels and lined up with lower UK Biobank plasma protein. Those are Illumina’s associations.

On rare disease, the paper restricts Genomics England analysis to 7,504 probands with monoallelic disorders. At a fixed odds ratio of 2 it reports 133 excess panel-matched variants for SpliceAI2 and 114 for the next-best model. The PR’s “17% more disease-relevant variants” is that ratio. The science note separately says SpliceAI2 found 33% more disease-relevant splice variants than legacy SpliceAI at a 2× confidence interval and 66% more at 4×. The abstract’s operational claim is that predicted splice-altering variants account for 15% of the excess genetic burden, with about half in deep intronic regions. We did not open the Genomics England research environment.

Author-reported splice-effect rows we opened in the manuscript, not an independent rerun
Benchmark in the PDFSpliceAI2Comparator the authors name
GTEx cryptic GT/AG auPRC0.770.66 SpliceAI
GTEx cryptic usage Spearman ρ0.630.47 Pangolin
GTEx sQTL auROC0.760.73 AlphaGenome
GEL excess variants at OR 2133114 next-best model

What can you download, and under what terms?

The GitHub README we opened on 9 October points to the manuscript PDF, the science note, pip install spliceai2, and Hugging Face for weights and precomputed scores. PyPI’s JSON for spliceai2 2.0 gives upload_time_iso_8601 2026-10-08T11:30:17.891664Z, author Kishore Jaganathan, and license “SpliceAI2 Model Terms of Use.” The README requires a CUDA GPU. Recommended summary-score thresholds are 0.1, 0.25, and 0.5.

The Hub model API for illumina-ai/SpliceAI2 lists createdAt 2026-09-24T15:48:23Z, lastModified 2026-10-06T22:42:02Z, gated auto, and two 13M checkpoint files. The dataset API for illumina-ai/SpliceAI2-data lists createdAt 2026-09-25T23:11:27Z, lastModified 2026-10-06T22:42:03Z, gated auto, and precomputed_scores_v2.0/GRCh38 trees. A GET of the model README without credentials returned a restricted-access notice. We did not accept the gate.

The LICENSE file in the GitHub tree is not an OSI open-source grant. It licenses SpliceAI2 to non-commercial organizations for non-commercial research, forbids redistribution of the materials, and was last updated 8 October 2026. That is the restricted-weight pattern in open weights versus open source, not a from-scratch public dump. Commercial use is directed to AI_licensing@illumina.com.

Where else can researchers reach it?

The science note lists BioInsight products: DRAGEN Annotation, Emedgene, and Illumina Connected Insights. The PR names the first two for customers. Those are availability claims, not scores we generated.

The PR and the science note place SpliceAI2 beside PromoterAI and PrimateAI-3D and say the suite can identify up to twice as many variants with predicted biological impact. The science note’s Genomics England inset says cryptic splice variants added 15% of candidates beyond clearer protein-disrupting mutations. “Up to twice” is Illumina’s combined-suite claim. Original SpliceAI’s ClinGen citation count in the PR is the company’s pedigree sentence, not a recount we ran.

What did we not run?

This is an evidence review of the 8 October press release and science note, the manuscript PDF linked from the README, the 9 October GitHub API record and README/LICENSE files, the PyPI 2.0 JSON, and the Hugging Face model and dataset API records. We did not accept the Hub gate, download checkpoints, score a VCF, or open Genomics England, GTEx, gnomAD, TOPMed, or UK Biobank. Treat the method as what the PDF describes, the percentages as what Illumina printed, and the public installer as a gated-weight research drop.

Common questions

Is SpliceAI2 open source?

The GitHub tree and the PyPI sdist are public. The LICENSE we opened is SpliceAI2 Model Terms of Use for non-commercial organizations, last updated 8 October 2026. Hub weights and precomputed scores are gated. That is not an Apache or MIT grant over the weights.

Did Illumina beat AlphaGenome in an independent league?

The science note says Oxford collaborators ran the AlphaGenome comparisons. The numbers we quote are from Illumina’s manuscript tables. We did not rerun AlphaGenome or contact those collaborators.

Does this replace RNA-seq on the affected tissue?

Illumina’s pitch is that DNA-only prediction helps when the relevant tissue is hard to sample. The paper and the science note still treat RNA-seq and proteomes as the checks. A predicted cryptic splice is not a diagnosis.

THE TAKEAWAY

What to remember

Use the 8 October note and the manuscript PDF for the splice-graph recipe and keep every GTEx, MPRA, and Genomics England percentage in Illumina’s column. Use GitHub and PyPI for the installer; use the Hub API for the gated weight and score drops. Do not invent an open-weight license.

Sources & further reading

  1. Introducing SpliceAI2: The Next Generation of Splicing and Transcript Isoform Prediction ↗
  2. Illumina releases SpliceAI2 to help advance rare disease research ↗
  3. A unified framework for quantitative splicing and transcript prediction ↗
  4. Illumina/SpliceAI2 README ↗
  5. SpliceAI2 Model Terms of Use ↗
  6. Illumina/SpliceAI2 repository API ↗
  7. spliceai2 2.0 package JSON ↗
  8. illumina-ai/SpliceAI2 model API ↗
  9. illumina-ai/SpliceAI2-data dataset API ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories