What did Hugging Face Bio publish?

The 8 October Hugging Face blog “Carbon-A: Finding genes in known and unknown genomes,” by Georgia Channing, says the lab is releasing the Carbon Annotation Database and Carbon-Annotator (Carbon-A). The model is described as 1.2 billion parameters, predicting protein-coding regions directly from DNA across mammals, other vertebrates, invertebrates, plants, fungi and protists, at nucleotide resolution on both strands, over a 98,304-base-pair context.

The same post says the current database holds 48,167 assemblies from 22,617 taxa, about 27 trillion base pairs and 566 million predicted protein-coding loci — about 11.0× more taxa and 9.0× more sequence than the RefSeq-derived training corpus. Roughly half of the target GenBank set is annotated, with another batch promised in three weeks. Fergal Martin of EMBL-EBI is quoted on the value of public genomes in the European Nucleotide Archive. Those totals and the quotation are in the blog; we did not recount assemblies. For a different biology-data build, see the Virtual Biology Initiative expansion.

The live Hub checkpoint is HuggingFaceBio/Carbon-A-1.2B. The Hugging Face model API we opened lists createdAt 2026-10-02T13:54:15Z, lastModified 2026-10-08T11:31:12Z, and license mit. The collection HuggingFaceBio/carbon-annotation-database, last updated 15:19 UTC on 8 October, lists the model, a database-explorer Space (id HuggingFaceBio/carbon-a-database-explorer, also advertised as HuggingFaceBio/genbank-annotation-explorer), training-data and genbank-annotations buckets, and a Carbon-A-Technical-Report bucket. That report bucket was not inspectable when we opened it, so we are not citing numbers from it.

How does Carbon-A work, and what scores are claimed?

The model card says Carbon-A uses a non-overlapping 6-mer tokenizer, 16,384 tokens, and two strand-specific classification heads that emit per-base non-coding versus CDS probabilities. Binary calls default to a 0.5 threshold. Training examples come from paired RefSeq GBFF annotations and FASTA on GCF assemblies. CDS intervals are unioned across isoforms independently on each strand. Windows without annotated CDS are kept. Transformers 4.56 or later is required, with trust_remote_code=True.

The blog’s headline comparison is a macro-averaged nucleotide F1 of 0.944 across 42 benchmark genomes, ahead of AUGUSTUS, Helixer, Tiberius, ANNEVO, OrionGeno, SegmentNT and NTv3 in that table. An animal-only checkpoint is said to annotate unseen plant genomes. Tetrahymena thermophila, which uses a nonstandard genetic code, is reported at 0.960 nucleotide F1. A pig-genome anecdote says the model scored better on a 2026 assembly than on the 2017 assembly it saw in training. All of those figures are Hugging Face’s.

The model card is explicit about the evaluation panel: it “was used during training; it is not a newly held-out test set.” Functional inference checks on an NVIDIA H100 are listed as code tests, “not a measure of agreement with reference CDS annotations.” That is the same distinction as reading a model card before treating a leaderboard as independent.

What did the wet lab show, and what did it not?

The interesting disagreements, the blog says, are genes Carbon-A calls that RefSeq does not. Hugging Face worked with ActiveSite and UCSD on PacBio Iso-Seq in cat, Syrian hamster, chicken and Arabidopsis. Iso-Seq reads full RNA molecules; if it supports a structure, that is evidence the transcript exists. Aggregate complete-CDS support is “roughly 0.62,” described as close to the reference and ahead of other ab initio systems in that comparison. That is transcription evidence, not a protein. The post says ribosome profiling and mass spectrometry are the next checks, and that further wet-lab results “we can’t share yet.” Compare that caution with how we treated AlphaProtein’s enzyme claims: a lab method is not a catalog of functions.

How do you run it, and what licence applies?

The card’s Python example uses AutoModelForTokenClassification and AutoTokenizer with BF16 and SDPA on CUDA. infer_genbank.py accepts GenBank or FASTA, including gzip, and writes per-base arrays in overlapping 98,304-base windows. Full-length windows need fused SDPA kernels; without them the card says attention matrices need more than 32 GB of extra GPU memory. CPU inference is described as slow. The model card and the card’s license field both say MIT, with Apache-2.0 notices retained in the model code. That is a permissive weight licence, not a dump of the genomes. See open weights versus open source.

A confidence score on each predicted gene is reported at AUROC 0.876 for distinguishing exact CDS matches from other predictions on Hugging Face’s benchmark. The blog points readers to the CADB Explorer at HuggingFaceBio/genbank-annotation-explorer; the collection still lists that Space as HuggingFaceBio/carbon-a-database-explorer with both hostnames live. We did not download the 4.6 GB weights or run either script.

What should readers not assume?

Carbon-A does not resolve isoforms, complete gene structures, bacteria or viruses. Ambiguous 6-mers are excluded. The evaluation panel is not a fresh hold-out. 566 million loci are model calls, not 566 million demonstrated proteins. The technical-report bucket in the collection was not inspectable when we checked, so we are not citing it. This is also not a replacement for maps such as AlphaGenome Atlas, which scores variant effects rather than calling CDS from raw eukaryotic assemblies.

Common questions

Did Hugging Face prove the new genes make proteins?

No. The Iso-Seq work is described as support for transcription and exon structure in four species. The blog says translation and function remain to be tested.

Can I run Carbon-A without a large GPU?

The card says full-length 98,304-base windows need fused SDPA kernels on a recent NVIDIA GPU, and that CPU inference is slow. We did not time a CPU run.

Is the comparison to AUGUSTUS an independent bake-off?

No. It is Hugging Face’s table on a panel the model card says was used during training.

THE TAKEAWAY

What to remember

Use Carbon-A if you want an MIT-licensed eukaryotic CDS caller and a large predicted GenBank layer, and if you can treat Iso-Seq support as transcription evidence only. Do not file the 566 million loci as new proteins, and do not treat 0.944 F1 as a held-out independent score.

Sources & further reading

  1. Carbon-A: Finding genes in known and unknown genomes ↗
  2. HuggingFaceBio/Carbon-A-1.2B model card ↗
  3. Carbon Annotation Database collection ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories