What landed on 8 and 9 October?

The Hub repository FINAL-Bench/Darwin-27B-ZTC-v2 is the dated object. The model API we queried lists createdAt 2026-10-08T15:28:29Z and lastModified 2026-10-09T03:28:23Z. Tags include darwin, ztc, zero-token-classifier, decision-engine, system-one, and s1mb. Likes were 36 and downloads 55 on the copy we opened. A download counter is not a quality signal.

The README title is “Darwin-27B-ZTC-v2: a zero-token decision engine.” The citation block names VIDRAFT and FINAL-Bench, year 2026. v2 is described as adding control and workflow decisions — games, drones, retrieval control, entity alignment, customer and incident workflows — on top of FINAL-Bench/Darwin-27B-ZTC. The S1MB results paragraph names a published results file of 2026-10-09 02:44 UTC and a maintainer merge into hotchpotch/s1mb-result under FINAL-Bench__Darwin-27B-ZTC-v2.

This is a typed decision checkpoint, not a chat model. It sits next to, and is not a test of, Celeris’s hosted decision API or OpenAI’s Decisions API beta.

What is a zero-token decision engine on this card?

The README says the model reads a piece of state and a typed question — noul for yes or no, choice for one of N labels, score for an ordered rubric — and returns a probability for every option. It uses one forward pass per question and generates no tokens. That is the same interface family as v1, which the v1 card we opened describes as a POST /v1/systemone request shape.

The v1 card, which we opened the same day, reports 0.743 accuracy, KL 0.204, and Brier 0.097 on the LocalLLaMA/typed-decisions test split: 400 cases, 2,000 decisions, zero-shot, with per-type accuracies 0.845 noul, 0.732 choice, and 0.675 score in that table. The v2 README says v2 was not trained on Typed Decisions data. Those v1 figures stay on the v1 card. We did not rerun that split.

A probability over named options is also the object of an 8 October USC preprint that treats typed models as allow-or-block gates. That paper is the option-channel attack note; it does not evaluate Darwin-27B-ZTC-v2. For why a Hub license field is not the same as a complete training dump, see open weights versus open source.

How do they say v2 was trained?

The README’s three-step recipe is: start from Darwin-27B-ZTC; continue full-weight training on the Open-Jev train split mixed with v1’s original training data, selecting the checkpoint on held-out development rows only; average the weights of v1 and the continued model 50/50. The average, they write, keeps v1’s skills on its original tasks and adds the new control and workflow skills.

Training overlap is disclosed on the same page. v2 was trained on the train split of ZefanCai/Open-Jev, marked CC0. S1MB includes 22 benchmarks built from the test split of the same Open-Jev tasks. The authors say they used no S1MB test case: an exact-match filter over every S1MB test state, query, and document removed zero training rows. They also say v2 was not trained on any S1MB data. That filter is theirs. We did not replay it.

The backbone class in the Hub config is Qwen3_5TextModel. The name on the repo is Darwin-27B. The safetensors total is about 26B parameters. Those are distinct publisher-facing representations. Do not collapse them into a newly counted 27 billion.

Which S1MB numbers are publisher-reported?

The README table is the authors’ application of the leaderboard’s viewer/src/lib/borda.ts and types.ts to the 9 October 02:44 UTC results file, over 102 models with complete results. Borda, they write, ranks models on every benchmark, gives 100 to first and 0 to last, and averages over 137 benchmarks. Task Avg is the mean of baseline-adjusted Noul, Choice, and Score scores.

On that table Darwin-27B-ZTC-v2 is first: Borda 89.58, Task Avg 66.46, Noul 67.66 (59 tasks), Choice 71.51 (57), Score 60.21 (21). Second is openjev/openjev at 87.50 / 62.60. Third is denis-pplx/AutoJev-27B at 87.07 / 60.80. The card says qasper-noul-test-v1 ran with --context-limit 32768 because 12 decisions exceed the default 8192 tokens, and that nothing was truncated. The evaluator adapter is named autojev at revision a8f283a.

The metadata JSON we opened at hotchpotch/s1mb-result/.../FINAL-Bench__Darwin-27B-ZTC-v2/metadata.json lists model_id FINAL-Bench__Darwin-27B-ZTC-v2, display_name Darwin-27B-ZTC-v2, and the Hub URL. It does not reprint the Borda row. The public S1MB Space is a JavaScript board; we used the card table and the results-tree metadata, not a screenshot of the Space.

Publisher-reported S1MB english-v1 rows from the v2 README, not a board we recomputed
ModelBordaTask AvgScore (21)
Darwin-27B-ZTC-v289.5866.4660.21
openjev/openjev87.5062.6054.83
denis-pplx/AutoJev-27B87.0760.8049.34
caiovicentino1/Eikos-27B85.4359.8647.83

What can you actually download?

The API siblings list eleven BF16 model-*-of-00011.safetensors shards, model.safetensors.index.json, readout.safetensors, decision_config.json, tokenizer.json, tokenizer_config.json, chat_template.jinja, autojev/model.py, autojev/types.py, autojev/LICENSE, ztc_server.py, and ztc_engine.py. The README says ztc_server.py is a POST /v1/systemone server and ztc_engine.py is an in-process engine for the Decision Index kit. autojev is credited to github.com/denis-pplx/autojev, MIT, with text-only backbone support added.

The usage snippet on the card calls snapshot_download("FINAL-Bench/Darwin-27B-ZTC-v2") and DecisionModel.predict. The v1 card, which v2 says it follows for layout, says the runtime needs torch, transformers with Qwen3.5 support, safetensors, pillow, and a GPU with space for about 54 GB of BF16 weights. usedStorage on the v2 API is about 51.3 GB. Those are Hub and card figures. We did not allocate that GPU.

Apache-2.0 is on the cardData.license field. That covers the files the Hub marks with that license. It is not a claim that every upstream Qwen tokenizer term disappeared, and it is not a hosted API.

What did we not verify?

This is an evidence review of the 8–9 October Hub API record and README, the v1 card, and the S1MB results metadata JSON. We did not load the backbone, call the systemone server, recompute Borda, or replay the Open-Jev overlap filter. Treat the method as what the README describes, the S1MB row as what the publisher printed after a maintainer merge, and the download as an Apache-2.0 27B-class decision checkpoint.

Common questions

Is this the same model as Darwin-27B-ZTC?

No. v2 starts from that checkpoint, continues training, and averages weights 50/50 with v1. The v1 typed-decisions 0.743 row stays on the v1 card. The v2 README says v2 was not trained on Typed Decisions data.

Did FINAL-Bench beat Celeris or OpenAI Decisions?

The measured board is S1MB. Those hosted APIs are not in the table we opened. A shared typed-question shape is not a transferred score.

Should this be an allow-or-block gate?

The card does not make that claim. An 8 October USC preprint argues that some open-weight typed gates fail open. That paper does not list Darwin-27B-ZTC-v2.

THE TAKEAWAY

What to remember

Use the 8–9 October Hub card for the zero-token interface, the 50/50 merge recipe, and the overlap disclosure. Keep the S1MB Borda row in FINAL-Bench’s column. Do not paste those ranks onto a hosted decision API, and do not treat “27B” and “26B parameters” as one verified count.

Sources & further reading

  1. Darwin-27B-ZTC-v2 model card ↗
  2. Darwin-27B-ZTC-v2 README ↗
  3. FINAL-Bench/Darwin-27B-ZTC-v2 model API ↗
  4. Darwin-27B-ZTC README ↗
  5. S1MB result metadata for Darwin-27B-ZTC-v2 ↗
  6. S1MB result tree for Darwin-27B-ZTC-v2 ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories