Announced 10 Oct 2026 · Sources checked
What changed on 10 October?
HAL-X’s DEV Community post dated 10 October 2026 is the access event for this story. Co-founder Farid Aghayev writes that the team open-sourced THX-01 “a couple of days” earlier and that the hosted API is now open with “no API key, no sign-up, and no payment required,” subject to rate limits. The same post cites vendor usage figures since launch—more than 25 million tokens and more than 211,000 decisions—which we treat as HAL-X’s claims, not independently audited traffic.
The weight files themselves are older than the free-API post. Hugging Face lists doofz/THX-01 as created on 23 September 2026, with a lastModified timestamp of 8 October 2026 on our 11 October check. The card and LICENSE file state Apache License 2.0. Pip install paths (pip install thx01) and a Space demo are documented on the same Hub page. So the fresh news is hosted access and the DEV framing of a free public endpoint, not a brand-new checkpoint date.
HAL-X’s HTML API reference at api.hal-x.ai/docs/thx-01 describes Authorization Bearer keys and per-key grants for /v1/systemone. That documentation and the DEV “no API key” claim sit side by side. On 11 October we sent an unauthenticated POST with a tiny choice question to https://api.hal-x.ai/v1/systemone and received HTTP 200 with a THX-01 answers object and a latency_ms field. Readers should still expect keys, grants, and rate limits as the durable production path; treat our single probe as a snapshot of endpoint behavior that day, not a permanent SLA.
What kind of model is THX-01?
THX-01 is a decision model, not a chat LLM. The Hub card lists about 322 million parameters—roughly 307M in an mmBERT-base encoder and 15M in a decision head—with up to 1,024 tokens of context per question after serialising the question, options, and state. Every question is packed as a single sequence with [MASK] slots over options; a two-layer head scores those positions and a temperature-fitted softmax returns probabilities. There is no token-by-token generation, so usage reports output_tokens as zero.
That design matches the broader October decision-model wave we have already covered for Microsoft-Decision-1 in Foundry and open weights such as Vega’s physics-style scorer. THX-01’s differentiator in HAL-X’s own framing is native number lookup, verbatim excerpt spans with character offsets, and citation payloads alongside TypeSafe-compatible choice, noul, and score types—capabilities the card says service layers otherwise emulate with repeated choice calls.
Post-training is described as eight Reinforcement Learning for Calibrated Decisions (RLCD) stages on about 2.1 million training decisions. The reward mixes logarithmic, spherical, and ranked-probability scores so that honest probabilities maximise expected reward. Soft targets teach “no good option” honesty and near-miss credit on ordinal scales. Languages emphasised in post-training include Azerbaijani, Russian, and English, with typed regional noise such as Azerbaijani without diacritics and Russian in Latin transliteration. Through the encoder, HAL-X claims coverage of more than 100 languages.

How do you call it, and what do the scores mean?
The main HTTP path is POST /v1/systemone (alias /v1/decide) with a state string or object and a questions map. Batch endpoints accept up to 1,024 items. Choice questions take criteria as key→description pairs; noul returns P(yes); score takes an ordered legend; number returns a value literally stated in the document (normalised, without unit conversion); excerpt returns a verbatim span; cite:true attaches supporting passages. Questions with more than 24 options run a two-round tournament according to the docs.
HAL-X publishes a TypeSafe migration note: change the base URL to api.hal-x.ai, set model thx-01, and keep the System One shape used by products built around TypeSafe’s Jev line. That compatibility claim is useful for routing experiments, but it is not an affiliation: THX-01 is an independent HAL-X release.
On HAL-X’s 15-category support-ticket suite (2,843 tickets across Clean, Corrupted, Messy, and Independent sets in Azerbaijani, Russian, English, and Turkish), the card reports THX-01 at 98.4% average accuracy with ECE 0.003 and about 10 ms latency, versus TypeSafe Jev 1.13 at 97.4% / 0.007 ECE / 331 ms in the same table. Claude Sonnet 5.5 appears only on a 200-ticket stratified subset. Extraction rows on held-out multilingual documents list number lookup at 93.4%, excerpt F1 at 84.1, and citation accuracy at 94.0%. All of these are author-reported; we did not rerun the suites or measure latency on matched hardware.

What should builders actually test?
Start with thresholds, not averages. Because probabilities are part of the API, a policy such as “auto-act above 0.8, otherwise escalate” matters more than a single leaderboard cell. HAL-X says calibrated confidence can automate about three quarters of its ticket traffic without an error on the published sets; that selective-automation curve is again the vendor’s plot. On your corpus, measure false automations separately for Clean versus noisy text.
Watch the native extraction contract. Number and excerpt locate text that is present; they do not convert currencies, convert miles to kilometres, or invent paraphrases. If your pipeline needs arithmetic or unit conversion, keep that in application code. Banking77-style fine intent sets and six-way emotion labels remain weak spots on the card (65.9% and 44.6% respectively), with correspondingly lower confidence—useful honesty if you were expecting LLM-style paraphrase robustness.
For local deployment, pip install thx01 pulls weights from the Hub on first use. Server extras expose the same REST shape. The free hosted path is attractive for prototypes; production still needs key management, concurrency limits, timeouts, and a fallback path when the gateway returns 429 or 5xx. Keep decision logs with request IDs from the X-Request-ID header if you escalate to humans.

How does this sit next to other decision models?
October’s decision-model shelf is crowded. Microsoft’s Foundry-hosted Decision-1 emphasises enterprise routing and judging; open engines such as Vega explore physics-inspired scoring on frozen backbones; Liquid and others ship edge-oriented variants. THX-01’s pitch is a small multilingual encoder with calibrated System One answers plus document-grounded number/excerpt tools. For contrast on open generative approaches to the same interface, see FuturixAI’s CARVE Jev-like family, which scores candidate tokens from a Qwen3.8-27B generative model and explicitly warns that its probabilities are not calibrated correctness estimates.
Practical selection still depends on your constraints: whether you need open weights, whether you need citation spans, whether 1,024 tokens of state is enough, and whether your evaluation set looks like HAL-X’s tickets or like long agent traces. Do not collapse “beats Jev on this table” into a general superiority claim—HAL-X itself cautions that latency rows use different serving setups.
Common questions
Is THX-01 new on 10 October, or only the free API?
The free hosted API announcement is dated 10 October 2026 on DEV. The doofz/THX-01 Hub repository was created 23 September 2026 and last showed an 8 October modification on our check. Treat weights and API access as separate dates.
Did Ai Lookout verify the free endpoint?
On 11 October we sent one unauthenticated POST to /v1/systemone and received HTTP 200 with a decision payload. HAL-X docs also describe API keys and grants. We did not load weights, run the ticket suite, or measure production rate limits.
Are the 98.4% ticket scores independent?
No. They come from HAL-X’s published tables on the Hub card and API docs. Competitor latency figures in those tables use each vendor’s own serving path.
What to remember
If you need typed, citation-aware decisions without generating prose, THX-01’s free API and Apache weights are worth a controlled pilot—with your thresholds and your data, not HAL-X’s ticket average alone.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





