Announced 6 Oct 2026 · Sources checked
What did Cohere release?
Cohere Labs released Tiny Aya L2-Thinker as an open-weight research model for multilingual reasoning. It has 3.35 billion parameters, accepts text, and can produce an explicit reasoning trace before its final answer. The design goal is practical: a question asked in Arabic, Hindi, Swahili or another supported language should be reasoned through in that language instead of silently switching to English.
That matters because a translated final answer can hide what happened inside the intermediate steps. A readable trace lets a user inspect terminology, cultural assumptions and logical mistakes in their own language. It does not make the answer automatically correct, and visible reasoning is not a complete account of a model’s internal computation. Our AI token guide explains why translation quality and task reasoning should be evaluated separately.
How does its multilingual reasoning work?
The project studies data mixing rather than building a separate reasoning model for every language. Cohere says the training recipe combines a large English reasoning set, smaller multilingual reasoning sets and multilingual non-reasoning instructions. English examples provide task-solving patterns; other-language examples teach the trace to follow the prompt; broader instructions help preserve language coverage.
Cohere reports that training only on English reasoning left the trace in the prompt language 12.8% of the time. Adding multilingual reasoning raised that measure to 86.1%, while the complete mixture reached 96.4% in the company’s ablation and 57.4% average task accuracy. The broader headline is above 93% in-language reasoning over 60 languages. These are creator-reported experiments, not an independent guarantee for every language or domain.
What can developers use today?
The model card lists a 32K combined input-and-output context window and provides examples for Transformers, vLLM, SGLang and Docker Model Runner. Thinking mode is enabled through the chat template and can be disabled for a direct answer. Cohere also released the multilingual reasoning dataset, making the training recipe easier to study.
The word open needs qualification. The Hugging Face repository is publicly listed but asks users to acknowledge conditions before obtaining files. The checkpoint is licensed CC BY-NC 4.0 and points to Cohere Labs’ acceptable-use policy. That permits research and other non-commercial use under the stated terms; it is not an Apache-style release. See our open weights versus open source explainer before planning a product around it.
How strong are the reported results?
Cohere evaluated six benchmarks covering mathematics, science and general reasoning, then measured both answer accuracy and the language of the reasoning trace. The company says performance changed little on average compared with its English-reasoning counterpart, except on competition mathematics, and that the model averaged fewer than 5,000 thinking tokens.
A 3.35B model is attractive because it needs less memory and compute than frontier systems, yet size does not remove the need for representative tests. Teams should sample the actual languages, scripts, dialects and tasks they serve; grade answers independently from reasoning-language compliance; and inspect code-switching, repetition and culturally specific errors. A model can reason in the requested language and still be wrong.
Where could it be useful?
The clearest uses are education, local-language assistance, evaluation research and tools that need inspectable intermediate work. A teacher could review whether a math explanation uses understandable terminology. A researcher could compare error patterns without translating every trace. A small local deployment could also be easier to control than a remote frontier API, depending on hardware and policy.
Developers should not expose long reasoning traces automatically in every product. Traces can reveal sensitive prompt material, create false confidence and add latency. A better interface may summarize the reasoning, show sources, or reveal steps only when a reviewer asks. Our AI answer verification checklist remains relevant: correctness must be checked against evidence, not inferred from a fluent explanation.
What are the important limitations?
This is a research release, not a claim of production readiness. Cohere notes variation across tasks and languages, and the model card includes an information cutoff in June 2024. The checkpoint can produce harmful, biased or inaccurate text, and a multilingual safety policy requires testing in every intended language rather than only in English.
The release is meaningful because it treats the language of reasoning as a first-class design target and publishes both weights and data. Its value will depend on independent replication, inference costs and whether users in low-resource languages find the traces genuinely clearer. Until then, the result is a strong research signal—not proof that multilingual reasoning is solved.
Common questions
Is Tiny Aya L2-Thinker free for commercial products?
Not under the published model license. The card lists CC BY-NC 4.0 plus Cohere Labs’ acceptable-use policy. Commercial users should obtain appropriate permission and review the final terms.
Does it support exactly 60 languages?
Cohere evaluates in-language reasoning across 60 languages. The model card says supervised reasoning covers 44 languages plus English, with broader coverage from additional multilingual instruction data. Quality can vary by language.
Can it hide the reasoning trace?
Yes. The provided chat template supports a no-thinking mode. Product teams should decide carefully whether a visible trace improves review or exposes unnecessary sensitive detail.
What to remember
Cohere Labs’ Tiny Aya L2-Thinker is a compact, open-weight research model trained to reason in the language of its prompt. Cohere reports more than 93% in-language reasoning across 60 languages and has released the model and supporting data. Developers can test it with common open inference tools, but should verify language-specific quality and note the gated, non-commercial license.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





