What did Kandinsky Lab open on 6 October?

The GitHub README for kandinskylab/kandinsky-6 is the release log we used. Under Project Updates dated 2026/10/06 it states that Kandinsky 6.0 and Kandinsky 6.0 Video Super-Resolution are open sourced, that a Hugging Face Space for Kandinsky 6.0 Pro Distill was added, and that the models are available in Diffusers, ComfyUI, vLLM-omni, SGLang, and FastVideo. The lead paragraph defines the family: Lite at 3 billion parameters and Pro at 29 billion, both producing 5-second clips with synchronized 44 kHz audio, including lip-sync, for text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes.

The technical report is arXiv:2610.05608, “Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation,” with the arXiv API record showing publication timestamp 2026-10-04T23:19:20Z. The abstract matches the README sizes and durations, names a dual-stream CrossDiT that joins a pretrained video stream to a newly trained audio stream with bidirectional cross-attention, and says the authors release code, checkpoints, and Diffusers integration under the MIT license. The LICENSE blob we fetched begins “The MIT License (MIT)” and “Copyright (c) 2026 Kandinsky Lab.” For readers tracking open media models alongside chat UIs, compare the license posture with earlier image drops such as Qwen-Image 2.1 Turbo and Google’s Nano Banana 2.1 card.

The open-source checklist in the README is about runnable software and weights, not about publishing a full training corpus. The abstract’s claim that Pro “clearly outperforms” Kandinsky 5.0 Video Pro in side-by-side human evaluation, and that it remains competitive on speech quality with leading audio-video systems, is the authors’ evaluation language. We did not rerun those comparisons.

Wooden black-and-white striped clapsticks resting on a film label marked Première. No people appear.
Film clapsticks and a Première reel label photographed 24 April 1992 by Bernd Schwabe in Hannover. CC BY-SA 4.0 via Wikimedia Commons (File:1992-04-24 Filmklappe Filmschindel Gedanken 2.jpg). No people appear. Archival production props; they do not depict Kandinsky 6.0 or a model demo. Cropped and converted to WebP. Photo: Bernd Schwabe in Hannover. CC BY-SA 4.0 · Cropped and resized.

How is the model built?

Architecturally, the report describes a dual-stream CrossDiT: a video stream carried forward from Kandinsky 5.0 capabilities and an audio stream trained from scratch, then joined through bidirectional cross-attention so timing and semantics stay aligned instead of being bolted on after the fact. Continuous pretraining first trains the audio stream on large-scale audio corpora, then trains both streams jointly on paired audio-video data while trying to preserve unimodal fidelity. Supervised fine-tuning, reinforcement-learning post-training, and distillation follow. Distilled checkpoints appear in the Hugging Face catalog names the README lists (for example Pro Distill and Lite Distill).

The generation contract is short and specific: five seconds of video plus synchronized 44 kHz audio in one pass, with lip-sync called out in both the README and the abstract. Resolution starts below Full HD in the serving defaults shown for vLLM-Omni (864×480, 125 frames at 24 fps, 50 steps, guidance 5.0 in the README’s Pro example) and can be raised with the separate kandinsky-6-sr super-resolution repository toward 1920×1080. Image-to-audio-video mode uses an input image as a masked tail frame in the vLLM-Omni example the README prints.

Because audio is generated with the video rather than attached later, failure modes differ from silent video models: speech sync, ambient sound, and music all sit inside the same sampling budget. That is the product claim. It is also why the authors emphasize speech quality in the human-evaluation summary. Independent listen tests are still required before anyone should treat that summary as a ranking.

How do you run it, and what does it cost in time?

The README’s quick start assumes an NVIDIA GPU and Python 3.13 or 3.14, plus `uv` and `just`. `just setup` picks a PyTorch build from GPU compute capability: Hopper (9.0) gets CUDA 13.0 with FlashAttention 3; Ampere, Ada, and consumer Blackwell get CUDA 13.0 and then compile SageAttention 2.2.0, which needs `nvcc` on PATH. The first generate downloads Kandinsky-6.0-Pro-distill-5s into `$KANDINSKY_HOME/weights` (defaulting to `~/.cache/kandinsky`) and writes an MP4 under a timestamped output directory. Catalog names listed for download include pro, pro-pretrain, lite, lite-distill, and lite-pretrain.

Integration paths named on 6 October are unusually broad for a lab video drop: Diffusers docs, a ComfyUI Manager node (plus a separate SR node), vLLM-Omni with an offline example and an online `/v1/videos` sketch, SGLang cookbook pages, and FastVideo. That breadth is useful for teams already standardized on one of those hosts. It is not evidence that every path exposes the same quality or memory profile. For adjacent voice/video product context see HeyGen’s voice-controlled Arena.

The README’s Performance table reports working time in seconds for a 5-second clip on the non-distilled model after warmup, excluding weight loading and MP4 encoding. Examples from that table: Pro Full HD is listed at 1247 s on an RTX 4090 and 402 s on an H100; Lite SD is 437 s on an RTX 4090 and 239 s on an H100. Consumer cards use block offload in the authors’ notes; several data-center cards use module offload. Those seconds are vendor measurements on named GPUs, not our stopwatch.

Author-reported working time (seconds) for a 5-second non-distilled clip (README Performance table)
SettingRTX 4090RTX 5090H100
Lite SD437309239
Lite Full HD578406284
Pro SD936754356
Pro Full HD1247854402
Vintage Western Electric mixing console and audio rack in a museum display. No people appear.
Western Electric 23C speech-input mixing console and studio rack at MIM Phoenix, photographed 4 December 2017 by bobistraveling. CC BY 2.0 via Wikimedia Commons. No people appear. Museum audio gear; it does not depict Kandinsky 6.0’s generated 44 kHz track or a GPU render. Cropped and converted to WebP. Photo: bobistraveling. CC BY 2.0 · Cropped and resized.

What should builders do with it?

Start with the license. MIT on code and, per the abstract, on checkpoints and Diffusers integration removes the territorial and revenue-tier hedges that have shown up in other “open” video releases this season. Still read the third-party notices inside the tree and confirm the specific Hugging Face repository card for the checkpoint you download before shipping a commercial pipeline.

For a first technical spike, the distilled Pro Space and `just generate` path are the lowest-friction checks that the stack installs on your GPU. Then pin one host—Diffusers or ComfyUI for interactive work, vLLM-Omni if you need an HTTP video API—and freeze the exact catalog name (pro-distill versus full pro) in your runbook. Keep a side-by-side folder of prompts that stress speech, music, and silent ambient sound; the authors’ speech-quality claim is the dimension most likely to disappoint if your use case is music-heavy.

Memory planning belongs next to quality. The README’s offload notes and the long Full HD times on consumer GPUs mean a “runs on 16 GB” marketing line from secondary writeups should be verified against the preset YAML you actually load. If you need longer than five seconds, that is outside the contract stated in the README lead—treat extensions as your own research.

  1. Clone kandinskylab/kandinsky-6, read LICENSE (MIT) and the 6 October Project Updates.
  2. Install with `just setup` on a supported NVIDIA GPU; try `just download pro-distill` then `just generate`.
  3. Pin Diffusers, ComfyUI, or vLLM-Omni as the production host and record the exact Hub revision.
  4. Score speech lip-sync and ambient audio on your own prompts before replacing a closed T2AV vendor.

What remains unverified?

We did not download multi-gigabyte weights, generate a clip, reproduce the Performance table, or re-run the paper’s side-by-side human study against Kandinsky 5.0 or proprietary systems. Hub download counts and Space uptime were not audited. Training-data composition beyond the abstract’s high-level recipe is not reconstructed here.

Secondary articles sometimes paraphrase memory presets and preference scores; when those numbers are not in the README or abstract paragraphs we opened, we leave them out rather than launder them into our voice.

Common questions

Is Kandinsky 6.0 Video really MIT-licensed?

The LICENSE file in kandinskylab/kandinsky-6 is the MIT License dated copyright 2026 Kandinsky Lab. The arXiv abstract also says code, model checkpoints, and Diffusers integration are released under MIT. Confirm the card on the specific Hugging Face checkpoint you use.

How long are the generated clips?

The README and abstract both describe 5-second clips with synchronized 44 kHz audio. A plug-in super-resolution model is documented for raising resolution toward Full HD.

Did Ai Lookout run Kandinsky 6.0?

No. This article is an evidence review of the GitHub README, MIT LICENSE, and arXiv 2610.05608 record. Latency and human-eval claims remain the authors’.

THE TAKEAWAY

What to remember

Kandinsky 6.0 Video’s 6 October open-source update puts MIT-licensed Lite 3B and Pro 29B joint audio-video models on GitHub and Hub with Diffusers and ComfyUI paths. Use the README for install and GPU-second tables; keep quality rankings attributed to the authors until you measure your own prompts.

Sources & further reading

  1. kandinskylab/kandinsky-6 README ↗
  2. LICENSE (MIT) ↗
  3. Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation ↗
  4. arXiv API record 2610.05608 ↗
  5. Kandinsky-6.0-Pro-distill-5s-Diffusers ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories