Kandinsky 6.0 was open-sourced on October 6, 2026, making it one of the freshest options for local and open AI video experimentation. The project combines video and synchronized audio generation instead of treating sound as a separate post-production step.
What is Kandinsky 6.0?
Kandinsky 6.0 Video is a family of diffusion models for synchronized text-to-audio-video and image-to-audio-video generation. The official project provides a 3B-parameter Lite model and a 29B-parameter Pro model. Both target five-second clips with synchronized 44 kHz audio and lip-sync support.
A separate super-resolution component can raise output to Full HD (1920×1080). The repository is published under the MIT license.
Kandinsky 6.0 in ComfyUI
The project includes an official ComfyUI path. According to its setup instructions, the kandinsky6 and kandinsky6-sr extensions can be installed through ComfyUI Manager, followed by a ComfyUI restart.
If you do not already have ComfyUI running, start with our ComfyUI installation guide. Then see our video workflow guide for the broader pipeline structure.
Hardware: read the benchmarks carefully
The upstream project publishes performance measurements across GPUs including RTX 4090, 5060 Ti, 5080, 5090, RTX PRO 6000, A100 and H100. These are project-reported benchmarks, not The Signal test results. The official quick start currently calls for an NVIDIA GPU and Python 3.13 or 3.14.
Why this release matters
The interesting part is not simply another text-to-video model. Kandinsky 6.0 puts synchronized audio, lip-sync, image-to-video, video generation and super-resolution into one open project with a ComfyUI route. That makes it particularly interesting for automated content pipelines.
What we want to test
The practical questions are GPU cost, VRAM behavior, generation time, lip-sync quality and whether the Lite model is useful on rented consumer GPUs. We will publish those as first-party results only after reproducing them.
Kandinsky 6.0 specifications
| Feature | Lite | Pro |
|---|---|---|
| Parameters | 3B | 29B |
| Modes | Text-to-audio-video, image-to-audio-video | Text-to-audio-video, image-to-audio-video |
| Clip length | 5 seconds | 5 seconds |
| Audio | Synchronized 44 kHz, including lip-sync | Synchronized 44 kHz, including lip-sync |
| Super-resolution | Separate Kandinsky 6 SR component; up to Full HD in the documented pipeline | |
| ComfyUI | Official extensions available through ComfyUI Manager | |
Kandinsky 6 vs a conventional text-to-video workflow
| Workflow concern | Kandinsky 6 approach | Conventional video-only model |
|---|---|---|
| Sound | Video and synchronized audio generated together | Often requires a separate audio/TTS/SFX stage |
| Lip-sync | Included in the project’s synchronized generation target | Often a separate post-process |
| Image-to-video | Supported with audio | Model-dependent |
| Upscaling | Dedicated Kandinsky 6 SR project | Usually external upscaler |
| Local setup | NVIDIA GPU; official Python/ComfyUI paths | Varies by model |
This is a capability comparison, not a claim that Kandinsky 6 has better visual quality than every competing model. Quality, speed and cost require controlled first-party tests.
Which Kandinsky 6 model should you start with?
Lite is the more logical first target when exploring the architecture because it is the smaller 3B model. Pro is 29B and targets the higher-capacity end of the family. The upstream repository also exposes distilled variants and device presets. Hardware feasibility depends on the exact variant, resolution, offloading strategy and GPU; do not infer a universal VRAM requirement from parameter count alone.
Kandinsky 6.0 FAQ
Is Kandinsky 6 open source?
The project states that Kandinsky 6.0 and its video super-resolution component were open-sourced on October 6, 2026. Check the repository and model-card licenses for the exact artifact you plan to redistribute or use commercially.
Does Kandinsky 6 work with ComfyUI?
Yes. The official project instructs users to install kandinsky6 and kandinsky6-sr through ComfyUI Manager and restart ComfyUI.
Can Kandinsky 6 generate audio with video?
Yes. Its core proposition is synchronized video and 44 kHz audio generation, including lip-sync, for both text-to-audio-video and image-to-audio-video modes.
Can Kandinsky 6 generate 1080p video?
The project provides a separate super-resolution component that can raise output to Full HD (1920×1080). This is an upscaling stage, not evidence that the base diffusion pass natively generates every clip at 1080p.
What hardware does Kandinsky 6 require?
The official quick start requires an NVIDIA GPU and Python 3.13 or 3.14. Actual memory and generation-time requirements vary by Lite/Pro variant, resolution, attention backend and offloading.


