Home/ MODELS/ Kandinsky 6.0 Is Open Source: Video + Audio Generation in ComfyUI
MODELS

Kandinsky 6.0 Is Open Source: Video + Audio Generation in ComfyUI

Kandinsky 6.0 is now open source with synchronized video and audio generation, Lite and Pro models, Full-HD super-resolution, and a ComfyUI integration.

Published Oct 7, 2026 · 4 min read

Kandinsky 6.0 was open-sourced on October 6, 2026, making it one of the freshest options for local and open AI video experimentation. The project combines video and synchronized audio generation instead of treating sound as a separate post-production step.

What is Kandinsky 6.0?

Kandinsky 6.0 Video is a family of diffusion models for synchronized text-to-audio-video and image-to-audio-video generation. The official project provides a 3B-parameter Lite model and a 29B-parameter Pro model. Both target five-second clips with synchronized 44 kHz audio and lip-sync support.

A separate super-resolution component can raise output to Full HD (1920×1080). The repository is published under the MIT license.

Kandinsky 6.0 in ComfyUI

The project includes an official ComfyUI path. According to its setup instructions, the kandinsky6 and kandinsky6-sr extensions can be installed through ComfyUI Manager, followed by a ComfyUI restart.

If you do not already have ComfyUI running, start with our ComfyUI installation guide. Then see our video workflow guide for the broader pipeline structure.

Hardware: read the benchmarks carefully

The upstream project publishes performance measurements across GPUs including RTX 4090, 5060 Ti, 5080, 5090, RTX PRO 6000, A100 and H100. These are project-reported benchmarks, not The Signal test results. The official quick start currently calls for an NVIDIA GPU and Python 3.13 or 3.14.

Why this release matters

The interesting part is not simply another text-to-video model. Kandinsky 6.0 puts synchronized audio, lip-sync, image-to-video, video generation and super-resolution into one open project with a ComfyUI route. That makes it particularly interesting for automated content pipelines.

What we want to test

The practical questions are GPU cost, VRAM behavior, generation time, lip-sync quality and whether the Lite model is useful on rented consumer GPUs. We will publish those as first-party results only after reproducing them.

Kandinsky 6.0 specifications

FeatureLitePro
Parameters3B29B
ModesText-to-audio-video, image-to-audio-videoText-to-audio-video, image-to-audio-video
Clip length5 seconds5 seconds
AudioSynchronized 44 kHz, including lip-syncSynchronized 44 kHz, including lip-sync
Super-resolutionSeparate Kandinsky 6 SR component; up to Full HD in the documented pipeline
ComfyUIOfficial extensions available through ComfyUI Manager

Kandinsky 6 vs a conventional text-to-video workflow

Workflow concernKandinsky 6 approachConventional video-only model
SoundVideo and synchronized audio generated togetherOften requires a separate audio/TTS/SFX stage
Lip-syncIncluded in the project’s synchronized generation targetOften a separate post-process
Image-to-videoSupported with audioModel-dependent
UpscalingDedicated Kandinsky 6 SR projectUsually external upscaler
Local setupNVIDIA GPU; official Python/ComfyUI pathsVaries by model

This is a capability comparison, not a claim that Kandinsky 6 has better visual quality than every competing model. Quality, speed and cost require controlled first-party tests.

Which Kandinsky 6 model should you start with?

Lite is the more logical first target when exploring the architecture because it is the smaller 3B model. Pro is 29B and targets the higher-capacity end of the family. The upstream repository also exposes distilled variants and device presets. Hardware feasibility depends on the exact variant, resolution, offloading strategy and GPU; do not infer a universal VRAM requirement from parameter count alone.

Kandinsky 6.0 FAQ

Is Kandinsky 6 open source?

The project states that Kandinsky 6.0 and its video super-resolution component were open-sourced on October 6, 2026. Check the repository and model-card licenses for the exact artifact you plan to redistribute or use commercially.

Does Kandinsky 6 work with ComfyUI?

Yes. The official project instructs users to install kandinsky6 and kandinsky6-sr through ComfyUI Manager and restart ComfyUI.

Can Kandinsky 6 generate audio with video?

Yes. Its core proposition is synchronized video and 44 kHz audio generation, including lip-sync, for both text-to-audio-video and image-to-audio-video modes.

Can Kandinsky 6 generate 1080p video?

The project provides a separate super-resolution component that can raise output to Full HD (1920×1080). This is an upscaling stage, not evidence that the base diffusion pass natively generates every clip at 1080p.

What hardware does Kandinsky 6 require?

The official quick start requires an NVIDIA GPU and Python 3.13 or 3.14. Actual memory and generation-time requirements vary by Lite/Pro variant, resolution, attention backend and offloading.

Sources

The Signal newsletter

Keep getting this

One edition a week on open models, local setups and the tools around them.

Read the latest issue

Email delivery opens once the newsletter platform is connected.

Scroll to Top