Sber has released Kandinsky 6.0 Video, introducing the first Russian artificial intelligence model capable of generating short video clips with synchronized sound. The system handles speech, ambient background noise, and automated lip synchronization directly from text prompts or image uploads. The tool is currently available for free in GigaChat, with open source code published under an MIT license for developers.
Users can create video clips up to 5 seconds in length with sound rendered at 44 kHz. Within the GigaChat interface, output resolution options include SD, HD, and Full HD. The generation process starts by submitting a text description or uploading an initial starter image, followed by audio prompts detailing speech, musical accompaniment, or environmental sounds like wind and vehicle traffic. Creators can also choose to generate silent clips if background audio is unnecessary.
Sber uses 2 connected neural modules to manage image sequences and acoustic tracks simultaneously. Training for the sound engine involved more than 20,000,000 video clips featuring clean audio, alongside specialized close up footage of speaking faces to refine mouth movements. The updated visual engine delivers realistic motion for humans, animals, and machinery during rapid camera angle shifts. If the prompt describes an unfamiliar object, the model retrieves external visual references to guide the final render.
Sber published the model under an open source MIT license to encourage developer experimentation. Beyond direct access in GigaChat, the engineering team plans to roll out compatibility with platforms including FAL, Diffusers, FastVideo, ComfyUI, SGLang, and vLLM frameworks. This open release lets independent engineers integrate synchronized audio generation into custom video pipelines without commercial restrictions.




