Skip to main content
The Grok Reference-to-Video node generates a video from a text prompt, using up to seven reference images to guide the style and content of the output. With the grok-imagine-video-1.5 model, you can also attach up to three preset voice references and refer to images and voices directly in the prompt using @ImageN and @AudioN tags. The node sends the request to an external API, waits for the generation to complete, and downloads the resulting video.

Inputs

Common Inputs

Grok Imagine Video 1.5 Inputs

Available when model is set to grok-imagine-video-1.5.

Grok Imagine Video Inputs

Available when model is set to grok-imagine-video.

Reference Inputs

Note: The sub-parameters shown depend on the selected model; grok-imagine-video-1.5 adds the voice_1, voice_2, and voice_3 inputs. At least one reference image is required, and the total is capped at 7 (a batched input counts once per image). With grok-imagine-video-1.5, the prompt can reference connected images as @Image1@Image7 and voice slots as @Audio1, @Audio2, @Audio3; an unnumbered @image or @audio refers to the first one. @AudioN refers to the voice_N widget, not the order of enabled voices. Referencing an image that is not connected or a voice slot set to none causes an error. The API supports only preset voices, not custom audio.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): e584c450563eaa7fcb7751d2325f9ef847fa34a4342df01f2bd9ce2e4ff8f2c3