Aws releases vLLM-Omni dlc for image and video generation on SageMaker AI
In this tutorial, aws shows how to turn a text prompt into an image with flux.2-klein-4b and then animate that image into a short video with wan2.1-vace-1.3b, both running on amazon sagemaker ai. The post explains that the two models are deployed from the same aws vllm‑omni deep learning container (dlc) image, with a real‑time endpoint for the image model and an asynchronous endpoint for the…
Key points
- aws vllm‑omni dlc supports flux.2-klein-4b image generation and wan2.1-vace-1.3b video generation.
- real‑time image endpoint uses ml.g6.xlarge; async video endpoint uses ml.g6e.xlarge.
- deployment script pinsomni‑sagemaker‑cuda‑v1.6 creates two sagemaker ai endpoints.
The workflow uses a text prompt to generate a png, resizes it, encodes it as a jpeg data url, and sends it along with a motion prompt to the video endpoint. Results are stored in amazon s3, and an optional streamlit interface lets users adjust prompts and view the mp4. The tutorial provides a command‑line script that pins the dlc image, creates the endpoints, and runs a sample workflow. It also lists prerequisites, instance types (ml.g6.xlarge for image, ml.g6e.xlarge for video), and performance checkpoints: 4.7‑second image latency and 8.9‑second video latency in a single run.
The article concludes that a shared container can support different model families while sagemaker ai applies the appropriate response pattern for each workload. It invites readers to try the sample code and adapt the pattern to other image and video models.
Generate images and video with vLLM-Omni on SageMaker AI
AWS Machine Learning Blog · 28 September 2026
Loading the full article…
This text was published by AWS Machine Learning Blog and written by Yadan Wei. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
2sources- Build real-time voice applications with vLLM-Omni on SageMaker AIPrimary source · AWS Machine Learning Blog ·
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- Modulate raises $25M to analyze voice tone and intent beyond transcripts · 2 src
- Meta Muse agent allegedly sold item and shared address without permission · 2 src
- AWS shows synthetic monitoring with Nova Act and AgentCore · 1 src
- TypeSafe AI launches Jev, a decision-focused AI 193x faster and 444x cheaper than ChatGPT · 3 src
- OpenAI says AI agents leaked 53 images from ChatGPT users · 9 src
Comments
via GitHub Discussions