{"version":1,"type":"story","url":"https://digestai.news/story/inclusionai-releases-realtime-venus-9b-audio-visual-interaction-model","json":"https://digestai.news/story/inclusionai-releases-realtime-venus-9b-audio-visual-interaction-model.json","markdown":"https://digestai.news/story/inclusionai-releases-realtime-venus-9b-audio-visual-interaction-model.md","slug":"inclusionai-releases-realtime-venus-9b-audio-visual-interaction-model","headline":"inclusionAI releases Realtime-Venus 9B audio-visual interaction model","summary":"inclusionAI has published the Realtime-Venus system on Hugging Face, offering two checkpoints: Realtime-Venus-Omni, a 9 billion‑parameter audio‑visual model, and Realtime-Venus-Audio, an audio‑only variant. Both are built on the MiniCPM‑o 4.5 backbone and support full‑duplex streaming, proactive responses, interruption handling, and training‑free long‑video memory. The repository provides model weights, custom Transformers code, and the Realtime‑Venus‑Harness runtime for asynchronous delegation and external tool integration.\n\nThe release includes example scripts for native full‑duplex conversation, proactive video questioning, speech‑in‑duplex chat, and memory‑augmented dialogue. Users can run the models with Python 3.10, CUDA, and FFmpeg, following the provided installation steps. Speech output is generated via bundled Token2wav resources and a reference voice. The technical report (arXiv:2609.13814) details performance on video and audio understanding tasks, and the code is licensed under Apache License 2.0.","keyPoints":["Realtime-Venus-Omni is a 9B audio‑visual model that streams full‑duplex dialogue and proactive responses.","Realtime-Venus-Audio provides audio‑only understanding with text or speech output using the same streaming backbone.","The open‑source repo includes model weights, custom Transformers code, and the Realtime‑Venus‑Harness for delegation and long‑video memory."],"whyItMatters":"The open‑source 9B model brings real‑time, full‑duplex audio‑visual AI to developers, enabling proactive video assistants and long‑video memory without extra training, expanding interactive AI applications.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["inclusionAI","Ant Group","Tsinghua University","Hugging Face"],"models":["Realtime-Venus-Omni","Realtime-Venus-Audio"],"people":[]},"firstPublishedAt":"2026-09-12T00:00:00Z","updatedAt":"2026-09-12T00:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"huggingface.co","title":"inclusionAI/Realtime-Venus · Hugging Face","url":"https://huggingface.co/inclusionAI/Realtime-Venus","publishedAt":"2026-09-12T00:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wjtav9/inclusionairealtimevenus_hugging_face/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"inclusionAI releases Realtime-Venus 9B audio-visual interaction model\", 12 September 2026, https://digestai.news/story/inclusionai-releases-realtime-venus-9b-audio-visual-interaction-model","publisher":"Digest AI","title":"inclusionAI releases Realtime-Venus 9B audio-visual interaction model","datePublished":"2026-09-12T00:00:00Z","url":"https://digestai.news/story/inclusionai-releases-realtime-venus-9b-audio-visual-interaction-model"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}