Video-Text-to-Text
Safetensors
qwen3_vl

JoyAI-VL-Interaction

The first open, vision-driven real-time interaction model — it watches a live video stream and decides on its own when to speak, stay silent, or delegate. While enabling online, real-time interaction, this release also delivers powerful offline video understanding, making it the most comprehensive open-source model for video-related capabilities in the 8B parameter class.

📄 Paper · 🌐 Project Page & Demos · 💻 GitHub · 🤗 Paper Page


Overview

Most large models today are turn-based: they answer only when you ask. But many moments in the real world don't wait for a question — a fire starts on a security feed, someone falls, a product flashes by in a livestream. Once missed, the moment is gone.

JoyAI-VL-Interaction is built for exactly these moments. It is an 8B-scale, vision-first interaction model that continuously watches a live video stream and, every second, decides on its own to take one of three actions:

  • Speak — respond when something is worth saying
  • Stay silent — keep watching when nothing warrants a response (a first-class, trained action)
  • Delegate — hand a hard subtask to a background model/agent, keep watching, and weave the result back in when it returns

The decision of when to act is learned inside the model (from second-by-second time-aligned data + RL), not bolted on by an external turn-detector or polling loop. Vision is the first-class driver; speech (ASR/TTS) is treated as pluggable I/O.

To our knowledge, this is the first open, vision-driven interaction model released together with its training recipe, data, and a complete deployable system.

Model Performance

Offline Video Evaluation

Across 26 standard video understanding benchmarks, JoyAI-VL-Interaction achieves an average score of 57.53, outperforming Qwen3-VL-8B-Instruct at 54.16 by 3.37 points.

Demo

https://github.com/user-attachments/assets/2853fc95-ad21-4972-8206-5f3d19798b14

Citation

If you find our work helpful, feel free to give us a cite.

@misc{yao2026joyaivlinteractionrealtimevisionlanguageinteraction,
      title={JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence}, 
      author={Dingyu Yao and Junhao Zhou and Chenxu Yang and Chuanyu Qin and Haowen Hou and Zheming Liang and Congcong Wang and Yuhang Cao and Shenglong Ye and Shuai Xie and Shuhuan Gu and Haoyang Huang and Qingyi Si and Nan Duan and Jiaqi Wang},
      year={2026},
      eprint={2606.14777},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.14777}, 
}
Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jdopensource/JoyAI-VL-Interaction

Quantizations
1 model

Paper for jdopensource/JoyAI-VL-Interaction