momo321654/LLaVA-OneVision-2-4B-p16m2-Instruct Image-Text-to-Text • 5B • Updated 11 days ago • 56
momo321654/LLaVA-OneVision-2-4B-p16m2-Instruct Image-Text-to-Text • 5B • Updated 11 days ago • 56
jdopensource/JoyAI-VL-Interaction-Preview Video-Text-to-Text • 9B • Updated Jun 22 • 3.33k • 77
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 216
Running Agents 13 LLaVA OneVision 1.5 📉 13 Interact with a multimodal chatbot using text and images