VLX-Go: Vision-Language Short-Horizon Waypoint Prediction for Embodied Navigation
• 12
Multimodal AI, VLM, VLA, VAM, etc
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
OmAI Lab is a passionate group building multimodal foundation models for physical AI that reshape our work and life.
Open Agent Leaderboard
Mark regions in images based on text descriptions
Process and answer questions about webpage videos
VLM-R1 model for Open-Vocabulary Object Detection