Public research demo. The model and runtime are pinned to released revisions. Uploaded videos are handled through Hugging Face/Gradio temporary storage and are not used for training. This 0.5B research model can still produce incorrect or hallucinated answers; do not use it for safety-critical decisions.
Length and compute limits are different. VQToken uniformly samples 2–8 frames across the complete uploaded video; there is no 30-second source-video limit. This public Space caps uploads at 50 MiB and 1920 × 1080 for resource safety. Its 30-second Hugging Face ZeroGPU setting limits one GPU function call—not the input video's duration.
Ready.
Upload a video and ask a question.
Project page · Public code · Released checkpoint
This is a bounded interactive smoke demo, not a formal benchmark result. The released learned VQ-Attention path is used with fixed, elbow, or silhouette budget selection; private experimental methods are not included.