Public research demo. The model and runtime are pinned to released revisions. Uploaded videos are handled through Hugging Face/Gradio temporary storage and are not used for training. This 0.5B research model can still produce incorrect or hallucinated answers; do not use it for safety-critical decisions.

Codebook-size selection
K is the learned VQ-Attention output budget; adaptive methods search up to 9 values.
2 8
12 64
12 32
12 32
8 64

Ready.

Upload a video and ask a question.

Project page · Public code · Released checkpoint

This is a bounded interactive smoke demo, not a formal benchmark result. The released learned VQ-Attention path is used with fixed, elbow, or silhouette budget selection; private experimental methods are not included.