Quo Vadis, Video Understanding with Vision-Language Foundation Models?

Published in NeurIPS Workshop on Video-Language Models, 2024, 2024