Publications
For the most up-to-date list, see my Google Scholar profile.
Selected Publications
Preprint, 2026
A hybrid long-video reasoning framework that proposes action-grounded hypotheses efficiently, then verifies only sparse RGB evidence with an expensive VLM.

Preprint, 2026
A unified egocentric encoder trained via hierarchical distillation across ego-exo viewpoints, modalities, and foundation models, using representation-specific proxies as mediators.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
The first Mamba-based architecture for temporal action detection in long untrimmed videos — only 17M parameters, trainable on an NVIDIA Jetson Nano.

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
A large language-vision model that incorporates 3D poses and object trajectories to understand the spatiotemporal relationships within daily living activities.

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
39th Annual AAAI Conference on Artificial Intelligence (AAAI), 2025
Introduces 3D skeletons into the vision-language embedding space to enable effective zero-shot learning for activities of daily living.

IEEE/IAPR INTERNATIONAL JOINT CONFERENCE ON BIOMETRICS (IJCB), 2026
3D latent-controlled diffusion for identity-preserving face swapping.

NeurIPS Workshop on Video-Language Models, 2024
A study of where video understanding stands with vision-language foundation models, and where it should go next.