Publications

For the most up-to-date list, see my Google Scholar profile.

Selected Publications

Arkaprava Sinha, Dominick Reilly, Siddharth Krishnan, Hieu Le, Srijan Das
Preprint, 2026
A hybrid long-video reasoning framework that proposes action-grounded hypotheses efficiently, then verifies only sparse RGB evidence with an expensive VLM.
Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le, Srijan Das
Preprint, 2026
A unified egocentric encoder trained via hierarchical distillation across ego-exo viewpoints, modalities, and foundation models, using representation-specific proxies as mediators.
Arkaprava Sinha, Monish Soundar Raj, Pu Wang, Ahmed Helmy, Hieu Le, Srijan Das
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
The first Mamba-based architecture for temporal action detection in long untrimmed videos — only 17M parameters, trainable on an NVIDIA Jetson Nano.
Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind, Pu Wang, Francois Bremond, Le Xue, Srijan Das
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
A large language-vision model that incorporates 3D poses and object trajectories to understand the spatiotemporal relationships within daily living activities.
Arkaprava Sinha, Dominick Reilly, Francois Bremond, Pu Wang, Srijan Das
39th Annual AAAI Conference on Artificial Intelligence (AAAI), 2025
Introduces 3D skeletons into the vision-language embedding space to enable effective zero-shot learning for activities of daily living.
Weston Bondurant, Arkaprava Sinha, Hieu Le, Srijan Das, Stephanie Schuckers
IEEE/IAPR INTERNATIONAL JOINT CONFERENCE ON BIOMETRICS (IJCB), 2026
3D latent-controlled diffusion for identity-preserving face swapping.
Mahmoud Ali, Di Yang, Arkaprava Sinha, Dominick Reilly, Srijan Das, Gianpiero Francesca, Francois Bremond
NeurIPS Workshop on Video-Language Models, 2024
A study of where video understanding stands with vision-language foundation models, and where it should go next.