Enhancing Hallucination Detection through Noise Injection
Improves inference-time hallucination detection by sampling under hidden-state or parameter perturbations.
Improves inference-time hallucination detection by sampling under hidden-state or parameter perturbations.
Presents QEVD, a benchmark and dataset for real-time situated fitness coaching.
Analyzes transformer limitations on random access and length generalization through controlled sequence tasks.
Studies locality, recurrent visual policies, and length generalization in visual reasoning tasks.
Studies state-tracking limitations and data efficiency differences between transformers and recurrent sequence models.
Introduces Qualcomm Interactive Cooking and LiveMamba for live, situated instructional guidance.
Introduces the Qualcomm Interactive Video Dataset for real-time audiovisual question answering.
Improves inference-time hallucination detection by sampling under hidden-state or parameter perturbations.
Presents QEVD, a benchmark and dataset for real-time situated fitness coaching.
Analyzes transformer limitations on random access and length generalization through controlled sequence tasks.
Studies end-to-end learning for pose-driven fitness activity recognition and repetition counting.
Trains language models with low-level visual skills for grounded compositional reasoning in videos.
Applies autoregressive language models to sketch generation through generated brush strokes.
Introduces an efficient real-time single-image super-resolution architecture for mobile devices.
Explores real-time situated interaction through a virtually embodied avatar.
Frames decision-making behavior through the lens of language generation.