paperJuly 2021

Spatio-Temporal Context for Action Detection

AuthorsManuel Sarmiento Calderó, David Varas, Elisenda Bou-Balust

Research in action detection has grown in the recent years, as it plays a key role in video understanding. Modelling the interactions (either spatial or temporal) between actors and their context has proven to be essential for this task. While recent works use spatial features with aggregated temporal information, this work proposes to use non-aggregated temporal information. This is done by adding an attention based method that leverages spatio-temporal interactions between elements in the scene along the clip. The main contribution of this work is the introduction of two cross attention blocks to effectively model the spatial relations and capture short range temporal interactions. Experiments on the AVA dataset show the advantages of the proposed approach that models spatio-temporal relations between relevant elements in the scene, outperforming other methods that model actor interactions with their context by +0.31 mAP.

Related readings and updates.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

October 27, 2025research area Computer Vision, research area Methods and AlgorithmsWorkshop at NeurIPS

This paper was accepted at the Evaluating the Evolving LLM Lifecycle Workshop at NeurIPS 2025.

Existing video understanding benchmarks often conflate knowledge-based and purely image-based questions, rather than clearly isolating a model’s temporal reasoning ability, which is the key aspect that distinguishes video understanding from other modalities. We identify two major limitations that obscure whether higher scores truly indicate stronger…

ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model

February 12, 2025research area Human-Computer Interaction, research area Speech and Natural Language Processingconference ICASSP

We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained to generate first-order ambisonics (FOA) audio, which is a conventional spatial audio format comprising four channels that can be rendered to multichannel spatial output. The proposed generative system is composed of a spatial…

Spatio-Temporal Context for Action Detection

Related readings and updates.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model

Discover opportunities in Machine Learning.