(CVPR2022) ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action Localization
-
Updated
Nov 14, 2022 - Python
(CVPR2022) ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action Localization
Official implementation for paper Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos
A visualization tool for temporal action localization (detection/segmentation).
The code for the paper "HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning" (ICCV'25).
A predictive Visual Place Recognition system using Event Forecasting Transformers and SNN encoders to enable robust robot localization during sensor dropout and total darkness.
Open-vocabulary temporal localization in endoscopic video with explainability & uncertainty.
Resource-aware X-CLIP baseline on Cholec80: FP16 + grad-accum training and evaluation for surgical video –text localization
Reference implementation of BLAST for causal, bounded-latency score attribution in streaming time-series anomaly detection (ICASSP 2027).
Two-stage pipeline on the Ego4D NLQ benchmark: finding the moment in a first-person video that answers a question, then answering it — VSLNet/VSLBase ablation for temporal localisation, Video-LLaVA for answer generation, scored against hand-written ground truth.
Temporal action localization in untrimmed videos with boundary detection
A clean, modular implementation of Moment-DETR for language-based temporal video grounding on custom datasets. Given a video and a natural-language query, the model predicts the start/end timestamps of the described moment, with a streamlined pipeline for preprocessing, feature extraction, training, and evaluation.
Three-stage dashcam accident understanding: recapture detection, collision/entry timing, driver action (DACON 236753, public LB 31st)
To associate your repository with the temporal-localization topic, visit your repo's landing page and select "manage topics."