Abstract
Understanding Human Actions by Skipping Through Old YouTube Based on CONCUR: Contrastive Observation Network with Cross-Video Understanding and Retrieval. We work on recognizing and localizing human actions in long, untrimmed videos—like the kind you’d randomly find on old YouTube. This is tricky when only the overall action label is known, without any info on when it happens. Our method, CONCUR, improves learning from such weak labels by comparing snippets both within the same video and across different ones. By aligning similar segments between videos and filtering out irrelevant parts, we help the model better understand where actions really occur. Tested on the THUMOS14 dataset, CONCUR outperforms existing methods by making smarter use of both temporal structure and video relationships.
About the speaker
Ulas Bingöl earned a Master’s degree in Electrical and Electronics Engineering from Middle East Technical University, Turkey, in 2022. During his master's studies, he worked in robotics perception - localization while working as a software lead in an international logistics robotics company. After completing his degree, he continued working in the industry before transitioning to academia. He is currently pursuing a PhD in Computer Science at the University of Konstanz, where his research focuses on computer vision and deep learning.
