April 12, 2024: Ismini Lourentzou

Speaker:
Ismini Lourentzou

Title:
Uni/Cross-modal Understanding with Limited Supervision

Abstract:
In recent years, significant advancements have been made in computer vision and natural language understanding, yielding various models tailored to diverse perception and reasoning tasks. However, in many cases, training perception models often requires large-scale labeled datasets. This talk will cover some of my lab’s research on learning with limited supervision for various applied unimodal and cross-modal tasks. Through case studies spanning cross-modal retrieval, video localization, and audio processing, the talk will cover techniques such as self-supervision and pseudo-supervision to address the challenge of data scarcity. Finally, I will outline open research directions in human-agent collaboration and embodied intelligence.