This poster proposes 'Visual Objectification in Films' as a new AI task. It details the ObyGaze12 dataset, annotation guidelines, baseline experiments with ViViT and X-CLIP, and future work on concept representation.
Paper title: Visual Objectification in Films: Towards a New AI Task for Video Interpretation Abstract: This poster proposes 'Visual Objectification in Films' as a new AI task. It details the ObyGaze12 dataset, annotation guidelines, baseline experiments with ViViT and X-CLIP, and future work on concept representation. Paper body (method & results): <!DOCTYPE html> <html lang="en"> <head> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <title>Visual Objectification in Films: Towards a New AI Task for Video Interpretation</title> <!--Generated on Tue Jan 23 16:23:42 2024 by LaTeXML (version 0.8.7) http://dlmf.nist.gov/LaTeXML/.--> <meta content="width=device-width, initial-scale=1, shrink-to-fit=no" name="viewport"/> <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv_0.7.4.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/latexml_styles.css" rel="stylesheet" type="text/css"/> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="/static/browse/0.3.4/js/addons.js"></script> <script src="/static/browse/0.3.4/js/feedbackOverlay.js"></script> <base href="/html/2401.13296v1/"/></head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="#S1" title="1 Introduction ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S2" title="2 Related works ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related works</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS1" title="2.1 Visual biases in film datasets ‣ 2 Related works ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.1 </span>Visual biases in film datasets</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS2" title="2.2 Interpretive-level tasks and dataset creation ‣ 2 Related works ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.2 </span>Interpretive-level tasks and dataset creation</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S2.SS3" title="2.3 Approaches to video and movie understanding ‣ 2 Related works ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.3 </span>Approaches to video and movie understanding</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S2.SS3.SSS0.Px1" title="Pre-trained models for video understanding ‣ 2.3 Approaches to video and movie understanding ‣ 2 Related works ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Pre-trained models for video understanding</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S2.SS3.SSS0.Px2" title="Movie-related tasks ‣ 2.3 Approaches to video and movie understanding ‣ 2 Related works ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Movie-related tasks</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S2.SS3.SSS0.Px3" title="Concept-based models ‣ 2.3 Approaches to video and movie understanding ‣ 2 Related works ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Concept-based models</span></a></li> </ol> </li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S3" title="3 Data and methods ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Data and methods</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S3.SS1" title="3.1 A thesaurus of objectification ‣ 3 Data and methods ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.1 </span>A thesaurus of objectification</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S3.SS2" title="3.2 Data selection ‣ 3 Data and methods ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.2 </span>Data selection</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S3.SS3" title="3.3 Data annotation ‣ 3 Data and methods ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3 </span>Data annotation</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S3.SS3.SSS0.Px1" title="Data processing and fusion ‣ 3.3 Data annotation ‣ 3 Data and methods ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Data processing and fusion</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S3.SS4" title="3.4 Analysis of the ObyGaze12 dataset ‣ 3 Data and methods ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.4 </span>Analysis of the <span class="ltx_text ltx_font_italic">ObyGaze12</span> dataset</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S4" title="4 Experiments ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>Experiments</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S4.SS1" title="4.1 Task accuracy ‣ 4 Experiments ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Task accuracy</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S4.SS2" title="4.2 Concept accuracy ‣ 4 Experiments ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Concept accuracy</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S4.SS2.SSS0.Px1" title="CAV computation ‣ 4.2 Concept accuracy ‣ 4 Experiments ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">CAV computation</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S4.SS2.SSS0.Px2" title="Interpretable classifier ‣ 4.2 Concept accuracy ‣ 4 Experiments ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Interpretable classifier</span></a></li> </ol> </li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S5" title="5 Discussion ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Discussion</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS0.SSS0.Px1" title="Ethical aspect ‣ 5 Discussion ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Ethical aspect</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS0.SSS0.Px2" title="Limitations and challenges ‣ 5 Discussion ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Limitations and challenges</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S6" title="6 Conclusion ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">6 </span>Conclusion</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S6.SS0.SSS0.Px1" title="Acknowledgements ‣ 6 Conclusion ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title">Acknowledgements</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S7" title="7 Dataset ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">7 </span>Dataset</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S7.SS1" title="7.1 List of films ‣ 7 Dataset ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">7.1 </span>List of films</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S7.SS2" title="7.2 Data annotation and processing ‣ 7 Dataset ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">7.2 </span>Data annotation and processing</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S7.SS3" title="7.3 Inter-annotator agreement calculation ‣ 7 Dataset ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">7.3 </span>Inter-annotator agreement calculation</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S8" title="8 Experiments on task accuracy ‣ Visual Objectification in Films: Towards a New AI Task for Video Interpretation"><span class="ltx_text ltx_ref_t