微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.…
- 21430v1 Announce Type: new Abstract: Multimodal language models increa…
- But the creation of stable benchmarks built around Hollywood films is …
- In this work, we address these concerns directly, by building a new co…
摘要引擎:抽取
正文提要
arXiv:2608.21430v1 Announce Type: new Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-language models struggle on this task (with many performing at near-chance levels of accuracy), while audio-visual models (including those that use audio in captioning scenes) reach a maximum accuracy of 61.1%, well below human-level performance.