Skip to main content
Aggregate arXiv cs.AI 人工智能 25 Aug 2026 - 13:00

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 21430v1 Announce Type: new Abstract: Multimodal language models increa…
  • But the creation of stable benchmarks built around Hollywood films is …
  • In this work, we address these concerns directly, by building a new co…

摘要引擎:抽取

正文提要

arXiv:2608.21430v1 Announce Type: new Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-language models struggle on this task (with many performing at near-chance levels of accuracy), while audio-visual models (including those that use audio in captioning scenes) reach a maximum accuracy of 61.1%, well below human-level performance.

来源:https://arxiv.org/abs/2608.21430

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表