Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
Summary
This paper introduces a new multimodal benchmark for evaluating AI models' narrative understanding of Hollywood films, addressing copyright issues, and shows that vision-language and audio-visual models perform below human-level accuracy on the task.
View Cached Full Text
Cached at: 08/25/26, 04:16 AM
# Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
Source: [https://arxiv.org/html/2608.21430](https://arxiv.org/html/2608.21430)
Kent K\. ChangAllison CooperJuishan HsuAffiliation:Cinema Studies, Bowdoin CollegeReina KushihashiAffiliation:UC BerkeleyAffiliation:UC BerkeleyMadison MarArnav PodichettyAffiliation:UC BerkeleyRachael SambergAffiliation:UC BerkeleyIpek Nil SancakAffiliation:UC BerkeleyAffiliation:Scholarly Communication and Information Policy, UC Berkeley\[0\.3em\]\{dbamman,kentkchang\}@berkeley\.eduYuhan ShaoAffiliation:UC Berkeley\[0\.5em\] School of InformationUC Berkeley
###### Abstract
Multimodal language models increasingly show promise for enabling the large\-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques\. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections\. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity \(where we publish the first large\-scale, open collection of weekly box office earnings reported by*Variety*magazine from 1922–1979\); and likely public domain status \(by researching copyright registrations and renewals in the US*Catalog of Copyright Entries*\)\. We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision\-language models struggle on this task \(with many performing at near\-chance levels of accuracy\), while audio\-visual models \(including those that use audio in captioning scenes\) reach a maximum accuracy of 61\.1%, well below human\-level performance\.
## 1Introduction
One of the most exciting consequences of advances in NLP, CV and AI has been in their application as analytical tools to shed light on questions of culture at scale\. While this use has long focused on the domain of text\([52](https://arxiv.org/html/2608.21430#bib.bib2);[39](https://arxiv.org/html/2608.21430#bib.bib7)\), we see increasing application for the analysis of film\. This work has shed light on changes in pacing and luminosity\([14](https://arxiv.org/html/2608.21430#bib.bib20)\), the representation of race and gender on screen\([26](https://arxiv.org/html/2608.21430#bib.bib18);[4](https://arxiv.org/html/2608.21430#bib.bib19);[7](https://arxiv.org/html/2608.21430#bib.bib22)\), and comedic timing\([60](https://arxiv.org/html/2608.21430#bib.bib9)\), along with many others\([46](https://arxiv.org/html/2608.21430#bib.bib6)\)\.
However, while progress in other application areas of these methods can be driven by the formation of open benchmarks\([15](https://arxiv.org/html/2608.21430#bib.bib3);[35](https://arxiv.org/html/2608.21430#bib.bib5);[32](https://arxiv.org/html/2608.21430#bib.bib4)\), film as an object of study presents distinct challenges due to the copyrighted nature of the underlying materials\. While many of the core tasks—from shot boundary segmentation\([56](https://arxiv.org/html/2608.21430#bib.bib16);[47](https://arxiv.org/html/2608.21430#bib.bib15)\)to character identification\([19](https://arxiv.org/html/2608.21430#bib.bib17);[37](https://arxiv.org/html/2608.21430#bib.bib12)\)—are often designed with application to the film industry in mind, restrictions on the digitization and republishing of copyrighted materials often leave researchers without direct access to the underlying data\. Even if extracting clips for research purposes is a fair use—an exception to copyright owners’ exclusive rights—film companies have a robust licensing market and often challenge uses that should be considered fair\. Datasets that instead rely on practitioners downloading videos from URLs on YouTube also find themselves with increasingly more restricted access over time, as videos are removed by users or through compliance with Digital Millennium Copyright Act \(“DMCA”\) takedown requests\. Benchmarks built on this data are unstable\.
One solution to this challenge is to build benchmarks on movies in the public domain\. The public domain includes not only works currently published before 1931 in the US,111Copyright protection extends for a fixed period of time\. Films are typically works of corporate authorship protected for 95 years from publication; as of 2026, films published prior to 1931 are in the public domain\. The matter is complicated, however, by the fact that elements of films may be remastered, and such changes to the original film would be protected by new copyright\.but also any film whose copyright owner \(typically the production studio\) did not renew its copyright during the period in US copyright law when such registration renewals were both available and required\. This, however, raises its own challenges: first, there is no known registry of public domain materials, and films that an online source might claim to be in the public domain can be contested; many of the feature films found on sources like the Internet Archive are still in copyright and could be subject to a DMCA takedown request\. Second, while thousands of movies are in the public domain, not all of them are equally notable—the public domain includes war propaganda released by the US government, films whose production studios failed to renew their copyright due to lack of interest in the film, and so on\. In building a benchmark around Hollywood films, we want to incorporate some measure of cultural significance as well\.
In this work we address these critiques through two interventions in data sourcing: first, determining the cultural impression of films in the United States \(measured by*popularity*at the US box office\); second, identifying films that are truly likely to be in the public domain—not through trusting the claims of a third party, but through researching that status with the US*Catalog of Copyright Entries*\. These two criteria—cultural significance and openness—differentiate our work from related efforts to build datasets that include elements of Hollywood films\([51](https://arxiv.org/html/2608.21430#bib.bib21);[41](https://arxiv.org/html/2608.21430#bib.bib14);[54](https://arxiv.org/html/2608.21430#bib.bib10)\), including historical ones\([57](https://arxiv.org/html/2608.21430#bib.bib1)\)\.
Given the dataset defined by these criteria, we build a benchmark around it for measuring the performance of multimodal models at the task of long\-form*narrative*understanding, focusing in particular on questions of temporality, plot, character, setting, perspective, representation, symbolism and object identification\. This benchmark lets us assess the performance of a range of models at measurement tasks increasingly driving scholarship in computational social science and the digital humanities\. While previous benchmarks have explored long\-form motion pictures\([57](https://arxiv.org/html/2608.21430#bib.bib1);[54](https://arxiv.org/html/2608.21430#bib.bib10)\)and interpretive tasks in other modalities like text\([50](https://arxiv.org/html/2608.21430#bib.bib11);[27](https://arxiv.org/html/2608.21430#bib.bib8)\), we bring these two paradigms together in this work to assess how such models can inform our analytical understanding of complex narrative phenomena in film\.
Our work therefore makes the following contributions:
- •We present the first systematic database of historical box office earnings \(extracted from*Variety*magazine\), spanning 1922–1979\. This database is publicly available at[https://github\.com/bamman\-group/variety\-boxoffice](https://github.com/bamman-group/variety-boxoffice)\.
- •We present a new dataset of popular films that are likely in the public domain in the United States, which can form the basis for benchmarks that others can trust will be stable over time\.
- •We present a new benchmark for narrative understanding in these films, and assess the performance of several multimodal models at this difficult task\. We find that many vision\-language models struggle on this task \(with many performing at near\-chance levels of accuracy\), while audio\-visual models \(including those that use audio in captioning scenes\) reach a maximum performance of 61\.1%, well below human\-level performance\. This Classical Hollywood Narrative Benchmark \(and code to support it\) is publicly available at[https://github\.com/bamman\-group/chnb](https://github.com/bamman-group/chnb)\.
## 2Defining the collection
### 2\.1Identifying popular movies
Our first goal is to identify movies that are popular\. Historical box office data in the United States is fragmentary, with major sources of information coming from ledgers kept by executives at individual studios\([24](https://arxiv.org/html/2608.21430#bib.bib27);[30](https://arxiv.org/html/2608.21430#bib.bib28);[25](https://arxiv.org/html/2608.21430#bib.bib26)\)and trade magazines such as*Variety*,*The Motion Picture Herald*and*Hollywood Reporter*\.*Variety*has the deepest historical collection of box office information: starting with its March 3, 1922 issue \(and persisting through the early 21st century\),*Variety*reports the weekly box office receipts for individual movies in specific theaters\. Figure[1](https://arxiv.org/html/2608.21430#S2.F1)gives two typical examples of this information\.


Figure 1:Weekly box office information from two issues of*Variety*, October 29, 1924 \(left\) and July 10, 1934 \(right\)\.Within a broader column noting the city \(Washington and Indianapolis\), each entry lists the theater, movie title\(s\) and box office estimate for that week\. We extract this information as a structured tuple—e\.g\.,⟨\\langleWashington, Columbia,*Feet of Clay*, $12,000⟩\\rangle—from all pages of*Variety*where the full text of at least one article on the page contains the phrase “estimates for this week”, “estimates for last week”, or where the title of an article contains the phrase “picture grosses”\. We prompt multimodal LLMs to extract all tuples from the scan of each page, providing two shots \(input image, output json\) to illustrate the desired behavior\.
Movies mentioned in*Variety*are often terse, relying on readers’ general familiarity with movies that year \(e\.g\., “Gone” to denote*Gone with the Wind*in 1940\)\. For each year, we manually create a mapping of aliases \(“Gone with the Wind”, “Gone”, “Gone with Wind”, etc\.\) to an IMDB record for that movie for all aliases appearing at least 3 times in a year\. This process allows us to generate a ranked list of movies each year by their total box office numbers reported by*Variety*\.
To evaluate the accuracy of this extraction method, we create a gold\-standard dataset by manually labeling 4,072 tuples in 21 issues of*Variety*spanning the years 1922–1979\. We use 7 of these issues for development and hold out 12 issues for held\-out evaluation\. We primarily assess validity through the Spearman rank correlation coefficient between the resulting ranked lists of movies \(comparing the ranks derived from human tuples vs\. model predicted tuples\), since the rank ultimately constitutes the decision criterion we use to define the collection below \(the top 100 movies per year\)\. We find that Gemini 3 Pro \(with high thinking and ultrahigh image resolution\) performs best on this held\-out evaluation, withρ=0\.961\\rho=0\.961\. The complete evaluation procedure can be found in Appendix[B](https://arxiv.org/html/2608.21430#A2)\.
We use this best performing model to extract this structured information for all*Variety*issues from 1922–1979\. By aggregating the total box office earnings for all movies, we are able to assemble ranked lists of the top grossing movies each year over this complete time frame, comprising 1\.4M weekly box office numbers for over 24,000 movies\. This represents, to our knowledge, the first comprehensive, openly available dataset of historical box office information for feature films\. All extractions, alias mapping, and aggregated yearly/weekly charts are openly available at[https://github\.com/bamman\-group/variety\-boxoffice](https://github.com/bamman-group/variety-boxoffice)\.
These top charts allow us to define a measure of popularity, including a film within the boundaries of our dataset if it is among the top 100 highest\-grossing films in a year\. Commercially distributed features form a logical corpus for a benchmark focusing on narrative understanding because they were produced in industrial conditions that prioritized narrative coherence and broad legibility—the properties we seek to evaluate\. TheVarietydataset is rich in films from the Classical Hollywood era, broadly defined as the period spanning 1917–1960\([8](https://arxiv.org/html/2608.21430#bib.bib37)\)and marked by industrial and aesthetic developments such as the studio system and the continuity system, created to improve the clarity and efficiency of storytelling onscreen\. Other dominant narrative trends of the period include character\-centered causality, goal\-oriented plots, and an emphasis on narrative closure\([10](https://arxiv.org/html/2608.21430#bib.bib38)\), along with the evolution of the star persona\([17](https://arxiv.org/html/2608.21430#bib.bib40)\), strategies to navigate the Hollywood Production Code\([29](https://arxiv.org/html/2608.21430#bib.bib39)\), and the formation and reinforcement of genre conventions\([2](https://arxiv.org/html/2608.21430#bib.bib44)\)\. The dataset spans several years of the Post\-Classical period as well; this era saw a loosening of character\-centered causality\([18](https://arxiv.org/html/2608.21430#bib.bib41)\), increasingly ambiguous endings\([11](https://arxiv.org/html/2608.21430#bib.bib43)\), the reworking of Classical\-era genres\([42](https://arxiv.org/html/2608.21430#bib.bib45)\), and the growing influence of European art cinema\([9](https://arxiv.org/html/2608.21430#bib.bib42)\)\.
Finally, we note that the limitations of using box office as a proxy for significance are well documented\([36](https://arxiv.org/html/2608.21430#bib.bib34);[48](https://arxiv.org/html/2608.21430#bib.bib29)\)\. Film scholars have cautioned against an over\-reliance on commercial metrics, arguing that too narrow of an approach misses aesthetically important works\([31](https://arxiv.org/html/2608.21430#bib.bib30);[44](https://arxiv.org/html/2608.21430#bib.bib31)\)or the broader historical and social context of cinema\([1](https://arxiv.org/html/2608.21430#bib.bib32);[53](https://arxiv.org/html/2608.21430#bib.bib33)\)\. Using box office gross as the main criterion in assembling our collection of popular films means our corpus excludes films that circulated outside mainstream distribution in the United States, from the “race” films of the silent era\([49](https://arxiv.org/html/2608.21430#bib.bib35)\)to transnational festival films\([55](https://arxiv.org/html/2608.21430#bib.bib36)\)\. It is worth emphasizing that box office is but one of many measures of significance that might be adopted to assemble a film dataset to evaluate multimodal understanding of narrative\. Our methodology for identifying the public domain status of popular films, described in detail in the following section, can easily be adapted for the development of benchmarks with different boundaries\.
### 2\.2Investigating public domain status
In the United States, most movies published prior to 1931 have fallen into the public domain \(as of the time of this writing\); as have most movies whose copyright was registered prior to 1964 but failed to be renewed 28 years later\. While the first criterion allows us to identify likely public domain movies relatively easily by considering their year of publication, the latter criterion is more difficult since it requires a\.\) identifying the year of copyright registration for a film and b\.\) identifying its*lack*of renewal \(including the renewal of any screenplays or expressive works from which the film was adapted, which may bear separate registration\)\.222Copyright renewal matters only for a certain time periods\. Beginning with the Copyright Act of 1976 \(effective January 1, 1978\), neither registration nor renewal is required for films to be protected by copyright\. Prior to the 1976 Act, however, there is a complex landscape of protection based on a combination of authorship \(individual vs\. corporate\), publication status, registration date, and renewal\.
We turn to two sources for identifying this information: the*Catalog of Copyright Entries*, published by the US Copyright Office through 1978; and the Copyright Public Records System \(CPRS\), an online database published by the US Copyright Office, which contains registrations \(and renewals\) from 1978 forward\. To identify candidate movies that may be in the public domain for lack of registration renewal, we extracted all registrations and renewals for movies recorded in the print*Catalog of Copyright Entries*using digitized versions on the Internet Archive \(which captures registrations/renewals until 1978\) and extracted all motion picture renewals from the Copyright Public Records System to capture renewals made after 1978\.
Using this information, we matched movies between theRegisteredset and theRenewedset using the original registration number \(which often appears as a canonical identifier in both sets\); any movie that was matched was automatically excluded, since this provides evidence that the movie was likely appropriately renewed\. Since registration numbers may change—whether through OCR mistakes, mistyping, or other factors—we also carried out a detailed manual review of the movies that did not match, attempting to match them based on the similarity of their title\.
We also check for one other situation that would lead a movie that has not been renewed to still be under copyright: even if the copyright for the film was not renewed, if the source \(such as a short story or novel\) was appropriately copyrighted and renewed, then the underlying story for the movie may still remain in copyright as well \(e\.g\., as is the case for*It’s a Wonderful Life*\)\. We draw on data from the American Film Institute, which provides information about whether a movie was based on some other original \(e\.g\., literary\) source; any movie described by AFI as being based on an additional source in copyright was removed from the collection\.
The two criteria laid out above—popularity \(among the top 100 movies per year by box office revenues\) and likely public domain status—provide the conceptual boundaries for this collection\. We adopt these stringent criteria in order to minimize DMCA takedown requests, and source the content of the films themselves from the Internet Archive, further limiting the collection to only sound films \(i\.e\., no silent\-era movies\)\. This results in dataset of 61 popular movies\.
## 3Building a benchmark
Given this collection of popular films whose copyright status are unlikely to be challenged, we build a benchmark around it to assess the long\-form narrative understanding capabilities of multimodal models\. We focus in particular on questions that illustrate the affordances of such models for work in cultural analytics of film\. For ease of evaluation, we frame the task as a multiple\-choice question format, with four answer options \(only one of which is correct\)\. We focus on eight narrative categories described below:
- •Temporality\.Questions that track attention to the order of events—both as they occur chronologically within the story world and as they are depicted to the viewer; these orderings are in tension in cases of anachrony\([23](https://arxiv.org/html/2608.21430#bib.bib25)\), such as flashbacks and flashforwards\. *Example question:*What is the order in which the four characters are arrested? A\.\) Countess de Mavon→\\rightarrowNurse Edith Cavell→\\rightarrowMme\. Moulin→\\rightarrowMme\. Rappard\. B\.\) …
- •Plot\.Questions that identify the narrative function of a scene, whether it introduces a complication, raises the stakes, resolves the central tension, and to track whether character goals stated early in the film are ultimately fulfilled\. *Example question:*Why didn’t Charles leave the crime scene right away after murdering Meinike? A\.\) He is setting up a paper trail; B\.\) …
- •Character\.Questions that identify roles and track character dynamics based on what the film shows: what characters do, say, how they are dressed, and how they are staged relative to one another\. *Example question:*Which character is shown smoking? A\.\) Jack; B\.\) …
- •Setting\.Questions that identify and distinguish between locations in the film, and observe how the physical staging of characters within a space \(e\.g\. social blocking\) conveys power, relationship, and intention, which require attention to mise\-en\-scène rather than plot\. *Example question:*Which of these locations do we not see the interior of? A\.\) King Little’s castle; B\.\) …
- •Perspective\.Questions that identify from whose vantage point events are presented, and whether the film ever gives the audience information the characters themselves do not have\. This tests attention to how the film is narrated—not what happens, but who knows what, and when\. *Example question:*When Norma tells Michael how much she loves him, who sees Dr\. Besant enter the room first? A\.\) Norma; B\.\) We do as viewers \(before any characters\); C\.\) …
- •Representation\.Questions that identify how the film depicts gender roles, social identity, and group membership, based on what is shown and said in the film\. *Example question:*Does this movie pass the Bechdel test? \(Two named women talking to each other about a topic that is not a man\.\) A\.\) Yes, between Irma and Phyllis repeatedly …B\.\) …
- •Symbolism\.Questions that identify objects, sounds, or visual actions that recur across the film and carry narrative or thematic weight\. *Example question:*In Marilyn’s performance where she is surrounded by dishes that she is washing, what does this chore symbolize, based on the lyrics to the song she sings? A\.\) Growing up; B\.\) …
- •Object identification\.Questions that identify a specific on\-screen object, track where it appears or what is done with it, or connect it to its function in the plot\. *Example question:*What object is repeatedly used by neighbors to cope with the heat? A\.\) Hand fans; B\.\) …



Figure 2:Example question \(*His Girl Friday*\): “How is the camera positioned as Hildy enters the room to talk to Earl Williams?”*A*\.\) On Hildy’s side of the bars to look through the grate at Earl\.*B*\.\) The camera is at eye level and moves parallel to her\.*C*\.\) On the inside of the cell to look out at Hildy\.*D*\.\) High angle above Hildy and Earl\.To create benchmark questions, seven co\-authors viewed the entirety of a movie and created an average of 12\.8 questions for each one, resulting in an initial set of 779 questions across 61 films\. We use a plain\-language scene description to refer to specific scenes \(not explicit timestamps\) and avoid distractors that are obviously off\-topic to avoid simply testing commonsense reasoning capability\. We assess expert\-level human performance by distributing questions for a sample of ten movies \(133 questions\) to co\-authors who did not write those questions, asking them to watch the movie and answer all questions \(going back and forth to the movie as needed\); we find human\-level accuracy to reach82\.0%82\.0\\%, reflecting in part the complex nature of narrative inferences\. Sources of error include ambiguity in the question/answer options \(where multiple choices could be argued to be correct\), but also reflects the natural difficulty of some information\-seeking questions \(e\.g\., where attention is required to a scene that is easily missed\)\. Table[4](https://arxiv.org/html/2608.21430#A4.T4)\(Appendix[D](https://arxiv.org/html/2608.21430#A4)\) lists the distribution of annotated categories, with greatest representation of questions around plot, character and setting\.
## 4Memorization
One of the challenges of working with popular movies is that they are frequently discussed online, and these discussions make their way into the pre\-training data for LLMs\. Past work has found this to be the case as well:[57](https://arxiv.org/html/2608.21430#bib.bib1)report an accuracy of66\.3%66\.3\\%\(compared to a random performance of50%50\\%\) when prompting Gemini 2\.5 Pro to answer questions based on the movie title and date of release alone \(with no access to the video\);[5](https://arxiv.org/html/2608.21430#bib.bib13)find this kind of “mirage reasoning” prevalent in multimodal medical benchmarks\. We see this as an example of test data contamination\([16](https://arxiv.org/html/2608.21430#bib.bib24);[13](https://arxiv.org/html/2608.21430#bib.bib23)\), where models use metapragmatic information*about*a movie—rather than the content of the movie itself—to make decisions\.
To account for this, we pass all questions through three frontier LLMs—Gemini Pro 3\.1, Claude Opus 4\.7 and GPT 5\.5—with the following prompt: “Based on your knowledge of the movie \{MOVIE\} \(\{YEAR\}\), answer the following question\.” All models exhibit similar rates of memorization \(Gemini39\.4%39\.4\\%, Opus40\.7%40\.7\\%, GPT38\.8%38\.8\\%\)\. To mitigate this effect, we subselect questions from the pool so that the performance across all models when prompted with the movie title and date alone is approximately25%25\\%\(reflecting a random guess\), detailed in Appendix[D](https://arxiv.org/html/2608.21430#A4)\. This yields a total of 628 benchmark questions \(discarding 151 from the original pool\)\. As Table[4](https://arxiv.org/html/2608.21430#A4.T4)\(Appendix[D](https://arxiv.org/html/2608.21430#A4)\) illustrates, the exclusion rate varies by category: models have internal knowledge of common analytical discussion topics about the film—including representation and symbolism—and much less so about questions that require access to specifics visuals\. For convenience, we letℬ\\mathcal\{B\}denote the post\-filter benchmark of 628 questions used in subsequent experiments\.
## 5Experiments
### 5\.1Setup
We evaluate five paradigms onℬ\\mathcal\{B\}\. For a film with videoVV\(frames and audio\) and a questionqq, each paradigm involves a modelMMthat predicts an answera^=M\(c,q\)\\hat\{a\}=M\(c,q\)from a contextccthat varies by howVVis compressed:
#### Closed\-book baseline:c=∅c=\\varnothing\.
The system answers fromqqand its parametric knowledge alone\. This establishes whether the benchmark is solvable without the movie at all, a precondition for any subsequent gain to be attributable to movie content rather than priors \(visual and textual\)\.
#### Subtitles\-only baseline:c=𝒮c=\\mathcal\{S\}\.
The QA backbone takes only the dialogue transcript𝒮\\mathcal\{S\}and the question, without visual or frame\-derived input\.𝒮\\mathcal\{S\}is generated by transcribing the audio track using Distil\-Whisperlarge\-v3\([21](https://arxiv.org/html/2608.21430#bib.bib56)\)\. Prior work on long\-movie comprehension has found subtitle\-only access to be a surprisingly strong baseline\([57](https://arxiv.org/html/2608.21430#bib.bib1)\)\. We include it as a baseline both for performance comparison and to isolate dialogue as a separate point for long\-video compression strategies for narrative understanding\.
#### End\-to\-end:c⊆Vc\\subseteq V\.
The QA model is itself a multimodal model and extracts a fixed sample of V directly\. We consider the following types of end\-to\-end models: a\.\)Long\-video modelsare architectures designed for long\-form video and extracts a uniform 64\-frame sample of the film\. b\.\)Vision–language modelsare general\-purpose VLMs and take a uniform 256\-frame sample\. c\.\)The audio\-visual model\(Gemini 3 Flash\) processes the video in its entirety, including the audio track\.
#### Socratic\([58](https://arxiv.org/html/2608.21430#bib.bib47)\):c=c=Captioner\(V\)\(V\)\.
We follow the two\-stage pipeline described in[12](https://arxiv.org/html/2608.21430#bib.bib46): First, aCaptionerproduces one of the following: a\.\)Frame captionsare generated from a 0\.5\-fps frame strip with no audio; the captioner sees roughly thirty stills per minute of film\. b\.\)Clip captionsare generated from video chunks \(target duration 60 seconds\) that respect shot boundaries\([47](https://arxiv.org/html/2608.21430#bib.bib15)\)and include the audio track\. These captions are then timestamped and concatenated chronologically into a world state history\([12](https://arxiv.org/html/2608.21430#bib.bib46)\)that the QA backbone receives \(instead ofVV\)\. This class tests whether textual compression preserves the audio\-visual signal for long\-video QA\. We adopt Gemini 3 Flash as the captioner\.
#### Agentic retrieval:c=π\(V,q\)c=\\pi\(V,q\)\.
To test whether query\-conditioned selectivity improves on the Socratic baseline \(which has access to full content\), we instantiate the retrieval policyπ\\piin two ways: a\.\)Frame image retrievalis the VideoAgent loop described in[20](https://arxiv.org/html/2608.21430#bib.bib48):π\\pistarts from 8 uniform frames in a 64\-frame pool and re\-fetches up to four more per iteration via CLIP\([40](https://arxiv.org/html/2608.21430#bib.bib50)\)cosine similarity on a confidence\-threshold loop \(≤3\\leq 3iterations\)\. b\.\)Caption retrievalutilizes the Letta agent\([38](https://arxiv.org/html/2608.21430#bib.bib49)\), whereπ\\piloads the Socratic captions \(both frame\- and clip\-based\) into archival memory and lets the QA backbone decide to trigger a semantic search\.333[https://docs\.letta\.com/api/python/resources/agents/subresources/passages/methods/search/](https://docs.letta.com/api/python/resources/agents/subresources/passages/methods/search/)\. We use the defaulttext\-embedding\-3\-smallmodel for semantic search\.
#### Pre\-processing pipeline\.
Several film\-level artifacts are precomputed once per film and shared across paradigms: a\.\) Uniform frame samples atN∈\{64,128,256\}N\\in\\\{64,128,256\\\}are extracted at indices⌊i⋅T/N⌋\\lfloor i\\cdot T/N\\rfloorwhereTTis the film’s frame count, and cached as JPEGs\. b\.\) Shot\-grouped chunks come from detected shot boundaries merged greedily to a target duration of 60 seconds, with a 10\-second minimum and the constraint that no shot is split\. c\.\) Per\-frame CLIP embeddings of the largest uniform pool,L2L\_\{2\}\-normalized and cached, drive the targeted frame\-retrieval step in the VideoAgent loop\. d\.\) Captions are produced by Gemini 3 Flash from either the 0\.5\-fps frame strip \(frame captions\) or the shot\-grouped video chunk with audio \(clip captions\)\.
Table 1:Accuracy onℬ\\mathcal\{B\}\. The largest 95% Wald confidence interval is±3\.9%\\pm 3\.9\\%\. Bold indicates best overall;underlinemarks the within\-subgroup leader significantly better than the runner\-up \(whose 95% CIs do not overlap\)\.ModelComputeAcc\.Closed\-book\(question only\)Qwen3\-VL\-8B≈\\approx0h 07m21\.5GLM\-4\.1V\-9B\-Thinking≈\\approx2h 46m20\.4GPT\-5\-mini≈\\approx0h 03m25\.2Claude Haiku 4\.5≈\\approx0h 09m22\.1Gemini 3 Flash≈\\approx2h 47m26\.6Subtitles onlyQwen3\-VL\-8B≈\\approx0h 20m32\.2GLM\-4\.1V\-9B\-Thinking≈\\approx31h 21m27\.7GPT\-5\-mini≈\\approx9h 22m39\.2Claude Haiku 4\.5≈\\approx0h 08m36\.9Gemini 3 Flash≈\\approx2h 04m48\.1End\-to\-endlong\-video\(64 frames\)VAMBA\-Qwen2\-VL\-7B≈\\approx1h 47m23\.7VideoChat\-Flash≈\\approx1h 04m27\.2HourLLaVA≈\\approx1h 02m23\.4LLaVA\-NeXT\-Video\-7B\-DPO≈\\approx0h 15m25\.0vision–language\(256 frames\)Qwen3\-VL\-8B≈\\approx4h 43m29\.1GLM\-4\.1V\-9B\-Thinking≈\\approx6h 54m24\.2GPT\-5\-mini≈\\approx7h 52m34\.2Gemini 3 Flash≈\\approx6h 22m39\.3audio\-visual\(full video\)Gemini 3 Flash≈\\approx37h 24m61\.1ModelComputeAcc\.Socraticframe\-basedQwen3\-VL\-8B≈\\approx1h 36m30\.3GPT\-5\-mini≈\\approx4h 42m36\.8Claude Haiku 4\.5≈\\approx0h 58m35\.0Gemini 3 Flash≈\\approx7h 15m46\.0clip\-basedQwen3\-VL\-8B≈\\approx1h 44m32\.8GPT\-5\-mini≈\\approx3h 30m48\.1Claude Haiku 4\.5≈\\approx1h 02m45\.9Gemini 3 Flash≈\\approx2h 53m58\.9Agentic retrievalframe image\(VideoAgent\)Qwen3\-VL\-8B≈\\approx16h 00m22\.9GPT\-5\-mini≈\\approx9h 22m24\.8Claude Haiku 4\.5≈\\approx1h 53m26\.1Gemini 3 Flash≈\\approx2h 04m28\.8frame caption\(Letta\)Qwen3\-VL\-8B≈\\approx1h 31m26\.1GPT\-5\-mini≈\\approx1h 57m33\.6Claude Haiku 4\.5≈\\approx3h 27m39\.8Gemini 3 Flash≈\\approx4h 48m26\.6clip caption\(Letta\)Qwen3\-VL\-8B≈\\approx1h 50m26\.8GPT\-5\-mini≈\\approx2h 13m26\.0Claude Haiku 4\.5≈\\approx1h 51m41\.2Gemini 3 Flash≈\\approx6h 51m30\.3
### 5\.2Results
We evaluateℬ\\mathcal\{B\}on the following: HourLLaVA\([34](https://arxiv.org/html/2608.21430#bib.bib52)\), VAMBA\-Qwen2\-VL\-7B\([43](https://arxiv.org/html/2608.21430#bib.bib53)\), VideoChat\-Flash\([33](https://arxiv.org/html/2608.21430#bib.bib57)\), and LLaVA\-NeXT\-Video\-DPO\([59](https://arxiv.org/html/2608.21430#bib.bib58)\)are long\-video models; Qwen3\-VL\-8B\([6](https://arxiv.org/html/2608.21430#bib.bib51)\)and GLM\-4\.1V\-9B\-Thinking\([28](https://arxiv.org/html/2608.21430#bib.bib55)\)are open\-weight VLMs; and finally, closed\-source models: GPT\-5\-mini\([45](https://arxiv.org/html/2608.21430#bib.bib54)\), Claude Haiku 4\.5\([3](https://arxiv.org/html/2608.21430#bib.bib59)\), and Gemini 3 Flash\([22](https://arxiv.org/html/2608.21430#bib.bib60)\)\.444Gemini 3 models are accessed via Vertex AI:[https://docs\.cloud\.google\.com/vertex\-ai/generative\-ai/docs/models](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models)\.Table[1](https://arxiv.org/html/2608.21430#S5.T1)summarizes model performance\. For each setup, we report accuracy and note the widest 95% Wald confidence intervals to facilitate testing the significance of direct model comparisons\.
In the closed\-book baseline, no lower CI bound exceeds 25%; models cannot answer questions inℬ\\mathcal\{B\}from their parametric knowledge alone\. Nor do the closed\-source backbones agree on which films they know better: per\-film closed\-book accuracies correlate atρ=\+0\.11\\rho=\+0\.11\(p=0\.40p=0\.40\) between Gemini Flash 3 and GPT\-5\-mini,\+0\.28\+0\.28\(p=0\.027p=0\.027\) between Gemini Flash 3 and Claude Haiku 4\.5, and\+0\.44\+0\.44\(p<0\.001p<0\.001\) between GPT\-5\-mini and Claude\. The subtitle\-only baseline shows that the questions are meaningfully answerable from real movie content; dialogue alone \(the ASR transcript\) raises accuracy to48\.148\.1for Gemini,39\.239\.2for GPT, and36\.936\.9for Claude\.
#### Clip\-based Socratic captioning is competitive with native long\-video processing\.
The strongest end\-to\-end configuration is Gemini 3 Flash on full video \(61\.161\.1±3\.8\\pm 3\.8\), followed by the same model on clip\-based captions \(58\.958\.9±3\.8\\pm 3\.8\)\. Given the overlapping CIs, there is no meaningful gap between the text\-mediated compression via clip captioning and Gemini’s native multimodal processing\. Clip\-based Socratic captioning is effective across models, which makes an empirical case for caption\-based compression as a viable alternative to video processing\.
#### Performance of agentic and Socratic methods differs across QA backbones\.
The Letta agent and Socratic pipelines take the same captions but differ in how they reach the model: Letta retrieves from archival memory, but Socratic includes the entire chronological description based on the captions in the prompt\. The Letta–Socratic gap is significant for Gemini on both caption types and for GPT on clip captions; for Claude the two paradigms are tied, and it is the strongest agentic\-retrieval backbone of the three\. In contrast, frame\-image retrieval barely exceeds chance on every backbone\. The agent’s iterative frame\-retrieval loop cannot compensate for what its modality omits \(audio and dialogue\)\. Its CLIP\-similarity is unsuitable for retrieving content in the audio stream that is relevant to the answer\.
#### Frame budget exerts limited impact\.
For the six end\-to\-end backbones run at multiple frame budgets, we report average accuracy in Table[5](https://arxiv.org/html/2608.21430#A5.T5)\(Appendix[E](https://arxiv.org/html/2608.21430#A5)\) forN∈\{64,128,256\}N\\in\\\{64,128,256\\\}\. The largest within\-backbone gain is\+3\.6\+3\.6pp fromN=64N\{=\}64toN=256N\{=\}256; within\-backbone CIs overlap heavily across the three budgets, and we therefore use a single canonical budget per sub\-paradigm in Table[1](https://arxiv.org/html/2608.21430#S5.T1)\(N=64N\{=\}64for long\-video models,N=256N\{=\}256for vision–language models\)\.
#### Vision is not sufficient for film narrative understanding\.
Of the four architectures designed for hour\-scale video at 64 frames \(HourLLaVA, VAMBA\-Qwen2\-VL\-7B, and VideoChat\-Flash, LLaVA\-NeXT\-Video\-DPO\), none are statistically indistinguishable from the closed\-book baseline\. The general\-purpose vision–language models with a larger frame budget do better, but in the case of Gemini 3 Flash, switching from a 256\-frame visual input \(39\.3%39\.3\\%\) to full video, including audio track \(61\.1%61\.1\\%\), gains\+21\.8\+21\.8pp\. We observe similar patterns inside Socratic: holding the QA backbone fixed, swapping the captions from frame\-based \(no audio\) to clip\-based introduces significant gains for all three models\.
#### Performance gains come from dialogue access and reasoning over video content\.
Per\-film accuracy in each backbone’s best configuration correlates positively with that backbone’s subtitles\-only accuracy \(Table[2](https://arxiv.org/html/2608.21430#S5.T2);ρsubtitles\\rho\_\{\\text\{subtitles\}\}ranges from\+0\.40\+0\.40to\+0\.52\+0\.52, allp≤0\.002p\\leq 0\.002\): the films where the strongest configurations win are in the films where dialogue alone is informative\. To assess whether memorization drives the performance gains onℬ\\mathcal\{B\}, we measure two proxies: web prevalence and title\-prediction accuracy\. Web prevalence is thelog10\(1\+hv\)\\log\_\{10\}\(1\+h\_\{v\}\)\-scaled count of Google Search results \(hvh\_\{v\}\) for the title query, and title\-prediction accuracy captures the fraction of 10 uniformly\-spaced 5\-frame window per film from which the backbone correctly predicts the canonical film title on IMDb, assessed by case\-insensitive exact string match \(28\.4%28\.4\\%for Gemini 3 Flash,3\.3%3\.3\\%for GPT\-5\-mini, and2\.1%2\.1\\%for Claude Haiku 4\.5\)\. A model whose performance partly relies on memorized content would perform better on films with greater web presence and whose visual iconography it can identify\. However, in Table[2](https://arxiv.org/html/2608.21430#S5.T2), we see that across all three closed models,ρhits\\rho\_\{\\text\{hits\}\}andρtitle\\rho\_\{\\text\{title\}\}are negative, suggesting films that are popular online and visually recognizable by the model are not easier even on the strongest configurations\.
Table 2:Spearman rank correlations across all 61 films, of per\-film accuracy in each backbone’s best\-performing configuration, compared with four predictors: closed\-book accuracy of Gemini 3 Flash, subtitles\-only accuracy, web prevalence, and title prediction accuracy\.ModelBestAcc\.Gemini closed\-bookSubtitlesWeb hitsTitle predictionρgcb\\rho\_\{\\text\{gcb\}\}ppρsubtitles\\rho\_\{\\text\{subtitles\}\}ppρhits\\rho\_\{\\text\{hits\}\}ppρtitle\\rho\_\{\\text\{title\}\}ppGPT\-5\-miniSocratic, clip captions48\.1\\mathbf\{48\.1\}\+0\.22\+0\.220\.0840\.084\+0\.40\+0\.400\.0020\.002−0\.25\-0\.250\.060\.06−0\.13\-0\.130\.310\.31Claude Haiku 4\.5Socratic, clip captions45\.9\\mathbf\{45\.9\}\+0\.37\+0\.370\.0030\.003\+0\.52\+0\.52<0\.001\{<\}0\.001−0\.45\-0\.45<0\.001\{<\}0\.001−0\.16\-0\.160\.230\.23Gemini 3 Flashend\-to\-end, full video61\.1\\mathbf\{61\.1\}\+0\.43\+0\.43<0\.001\{<\}0\.001\+0\.42\+0\.42<0\.001\{<\}0\.001−0\.27\-0\.270\.040\.04−0\.26\-0\.260\.050\.05
Gemini’sρgcb=\+0\.43\\rho\_\{\\text\{gcb\}\}=\+0\.43is a within\-model consistency effect, but for Claude and GPT\-5\-mini, their clip\-captioned Socratic routes the audio\-visual content through Gemini Flash as a captioner before reaching the QA backbone\. The\+0\.37\+0\.37and\+0\.22\+0\.22we observe for those two against the closed\-book ranking of Gemini Flash, then, show that Gemini’s parametric film knowledge influences its caption output enough to leave a per\-film signal in the performance of the downstream models\.
## 6Conclusion
We present in this work a new benchmark of narrative questions built around popular movies from the Classical Hollywood era, defining the boundaries of that collection by movies that are popular \(as measured by box office numbers reported by*Variety*magazine\) and whose copyright status is unlikely to be challenged \(either by being released prior to 1931 or by registering their copyright but failing to renew it\)\. The questions require attention to complex narrative elements involving plot, character, setting, perspective, and more, and prove challenging for frontier multimodal language models\. As more research leverages such models for the large\-scale computational analysis of film, we expect this benchmark to provide a proving ground for assessing comparative model performance\. Data and code to support this work are available at[https://github\.com/bamman\-group/variety\-boxoffice](https://github.com/bamman-group/variety-boxoffice)and[https://github\.com/bamman\-group/chnb](https://github.com/bamman-group/chnb)\.
## 7Limitations
While this work aims to address a gap in long\-form multimodal benchmarks for assessing narrative understanding abilities of contemporary models, it is limited in several ways\. By selecting movies based on popularity alone, we encode only one of the many possible forms of cultural significance, and omit movies from the time period that circulated outside of major metropolitan cities in the United States\. The process we describe for investigating public domain status, however, could be applied to define new collections under alternative criteria\. Additionally, all questions in the benchmark were created by researchers \(of varying disciplinary backgrounds\) at U\.S\. universities, which influences the narrative aspects in a film we find salient\. Finally, the benchmark specifically covers the era of Classical Hollywood cinema \(through 1963\); while we expect models with good long\-form narrative understanding to be able to perform well on this data, performance may not generalize to films outside of this time period \(both older and newer\); we see this as a necessary trade\-off for defining a collection of movies on which additional stable benchmarks can be built\.
## Acknowledgments
The research reported in this article was supported by the Humanities and AI Virtual Institute \(HAVI\), a program of Schmidt Sciences, and by Google\. This research used the Savio computational cluster resource provided by the Berkeley Research Computing program at the University of California, Berkeley\.
## References
- Allen and Gomery \(1985\)R\. C\. Allen and D\. GomeryFilm history: theory and practice\.Knopf\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Altman \(1999\)R\. AltmanFilm/genre\.BFI Publishing\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Anthropic \(2025\)AnthropicClaude Haiku 4\.5 system card\.Technical reportAnthropic\.Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Arnoldet al\.\(2019\)T\. Arnold, L\. Tilton, and A\. BerkeVisual style in two network era sitcoms\.Journal of Cultural Analytics4\(2\)\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
- Asadiet al\.\(2026\)M\. Asadi, J\. W\. O’Sullivan, F\. Cao, T\. Nedaee, K\. Fardi, F\. Li, E\. Adeli, and E\. AshleyMirage the illusion of visual understanding\.arXiv preprint arXiv:2603\.21687\.Cited by:[§4](https://arxiv.org/html/2608.21430#S4.p1.1)\.
- Baiet al\.\(2025\)S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge, W\. Ge, Z\. Guo, Q\. Huang, J\. Huang, F\. Huang, B\. Hui, S\. Jiang, Z\. Li, M\. Li, M\. Li, K\. Li, Z\. Lin, J\. Lin, X\. Liu, J\. Liu, C\. Liu, Y\. Liu, D\. Liu, S\. Liu, D\. Lu, R\. Luo, C\. Lv, R\. Men, L\. Meng, X\. Ren, X\. Ren, S\. Song, Y\. Sun, J\. Tang, J\. Tu, J\. Wan, P\. Wang, P\. Wang, Q\. Wang, Y\. Wang, T\. Xie, Y\. Xu, H\. Xu, J\. Xu, Z\. Yang, M\. Yang, J\. Yang, A\. Yang, B\. Yu, F\. Zhang, H\. Zhang, X\. Zhang, B\. Zheng, H\. Zhong, J\. Zhou, F\. Zhou, J\. Zhou, Y\. Zhu, and K\. ZhuQwen3\-VL technical report\.arXiv preprint arXiv:2511\.21631\.Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Bammanet al\.\(2024\)D\. Bamman, R\. Samberg, R\. J\. So, and N\. ZhouMeasuring diversity in Hollywood through the large\-scale computational analysis of film\.Proceedings of the National Academy of Sciences121\(46\),pp\. e2409770121\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2409770121),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2409770121),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.2409770121Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
- Bordwellet al\.\(1985\)D\. Bordwell, J\. Staiger, and K\. ThompsonThe classical hollywood cinema: film style & mode of production to 1960\.Columbia University Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Bordwell \(1979\)D\. BordwellThe art cinema as a mode of film practice\.Film Criticism4\(1\),pp\. 56–64\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Bordwell \(1985\)D\. BordwellNarration in the fiction film\.Univ of Wisconsin Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Bordwell \(2006\)D\. BordwellThe way hollywood tells it: story and style in modern movies\.Univ of California Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Chandrasegaranet al\.\(2024\)K\. Chandrasegaran, A\. Gupta, L\. M\. Hadzic, T\. Kota, J\. He, C\. Eyzaguirre, Z\. Durante, M\. Li, J\. Wu, and L\. Fei\-FeiHourVideo: 1\-hour video\-language understanding\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 53168–53197\.External Links:[Document](https://dx.doi.org/10.52202/079017-1684),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/5f2809607f692d79a01c05c43d702883-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px4.p1.1)\.
- Changet al\.\(2023\)K\. K\. Chang, M\. Cramer, S\. Soni, and D\. BammanSpeak, memory: an archaeology of books known to ChatGPT/GPT\-4\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 7312–7327\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.453/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.453)Cited by:[§4](https://arxiv.org/html/2608.21430#S4.p1.1)\.
- Cuttinget al\.\(2011\)J\. E\. Cutting, K\. L\. Brunick, J\. E\. DeLong, C\. Iricinschi, and A\. CandanQuicker, faster, darker: changes in Hollywood film over 75 years\.i\-Perception2\(6\),pp\. 569–576\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
- Denget al\.\(2009\)J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-FeiImagenet: a large\-scale hierarchical image database\.In2009 IEEE conference on computer vision and pattern recognition,pp\. 248–255\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p2.1)\.
- Dodgeet al\.\(2021\)J\. Dodge, M\. Sap, A\. Marasović, W\. Agnew, G\. Ilharco, D\. Groeneveld, M\. Mitchell, and M\. GardnerDocumenting large webtext corpora: a case study on the colossal clean crawled corpus\.InProceedings of the 2021 conference on empirical methods in natural language processing,pp\. 1286–1305\.Cited by:[§4](https://arxiv.org/html/2608.21430#S4.p1.1)\.
- Dyer \(1998\)R\. DyerStars\. 1979\.London: BFI\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Elsaesseret al\.\(1975\)T\. Elsaesseret al\.The pathos of failure: the unmotivated hero\.Monogram6\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Everinghamet al\.\(2006\)M\. Everingham, J\. Sivic, and A\. Zisserman“Hello\! my name is… Buffy” – Automatic naming of characters in TV video\.InBMVC,Vol\.2,pp\. 6\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p2.1)\.
- Fanet al\.\(2025\)Y\. Fan, X\. Ma, R\. Wu, Y\. Du, J\. Li, Z\. Gao, and Q\. LiVideoagent: a memory\-augmented multimodal agent for video understanding\.InEuropean Conference on Computer Vision,pp\. 75–92\.Cited by:[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px5.p1.1)\.
- Gandhiet al\.\(2023\)S\. Gandhi, P\. von Platen, and A\. M\. RushDistil\-whisper: robust knowledge distillation via large\-scale pseudo labelling\.External Links:2311\.00430Cited by:[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px2.p1.1)\.
- Gemini Teamet al\.\(2023\)Gemini Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican, D\. Silver, M\. Johnson, I\. Antonoglou, J\. Schrittwieser, A\. Glaese, J\. Chen, E\. Pitler, T\. Lillicrap, A\. Lazaridou, O\. Firat, J\. Molloy, M\. Isard, P\. R\. Barham, T\. Hennigan, B\. Lee, F\. Viola, M\. Reynolds, Y\. Xu, R\. Doherty, E\. Collins, C\. Meyer, E\. Rutherford, E\. Moreira, K\. Ayoub, M\. Goel, J\. Krawczyk, C\. Du, E\. Chi, H\. Cheng, E\. Ni, P\. Shah, P\. Kane, B\. Chan, M\. Faruqui, A\. Severyn, H\. Lin, Y\. Li, Y\. Cheng, A\. Ittycheriah, M\. Mahdieh, M\. Chen, P\. Sun, D\. Tran, S\. Bagri, B\. Lakshminarayanan, J\. Liu, A\. Orban, F\. Güra, H\. Zhou, X\. Song, A\. Boffy, H\. Ganapathy, S\. Zheng, H\. Choe, Á\. Weisz, T\. Zhu, Y\. Lu, S\. Gopal, J\. Kahn, M\. Kula, J\. Pitman, R\. Shah, E\. Taropa, M\. A\. Merey, M\. Baeuml, Z\. Chen, L\. E\. Shafey, Y\. Zhang, O\. Sercinoglu, G\. Tucker, E\. Piqueras, M\. Krikun, I\. Barr, N\. Savinov, I\. Danihelka, B\. Roelofs, A\. White, A\. Andreassen, T\. von Glehn, L\. Yagati, M\. Kazemi, L\. Gonzalez, M\. Khalman, J\. Sygnowski, A\. Frechette, C\. Smith, L\. Culp, L\. Proleev, Y\. Luan, X\. Chen, J\. Lottes, N\. Schucher, F\. Lebron, A\. Rrustemi, N\. Clay, P\. Crone, T\. Kocisky, J\. Zhao, B\. Perz, D\. Yu, H\. Howard, A\. Bloniarz, J\. W\. Rae, H\. Lu, L\. Sifre, M\. Maggioni, F\. Alcober, D\. Garrette, M\. Barnes, S\. Thakoor, J\. Austin, G\. Barth\-Maron, W\. Wong, R\. Joshi, R\. Chaabouni, D\. Fatiha, A\. Ahuja, G\. S\. Tomar, E\. Senter, M\. Chadwick, I\. Kornakov, N\. Attaluri, I\. Iturrate, R\. Liu, Y\. Li, S\. Cogan, J\. Chen, C\. Jia, C\. Gu, Q\. Zhang, J\. Grimstad, A\. J\. Hartman, X\. Garcia, T\. S\. Pillai, J\. Devlin, M\. Laskin, D\. d\. L\. Casas, D\. Valter, C\. Tao, L\. Blanco, A\. P\. Badia, D\. Reitter, M\. Chen, J\. Brennan, C\. Rivera, S\. Brin, S\. Iqbal, G\. Surita, J\. Labanowski, A\. Rao, S\. Winkler, E\. Parisotto, Y\. Gu, K\. Olszewska, R\. Addanki, A\. Miech, A\. Louis, D\. Teplyashin, G\. Brown, E\. Catt, J\. Balaguer, J\. Xiang, P\. Wang, Z\. Ashwood, A\. Briukhov, A\. Webson, S\. Ganapathy, S\. Sanghavi, A\. Kannan, M\. Chang, A\. Stjerngren, J\. Djolonga, Y\. Sun, A\. Bapna, M\. Aitchison, P\. Pejman, H\. Michalewski, T\. Yu, C\. Wang, J\. Love, J\. Ahn, D\. Bloxwich, K\. Han, P\. Humphreys, T\. Sellam, J\. Bradbury, V\. Godbole, S\. Samangooei, B\. Damoc, A\. Kaskasoli, S\. M\. R\. Arnold, V\. Vasudevan, S\. Agrawal, J\. Riesa, D\. Lepikhin, R\. Tanburn, S\. Srinivasan, H\. Lim, S\. Hodkinson, P\. Shyam, J\. Ferret, S\. Hand, A\. Garg, T\. L\. Paine, J\. Li, Y\. Li, M\. Giang, A\. Neitz, Z\. Abbas, S\. York, M\. Reid, E\. Cole, A\. Chowdhery, D\. Das, D\. Rogozińska, V\. Nikolaev, P\. Sprechmann, Z\. Nado, L\. Zilka, F\. Prost, L\. He, M\. Monteiro, G\. Mishra, C\. Welty, J\. Newlan, D\. Jia, M\. Allamanis, C\. H\. Hu, R\. de Liedekerke, J\. Gilmer, C\. Saroufim, S\. Rijhwani, S\. Hou, D\. Shrivastava, A\. Baddepudi, A\. Goldin, A\. Ozturel, A\. Cassirer, Y\. Xu, D\. Sohn, D\. Sachan, R\. K\. Amplayo, C\. Swanson, D\. Petrova, S\. Narayan, A\. Guez, S\. Brahma, J\. Landon, M\. Patel, R\. Zhao, K\. Villela, L\. Wang, W\. Jia, M\. Rahtz, M\. Giménez, L\. Yeung, J\. Keeling, P\. Georgiev, D\. Mincu, B\. Wu, S\. Haykal, R\. Saputro, K\. Vodrahalli, J\. Qin, Z\. Cankara, A\. Sharma, N\. Fernando, W\. Hawkins, B\. Neyshabur, S\. Kim, A\. Hutter, P\. Agrawal, A\. Castro\-Ros, G\. van den Driessche, T\. Wang, F\. Yang, S\. Chang, P\. Komarek, R\. McIlroy, M\. Lučić, G\. Zhang, W\. Farhan, M\. Sharman, P\. Natsev, P\. Michel, Y\. Bansal, S\. Qiao, K\. Cao, S\. Shakeri, C\. Butterfield, J\. Chung, P\. K\. Rubenstein, S\. Agrawal, A\. Mensch, K\. Soparkar, K\. Lenc, T\. Chung, A\. Pope, L\. Maggiore, J\. Kay, P\. Jhakra, S\. Wang, J\. Maynez, M\. Phuong, T\. Tobin, A\. Tacchetti, M\. Trebacz, K\. Robinson, Y\. Katariya, S\. Riedel, P\. Bailey, K\. Xiao, N\. Ghelani, L\. Aroyo, A\. Slone, N\. Houlsby, X\. Xiong, Z\. Yang, E\. Gribovskaya, J\. Adler, M\. Wirth, L\. Lee, M\. Li, T\. Kagohara, J\. Pavagadhi, S\. Bridgers, A\. Bortsova, S\. Ghemawat, Z\. Ahmed, T\. Liu, R\. Powell, V\. Bolina, M\. Iinuma, P\. Zablotskaia, J\. Besley, D\. Chung, T\. Dozat, R\. Comanescu, X\. Si, J\. Greer, G\. Su, M\. Polacek, R\. L\. Kaufman, S\. Tokumine, H\. Hu, E\. Buchatskaya, Y\. Miao, M\. Elhawaty, A\. Siddhant, N\. Tomasev, J\. Xing, C\. Greer, H\. Miller, S\. Ashraf, A\. Roy, Z\. Zhang, A\. Ma, A\. Filos, M\. Besta, R\. Blevins, T\. Klimenko, C\. Yeh, S\. Changpinyo, J\. Mu, O\. Chang, M\. Pajarskas, C\. Muir, V\. Cohen, C\. L\. Lan, K\. Haridasan, A\. Marathe, S\. Hansen, S\. Douglas, R\. Samuel, M\. Wang, S\. Austin, C\. Lan, J\. Jiang, J\. Chiu, J\. A\. Lorenzo, L\. L\. Sjösund, S\. Cevey, Z\. Gleicher, T\. Avrahami, A\. Boral, H\. Srinivasan, V\. Selo, R\. May, K\. Aisopos, L\. Hussenot, L\. B\. Soares, K\. Baumli, M\. B\. Chang, A\. Recasens, B\. Caine, A\. Pritzel, F\. Pavetic, F\. Pardo, A\. Gergely, J\. Frye, V\. Ramasesh, D\. Horgan, K\. Badola, N\. Kassner, S\. Roy, E\. Dyer, V\. C\. Campos, A\. Tomala, Y\. Tang, D\. E\. Badawy, E\. White, B\. Mustafa, O\. Lang, A\. Jindal, S\. Vikram, Z\. Gong, S\. Caelles, R\. Hemsley, G\. Thornton, F\. Feng, W\. Stokowiec, C\. Zheng, P\. Thacker, Ç\. Ünlü, Z\. Zhang, M\. Saleh, J\. Svensson, M\. Bileschi, P\. Patil, A\. Anand, R\. Ring, K\. Tsihlas, A\. Vezer, M\. Selvi, T\. Shevlane, M\. Rodriguez, T\. Kwiatkowski, S\. Daruki, K\. Rong, A\. Dafoe, N\. FitzGerald, K\. Gu\-Lemberg, M\. Khan, L\. A\. Hendricks, M\. Pellat, V\. Feinberg, J\. Cobon\-Kerr, T\. Sainath, M\. Rauh, S\. H\. Hashemi, R\. Ives, Y\. Hasson, E\. Noland, Y\. Cao, N\. Byrd, L\. Hou, Q\. Wang, T\. Sottiaux, M\. Paganini, J\. Lespiau, A\. Moufarek, S\. Hassan, K\. Shivakumar, J\. van Amersfoort, A\. Mandhane, P\. Joshi, A\. Goyal, M\. Tung, A\. Brock, H\. Sheahan, V\. Misra, C\. Li, N\. Rakićević, M\. Dehghani, F\. Liu, S\. Mittal, J\. Oh, S\. Noury, E\. Sezener, F\. Huot, M\. Lamm, N\. De Cao, C\. Chen, S\. Mudgal, R\. Stella, K\. Brooks, G\. Vasudevan, C\. Liu, M\. Chain, N\. Melinkeri, A\. Cohen, V\. Wang, K\. Seymore, S\. Zubkov, R\. Goel, S\. Yue, S\. Krishnakumaran, B\. Albert, N\. Hurley, M\. Sano, A\. Mohananey, J\. Joughin, E\. Filonov, T\. Kępa, Y\. Eldawy, J\. Lim, R\. Rishi, S\. Badiezadegan, T\. Bos, J\. Chang, S\. Jain, S\. G\. S\. Padmanabhan, S\. Puttagunta, K\. Krishna, L\. Baker, N\. Kalb, V\. Bedapudi, A\. Kurzrok, S\. Lei, A\. Yu, O\. Litvin, X\. Zhou, Z\. Wu, S\. Sobell, A\. Siciliano, A\. Papir, R\. Neale, J\. Bragagnolo, T\. Toor, T\. Chen, V\. Anklin, F\. Wang, R\. Feng, M\. Gholami, K\. Ling, L\. Liu, J\. Walter, H\. Moghaddam, A\. Kishore, J\. Adamek, T\. Mercado, J\. Mallinson, S\. Wandekar, S\. Cagle, E\. Ofek, G\. Garrido, C\. Lombriser, M\. Mukha, B\. Sun, H\. R\. Mohammad, J\. Matak, Y\. Qian, V\. Peswani, P\. Janus, Q\. Yuan, L\. Schelin, O\. David, A\. Garg, Y\. He, O\. Duzhyi, A\. Älgmyr, T\. Lottaz, Q\. Li, V\. Yadav, L\. Xu, A\. Chinien, R\. Shivanna, A\. Chuklin, J\. Li, C\. Spadine, T\. Wolfe, K\. Mohamed, S\. Das, Z\. Dai, K\. He, D\. von Dincklage, S\. Upadhyay, A\. Maurya, L\. Chi, S\. Krause, K\. Salama, P\. G\. Rabinovitch, P\. K\. R\. M, A\. Selvan, M\. Dektiarev, G\. Ghiasi, E\. Guven, H\. Gupta, B\. Liu, D\. Sharma, I\. H\. Shtacher, S\. Paul, O\. Akerlund, F\. Aubet, T\. Huang, C\. Zhu, E\. Zhu, E\. Teixeira, M\. Fritze, F\. Bertolini, L\. Marinescu, M\. Bölle, D\. Paulus, K\. Gupta, T\. Latkar, M\. Chang, J\. Sanders, R\. Wilson, X\. Wu, Y\. Tan, L\. N\. Thiet, T\. Doshi, S\. Lall, S\. Mishra, W\. Chen, T\. Luong, S\. Benjamin, J\. Lee, E\. Andrejczuk, D\. Rabiej, V\. Ranjan, K\. Styrc, P\. Yin, J\. Simon, M\. R\. Harriott, M\. Bansal, A\. Robsky, G\. Bacon, D\. Greene, D\. Mirylenka, C\. Zhou, O\. Sarvana, A\. Goyal, S\. Andermatt, P\. Siegler, B\. Horn, A\. Israel, F\. Pongetti, C\. “\. Chen, M\. Selvatici, P\. Silva, K\. Wang, J\. Tolins, K\. Guu, R\. Yogev, X\. Cai, A\. Agostini, M\. Shah, H\. Nguyen, N\. Ó\. Donnaile, S\. Pereira, L\. Friso, A\. Stambler, A\. Kurzrok, C\. Kuang, Y\. Romanikhin, M\. Geller, Z\. J\. Yan, K\. Jang, C\. Lee, W\. Fica, E\. Malmi, Q\. Tan, D\. Banica, D\. Balle, R\. Pham, Y\. Huang, D\. Avram, H\. Shi, J\. Singh, C\. Hidey, N\. Ahuja, P\. Saxena, D\. Dooley, S\. P\. Potharaju, E\. O’Neill, A\. Gokulchandran, R\. Foley, K\. Zhao, M\. Dusenberry, Y\. Liu, P\. Mehta, R\. Kotikalapudi, C\. Safranek\-Shrader, A\. Goodman, J\. Kessinger, E\. Globen, P\. Kolhar, C\. Gorgolewski, A\. Ibrahim, Y\. Song, A\. Eichenbaum, T\. Brovelli, S\. Potluri, P\. Lahoti, C\. Baetu, A\. Ghorbani, C\. Chen, A\. Crawford, S\. Pal, M\. Sridhar, P\. Gurita, A\. Mujika, I\. Petrovski, P\. Cedoz, C\. Li, S\. Chen, N\. D\. Santo, S\. Goyal, J\. Punjabi, K\. Kappaganthu, C\. Kwak, P\. Lv, S\. Velury, H\. Choudhury, J\. Hall, P\. Shah, R\. Figueira, M\. Thomas, M\. Lu, T\. Zhou, C\. Kumar, T\. Jurdi, S\. Chikkerur, Y\. Ma, A\. Yu, S\. Kwak, V\. Ähdel, S\. Rajayogam, T\. Choma, F\. Liu, A\. Barua, C\. Ji, J\. H\. Park, V\. Hellendoorn, A\. Bailey, T\. Bilal, H\. Zhou, M\. Khatir, C\. Sutton, W\. Rzadkowski, F\. Macintosh, K\. Shagin, P\. Medina, C\. Liang, J\. Zhou, P\. Shah, Y\. Bi, A\. Dankovics, S\. Banga, S\. Lehmann, M\. Bredesen, Z\. Lin, J\. E\. Hoffmann, J\. Lai, R\. Chung, K\. Yang, N\. Balani, A\. Bražinskas, A\. Sozanschi, M\. Hayes, H\. F\. Alcalde, P\. Makarov, W\. Chen, A\. Stella, L\. Snijders, M\. Mandl, A\. Kärrman, P\. Nowak, X\. Wu, A\. Dyck, K\. Vaidyanathan, R\. R, J\. Mallet, M\. Rudominer, E\. Johnston, S\. Mittal, A\. Udathu, J\. Christensen, V\. Verma, Z\. Irving, A\. Santucci, G\. Elsayed, E\. Davoodi, M\. Georgiev, I\. Tenney, N\. Hua, G\. Cideron, E\. Leurent, M\. Alnahlawi, I\. Georgescu, N\. Wei, I\. Zheng, D\. Scandinaro, H\. Jiang, J\. Snoek, M\. Sundararajan, X\. Wang, Z\. Ontiveros, I\. Karo, J\. Cole, V\. Rajashekhar, L\. Tumeh, E\. Ben\-David, R\. Jain, J\. Uesato, R\. Datta, O\. Bunyan, S\. Wu, J\. Zhang, P\. Stanczyk, Y\. Zhang, D\. Steiner, S\. Naskar, M\. Azzam, M\. Johnson, A\. Paszke, C\. Chiu, J\. S\. Elias, A\. Mohiuddin, F\. Muhammad, J\. Miao, A\. Lee, N\. Vieillard, J\. Park, J\. Zhang, J\. Stanway, D\. Garmon, A\. Karmarkar, Z\. Dong, J\. Lee, A\. Kumar, L\. Zhou, J\. Evens, W\. Isaac, G\. Irving, E\. Loper, M\. Fink, I\. Arkatkar, N\. Chen, I\. Shafran, I\. Petrychenko, Z\. Chen, J\. Jia, A\. Levskaya, Z\. Zhu, P\. Grabowski, Y\. Mao, A\. Magni, K\. Yao, J\. Snaider, N\. Casagrande, E\. Palmer, P\. Suganthan, A\. Castaño, I\. Giannoumis, W\. Kim, M\. Rybiński, A\. Sreevatsa, J\. Prendki, D\. Soergel, A\. Goedeckemeyer, W\. Gierke, M\. Jafari, M\. Gaba, J\. Wiesner, D\. G\. Wright, Y\. Wei, H\. Vashisht, Y\. Kulizhskaya, J\. Hoover, M\. Le, L\. Li, C\. Iwuanyanwu, L\. Liu, K\. Ramirez, A\. Khorlin, A\. Cui, T\. Lin, M\. Wu, R\. Aguilar, K\. Pallo, A\. Chakladar, G\. Perng, E\. A\. Abellan, M\. Zhang, I\. Dasgupta, N\. Kushman, I\. Penchev, A\. Repina, X\. Wu, T\. van der Weide, P\. Ponnapalli, C\. Kaplan, J\. Simsa, S\. Li, O\. Dousse, F\. Yang, J\. Piper, N\. Ie, R\. Pasumarthi, N\. Lintz, A\. Vijayakumar, D\. Andor, P\. Valenzuela, M\. Lui, C\. Paduraru, D\. Peng, K\. Lee, S\. Zhang, S\. Greene, D\. D\. Nguyen, P\. Kurylowicz, C\. Hardin, L\. Dixon, L\. Janzer, K\. Choo, Z\. Feng, B\. Zhang, A\. Singhal, D\. Du, D\. McKinnon, N\. Antropova, T\. Bolukbasi, O\. Keller, D\. Reid, D\. Finchelstein, M\. A\. Raad, R\. Crocker, P\. Hawkins, R\. Dadashi, C\. Gaffney, K\. Franko, A\. Bulanova, R\. Leblond, S\. Chung, H\. Askham, L\. C\. Cobo, K\. Xu, F\. Fischer, J\. Xu, C\. Sorokin, C\. Alberti, C\. Lin, C\. Evans, A\. Dimitriev, H\. Forbes, D\. Banarse, Z\. Tung, M\. Omernick, C\. Bishop, R\. Sterneck, R\. Jain, J\. Xia, E\. Amid, F\. Piccinno, X\. Wang, P\. Banzal, D\. J\. Mankowitz, A\. Polozov, V\. Krakovna, S\. Brown, M\. Bateni, D\. Duan, V\. Firoiu, M\. Thotakuri, T\. Natan, M\. Geist, S\. T\. Girgin, H\. Li, J\. Ye, O\. Roval, R\. Tojo, M\. Kwong, J\. Lee\-Thorp, C\. Yew, D\. Sinopalnikov, S\. Ramos, J\. Mellor, A\. Sharma, K\. Wu, D\. Miller, N\. Sonnerat, D\. Vnukov, R\. Greig, J\. Beattie, E\. Caveness, L\. Bai, J\. Eisenschlos, A\. Korchemniy, T\. Tsai, M\. Jasarevic, W\. Kong, P\. Dao, Z\. Zheng, F\. Liu, F\. Yang, R\. Zhu, T\. H\. Teh, J\. Sanmiya, E\. Gladchenko, N\. Trdin, D\. Toyama, E\. Rosen, S\. Tavakkol, L\. Xue, C\. Elkind, O\. Woodman, J\. Carpenter, G\. Papamakarios, R\. Kemp, S\. Kafle, T\. Grunina, R\. Sinha, A\. Talbert, D\. Wu, D\. Owusu\-Afriyie, C\. Du, C\. Thornton, J\. Pont\-Tuset, P\. Narayana, J\. Li, S\. Fatehi, J\. Wieting, O\. Ajmeri, B\. Uria, Y\. Ko, L\. Knight, A\. Héliou, N\. Niu, S\. Gu, C\. Pang, Y\. Li, N\. Levine, A\. Stolovich, R\. Santamaria\-Fernandez, S\. Goenka, W\. Yustalim, R\. Strudel, A\. Elqursh, C\. Deck, H\. Lee, Z\. Li, K\. Levin, R\. Hoffmann, D\. Holtmann\-Rice, O\. Bachem, S\. Arora, C\. Koh, S\. H\. Yeganeh, S\. Põder, M\. Tariq, Y\. Sun, L\. Ionita, M\. Seyedhosseini, P\. Tafti, Z\. Liu, A\. Gulati, J\. Liu, X\. Ye, B\. Chrzaszcz, L\. Wang, N\. Sethi, T\. Li, B\. Brown, S\. Singh, W\. Fan, A\. Parisi, J\. Stanton, V\. Koverkathu, C\. A\. Choquette\-Choo, Y\. Li, T\. J\. Lu, A\. Ittycheriah, P\. Shroff, M\. Varadarajan, S\. Bahargam, R\. Willoughby, D\. Gaddy, G\. Desjardins, M\. Cornero, B\. Robenek, B\. Mittal, B\. Albrecht, A\. Shenoy, F\. Moiseev, H\. Jacobsson, A\. Ghaffarkhah, M\. Rivière, A\. Walton, C\. Crepy, A\. Parrish, Z\. Zhou, C\. Farabet, C\. Radebaugh, P\. Srinivasan, C\. van der Salm, A\. Fidjeland, S\. Scellato, E\. Latorre\-Chimoto, H\. Klimczak\-Plucińska, D\. Bridson, D\. de Cesare, T\. Hudson, P\. Mendolicchio, L\. Walker, A\. Morris, M\. Mauger, A\. Guseynov, A\. Reid, S\. Odoom, L\. Loher, V\. Cotruta, M\. Yenugula, D\. Grewe, A\. Petrushkina, T\. Duerig, A\. Sanchez, S\. Yadlowsky, A\. Shen, A\. Globerson, L\. Webb, S\. Dua, D\. Li, S\. Bhupatiraju, D\. Hurt, H\. Qureshi, A\. Agarwal, T\. Shani, M\. Eyal, A\. Khare, S\. R\. Belle, L\. Wang, C\. Tekur, M\. S\. Kale, J\. Wei, R\. Sang, B\. Saeta, T\. Liechty, Y\. Sun, Y\. Zhao, S\. Lee, P\. Nayak, D\. Fritz, M\. R\. Vuyyuru, J\. Aslanides, N\. Vyas, M\. Wicke, X\. Ma, E\. Eltyshev, N\. Martin, H\. Cate, J\. Manyika, K\. Amiri, Y\. Kim, X\. Xiong, K\. Kang, F\. Luisier, N\. Tripuraneni, D\. Madras, M\. Guo, A\. Waters, O\. Wang, J\. Ainslie, J\. Baldridge, H\. Zhang, G\. Pruthi, J\. Bauer, F\. Yang, R\. Mansour, J\. Gelman, Y\. Xu, G\. Polovets, J\. Liu, H\. Cai, W\. Chen, X\. Sheng, E\. Xue, S\. Ozair, C\. Angermueller, X\. Li, A\. Sinha, W\. Wang, J\. Wiesinger, E\. Koukoumidis, Y\. Tian, A\. Iyer, M\. Gurumurthy, M\. Goldenson, P\. Shah, M\. K\. Blake, H\. Yu, A\. Urbanowicz, J\. Palomaki, C\. Fernando, K\. Durden, H\. Mehta, N\. Momchev, E\. Rahimtoroghi, M\. Georgaki, A\. Raul, S\. Ruder, M\. Redshaw, J\. Lee, D\. Zhou, K\. Jalan, D\. Li, B\. Hechtman, P\. Schuh, M\. Nasr, K\. Milan, V\. Mikulik, J\. Franco, T\. Green, N\. Nguyen, J\. Kelley, A\. Mahendru, A\. Hu, J\. Howland, B\. Vargas, J\. Hui, K\. Bansal, V\. Rao, R\. Ghiya, E\. Wang, K\. Ye, J\. M\. Sarr, M\. M\. Preston, M\. Elish, S\. Li, A\. Kaku, J\. Gupta, I\. Pasupat, D\. Juan, M\. Someswar, T\. M\., X\. Chen, A\. Amini, A\. Fabrikant, E\. Chu, X\. Dong, A\. Muthal, S\. Buthpitiya, S\. Jauhari, N\. Hua, U\. Khandelwal, A\. Hitron, J\. Ren, L\. Rinaldi, S\. Drath, A\. Dabush, N\. Jiang, H\. Godhia, U\. Sachs, A\. Chen, Y\. Fan, H\. Taitelbaum, H\. Noga, Z\. Dai, J\. Wang, C\. Liang, J\. Hamer, C\. Ferng, C\. Elkind, A\. Atias, P\. Lee, V\. Listík, M\. Carlen, J\. van de Kerkhof, M\. Pikus, K\. Zaher, P\. Müller, S\. Zykova, R\. Stefanec, V\. Gatsko, C\. Hirnschall, A\. Sethi, X\. F\. Xu, C\. Ahuja, B\. Tsai, A\. Stefanoiu, B\. Feng, K\. Dhandhania, M\. Katyal, A\. Gupta, A\. Parulekar, D\. Pitta, J\. Zhao, V\. Bhatia, Y\. Bhavnani, O\. Alhadlaq, X\. Li, P\. Danenberg, D\. Tu, A\. Pine, V\. Filippova, A\. Ghosh, B\. Limonchik, B\. Urala, C\. K\. Lanka, D\. Clive, Y\. Sun, E\. Li, H\. Wu, K\. Hongtongsak, I\. Li, K\. Thakkar, K\. Omarov, K\. Majmundar, M\. Alverson, M\. Kucharski, M\. Patel, M\. Jain, M\. Zabelin, P\. Pelagatti, R\. Kohli, S\. Kumar, J\. Kim, S\. Sankar, V\. Shah, L\. Ramachandruni, X\. Zeng, B\. Bariach, L\. Weidinger, T\. Vu, A\. Andreev, A\. He, K\. Hui, S\. Kashem, A\. Subramanya, S\. Hsiao, D\. Hassabis, K\. Kavukcuoglu, A\. Sadovsky, Q\. Le, T\. Strohman, Y\. Wu, S\. Petrov, J\. Dean, and O\. VinyalsGemini: a family of highly capable multimodal models\.arXiv \[cs\.CL\]\.Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Genette \(1980\)G\. GenetteNarrative discourse: an essay in method\.Vol\.3,Cornell University Press\.Cited by:[1st item](https://arxiv.org/html/2608.21430#S3.I1.i1.p1.1)\.
- Glancy \(1992\)H\. M\. GlancyMGM film grosses, 1924–1948: The Eddie Mannix ledger\.Historical Journal of Film, Radio and Television12\(2\),pp\. 127–144\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p1.1)\.
- Glancy \(1995\)H\. M\. GlancyWarner Bros film grosses, 1921–51: The William Schaefer ledger\.Historical Journal of Film, Radio and Television15\(1\),pp\. 55–73\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p1.1)\.
- Guhaet al\.\(2015\)T\. Guha, C\. Huang, N\. Kumar, Y\. Zhu, and S\. S\. NarayananGender representation in cinematic content: a multimodal approach\.InProceedings of the 2015 ACM on International Conference on Multimodal Interaction,pp\. 31–34\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
- Hamiltonet al\.\(2026\)S\. Hamilton, M\. Wilkens, and A\. PiperNarraBench: a comprehensive framework for narrative benchmarking\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 3786–3801\.External Links:[Link](https://aclanthology.org/2026.eacl-long.176/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.176),ISBN 979\-8\-89176\-380\-7Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p5.1)\.
- Honget al\.\(2026\)W\. Hong, W\. Yu, X\. Gu, G\. Wang, G\. Gan, H\. Tang, J\. Cheng, J\. Qi, J\. Ji, L\. Pan, S\. Duan, W\. Wang, Y\. Wang, Y\. Cheng, Z\. He, Z\. Su, Z\. Yang, Z\. Pan, A\. Zeng, B\. Wang, B\. Chen, B\. Shi, C\. Pang, C\. Zhang, D\. Yin, F\. Yang, G\. Chen, H\. Li, J\. Zhu, J\. Chen, J\. Xu, J\. Xu, J\. Chen, J\. Lin, J\. Chen, J\. Wang, J\. Chen, L\. Lei, L\. Gong, L\. Pan, M\. Liu, M\. Xu, M\. Zhang, Q\. Zheng, R\. Lyu, S\. Tu, S\. Yang, S\. Meng, S\. Zhong, S\. Huang, S\. Zhao, S\. Xue, T\. Zhang, T\. Luo, T\. Hao, T\. Tong, W\. Jia, W\. Li, X\. Liu, X\. Zhang, X\. Lyu, X\. Zhang, X\. Fan, X\. Huang, Y\. Xue, Y\. Wang, Y\. Wang, Y\. Wang, Y\. An, Y\. Du, Y\. Huang, Y\. Niu, Y\. Shi, Y\. Wang, Y\. Wang, Y\. Yue, Y\. Li, Y\. Liu, Y\. Zhang, Y\. Wang, Y\. Zhang, Z\. Xue, Z\. Du, Z\. Hou, Z\. Wang, P\. Zhang, D\. Liu, B\. Xu, J\. Li, M\. Huang, Y\. Dong, and J\. TangGLM\-4\.5V and GLM\-4\.1V\-thinking: towards versatile multimodal reasoning with scalable reinforcement learning\.Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Jacobs \(1997\)L\. JacobsThe wages of sin: censorship and the fallen woman film, 1928\-1942\.Univ of California Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Jewell \(1994\)R\. B\. JewellRKO film grosses, 1929–1951: The CJ Tevlin ledger\.Historical Journal of Film, Radio and Television14\(1\),pp\. 37–49\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p1.1)\.
- Johnston \(1973\)C\. JohnstonWomen’s cinema as counter\-cinema\.SEFT,London\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Kayet al\.\(2017\)W\. Kay, J\. Carreira, K\. Simonyan, B\. Zhang, C\. Hillier, S\. Vijayanarasimhan, F\. Viola, T\. Green, T\. Back, P\. Natsev,et al\.The Kinetics human action video dataset\.arXiv preprint arXiv:1705\.06950\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p2.1)\.
- Liet al\.\(2026\)X\. Li, Y\. Wang, J\. Yu, X\. Zeng, Y\. Zhu, H\. Huang, J\. Gao, K\. Li, Y\. He, C\. Wang, Y\. Qiao, Y\. Wang, and L\. WangVideoChat\-Flash: hierarchical compression for long\-context video modeling\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MUjdNcfNPv)Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Linet al\.\(2025\)J\. Lin, J\. Wu, X\. Sun, Z\. Wang, J\. Liu, Y\. Su, X\. Yu, H\. Chen, J\. Luo, Z\. Liu, and E\. BarsoumUnleashing hour\-scale video training for long video\-language understanding\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38,pp\. 17523–17552\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/19741c617c87a25627682e5714af1501-Paper-Conference.pdf)Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Linet al\.\(2014\)T\. Lin, M\. Maire, S\. Belongie, J\. Hays, P\. Perona, D\. Ramanan, P\. Dollár, and C\. L\. ZitnickMicrosoft COCO: common objects in context\.InEuropean conference on computer vision,pp\. 740–755\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p2.1)\.
- Maltby \(2006\)R\. MaltbyOn the prospect of writing cinema history from below\.TMG Journal for Media History9\(2\)\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Nagraniet al\.\(2018\)A\. Nagrani, S\. Albanie, and A\. ZissermanLearnable pins: cross\-modal embeddings for person identity\.InProceedings of the European Conference on Computer Vision \(ECCV\),Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p2.1)\.
- Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards LLMs as operating systems\.arXiv \[cs\.AI\]\.Cited by:[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px5.p1.1)\.
- Piper \(2019\)A\. PiperEnumerations: data and literary study\.University of Chicago Press\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. SutskeverLearning transferable visual models from natural language supervision\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 8748–8763\.External Links:[Link](https://proceedings.mlr.press/v139/radford21a.html)Cited by:[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px5.p1.1)\.
- Rawalet al\.\(2024\)R\. Rawal, K\. Saifullah, M\. Farré, R\. Basri, D\. Jacobs, G\. Somepalli, and T\. GoldsteinCinepile: a long video question answering dataset and benchmark\.arXiv preprint arXiv:2405\.08813\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p4.1)\.
- Ray \(2020\)R\. B\. RayA certain tendency of the hollywood cinema, 1930\-1980\.Princeton University Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p6.1)\.
- Renet al\.\(2025\)W\. Ren, W\. Ma, H\. Yang, C\. Wei, G\. Zhang, and W\. ChenVamba: understanding hour\-long videos with hybrid mamba\-transformers\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 21197–21208\.Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Rich \(2013\)B\. R\. RichNew queer cinema: the director’s cut\.Duke University Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Singhet al\.\(2025\)A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. J\. Ostrow, A\. Ananthram, A\. Nathan, A\. Luo, A\. Helyar, A\. Madry, A\. Efremov, A\. Spyra, A\. Baker\-Whitcomb, A\. Beutel, A\. Karpenko, A\. Makelov, A\. Neitz, A\. Wei, A\. Barr, A\. Kirchmeyer, A\. Ivanov, A\. Christakis, A\. Gillespie, A\. Tam, A\. Bennett, A\. Wan, A\. Huang, A\. M\. Sandjideh, A\. Yang, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Gheorghe, A\. G\. Garcia, A\. Braunstein, A\. Liu, A\. Schmidt, A\. Mereskin, A\. Mishchenko, A\. Applebaum, A\. Rogerson, A\. Rajan, A\. Wei, A\. Kotha, A\. Srivastava, A\. Agrawal, A\. Vijayvergiya, A\. Tyra, A\. Nair, A\. Nayak, B\. Eggers, B\. Ji, B\. Hoover, B\. Chen, B\. Chen, B\. Barak, B\. Minaiev, B\. Hao, B\. Baker, B\. Lightcap, B\. McKinzie, B\. Wang, B\. Quinn, B\. Fioca, B\. Hsu, B\. Yang, B\. Yu, B\. Zhang, B\. Brenner, C\. R\. Zetino, C\. Raymond, C\. Lugaresi, C\. Paz, C\. Hudson, C\. Whitney, C\. Li, C\. Chen, C\. Cole, C\. Voss, C\. Ding, C\. Shen, C\. Huang, C\. Colby, C\. Hallacy, C\. Koch, C\. Lu, C\. Kaplan, C\. Kim, C\. J\. Minott\-Henriques, C\. Frey, C\. Yu, C\. Czarnecki, C\. Reid, C\. Wei, C\. Decareaux, C\. Scheau, C\. Zhang, C\. Forbes, D\. Tang, D\. Goldberg, D\. Roberts, D\. Palmie, D\. Kappler, D\. Levine, D\. Wright, D\. Leo, D\. Lin, D\. Robinson, D\. Grabb, D\. Chen, D\. Lim, D\. Salama, D\. Bhattacharjee, D\. Tsipras, D\. Li, D\. Yu, D\. J\. Strouse, D\. Williams, D\. Hunn, E\. Bayes, E\. Arbus, E\. Akyurek, E\. Y\. Le, E\. Widmann, E\. Yani, E\. Proehl, E\. Sert, E\. Cheung, E\. Schwartz, E\. Han, E\. Jiang, E\. Mitchell, E\. Sigler, E\. Wallace, E\. Ritter, E\. Kavanaugh, E\. Mays, E\. Nikishin, F\. Li, F\. P\. Such, F\. d\. A\. B\. Peres, F\. Raso, F\. Bekerman, F\. Tsimpourlas, F\. Chantzis, F\. Song, F\. Zhang, G\. Raila, G\. McGrath, G\. Briggs, G\. Yang, G\. Parascandolo, G\. Chabot, G\. Kim, G\. Zhao, G\. Valiant, G\. Leclerc, H\. Salman, H\. Wang, H\. Sheng, H\. Jiang, H\. Wang, H\. Jin, H\. Sikchi, H\. Schmidt, H\. Aspegren, H\. Chen, H\. Qiu, H\. Lightman, I\. Covert, I\. Kivlichan, I\. Silber, I\. Sohl, I\. Hammoud, I\. Clavera, I\. Lan, I\. Akkaya, I\. Kostrikov, I\. Kofman, I\. Etinger, I\. Singal, J\. Hehir, J\. Huh, J\. Pan, J\. Wilczynski, J\. Pachocki, J\. Lee, J\. Quinn, J\. Kiros, J\. Kalra, J\. Samaroo, J\. Wang, J\. Wolfe, J\. Chen, J\. Wang, J\. Harb, J\. Han, J\. Wang, J\. Zhao, J\. Chen, J\. Yang, J\. Tworek, J\. Chand, J\. Landon, J\. Liang, J\. Lin, J\. Liu, J\. Wang, J\. Tang, J\. Yin, J\. Jang, J\. Morris, J\. Flynn, J\. Ferstad, J\. Heidecke, J\. Fishbein, J\. Hallman, J\. Grant, J\. Chien, J\. Gordon, J\. Park, J\. Liss, J\. Kraaijeveld, J\. Guay, J\. Mo, J\. Lawson, J\. McGrath, J\. Vendrow, J\. Jiao, J\. Lee, J\. Steele, J\. Wang, J\. Mao, K\. Chen, K\. Hayashi, K\. Xiao, K\. Salahi, K\. Wu, K\. Sekhri, K\. Sharma, K\. Singhal, K\. Li, K\. Nguyen, K\. Gu\-Lemberg, K\. King, K\. Liu, K\. Stone, K\. Yu, K\. Ying, K\. Georgiev, K\. Lim, K\. Tirumala, K\. Miller, L\. Ahmad, L\. Lv, L\. Clare, L\. Fauconnet, L\. Itow, L\. Yang, L\. Romaniuk, L\. Anise, L\. Byron, L\. Pathak, L\. Maksin, L\. Lo, L\. Ho, L\. Jing, L\. Wu, L\. Xiong, L\. Mamitsuka, L\. Yang, L\. McCallum, L\. Held, L\. Bourgeois, L\. Engstrom, L\. Kuhn, L\. Feuvrier, L\. Zhang, L\. Switzer, L\. Kondraciuk, L\. Kaiser, M\. Joglekar, M\. Singh, M\. Shah, M\. Stratta, M\. Williams, M\. Chen, M\. Sun, M\. Cayton, M\. Li, M\. Zhang, M\. Aljubeh, M\. Nichols, M\. Haines, M\. Schwarzer, M\. Gupta, M\. Shah, M\. Huang, M\. Dong, M\. Wang, M\. Glaese, M\. Carroll, M\. Lampe, M\. Malek, M\. Sharman, M\. Zhang, M\. Wang, M\. Pokrass, M\. Florian, M\. Pavlov, M\. Wang, M\. Chen, M\. Wang, M\. Feng, M\. Bavarian, M\. Lin, M\. Abdool, M\. Rohaninejad, N\. Soto, N\. Staudacher, N\. LaFontaine, N\. Marwell, N\. Liu, N\. Preston, N\. Turley, N\. Ansman, N\. Blades, N\. Pancha, N\. Mikhaylin, N\. Felix, N\. Handa, N\. Rai, N\. Keskar, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, O\. Gleeson, P\. Mishkin, P\. Lesiewicz, P\. Baltescu, P\. Belov, P\. Zhokhov, P\. Pronin, P\. Guo, P\. Thacker, Q\. Liu, Q\. Yuan, Q\. Liu, R\. Dias, R\. Puckett, R\. Arora, R\. T\. Mullapudi, R\. Gaon, R\. Miyara, R\. Song, R\. Aggarwal, R\. J\. Marsan, R\. Yemiru, R\. Xiong, R\. Kshirsagar, R\. Nuttall, R\. Tsiupa, R\. Eldan, R\. Wang, R\. James, R\. Ziv, R\. Shu, R\. Nigmatullin, S\. Jain, S\. Talaie, S\. Altman, S\. Arnesen, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Yoo, S\. Heon, S\. Ethersmith, S\. Grove, S\. Taylor, S\. Bubeck, S\. Banesiu, S\. Amdo, S\. Zhao, S\. Wu, S\. Santurkar, S\. Zhao, S\. R\. Chaudhuri, S\. Krishnaswamy, Shuaiqi, Xia, S\. Cheng, S\. Anadkat, S\. P\. Fishman, S\. Tobin, S\. Fu, S\. Jain, S\. Mei, S\. Egoian, S\. Kim, S\. Golden, S\. Q\. Mah, S\. Lin, S\. Imm, S\. Sharpe, S\. Yadlowsky, S\. Choudhry, S\. Eum, S\. Sanjeev, T\. Khan, T\. Stramer, T\. Wang, T\. Xin, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Degry, T\. Shadwell, T\. Fu, T\. Gao, T\. Garipov, T\. Sriskandarajah, T\. Sherbakov, T\. Kaftan, T\. Hiratsuka, T\. Wang, T\. Song, T\. Zhao, T\. Peterson, V\. Kharitonov, V\. Chernova, V\. Kosaraju, V\. Kuo, V\. Pong, V\. Verma, V\. Petrov, W\. Jiang, W\. Zhang, W\. Zhou, W\. Xie, W\. Zhan, W\. McCabe, W\. DePue, W\. Ellsworth, W\. Bain, W\. Thompson, X\. Chen, X\. Qi, X\. Xiang, X\. Shi, Y\. Dubois, Y\. Yu, Y\. Khakbaz, Y\. Wu, Y\. Qian, Y\. T\. Lee, Y\. Chen, Y\. Zhang, Y\. Xiong, Y\. Tian, Y\. Cha, Y\. Bai, Y\. Yang, Y\. Yuan, Y\. Li, Y\. Zhang, Y\. Yang, Y\. Jin, Y\. Jiang, Y\. Wang, Y\. Wang, Y\. Liu, Z\. Stubenvoll, Z\. Dou, Z\. Wu, and Z\. WangOpenAI GPT\-5 system card\.arXiv \[cs\.CL\]\.Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Somandepalliet al\.\(2021\)K\. Somandepalli, T\. Guha, V\. R\. Martinez, N\. Kumar, H\. Adam, and S\. NarayananComputational media intelligence: human\-centered machine analysis of media\.Proceedings of the IEEE109\(5\),pp\. 891–910\.External Links:[Document](https://dx.doi.org/10.1109/JPROC.2020.3047978)Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
- Soucek and Lokoc \(2024\)T\. Soucek and J\. LokocTransnet v2: an effective deep network architecture for fast shot transition detection\.InProceedings of the 32nd ACM International Conference on Multimedia,pp\. 11218–11221\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px4.p1.1)\.
- Staiger \(1992\)J\. StaigerInterpreting films: studies in the historical reception of american cinema\.Princeton University Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Stewart \(2005\)J\. N\. StewartMigrating to the movies: cinema and black urban modernity\.Univ of California Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Suiet al\.\(2025\)P\. Sui, J\. D\. Rodriguez, P\. Laban, J\. D\. Murphy, J\. P\. Dexter, R\. J\. So, S\. Baker, and P\. ChaudhuriKRISTEVA: close reading as a novel task for benchmarking interpretive reasoning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 32829–32849\.External Links:[Link](https://aclanthology.org/2025.acl-long.1577/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1577),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p5.1)\.
- Tapaswiet al\.\(2016\)M\. Tapaswi, Y\. Zhu, R\. Stiefelhagen, A\. Torralba, R\. Urtasun, and S\. FidlerMovieQA: understanding stories in movies through question\-answering\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 4631–4640\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p4.1)\.
- Underwood \(2019\)T\. UnderwoodDistant horizons: digital evidence and literary change\.University of Chicago Press\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
- Verhoevenet al\.\(2019\)D\. Verhoeven, B\. Coate, and V\. ZemaityteRe\-distributing gender in the global film industry: beyond \#MeToo and \#MeThree\.Media Industries\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Wanget al\.\(2025\)W\. Wang, Z\. He, W\. Hong, Y\. Cheng, X\. Zhang, J\. Qi, M\. Ding, X\. Gu, S\. Huang, B\. Xu,et al\.Lvbench: an extreme long video understanding benchmark\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 22958–22967\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p4.1),[§1](https://arxiv.org/html/2608.21430#S1.p5.1)\.
- White \(2015\)P\. WhiteWomen’s cinema, world cinema: projecting contemporary feminisms\.Duke University Press\.Cited by:[§2\.1](https://arxiv.org/html/2608.21430#S2.SS1.p7.1)\.
- Zabihet al\.\(1995\)R\. Zabih, J\. Miller, and K\. MaiA feature\-based algorithm for detecting and classifying scene breaks\.InProceedings of the third ACM international conference on Multimedia,pp\. 189–200\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p2.1)\.
- Zaraniset al\.\(2025\)E\. Zaranis, A\. Farinhas, S\. Santos, B\. Canaverde, M\. M\. Ramos, A\. K\. Surikuchi, A\. Viveiros, B\. Liao, E\. Bueno\-Benito, N\. Sivakumaran, P\. Vasylenko, S\. Yu, S\. Sannigrahi, W\. Mohammed, B\. Peters, D\. S\. Villegas, E\. Stengel\-Eskin, G\. Attanasio, J\. Yoon, S\. Frank, A\. Suglia, C\. Zerva, D\. Elliott, M\. Dimiccoli, M\. Bansal, O\. Lanz, R\. Bernardi, R\. Fernández, S\. Pezzelle, V\. Niculae, and A\. F\. T\. MartinsMovie Facts and Fibs \(MF2\{\}^\{2\}\): a benchmark for long movie understanding\.External Links:2506\.06275,[Link](https://arxiv.org/abs/2506.06275)Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p4.1),[§1](https://arxiv.org/html/2608.21430#S1.p5.1),[§4](https://arxiv.org/html/2608.21430#S4.p1.1),[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px2.p1.1)\.
- Zenget al\.\(2023\)A\. Zeng, M\. Attarian, B\. Ichter, K\. M\. Choromanski, A\. Wong, S\. Welker, F\. Tombari, A\. Purohit, M\. S\. Ryoo, V\. Sindhwani, J\. Lee, V\. Vanhoucke, and P\. FlorenceSocratic models: composing zero\-shot multimodal reasoning with language\.InThe Eleventh International Conference on Learning Representations,Cited by:[§5\.1](https://arxiv.org/html/2608.21430#S5.SS1.SSS0.Px4)\.
- Zhanget al\.\(2024\)Y\. Zhang, B\. Li, h\. Liu, Y\. j\. Lee, L\. Gui, D\. Fu, J\. Feng, Z\. Liu, and C\. LiLLaVA\-NeXT: a strong zero\-shot video understanding model\.External Links:[Link](https://llava-vl.github.io/blog/2024-04-30-llava-next-video/)Cited by:[§5\.2](https://arxiv.org/html/2608.21430#S5.SS2.p1.1)\.
- Zribiet al\.\(2026\)Y\. Zribi, F\. Cafiero, V\. Lépinay, and C\. Vidal\-GorèneTiming in stand\-up comedy: text, audio, laughter, kinesics \(TIC\-TALK\): pipeline and database for the multimodal study of comedic timing\.arXiv preprint arXiv:2603\.21803\.Cited by:[§1](https://arxiv.org/html/2608.21430#S1.p1.1)\.
## Appendix AAuthor contributions
- •Conceptualization of narrative categories: DB, KC, AC, JH, RK, MM, AP, INS, YS
- •Benchmark question annotation: JH, RK, MM, AP, INS, YS \(major\); DB \(minor\)
- •Benchmark answer annotation \(evaluation\): DB, KC, AC, JH, RK, MM, AP, RS, INS, YS
- •*Variety*box office annotation: MM, DB
- •Data curation \(*Variety*,*Catalog of Copyright Entries*, Internet Archive\): DB
- •Statistical analysis: DB, KC
- •Computational methodology: KC, DB
- •Validation: DB, KC
- •Writing: DB, AC, KC, RS
## Appendix BBox office extraction accuracy
Our procedure for generating ranked lists of movies per year by their total box office numbers involves using a multimodal LLM to extract individual⟨\\langlecity, theater, movie, gross⟩\\rangletuples from pages of*Variety*magazine, and map movie titles to IMDB identifiers \(as described in the main body above\)\. We evaluate the accuracy of this process by comparing the extractions derived from several models with those manually created by visually inspecting 12 issues of*Variety*\.
In order to most closely measure the performance on our final goal \(identifying the topnnmovies by box office per year\), we apply the same process of mapping titles to IMDB identifiers for both the gold and extracted sets\. If a movie does not exist in that mapping, we identify it by its lowercased title\. We sum all gross values for the same movie across all tuples to yield a⟨\\langlemovie, $⟩\\rangleset and generate a ranking over movies by their total gross\. We do so over the gold annotations and the predicted values to generate two ranked lists\.
We measure accuracy with three metrics:
1. 1\.For all movies that exist in the gold*and*predicted sets, we measure the Spearman rank correlation coefficientρ\\rhobetween the gross dollar amounts in the two lists\. This captures the degree to which the predicted ranks correspond with the true ranks, when we have a gross dollar amount for all movies\.
2. 2\.To capture the degree to which the predicted ranks invent movies that do not exist, or fail to extract ones that do, we report the “Movie F1” score, where precision is defined as the fraction of predicted movies that exist in gold annotations and recall as the fraction of movies in the gold annotations that exist in the predicted set\. Low precision would mean that models are hallucinating movies that do not exist; low recall would mean that models are failing to identify the ones that do\.
3. 3\.While the two metrics above evaluate our final goal \(assessing the ranking of movie titles by reported gross\), we also evaluate the accuracy of exact tuples \(i\.e\., the degree to which a tuple gets the movie title, theater, city and gross exactly correct\)\.
Table[3](https://arxiv.org/html/2608.21430#A2.T3)reports those metrics over several commercial model variants; each number represents the average of that metric over all*Variety*issues \(so that each issue carries equal weight in assessment\)\.
Table 3:Box office extraction accuracy, with 95% bootstrap confidence intervals\.Modelρ\\rhoMovie F1Tuple F1Gemini 3 Pro0\.966\[0\.953\-0\.977\]0\.921\[0\.909\-0\.934\]0\.825\[0\.800\-0\.851\]\- \(high→\\rightarrowlow thinking level\)0\.934\[0\.912\-0\.950\]0\.858\[0\.802\-0\.896\]0\.711\[0\.636\-0\.770\]\- \(ultrahigh→\\rightarrowhigh res\)0\.916\[0\.887\-0\.938\]0\.847\[0\.800\-0\.886\]0\.609\[0\.533\-0\.680\]\- \(pro→\\rightarrowflash\)0\.930\[0\.903\-0\.951\]0\.858\[0\.826\-0\.886\]0\.675\[0\.633\-0\.728\]GPT 5\.4 \(original res\)0\.916\[0\.893\-0\.937\]0\.809\[0\.774\-0\.845\]0\.684\[0\.645\-0\.725\]\- \(mini\)0\.605\[0\.493\-0\.699\]0\.388\[0\.329\-0\.451\]0\.083\[0\.048\-0\.128\]
## Appendix CSample registrations
Figure[3](https://arxiv.org/html/2608.21430#A3.F3)illustrates a copyright registration notice for the movie*The Bells of St\. Mary’s*take from the*Catalog of Copyright Entries, Cumulative Series 1940–1949*; figure[4](https://arxiv.org/html/2608.21430#A3.F4)shows the copyright registration renewal for that same film exactly 28 years later\. In this case, the renewal references the registration date \(6Dec45\) and number \(L81\) of the original\.
Figure 3:Registration for*Bells of St\. Mary’s*recorded in the*Catalog of Copyright Entries, Cumulative Series 1940–1949*Figure 4:Registration renewal for*Bells of St\. Mary’s*recorded in the*Catalog of Copyright Entries, Third Series, Volume 27, Parts 12–13, Number 1: Motion Pictures \(January–June 1973\)*We use Gemini Pro 2\.5 to extract the movie title, registration identifier and copyright date from each entry from the “Registrations” section of the CCE—e\.g\.,⟨\\langleBells of St\. Mary’s, LP81, 6Dec45⟩\\rangle—and movie title, original registration number, original copyright date, and renewal number and renewal date from the “Renewals” section—e\.g\.,⟨\\langleBells of St\. Mary’s, LP81, 6Dec45, R542408, 3Jan73⟩\\ranglefor volumes from 1957–1978 \(corresponding to an earliest original registration date of 1929\)\. We further extracted all motion picture renewals from the Copyright Public Records System to capture renewals made after 1978\.
## Appendix DMemorization
As noted in the main text, to identify questions that are highly memorized by models \(answerable without access to the movie\), we pass all questions through three frontier LLMs—Gemini Pro 3\.1, Claude Opus 4\.7 and GPT 5\.5—with the following prompt: “Based on your knowledge of the movie \{MOVIE\} \(\{YEAR\}\), answer the following question\.” All models exhibit similar rates of memorization \(Gemini39\.4%39\.4\\%, Opus40\.7%40\.7\\%, GPT38\.8%38\.8\\%\)\. To mitigate this effect, we subselect questions from the benchmark pool so that the average performance across all models when prompted with the movie title and date alone is approximately25%25\\%\(reflecting a true random guess\)\. The selection process ranks all questions in ascending order by the total number of the three models that correctly answer it from metadata alone, and adds questions sequentially to the benchmark until an average accuracy of25%25\\%is reached\. This yields a total of 628 benchmark questions \(discarding 151 from the original pool\)\. As Table[4](https://arxiv.org/html/2608.21430#A4.T4)illustrates, the exclusion rate varies widely by category\. For convenience, we let𝒟\\mathcal\{D\}denote the full pool of 779 candidate questions, andℬ\\mathcal\{B\}the post\-filter benchmark of 628 questions used in the experiments described §[5](https://arxiv.org/html/2608.21430#S5)\.
Table 4:Exclusion rate by category\.CategoryExclusion rateℬ\\mathcal\{B\}count𝒟\\mathcal\{D\}countRepresentation0\.3892236Symbolism0\.2412229Temporality0\.2215368Plot0\.209155196Object Identification0\.20092115Character0\.167135162Perspective0\.1554958Setting0\.130100115Total628779
## Appendix EFrame budget sweep
For the six end\-to\-end backbones run at multiple frame budgets \(GPT\-5\-mini, Qwen3\-VL\-8B, and GLM\-4\.1V\-9B\-Thinking on the vision–language side; HourLLaVA, VAMBA\-Qwen2\-VL\-7B, and VideoChat\-Flash on the long\-video side\), average accuracy is reported in Table[5](https://arxiv.org/html/2608.21430#A5.T5)\(appendix[E](https://arxiv.org/html/2608.21430#A5)\) forN∈\{64,128,256\}N\\in\\\{64,128,256\\\}\. The largest within\-backbone gain over the full sweep is GPT\-5\-mini’s\+3\.6\+3\.6pp fromN=64N\{=\}64toN=256N\{=\}256; the next largest is Qwen3\-VL\-8B’s\+4\.5\+4\.5pp fromN=64N\{=\}64toN=128N\{=\}128\. Within\-backbone CIs overlap heavily across the three budgets, and three of the six backbones \(Qwen3\-VL\-8B, GLM\-4\.1V\-9B\-Thinking, VideoChat\-Flash, VAMBA\-Qwen2\-VL\-7B in the non\-monotone direction\) do not improve monotonically withNN\. The frame\-budget effect is an order of magnitude smaller than the audio\-access gap shown in §[5\.2](https://arxiv.org/html/2608.21430#S5.SS2), and we therefore use a single canonical budget per sub\-paradigm in Table[1](https://arxiv.org/html/2608.21430#S5.T1)\(N=64N\{=\}64for long\-video models,N=256N\{=\}256for vision–language models\)\.
Table 5:Frame\-budget sweepN∈\{64,128,256\}N\\in\\\{64,128,256\\\}for end\-to\-end vision\-language and long\-video backbones run at multiple input budgets\.BackboneN=64N\{=\}64N=128N\{=\}128N=256N\{=\}256End\-to\-end vision\-language modelsGPT\-5\-mini30\.6\[27\.2–33\.9\]32\.8\[29\.3–36\.3\]34\.2\[30\.7–37\.7\]Qwen3\-VL\-8B25\.3\[22\.0–28\.5\]29\.8\[26\.4–33\.1\]28\.8\[25\.5–32\.2\]GLM\-4\.1V\-9B\-Thinking25\.8\[22\.6–29\.0\]25\.6\[22\.5–28\.8\]24\.2\[21\.2–27\.4\]End\-to\-end long\-video modelsHourLLaVA23\.6\[20\.4–26\.8\]24\.7\[21\.5–27\.9\]26\.3\[23\.1–29\.5\]VAMBA\-Qwen2\-VL\-7B23\.7\[20\.5–26\.9\]25\.2\[22\.0–28\.5\]23\.2\[20\.1–26\.4\]VideoChat\-Flash27\.4\[24\.0–30\.7\]27\.2\[23\.9–30\.6\]27\.1\[23\.7–30\.4\]
## Appendix FMovies in Benchmark
TitleYearIMDB GenresDirectorProduction CompanyMcLintock\!1963Comedy, WesternAndrew V\. McLaglenBatjac ProductionsBeneath the 12\-Mile Reef1953Adventure, Drama, RomanceRobert D\. Webb20th Century FoxCyrano de Bergerac1951Adventure, Drama, RomanceMichael GordonStanley Kramer ProductionsGo for Broke\!1951Drama, History, WarRobert PiroshLoew’sRoyal Wedding1951Comedy, Musical, RomanceStanley DonenLoew’sThree Guys Named Mike1951Comedy, RomanceCharles WaltersMetro\-Goldwyn\-Mayer \(MGM\)The Inspector General1949Comedy, Musical, RomanceHenry KosterWarner Bros\.Tulsa1949Drama, WesternStuart HeislerWalter Wanger ProductionsHe Walked by Night1948Crime, Drama, Film\-NoirAlfred L\. WerkerBryan Foy ProductionsMy Favorite Brunette1947Comedy, Crime, MysteryElliott NugentHope EnterprisesThe Perils of Pauline1947Drama, RomanceGeorge MarshallParamount PicturesSmash Up: The Story of a Woman1947Comedy, Crime, DramaStuart HeislerWalter Wanger ProductionsTill the Clouds Roll By1947Biography, MusicalRichard WhorfMetro\-Goldwyn\-Mayer \(MGM\)Angel on My Shoulder1946Adventure, Comedy, FantasyArchie MayoCharles R\. Rogers ProductionsThe Strange Love of Martha Ivers1946Drama, Film\-Noir, RomanceLewis MilestoneHal Wallis ProductionsThe Stranger1946Crime, Drama, Film\-NoirOrson WellesInternational Pictures, The Haig CorporationBlood on the Sun1945Drama, Romance, ThrillerFrank LloydWilliam Cagney ProductionsCaptain Kidd1945Adventure, Biography, DramaRowland V\. LeeBenedict Bogeaus ProductionThe Stork Club1945Comedy, Musical, RomanceHal WalkerB\.G\. DeSylva Productions Inc\.Stage Door Canteen1943Comedy, Music, RomanceFrank BorzageSol Lesser ProductionsRudyard Kipling’s Jungle Book1942Action, Adventure, FamilyZoltan KordaAlexander Korda FilmsBilly the Kid in Santa Fe1941Drama, WesternSam NewfieldSigmund Neufeld ProductionsMeet John Doe1941Comedy, Drama, RomanceFrank CapraFrank Capra ProductionsSecond Chorus1941Comedy, Musical, RomanceH\.C\. PotterBoris Morros ProductionsHis Girl Friday1940Comedy, Drama, RomanceHoward HawksColumbia PicturesGulliver’s Travels1939Adventure, Animation, ComedyDave FleischerFleischer StudiosThe Little Princess1939Comedy, Drama, FamilyWalter Lang20th Century FoxLove Affair1939Comedy, Drama, RomanceLeo McCareyRKO Radio PicturesMade for Each Other1939Comedy, Drama, RomanceJohn CromwellSelznick International PicturesNurse Edith Cavell1939Biography, Drama, WarHerbert WilcoxImperadio Pictures Ltd\.Letter of Introduction1938Comedy, Drama, MysteryJohn M\. StahlUniversal PicturesA Star Is Born1937Drama, RomanceWilliam A\. WellmanSelznick International PicturesSwing High, Swing Low1937Comedy, Drama, MusicalMitchell LeisenParamount PicturesLittle Lord Fauntleroy1936Drama, FamilyJohn CromwellSelznick International PicturesBecky Sharp1935Drama, Romance, WarRouben MamoulianPioneer Pictures CorporationOf Human Bondage1934Drama, Film\-Noir, RomanceJohn CromwellRKO Radio PicturesA Farewell to Arms1932Drama, Romance, WarFrank BorzageParamount PicturesBird of Paradise1932Adventure, Drama, RomanceKing VidorRKO Radio PicturesRain1932DramaLewis MilestoneFeature ProductionsThe Front Page1931Comedy, Crime, DramaLewis MilestoneThe Caddo CompanyParlor, Bedroom and Bath1931ComedyEdward SedgwickMetro\-Goldwyn\-Mayer \(MGM\)Street Scene1931Drama, RomanceKing VidorThe Samuel Goldwyn Company, Feature ProductionsAll Quiet on the Western Front1930Drama, WarLewis MilestoneUniversal PicturesAnimal Crackers1930Comedy, Family, MusicalVictor HeermanParamount PicturesAnybody’s Woman1930Drama, RomanceDorothy ArznerParamount PicturesThe Big House1930Crime, Drama, ThrillerGeorge W\. HillMetro\-Goldwyn\-Mayer \(MGM\), Cosmopolitan ProductionsThe Big Trail1930Adventure, Drama, RomanceRaoul WalshFox Film CorporationCheck and Double Check1930ComedyMelville W\. BrownRKO Radio PicturesThe Divorcee1930Drama, RomanceRobert Z\. LeonardMetro\-Goldwyn\-Mayer \(MGM\)Feet First1930Adventure, Comedy, FamilyClyde BruckmanThe Harold Lloyd CorporationMin and Bill1930Comedy, DramaGeorge W\. HillMetro\-Goldwyn\-Mayer \(MGM\)Reaching for the Moon1930Comedy, RomanceEdmund GouldingFeature ProductionsSong o’ My Heart1930Drama, Music, RomanceFrank BorzageFox Film CorporationBulldog Drummond1929Crime, Drama, MysteryF\. Richard JonesThe Samuel Goldwyn CompanyThe Canary Murder Case1929Crime, Drama, MysteryMalcolm St\. ClairParamount PicturesCoquette1929Drama, RomanceSam TaylorPickford CorporationHappy Days1929Comedy, Musical, RomanceBenjamin StoloffFox Film CorporationSally1929MusicalJohn Francis DillonFirst National PicturesThe Trespasser1929Drama, RomanceEdmund GouldingGloria ProductionsWeary River1929Drama, RomanceFrank LloydFirst National PicturesThe Wild Party1929Comedy, Drama, RomanceDorothy ArznerParamount PicturesSimilar Articles
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
MuseBench is a comprehensive benchmark introduced to evaluate multimodal large language models on nuanced, intent-level understanding of audiovisual arts, revealing that even the best model achieves only 48.29% accuracy compared to 87.18% for human experts.
Detecting AI-Generated Content on Social Media with Multi-modal Language Models
This paper from Meta and Carnegie Mellon presents a multi-modal vision-language model pipeline for detecting AI-generated content on social media, achieving state-of-the-art performance and positive downstream impacts on user engagement.
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Artifact-Bench is a comprehensive benchmark that evaluates multimodal large language models on detecting and analyzing artifacts in AI-generated videos, revealing significant limitations and misalignment with human perception.
ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models
This paper introduces ArtECulture, a benchmark for culture-conditioned visual emotion understanding in multimodal large language models, covering English, Chinese, and Arabic cultures with balanced Western and non-Western artwork. Evaluations reveal the task remains challenging, and the authors propose a retrieval-augmented framework to inject cultural knowledge into MLLMs.
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Introduces MultivationBench, a benchmark for evaluating multimodal large language models' sequential motivation reasoning using story-driven visual narratives based on Maslow's hierarchy and Reiss's desires. Results show all tested models struggle with dynamic motivation inference.