<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2023.1125533</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Hypothesis and Theory</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Understanding image-text relations and news values for multimodal news analysis</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Cheema</surname> <given-names>Gullal S.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2100782/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Hakimov</surname> <given-names>Sherzod</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2262291/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>M&#x000FC;ller-Budack</surname> <given-names>Eric</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2262348/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Otto</surname> <given-names>Christian</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1659529/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Bateman</surname> <given-names>John A.</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1235093/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ewerth</surname> <given-names>Ralph</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2186846/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>TIB &#x02013; Leibniz Information Centre for Science and Technology</institution>, <addr-line>Hannover</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Computational Linguistics, University of Potsdam</institution>, <addr-line>Potsdam</addr-line>, <country>Germany</country></aff>
<aff id="aff3"><sup>3</sup><institution>L3S Research Center, Leibniz University Hannover</institution>, <addr-line>Hannover</addr-line>, <country>Germany</country></aff>
<aff id="aff4"><sup>4</sup><institution>Department of English and Linguistics, Universit&#x000E4;t Bremen</institution>, <addr-line>Bremen</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Andy L&#x000FC;cking, Universit&#x000E9; Paris Cit&#x000E9;, France</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Marco Polignano, University of Bari Aldo Moro, Italy; Everardo Reyes, Universit&#x000E9; Paris 8, France</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Gullal S. Cheema <email>gullal.cheema&#x00040;tib.eu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Language and Computation, a section of the journal Frontiers in Artificial Intelligence</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>05</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>6</volume>
<elocation-id>1125533</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>03</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Cheema, Hakimov, M&#x000FC;ller-Budack, Otto, Bateman and Ewerth.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Cheema, Hakimov, M&#x000FC;ller-Budack, Otto, Bateman and Ewerth</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The analysis of news dissemination is of utmost importance since the credibility of information and the identification of disinformation and misinformation affect society as a whole. Given the large amounts of news data published daily on the Web, the empirical analysis of news with regard to research questions and the detection of problematic news content on the Web require computational methods that work at scale. Today&#x00027;s online news are typically disseminated in a multimodal form, including various presentation modalities such as text, image, audio, and video. Recent developments in multimodal machine learning now make it possible to capture basic &#x0201C;descriptive&#x0201D; relations between modalities&#x02013;such as correspondences between words and phrases, on the one hand, and corresponding visual depictions of the verbally expressed information on the other. Although such advances have enabled tremendous progress in tasks like image captioning, text-to-image generation and visual question answering, in domains such as news dissemination, there is a need to go further. In this paper, we introduce a novel framework for the computational analysis of multimodal news. We motivate a set of more complex image-text relations as well as multimodal news values based on real examples of news reports and consider their realization by computational approaches. To this end, we provide (a) an overview of existing literature from <italic>semiotics</italic> where detailed proposals have been made for taxonomies covering diverse image-text relations generalisable to any domain; (b) an overview of computational work that derives models of image-text relations from data; and (c) an overview of a particular class of news-centric attributes developed in journalism studies called news values. The result is a novel framework for multimodal news analysis that closes existing gaps in previous work while maintaining and combining the strengths of those accounts. We assess and discuss the elements of the framework with real-world examples and use cases, setting out research directions at the intersection of multimodal learning, multimodal analytics and computational social sciences that can benefit from our approach.</p>
</abstract>
<kwd-group>
<kwd>multimodality</kwd>
<kwd>news analysis</kwd>
<kwd>news values</kwd>
<kwd>computational analytics</kwd>
<kwd>machine learning</kwd>
<kwd>image-text relations</kwd>
<kwd>semiotics</kwd>
<kwd>journalism</kwd>
</kwd-group>
<counts>
<fig-count count="13"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="151"/>
<page-count count="29"/>
<word-count count="21748"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>News media today convey information about events worldwide in a broad variety of formats, including print, television, news websites, and social media platforms. News websites have evolved to use a broad range of presentation modalities, including text, photos, diagrams, and videos. With an unprecedented increase of information and ease of access and consumption due to the Internet, it has become increasingly important to analyse news content (Karlsson and Sj&#x000F8;vaag, <xref ref-type="bibr" rid="B57">2016</xref>), evaluate its correctness, and understand the spread of information, regardless of the presentation modalities involved. Given the vast quantity of news articles, however, such analysis cannot be accomplished without computational methods, be that for empirical analysis of multimodal news for research or as a response to the highly-pressing need for (software) tools that help to identify and filter problematic news content on the Web and in social media.</p>
<p>Multimodal news analysis involves several challenges, as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. There, the two pairs contain different cross-modal relations, which add complexity to the conveyed message. The samples are both taken from sports news but exhibit different relations between image and text. The sample on the right shows a stock photograph of the <italic>Tokyo Olympics 2020</italic>. But this image is either uncorrelated or contradictory to the surrounding text because the text mentions an investigation resulting from the death of workers at Olympics construction sites. The image, in contrast, simply depicts two persons on stage celebrating Olympics torch relay. The information in both modalities does not convey a similar message and, indeed, taken at face value, may well suggest unwarranted connections &#x02013; for example: are the people shown those failing to investigate? The sample on the left of the figure is very different in that the caption and image, in this case, are congruent. They exhibit a complementary and additive relation because the caption provides more information (opponent team and date) to that of the image. At the same time, the image complements the text by showing the mentioned football player.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Examples of image-text pairs from news. <bold>Left</bold>: An additive relationship between image and text. <bold>Right</bold>: A stock photograph uncorrelated or contradictory to the text. Highlighted text shows the contradicting and overlapping aspects within the images respectively. In general, an image can have relations to text in different parts of an article, but here we focus on the headline (bold) and caption of the image as these already exhibit some of the complexity involved.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0001.tif"/>
</fig>
<p>These examples highlight the difficulty of analysing multimodal news articles where information from different modalities may stand in a broad range of interrelationships. Current computational approaches for news analysis mostly focus on <italic>unimodal</italic> and <italic>data-driven</italic> aspects like sentiment analysis (Godbole et al., <xref ref-type="bibr" rid="B40">2007</xref>; Taj et al., <xref ref-type="bibr" rid="B120">2019</xref>), stance detection (Hanselowski et al., <xref ref-type="bibr" rid="B47">2018</xref>), and focus location estimation (D&#x00027;Ignazio et al., <xref ref-type="bibr" rid="B35">2014</xref>; Imani et al., <xref ref-type="bibr" rid="B54">2017</xref>). Moreover, hardly any computational work has adopted relevant taxonomies of text-image relationships developed within semiotics (Barthes, <xref ref-type="bibr" rid="B9">1977</xref>; Marsh and White, <xref ref-type="bibr" rid="B74">2003</xref>; Martinec and Salway, <xref ref-type="bibr" rid="B77">2005</xref>) or media studies (Bednarek, <xref ref-type="bibr" rid="B11">2016</xref>; Caple et al., <xref ref-type="bibr" rid="B23">2020</xref>), sometimes specifically for analyzing news articles. As a result, computational methods do not model cross-modal relations between text and images and cannot provide interpretations of the (overall) multimodal message, which can be caused by complex mechanisms known under the notion of meaning multiplication (Lemke, <xref ref-type="bibr" rid="B66">1998</xref>; Bateman, <xref ref-type="bibr" rid="B10">2014</xref>). In addition, the text-image relations described in the literature require closer connections with other aspects of news analysis. In this context, news values (Caple and Bednarek, <xref ref-type="bibr" rid="B21">2016</xref>; Harcup and O&#x00027;Neill, <xref ref-type="bibr" rid="B49">2017</xref>) can play crucial roles in modulating the text-image relations that apply. News values are defined as aspects of events that make them &#x0201C;newsworthy&#x0201D;, such as the involvement of elite personalities (e.g., celebrities, politicians) or the impact of an event and its negative consequences. Bednarek and Caple (<xref ref-type="bibr" rid="B13">2017</xref>) suggest that the concept of news values can equally well be applied to text and image content, which needs to be considered from the perspective of automatic analysis, particularly with respect to interactions with text-image relations.</p>
<p>In this paper, therefore, we introduce a framework for the scalable analysis of multimodal news articles that covers a set of <italic>computable</italic> image-text relations and news values. First, we adapt existing models and taxonomies for image-text relations (Otto et al., <xref ref-type="bibr" rid="B93">2019b</xref>) and expand them for multimodal news analysis, adding further semantic relations inspired by semiotics. Second, we adopt news values (Bednarek and Caple, <xref ref-type="bibr" rid="B13">2017</xref>) from journalism studies, modify them regarding specific <italic>multimodal</italic> aspects, and add distinctive sub-categories. This set of defined relations and news values aims at enabling computational methods to provide more interpretable characterizations of articles, thus allowing for empirical analysis of multimodal news articles at scale. We also develop two additional aspects that respond to the two main stakeholder groups in the news process: publishers and readers. Therefore, we integrate <italic>author intent</italic> in our framework to capture the primary purpose of the news piece and incorporate <italic>subjective interpretation</italic> to allow for variation with respect to image-text relations and news values based on the target audience&#x00027;s background, experience, and demographics.</p>
<p>Overall, the framework presented in this paper covers interdisciplinary topics and contributes to research in multimodal analytics, machine learning, communication sciences and media studies. The principal contributions of this paper can be summarized as follows:</p>
<list list-type="bullet">
<list-item><p>We review and discuss a broad range of literature from semiotics, computational science, and journalism studies concerning news factors and consider methods for combining them.</p></list-item>
<list-item><p>We develop a framework based on multimodal news content, author intent, news values, and subjective interpretation that addresses the entire process of news communication from production to consumption.</p></list-item>
<list-item><p>We critically compare the taxonomy we develop with previously proposed taxonomies, and provide a large number of real-world news examples to motivate and explain different parts of the framework.</p></list-item>
<list-item><p>Finally, we discuss applications and tools that can benefit from the framework. Hopefully, this will inspire new interdisciplinary research directions.</p></list-item>
</list>
<p>The remainder of the paper is structured as follows. In Section 2, we describe the related work from communication and computational science with respect to image-text relations to provide the context for our framework and identify research gaps. In Section 3, we explain our proposed framework with example use cases and discuss their computational aspects and implications. And lastly, in Section 4, we assess the framework in terms of applications and use-cases before concluding the paper with some directions for future research.</p>
</sec>
<sec id="s2">
<title>2. Related work and background</title>
<p>A wide variety of literature exists in multiple fields that have explored the relationship between different modalities and analyzed news from popular media outlets and social media. For current purposes, however, we focus in this section on relevant work for image and text relations in communication science and computational science, as these are equally important for understanding the content and importance of our proposals below. We briefly describe relevant previous work and provide a detailed review of that work motivating our own. In the case of computational approaches, we focus for the present paper on proposed taxonomies and applications of computational methods built using various machine learning models for analysing different forms of digital media.</p>
<sec>
<title>2.1. Communication science</title>
<p>We first describe those approaches that can be generally applied to image-text pairs regardless of the domain and then discuss news values central to analysing and understanding news articles. The discussion follows a chronological order signifying the evolution of treatments of image-text relations over the years.</p>
<sec>
<title>2.1.1. Image-text relations</title>
<p><italic><bold>Early and main taxonomies:</bold></italic> In the pioneering work of Barthes (e.g., Barthes, <xref ref-type="bibr" rid="B9">1977</xref>), three relations are introduced to capture the relative importance of a modality in an image-text pair and to characterize the functions (purposes) served by each modality. These relations are: (1) <italic>Anchorage</italic>: text supporting image in a way that the text guides the viewer in describing and interpreting an image, restricting the commonly assumed &#x0201C;polysemy&#x0201D; of the image; (2) <italic>Illustration</italic>: an image supporting text such that the image recasts in a pictorial form information largely already present in the text; and (3) <italic>Relay</italic>: where text and image exhibit a bidirectional and equal relationship such as complementarity or interdependence. Many semioticians have subsequently explored the integration of meaning in visual and verbal representation in specific domains, although without always specifying a system of relations such as that of Barthes. Most notable here are a study of verbal-visual relations in film documentaries by van Leeuwen (<xref ref-type="bibr" rid="B128">1991</xref>), a study of relating abstract images and text segments by Martin (<xref ref-type="bibr" rid="B75">1994</xref>), an analysis of scientific articles that combine tables, diagrams and text by Lemke (<xref ref-type="bibr" rid="B66">1998</xref>), and detailed analyses of inter-semiotic relations between images and text in advertisements by St&#x000F6;ckl (<xref ref-type="bibr" rid="B115">1997</xref>) and Royce (<xref ref-type="bibr" rid="B106">1998</xref>). A detailed overview of these and several other approaches discussed below is given in Bateman (<xref ref-type="bibr" rid="B10">2014</xref>).</p>
<p>Some later works combine Barthes&#x00027; taxonomy with notable input from linguistics, in particular by adopting accounts of semantic relationships between clauses in grammar (Halliday, <xref ref-type="bibr" rid="B45">1985</xref>; Halliday et al., <xref ref-type="bibr" rid="B46">2014</xref>) as models for mapping image-text pairs into meaningful relations as well. Martinec and Salway (<xref ref-type="bibr" rid="B77">2005</xref>), for example, propose a taxonomy drawing on both Barthes and Halliday that characterizes relations by systematically distinguishing relative importance (status) and semantic relations (logico-semantic). Unique relations are then formed by cross-classifying along the two dimensions to provide detailed classes of image-text pairs combining <italic>status</italic> and <italic>logico-semantic</italic> information. For instance, they divide <italic>equal status</italic> further into <italic>complementary</italic> and <italic>independent</italic> to distinguish when both modalities are necessary to convey the message or when they contain overlapping information, respectively. Under semantic relations, they re-purpose Halliday&#x00027;s main types of <italic>elaboration, extension</italic> and <italic>enhancement</italic> to relate images and texts more generally, adding new information and enrichment of certain attributes (like time or place). In a related approach developed in the educational context, Unsworth (<xref ref-type="bibr" rid="B126">2007</xref>) also extends Martinec and Salway (<xref ref-type="bibr" rid="B77">2005</xref>) <italic>logico-semantic</italic> relations specifically for science textbooks but with different labeling of some relations suited for textbooks, articulating them further for more fine-grained dimensions or sub-categories. For instance, they add a <italic>clarification</italic> relation under <italic>elaboration</italic> suited to diagrams in textbooks, where a diagram provides clarification for the surrounding text. Some other works by Wunderli (<xref ref-type="bibr" rid="B137">1995</xref>) and van Leeuwen (<xref ref-type="bibr" rid="B129">2005</xref>) also specify cases of negative relations such as <italic>contradiction</italic> between image and text, e.g., by creating a contrast to draw attention to certain aspects or elements in each modality.</p>
<p><italic><bold>Other taxonomies:</bold></italic> Another relevant approach drawing on quite different considerations from communication studies is that of Marsh and White (<xref ref-type="bibr" rid="B74">2003</xref>). These authors develop contextual ties between image and text to improve information retrieval and document design. By analyzing articles in several subject areas, they identified 49 relationships grouped into three higher-level categories according to the closeness of the conceptual relationship between image and text. The three main categories are centered around the idea of an image as an <italic>illustration</italic>, and how this illustration relates to the surrounding text. Their categories are: (1) <italic>Minimal</italic>: illustration expressing little relation to the text; (2) <italic>Close</italic>: illustration expressing a close (highly related) relation to the text; and (3) <italic>Transcendental</italic>: illustration that is closely related, but also going beyond the text.</p>
<p><italic><bold>Use-case studies:</bold></italic> Several use-case studies have manually applied taxonomies of text-image relation to real-world problems and data analysis. Mehmet et al. (<xref ref-type="bibr" rid="B78">2014</xref>) work on social media semantics applies Unsworth&#x00027;s taxonomy to understand how messages and online conversations construct and convey meaning. Interestingly, the image-text pairs in this work are not necessarily co-present within single units but can exist on different platforms at different locations and times. In addition, both Wu (<xref ref-type="bibr" rid="B136">2014</xref>) and Nhat and Pha (<xref ref-type="bibr" rid="B89">2019</xref>) apply Unsworth&#x00027;s taxonomy to picture books and English comics for children, respectively. Several broader annotation-based studies then present results in terms of quantitative distributions of the relations found in selected texts (Moya Guijarro, <xref ref-type="bibr" rid="B83">2014</xref>; Nhat and Pha, <xref ref-type="bibr" rid="B89">2019</xref>).</p>
</sec>
<sec>
<title>2.1.2. News values</title>
<p><italic><bold>Seminal works and different perspectives:</bold></italic> The systematic consideration of news values is generally traced back to the seminal study of Galtung and Ruge (<xref ref-type="bibr" rid="B37">1965</xref>) on Scandinavian news discourse. These authors introduce a list of news values divided into two broad categories: culture-free and culture-bound. Culture-free news values are those that are based solely on perception, such as <italic>frequency</italic> and <italic>threshold</italic> (e.g., a large or impactful event), whereas culture-bound news values involve judgements specific to a target culture, such as the presence of <italic>elite persons or nations</italic> and <italic>negativity</italic>, i.e., an event judged negatively by the target culture. Since then, several attempts have been made at redefining and extending these categories with more news values from different perspectives (Bell, <xref ref-type="bibr" rid="B15">1991</xref>; Harcup and O&#x00027;Neill, <xref ref-type="bibr" rid="B48">2001</xref>; Brighton and Foy, <xref ref-type="bibr" rid="B18">2007</xref>). For an event to be constructed as news, then, several factors are at play reflecting various aspects of the news production process. Caple and Bednarek (<xref ref-type="bibr" rid="B21">2016</xref>) usefully differentiate these factors into three categories as follows:</p>
<list list-type="bullet">
<list-item><p><bold>News writing objectives:</bold> general goals associated with news writing, such as <italic>clarity of expression, brevity, color, accuracy</italic> and so on.</p></list-item>
<list-item><p><bold>Selection factors:</bold> any factors or criteria impacting whether or not a story becomes published that are not intrinsic to the presented news item itself, e.g., <italic>commercial pressures, availability of reporters, deadlines</italic> and so on.</p></list-item>
<list-item><p><bold>News values:</bold> the &#x0201C;newsworthy&#x0201D; aspects of actors, happenings and issues as established by a set of recognized values such as <italic>negativity, proximity, timeliness</italic> and so on.</p></list-item>
</list>
<p>Some researchers consider factors in news writing objectives and selection factors to be part of news values as well, but this blurs the line differentiating them from news values that are event-dependent and audience-centric. Furthermore, news values themselves can be produced from different perspectives defined succinctly by Bednarek and Caple (<xref ref-type="bibr" rid="B12">2012</xref>) and Caple and Bednarek (<xref ref-type="bibr" rid="B21">2016</xref>) as below:</p>
<list list-type="bullet">
<list-item><p><bold>Material perspective:</bold> News values that exist in the actual events and people who are reported on in the news, that is, in events in their material reality.</p></list-item>
<list-item><p><bold>Cognitive perspective:</bold> News values that exist in the minds of journalists.</p></list-item>
<list-item><p><bold>Discursive perspective:</bold> News values that are constructed in the discourses involved in the production of news using language and image.</p></list-item>
</list>
<p>Although many researchers define news values from a <italic>cognitive</italic> perspective, Caple (<xref ref-type="bibr" rid="B20">2013</xref>) takes a <italic>discursive</italic> perspective, stating that the other process is highly subjective and may lead to a different interpretation because of every journalist&#x00027;s or news worker&#x00027;s assumptions about news values. For our work and analysis of multimodal news articles, we are then interested particularly in the discursive event-dependent approach (Caple and Bednarek, <xref ref-type="bibr" rid="B21">2016</xref>) . The discursive approach gives a view of how news values are constructed from different modalities and provides crucial insights into how news media package events as &#x0201C;news&#x0201D;. From this discursive perspective, Bednarek and Caple (<xref ref-type="bibr" rid="B12">2012</xref>) define nine news values present in both language and image, which are <italic>Negativity, Proximity, Timeliness, Prominence, Novelty, Consonance, Impact, Superlativeness</italic> and <italic>Aesthetics</italic>. These news values and relevant changes in our framework are discussed in detail in Section 3.3.</p>
<p><italic><bold>Use-case studies:</bold></italic> While these news values are particularly defined for news stories from media channels, they also apply to social media news stories shared across platforms such as Facebook (Bednarek, <xref ref-type="bibr" rid="B11">2016</xref>). Some other works consequently apply news value theory to news articles shared on Facebook (Park and Kaye, <xref ref-type="bibr" rid="B96">2021</xref>), Twitter (Araujo and van der Meer, <xref ref-type="bibr" rid="B6">2020</xref>) and the Russian social media site Vkontakte (Judina and Platonov, <xref ref-type="bibr" rid="B56">2019</xref>). Park and Kaye (<xref ref-type="bibr" rid="B96">2021</xref>) studied the correlation between news values of social significance and deviance with social media users&#x00027; tendency to like, comment and share mainstream news stories on Facebook. Araujo and van der Meer (<xref ref-type="bibr" rid="B6">2020</xref>) conducted a large study on 1.8 million tweets and showed that the news values of social impact, geographical closeness and facticity, among other factors, can explain the intensity of online activities. Harcup and O&#x00027;Neill (<xref ref-type="bibr" rid="B49">2017</xref>) also investigate news values and adapt them for both actual news stories and news on social media by adding values like <italic>shareability, follow-up</italic>, and <italic>entertainment</italic>. Tandoc Jr et al. (<xref ref-type="bibr" rid="B121">2021</xref>) studied the &#x0201C;newsness&#x0201D; of fake news by examining fake articles based on news values, topic and format. Interestingly, this study found that most fake news articles included the news values of timeliness, negativity and prominence. The only difference between real and fake news that was found concerned objectivity: fake news tended to include the author&#x00027;s personal opinion (i.e., not objective) through adjectives and judgements not attributed to any source.</p>
</sec>
</sec>
<sec>
<title>2.2. Computational science</title>
<p><italic><bold>Multimodal machine learning:</bold></italic> In computational sciences, multimodal learning is a well-established research area whereby meaningful information is extracted from two or more modalities and combined to achieve a specific task. The aim is to reduce the semantic gap between more directly accessible features and semantic interpretations (Smeulders et al., <xref ref-type="bibr" rid="B110">2000</xref>) so that the derived semantics can serve as an effective bridge between contributions made in quite different presentation modalities, such as text and image. Bridging the semantic gap is expected to lead to better performance (often involving prediction) in tasks using multiple modalities instead of just one, e.g., sentiment prediction (Poria et al., <xref ref-type="bibr" rid="B99">2017</xref>) from videos using sequences of image frames and audio. With the recent groundbreaking advances in machine learning and the availability of large multimodal datasets, representation learning models can be trained automatically (Ngiam et al., <xref ref-type="bibr" rid="B88">2011</xref>; Chen et al., <xref ref-type="bibr" rid="B26">2020</xref>; Jia et al., <xref ref-type="bibr" rid="B55">2021</xref>; Radford et al., <xref ref-type="bibr" rid="B102">2021</xref>), and their meaningfulness assessed via standard prediction testing benchmarks. A large body of work exists (Baltrusaitis et al., <xref ref-type="bibr" rid="B8">2019</xref>) targeting areas spanning learning algorithms, the fusion of multiple modalities, evaluation metrics and applications that go beyond engineering to domains like medicine and arts. Some noticeable areas and applications with image and text as modalities that have made significant progress in the last decade are image captioning (Kiros et al., <xref ref-type="bibr" rid="B61">2014</xref>; Karpathy and Fei-Fei, <xref ref-type="bibr" rid="B58">2015</xref>; Hossain et al., <xref ref-type="bibr" rid="B53">2019</xref>), cross-modal retrieval (Socher et al., <xref ref-type="bibr" rid="B111">2014</xref>; Xu et al., <xref ref-type="bibr" rid="B141">2015</xref>; Zhen et al., <xref ref-type="bibr" rid="B148">2019</xref>), and text to image generation (Mansimov et al., <xref ref-type="bibr" rid="B73">2016</xref>; Qiao et al., <xref ref-type="bibr" rid="B101">2019</xref>; Ramesh et al., <xref ref-type="bibr" rid="B103">2021</xref>).</p>
<p><italic><bold>Beyond descriptive and literal relations:</bold></italic> In the case of most vision and language model learning and benchmark tasks, the surrounding text is always semantically related at a descriptive and literal level to the image. However, this assumption hardly holds for examples commonly found in online news or advertisements. In such media, the presentation modalities can be weakly linked or abstractly related to one another so as to emotionally engage and influence the reader. Social media offer another example where posts are mostly multimodal, with users uploading billions of photographs<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> with opinions every day on sites such as Facebook and Twitter. Consequently, some researchers (Chen et al., <xref ref-type="bibr" rid="B25">2013</xref>) have investigated Twitter data to distinguish between visually relevant and visually-irrelevant tweets and have built models to predict the two categories. Taking this one step further, Vempala and Preotiuc-Pietro (<xref ref-type="bibr" rid="B130">2019</xref>) take inspiration from Marsh and White (<xref ref-type="bibr" rid="B74">2003</xref>) and have studied image-text tweets for modeling relations that capture which modality adds information (or not) to the other. There is also work on estimating the semantic correlation (Zhang et al., <xref ref-type="bibr" rid="B146">2008</xref>; Xue et al., <xref ref-type="bibr" rid="B142">2015</xref>) and concept level relatedness between modalities (Yanai and Barnard, <xref ref-type="bibr" rid="B143">2005</xref>) (e.g., between local image regions and words), but researchers have only recently (Chinnappa et al., <xref ref-type="bibr" rid="B28">2019</xref>) started combining inter-modal relationships with computational modeling to investigate relations at a deeper level.</p>
<p>Zhang et al. (<xref ref-type="bibr" rid="B145">2018</xref>) focus on investigating non-literal relations between visual and textual persuasion for the automatic analysis of advertisements. They divide the relations into <italic>parallel</italic> vs. <italic>non-parallel</italic>, e.g., when the image and text convey the same message independently (but not necessarily exhibiting literal overlap) and when the meaning is ambiguous or incoherent if the image and text are combined, respectively. However, one difficulty with communication science taxonomies is that their level of detail sometimes makes it difficult to assign a particular class to an image-text pair, especially for an untrained analyst. Recently, to address such issues, both Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>) and Otto et al. (<xref ref-type="bibr" rid="B93">2019b</xref>) have extensively combined image-text taxonomies from semiotics with metrics from computational science and proposed interpretable and computable categories for image-text relations. Inspired from semiotics, Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>) focused on investigating Instagram posts from three perspectives: author intent, contextual or literal meaning overlap, and signified meanings (meaning multiplication) between the image and text. Their contextual and signified relations are inspired by the semiotically-inflected proposals of Kloepfer (<xref ref-type="bibr" rid="B62">1976</xref>) and Marsh and White (<xref ref-type="bibr" rid="B74">2003</xref>), respectively. The authors report noticeable gains in multimodal intent classification when the interpretation of image and text diverges. Alikhani et al. (<xref ref-type="bibr" rid="B4">2020</xref>) introduce a new dataset and study different types of cross-modal coherence relations (such as subjective, story, meta) between image-text pairs. Most recently, Utescher and Zarrie&#x000DF; (<xref ref-type="bibr" rid="B127">2021</xref>) studied types of referential relations between image and text in long text documents such as Wikipedia articles and posed a new research direction for vision and language modeling frameworks. Similarly, Sosea et al. (<xref ref-type="bibr" rid="B113">2021</xref>) show the use case of modeling image-text relationships (<italic>unrelated, similar, complementary</italic>) to improve multimodal disaster tweet classification.</p>
<p><italic><bold>Image-text computational metrics:</bold></italic> Whereas the approaches above generally attempt to provide classification in terms of text-image relations directly, it has also been proposed that more widely applicable and accurate characterizations might be gained by drawing on more robust metrics derived from the text, image, and their inter-relations from which more specific text-image relations can be derived. Henning and Ewerth (<xref ref-type="bibr" rid="B51">2018</xref>), for example, propose two metrics to characterize image-text relations: <italic>cross-modal mutual information</italic> (CMI) and <italic>semantic correlation</italic> (SC). For a given image-text pair, CMI (with values ranging between 0 and 1) measures the number of shared real-world objects, entities and concepts, while SC (ranging between -1 and 1) measures how much meaning is shared between the two modalities. In contrast to the common assumptions (Parekh et al., <xref ref-type="bibr" rid="B95">2021</xref>) made in several works including image captioning and image-text synthesis where image and text are always semantically related, the semantic correlation here can also be negative, which means relations where the overall meaning is incoherent can also be captured. Similarly, Otto et al. (<xref ref-type="bibr" rid="B92">2019a</xref>) propose a further metric called the <italic>abstractness level</italic> (ABS), which measures whether the image is an abstraction of the text or vice versa.</p>
<p>Otto et al. (<xref ref-type="bibr" rid="B93">2019b</xref>) also propose a taxonomy of eight image-text relations derived from Barthes (<xref ref-type="bibr" rid="B9">1977</xref>) and Martinec and Salway (<xref ref-type="bibr" rid="B77">2005</xref>), mapping the three metrics SC, CMI and STATUS to distinctive image-text classes. While CMI and SC measure the information and meaning overlap, STATUS, as introduced in Section 2.1.1, measures the relative importance of an image and text. In addition to the three positively correlated categories termed as <italic>complementary, illustration</italic> and <italic>anchorage</italic>, there are three negatively correlated (<italic>contrasting, bad illustration, bad anchorage</italic>), and two image-text relations with no information overlap (<italic>interdependent, uncorrelated</italic>), which may be deployed in domains such as advertisements and social media to catch the user&#x00027;s attention.</p>
<p><italic><bold>Computational news values and analysis:</bold></italic> Recent efforts toward estimating and extracting news values have focused almost exclusively on the text. Potts et al. (<xref ref-type="bibr" rid="B100">2015</xref>) investigates linguistic techniques such as tagged lemma frequencies, collocation, part-of-speech tags, and semantic tagging to extract news values. The study used a 36-million-word corpus of news reporting on Hurricane Katrina to explore the usefulness of computer-based methods. Bednarek et al. (<xref ref-type="bibr" rid="B14">2021</xref>) conduct a similar empirical study of news articles concerning the Australian national holiday &#x0201C;Australia Day&#x0201D; in the Australian press. In addition, di Buono et al. (<xref ref-type="bibr" rid="B33">2017</xref>) and Piotrkowicz et al. (<xref ref-type="bibr" rid="B97">2017</xref>) focused on developing fully automatic methods for extracting and classifying news values from headline text. di Buono et al. (<xref ref-type="bibr" rid="B33">2017</xref>) relied on news value specific statistical features and Piotrkowicz et al. (<xref ref-type="bibr" rid="B97">2017</xref>) additionally experimented with word embeddings and emotion labels to train machine learning algorithms such as Support Vector Machines (SVN: Cortes and Vapnik, <xref ref-type="bibr" rid="B30">1995</xref>) and Convolutional Neural Networks (CNN: LeCun et al., <xref ref-type="bibr" rid="B65">1998</xref>). Similarly, Belyaeva et al. (<xref ref-type="bibr" rid="B16">2018</xref>) address the problem of estimating news values like <italic>frequency, threshold</italic>, and <italic>proximity</italic> by applying various text mining methods.</p>
<p>For news media research in general, Motta et al. (<xref ref-type="bibr" rid="B82">2020</xref>) describe a framework of newsworthy aspects (called news angles) and a data schema to bridge the gap between formal journalism literature and computational capabilities (named entity recognition, temporal reasoning, opinion mining, and statistical analysis). News angles, in addition to signifying the newsworthiness of an event, also provide a predefined theme/structure to report the event. Similarly, some very recent works address the use of technology toward supporting journalistic practices using semi-automated news discovery tools (Diakopoulos et al., <xref ref-type="bibr" rid="B34">2021</xref>), discuss responsible media technology use for personalized media experiences and news recommendation with automatic media content creation (Trattner et al., <xref ref-type="bibr" rid="B125">2021</xref>), and propose the development of AI-augmented software to support citizen involvement in local journalism (Tessem et al., <xref ref-type="bibr" rid="B122">2022</xref>). All these works position the use of recent breakthroughs in image recognition, natural language processing and other AI technologies for automated or semi-automated news discovery, content creation, and improving user interaction and accessibility. In addition, problems of bias in automated systems, fake news, privacy, news literacy and societal challenges are also discussed and considered important challenges in the ever-changing landscape of technology and information accessibility. However, the role of news angles and news in general in the context of modalities other than text has yet to be explored or discussed.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Proposed framework for multimodal news analysis</title>
<p>In this section, we introduce and discuss our framework for multimodal news analysis drawing motivation from the related work. We bring together different analytical perspectives that have either been used (such as news values) or not yet explored (image-text relations) to analyse news. From computer science perspective, most tasks in multimodal machine learning with respect to news revolve around fake news detection (Giachanou et al., <xref ref-type="bibr" rid="B39">2020</xref>; Singh et al., <xref ref-type="bibr" rid="B109">2021</xref>), news image captioning (Liu et al., <xref ref-type="bibr" rid="B69">2021</xref>), geo-location estimation (Zhou and Luo, <xref ref-type="bibr" rid="B150">2012</xref>; Tahmasebzadeh et al., <xref ref-type="bibr" rid="B119">2022</xref>), and source or popularity prediction (Ramisa, <xref ref-type="bibr" rid="B104">2017</xref>).</p>
<p><xref ref-type="fig" rid="F2">Figure 2</xref> offers an overview of our news analysis framework, covering:</p>
<list list-type="simple">
<list-item><p>(a) The news production aspects of author intent and news values signifying communicative purpose and newsworthy elements in the news article, respectively,</p></list-item>
<list-item><p>(b) Multimodality aspects of different image-text relations that signify the interplay of image and text and their use in news, and</p></list-item>
<list-item><p>(c) The news consumption aspect of subjective interpretation, which explains how user modeling and the interpretation connect to the other two aspects given in (a) and (b).</p></list-item>
</list>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Proposed framework with different perspectives and classes for multimodal news article analysis.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0002.tif"/>
</fig>
<p>With our framework, we intend to enable new tasks on large-scale data that can aid in discourse analysis and be of interest to researchers in computational science, communication and media studies. Recent work on news values (Caple, <xref ref-type="bibr" rid="B20">2013</xref>; Caple et al., <xref ref-type="bibr" rid="B23">2020</xref>) has progressed from text-only to image and text both, which ultimately begs the question of how these modalities need to be related to one another (image-text relations). However, this aspect has not yet been considered in combination with news values. Therefore, a framework that can provide a view of all these aspects is going to be far more revealing in discourse analysis than any approach restricted to just one aspect. In the next sections, this is demonstrated by the discussion of several examples and case studies.</p>
<p>Next, we first summarize the individual parts of the framework and then explain them in detail with illustrative examples. In this section, we dig deeper into the individual aspects of (a) and (b) while providing detailed examples, and briefly touch upon the interplay between these different aspects and the news consumption aspect of subjective interpretation. The four main areas of the framework are the following:</p>
<list list-type="bullet">
<list-item><p><bold>Author intent</bold> captures the author&#x00027;s primary purpose (communicative goal) behind writing the article. A news article has some facts and information, may have other people&#x00027;s opinions, and the author&#x00027;s own opinion and analysis on an issue. For example, the authors&#x00027; intent could be just to inform readers about an event or persuade them by making them align with an idea in the article. In the field of rhetorical criticism and genre analysis, intent corresponds to the communicative purpose of a text (Swales, <xref ref-type="bibr" rid="B117">1990</xref>; Miller, <xref ref-type="bibr" rid="B81">1994</xref>). The intent classes in our framework are drawn specifically from recent work on intent taxonomy in Instagram posts (Kruk et al., <xref ref-type="bibr" rid="B64">2019</xref>); broader classification systems are discussed, for example, by Martin and Rose (<xref ref-type="bibr" rid="B76">2008</xref>).</p></list-item>
<list-item><p><bold>Cross-modal relations</bold> capture the interplay between image and text, capturing their use in news articles. In other words, this primarily indicates how image and text are tied together for the story in an article. Various image-text relation taxonomies from semiotics were reviewed above. However, even though these often have detailed definitions and divisions of classes, their categories (Marsh and White, <xref ref-type="bibr" rid="B74">2003</xref>; Martinec and Salway, <xref ref-type="bibr" rid="B77">2005</xref>; Unsworth, <xref ref-type="bibr" rid="B126">2007</xref>) sometimes overlap and are hard to interpret. Conversely, categories proposed in computational science are often distinctive and informative, but are either proposed for a particular domain [like Instagram posts from Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>)] or mainly to capture the role of modalities (Zhang et al., <xref ref-type="bibr" rid="B145">2018</xref>; Vempala and Preotiuc-Pietro, <xref ref-type="bibr" rid="B130">2019</xref>) and generalized image-text relationships (Otto et al., <xref ref-type="bibr" rid="B93">2019b</xref>). In our framework, we both re-purpose existing relations with clearer definitions of classes and scope and introduce new relations specifically for news.</p></list-item>
<list-item><p><bold>News values</bold> capture the news-centric attributes defined in the journalism literature that we introduced in Section 2.1.2. As explained earlier, news values provide crucial insights into how some events are packaged as &#x0201C;news&#x0201D; by news media. We mainly adopt the news values from Caple et al. (<xref ref-type="bibr" rid="B23">2020</xref>) and, when needed, extend them with further categorization to make them more inclusive and discrete. In particular, we also review computational approaches that can be used to detect news values from image and text.</p></list-item>
<list-item><p><bold>Subjective interpretation</bold> capture the user-centric changes in the aspects discussed above. As any article is written for a target audience, the demographics and their characteristics certainly have some influence on the use of imagery and language in a news piece. For this reason, we provide an additional dimension of classification that allows us to link cross-modal relations and news values with certain codified aspects of the user&#x00027;s background. We term this the &#x02018;subjective interpretation&#x00027;. Although we do not delve into this aspect in detail, it allows for the possibility to do comparative studies given analytical aspects above and user characteristics.</p></list-item>
</list>
<p>In the next sections, we elaborate on each of these aspects in the same order as introduced above, propose additions and modifications, explain the difference to existing literature, and provide real-world examples to motivate the use of the proposed categories.</p>
<sec>
<title>3.1. Author intent</title>
<p>As briefly described before, <italic>Author Intent</italic> represents the author&#x00027;s main purpose behind the news article. Recently, Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>) propose an intent taxonomy for Instagram posts centered around the presentation of self (Hogan, <xref ref-type="bibr" rid="B52">2010</xref>; Mahoney et al., <xref ref-type="bibr" rid="B72">2016</xref>). They propose eight categories, which are: <italic>advocative, promotive, exhibitionist, expressive, informative, entertainment, provocative:discrimination</italic>, and <italic>provocative:controversial</italic>. While independent journalists&#x00027; and authors&#x00027; intent can be placed under some of these categories, we are more interested in intent arising as a combination of author, the idea of the article and the publisher. Thus, the categories that strictly represent presentation of self (<italic>exhibitionist, expressive</italic>) are in general less significant intents for news. Categories under provocation, on the other hand, are related to hate-speech and propaganda, which is certainly an interesting area on its own right in computational analytics.</p>
<p>Building on this, we propose a more structured two-level taxonomy for capturing author intent. <italic>Informative, Entertaining</italic> and <italic>Persuasive</italic> are the top-level categories. <italic>Informative</italic> means that the author&#x00027;s main purpose is to inform the reader about news via unbiased reporting of facts and information. The articles under this category are descriptive and expository focused on details, facts, eye-witness accounts and opinions/arguments of others (such as public figures or people in general) and linguistically distinguishable as aligning along some of the distinct dimensions of registerial variation described in detail by Biber (<xref ref-type="bibr" rid="B17">1988</xref>). <italic>Entertaining</italic> articles, in contrast, refer to satirical or spoof pieces that often take a humorous jab at organizations or people that hold power. Finally, <italic>Persuasive</italic> refers to an author supporting or refuting a policy, person, organization or broadly an idea in the article <italic>via</italic> personal opinions, interpretations, point of view or judgements. Not all the intent categories are mutually exclusive, and so an article can have more than one author intent. These diverse properties naturally suggest that corresponding computational language models differentiate among the text types and registers represented.</p>
<p>Developing the account further, <italic>Informative</italic> is then divided into two categories:</p>
<list list-type="bullet">
<list-item><p><italic>Report facts</italic>: When the article mainly focuses on what (description of the event/news), when (time, date) and where (location) of the news/event. For example, a news article about a hurricane with the description of wind speed, amount of rain, time and duration, location of landfall and the details of destruction and damage it caused. A real news factual example is shown in <xref ref-type="fig" rid="F3">Figure 3B</xref>, where the article is about a new chip, its components and the benefits. The articles under this category can also be an introduction of a problem and possible solutions to the readers. For example, a news about water contamination in the region and the solutions recommended by health authorities.</p></list-item>
<list-item><p><italic>Report opinions</italic>: In this case, articles contain references to other people&#x00027;s opinions, arguments and comments, which sometimes in text refers to quoted snippets. For example, a news article with references to people questioning the government about their ill-preparedness for the hurricane. In <xref ref-type="fig" rid="F3">Figure 3A</xref>, another news example can be seen which is scientific and factual, but also consists of comments and opinions of another person, e.g., a professor, given in quotes. The examples show that not all quoted snippets are opinions, as the first quoted snippet is a summary of the research rather than an opinion.</p></list-item>
</list>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><italic>Author Intent</italic> examples for <italic>Report Facts</italic> and <italic>Opinions</italic>. Left: the portions of the text (quotes) that lead to the assignment of <italic>Opinions</italic> are highlighted (yellow). Both articles include factual information, and thus the category <italic>Facts</italic> is assigned. <bold>(A)</bold> Facts and Opinions. <bold>(B)</bold> Facts.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0003.tif"/>
</fig>
<p><italic>Entertaining</italic> news articles rely heavily on irony and humor, and it is almost impossible to mistake them as serious news. This is different from fake news, which intentionally misleads people into believing that the story is real (Golbeck et al., <xref ref-type="bibr" rid="B41">2018</xref>). The articles here mimic real news but still cue the reader that they should not be taken seriously. Examples of this category are news articles from satire websites like <italic><ext-link ext-link-type="uri" xlink:href="https://thespoof.com">thespoof.com</ext-link>, <ext-link ext-link-type="uri" xlink:href="https://thecivilian.co.nz">thecivilian.co.nz</ext-link></italic> and <italic><ext-link ext-link-type="uri" xlink:href="https://thedailymash.co.uk/">thedailymash.co.uk</ext-link></italic>. Two examples from such websites are shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, where highlighted text shows informal language that is not used in actual news articles. Interestingly, the images used in these examples are neutral and do not reflect the entertaining intent of the article.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><italic>Author Intent</italic> examples: Entertaining and Opinions. The portions of the text (entertaining, satirical parts, or quotes) that lead to the assignment of corresponding categories are highlighted (yellow). <bold>(A)</bold> Entertaining. <bold>(B)</bold> Entertaining &#x00026; Opinions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0004.tif"/>
</fig>
<p><italic>Persuasive</italic> articles are written with the purpose to persuade and convince the readers to consider or accept author&#x00027;s opinion, position and point of view. Opinion pieces and articles fall under this category. It is also a possibility that an article holds no position and so only falls under the <italic>informative</italic> category. If not, <italic>persuasive</italic> is divided into two opposing sub-categories: <italic>Promote</italic> and <italic>Demote</italic>, defined with respect to the core theme or idea of the article. <italic>Promote</italic> refers to when an author aligns with or supports a policy, person or organization and, in so doing, aims to persuade the reader to do the same. <italic>Demote</italic> is the converse of this. We show two examples of such news articles in <xref ref-type="fig" rid="F5">Figure 5</xref>, where article (a) is optimistic about the Internet for opinion journalism, while article (b) is about press culture in the US and the lack of gender diversity. Additionally, use of harsh language or tonality, attacks toward someone, presence of controversy and propaganda are all possibilities within such articles. These establish important targets for further research but are not yet explicitly categorized under our taxonomy.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><italic>Author Intent</italic> examples for Persuasive&#x02013;Promotive and Persuasive&#x02013;Demotive. The portions of the text that lead to the assignment of corresponding categories are highlighted (yellow). <bold>(A)</bold> Promote. <bold>(B)</bold> Demote.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0005.tif"/>
</fig>
</sec>
<sec>
<title>3.2. Cross-modal relations</title>
<p>Our proposed taxonomy of cross-modal relations (last column) is inspired from multiple taxonomies. <xref ref-type="table" rid="T1">Table 1</xref> shows a comparison of taxonomies proposed in the computational science literature that draw on semiotics. The top three rows in the table represent relations where the image and text are coherent and belong together, and relations in the next three rows are where the image and text differ in meaning and are incoherent. While the top three relations can be represented in both Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>) and Vempala and Preotiuc-Pietro (<xref ref-type="bibr" rid="B130">2019</xref>), the differences between relations based on modality&#x00027;s importance cannot be established uniquely. Moreover, Vempala and Preotiuc-Pietro (<xref ref-type="bibr" rid="B130">2019</xref>) cannot represent any of the relations where either image and text are incoherent or related on an abstract level (row 7 in <xref ref-type="table" rid="T1">Table 1</xref>) without sharing any information. Whereas, Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>) can represent the incoherent relations, all ambiguously under one combination of its contextual and semiotic classes. As our relation taxonomy is inspired from Otto et al. (<xref ref-type="bibr" rid="B93">2019b</xref>), all the relations in the leftmost column can be uniquely represented by our taxonomy (rightmost column). In addition, we have extensions of CMI and SC to explicitly model the presence of additional information and fine-grained relations, respectively.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>A comparison of taxonomies proposed in computational sciences to build multimodal models.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left" colspan="2"><bold>Otto et al. (</bold><xref ref-type="bibr" rid="B93"><bold>2019b</bold></xref><bold>)</bold></th>
<th valign="top" align="left"><bold>Vempala and Preotiuc-Pietro (<xref ref-type="bibr" rid="B130">2019</xref>)</bold></th>
<th valign="top" align="left" colspan="2"><bold>Kruk et al. (</bold><xref ref-type="bibr" rid="B64"><bold>2019</bold></xref><bold>)</bold></th>
<th valign="top" align="left" colspan="2"><bold>Our taxonomy</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="2">Anchorage</td>
<td valign="top" align="left" rowspan="4">Text is represented<break/> Image adds</td>
<td valign="top" align="left" rowspan="8">Contextual :<break/> Semiotic :<break/>Transcendent<break/> Parallel/Additive</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">Partial - High</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">Image extends text / Both</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">High-positive</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Image</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Image</td>
</tr>
 <tr>
<td valign="top" align="left" colspan="2">Illustration</td>
<td valign="top" align="left" rowspan="8">Text is represented<break/> Image does not add</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">Partial - High</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">Text extends image / Both</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">High-positive</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Text</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Text</td>
</tr>
 <tr>
<td valign="top" align="left" colspan="2">Complementary</td>
<td valign="top" align="left" rowspan="4">Contextual :<break/> Semiotic :<break/>Close<break/> Parallel</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">High</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">Both extend</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">High-positive</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
</tr> <tr>
<td valign="top" align="left" colspan="2">Bad Anchorage</td>
<td valign="top" align="left" rowspan="16">Not represented</td>
<td valign="top" align="left" rowspan="12">Contextual :<break/> Semiotic :<break/>Minimal<break/> Divergent</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">Partial - High</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">Image extends text/both</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">&#x0003C; 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">Low-positive/negative</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Image</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Image</td>
</tr>
 <tr>
<td valign="top" align="left" colspan="2">Bad Illustration</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">Partial&#x02013;high</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">Text extends image/both</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">&#x0003C; 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">Low-positive/negative</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Text</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Text</td>
</tr>
 <tr>
<td valign="top" align="left" colspan="2">Contrasting</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">High</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">Both extend</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">&#x0003C; 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">Low-positive/negative</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
</tr>
 <tr>
<td valign="top" align="left" colspan="2">Interdependent</td>
<td valign="top" align="left" rowspan="4">Contextual :<break/> Semiotic :<break/>Minimal<break/> Additive</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">No Overlap</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">= 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">No extension</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">&#x0003E; 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">High-positive</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
</tr> <tr>
<td valign="top" align="left" colspan="2">Uncorrelated</td>
<td valign="top" align="left" rowspan="4">Text is not represented<break/> Image does not add</td>
<td valign="top" align="left" rowspan="4">Contextual :<break/> Semiotic :<break/>Minimal<break/> Divergent</td>
<td valign="top" align="left">OO</td>
<td valign="top" align="left">No Overlap</td>
</tr>
 <tr>
<td valign="top" align="left">CMI</td>
<td valign="top" align="left">= 0</td>
<td valign="top" align="left">OE</td>
<td valign="top" align="left">No extension</td>
</tr>
<tr>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">= 0</td>
<td valign="top" align="left">SC</td>
<td valign="top" align="left">Zero correlation</td>
</tr>
<tr>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
<td valign="top" align="left">STATUS</td>
<td valign="top" align="left">Equal</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>OO, Objective Overlap, OE, Objective Extension, CMI, Cross-modal Mutual Information, SC, Semantic Correlation.</p>
</table-wrap-foot>
</table-wrap>
<p>We suggest additions and modifications for object- and semantic-level relations along with news-centric attributes for richer news analysis. We define these relations building on measures established in previous work. However, while relations in other taxonomies can be broken down (sometimes not distinctively) at conceptual, semantic and relevance levels, our framework is more fine-grained. Furthermore, our framework can be applied to any image-text pair, while at the same time offering relevant relations and attributes specifically for news analysis.</p>
<p>There are then two types of high-level cross-modal relations in the framework, Objective relations and Semantic relations, which are explained in detail in following two Sections 3.2.1 and 3.2.2.</p>
<sec>
<title>3.2.1. Objective relations</title>
<p><bold>Objective relations</bold> measure the amount of shared real-world entities or concepts between the two modalities (image and text). Information for objective relations refers to real-world objects and named entities, such as persons, locations and landmarks, actions, events, scenery, organizations, products. We use the idea of imageability (Kastner et al., <xref ref-type="bibr" rid="B59">2020</xref>) here to differentiate between abstract and concrete entities/concepts. Under objective relations, entities or concepts have either concrete iconic depictions or give a clear mental image. In addition, information can also refer to concepts like scenes and actions, which have a concrete mental image. Rest of the named entities and concepts which have no concrete depiction are considered in semantic relations. For example, the term religion can have various possible visuals (symbol, church/temple), but is still an abstract concept. In contrast, a specific religious ceremony like baptism has a far clearer &#x02013; i.e., more reliably associated &#x02013; mental image and concrete depiction. This is specifically true for certain named entities in text such as large numbers, season and time, for which there is no clear depiction. Previous works mix concrete and abstract concepts under one relation, such as &#x0201C;text is represented&#x0201D; relation in Vempala and Preotiuc-Pietro (<xref ref-type="bibr" rid="B130">2019</xref>), &#x0201C;contextual&#x0201D; relation in Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>), and Otto et al. (<xref ref-type="bibr" rid="B93">2019b</xref>) referring to entities as &#x0201C;main objects&#x0201D; in foreground of an image. From the computational perspective, object (Deng et al., <xref ref-type="bibr" rid="B31">2009</xref>; Thomee et al., <xref ref-type="bibr" rid="B124">2016</xref>), scene (Xiao et al., <xref ref-type="bibr" rid="B138">2010</xref>; Zhou et al., <xref ref-type="bibr" rid="B149">2018</xref>), event (Xiong et al., <xref ref-type="bibr" rid="B139">2015</xref>; M&#x000FC;ller-Budack et al., <xref ref-type="bibr" rid="B86">2021a</xref>), and action recognition (Heilbron et al., <xref ref-type="bibr" rid="B50">2015</xref>; Gu et al., <xref ref-type="bibr" rid="B43">2018</xref>) models and their classes provide the list of all concepts and entities that have clear depictions and are thus identifiable in an image.</p>
<p>It is further divided into two classes: <italic>Objective Overlap</italic> that signifies mutual information between two modalities (similar to CMI), and <italic>Objective Extension</italic>, a new aspect, that indicates the modality contributing more entities and concepts.</p>
<p>In practice, to model these relations and to identify the correct sub-categories, we need some specific ranges of values: (i) the set of entities or concepts in an image (<italic>N</italic><sub><italic>I</italic></sub>), (ii) the set of entities or concepts in the text (<italic>N</italic><sub><italic>T</italic></sub>), and, (iii) the set of entities or concepts that overlap between image and text (<italic>N</italic><sub><italic>O</italic></sub>). Examples of image-text pairs from news articles are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. They include image-headlines and image-captions pairs, respectively, to explain the categories. Both categories of objective relations should also be beneficial and informative in the case of informal news sources such as Twitter, where text has a character limit resulting in much information being packed into a few sentences.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>News articles annotated with the sub-categories of <italic>Objective relations</italic>. Each sample includes the identified textual (<bold>T</bold>) and visual (<bold>I</bold>) concepts, entities, etc., along with the assigned sub-categories of <italic>Objective Overlap (OO)</italic> and <italic>Objective Extension (OE)</italic>. &#x0002A;Used ControlNet (Zhang and Agrawala, <xref ref-type="bibr" rid="B147">2023</xref>) to re-create images because of licensing issues. <bold>(A)</bold> OO: Uncorrelated, OE: No Extension. <bold>(B)</bold> OO: Partial Ovelap, OE: Both Extend. <bold>(C)</bold> OO: High Overlap, OE: Image Extends Text. <bold>(D)</bold> OO: Partial Overlap, OE: Text Extends Image. <bold>(E)</bold> OO: High Overlap, OE: Both Extend.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0006.tif"/>
</fig>
<p><italic><bold>Objective overlap (OO)</bold></italic> measures the overlap between a multimodal news article&#x00027;s image and text components. The purpose is to establish material and conceptual links between one or multiple images and text in a news article. It measures the intersection of information from two presentation modalities (image and text) and can be an actual number or a quantitative value signifying the extent of overlap (as in CMI). In our framework, this is divided into three further sub-categories:</p>
<list list-type="bullet">
<list-item><p><italic>Uncorrelated</italic> stands for cases where none of the defined material or conceptual ties co-exist in the textual and visual content (as when CMI &#x0003D; 0). According to the introduced notation, this is when <italic>N</italic><sub><italic>O</italic></sub> &#x0003D; &#x02205;. Example (a) in <xref ref-type="fig" rid="F6">Figure 6</xref> shows an uncorrelated pair, where pills in the image have no shared word or phrase in the text.</p></list-item>
<list-item><p><italic>Partial overlap</italic> is when there is an overlap between the defined concepts in image and text, but not every aspect exists in both modalities (similar to CMI &#x0007E;0.5), i.e., <italic>N</italic><sub><italic>O</italic></sub> &#x0003D; (<italic>N</italic><sub><italic>I</italic></sub>&#x02229;<italic>N</italic><sub><italic>T</italic></sub>) and 0 &#x0003C; |<italic>N</italic><sub><italic>O</italic></sub>| &#x0003C; <italic>min</italic>(|<italic>N</italic><sub><italic>I</italic></sub>|, |<italic>N</italic><sub><italic>T</italic></sub>|). Examples (b) and (d) in <xref ref-type="fig" rid="F6">Figure 6</xref> show this overlap, where in both modalities additional concepts are present besides the overlapping concepts.</p></list-item>
<list-item><p><italic>High overlap</italic> is when all the entities or concepts from either modality are present in <italic>N</italic><sub><italic>O</italic></sub>. With our notation, it can be computed as <italic>N</italic><sub><italic>O</italic></sub> &#x0003D; (<italic>N</italic><sub><italic>I</italic></sub>&#x02229;<italic>I</italic><sub><italic>T</italic></sub>) &#x0003D; <italic>min</italic>(|<italic>N</italic><sub><italic>I</italic></sub>|, |<italic>N</italic><sub><italic>T</italic></sub>|)), where <italic>N</italic><sub><italic>O</italic></sub> &#x0003D; <italic>N</italic><sub><italic>I</italic></sub> when high overlap is because of image&#x00027;s contribution and <italic>N</italic><sub><italic>O</italic></sub> &#x0003D; <italic>N</italic><sub><italic>T</italic></sub> when high overlap is because of text&#x00027;s contribution. Example (c) shows high overlap with all aspects (majesty, guy, say) in text seen in the image. In the third case, when <italic>N</italic><sub><italic>O</italic></sub> &#x0003D; <italic>N</italic><sub><italic>I</italic></sub> &#x0003D; <italic>N</italic><sub><italic>T</italic></sub>, there is close to a one-to-one correspondence between concepts in image and text (CMI &#x0007E;1). A prominent example of this is given by image captioning datasets, where the text exclusively focuses on image content. <xref ref-type="fig" rid="F6">Figure 6E</xref> shows the high overlap between concepts and words like kitchen, tofu, plastic, boilers and workers. Although an high overlap is unlikely when comparing an news text article and an image, there is a possibility of high overlap between the image-caption (or parts of the text) and image-headline pairs &#x02013; when, for instance, an image acts as strong evidence (reference) for the text, e.g., news about a crime scene or a natural disaster.</p></list-item>
</list>
<p><italic><bold>Objective extension (OE)</bold></italic> indicates which modality extends another based on the amount of information contributed by each modality. The relation is inspired from <italic>Extension</italic> described by Martinec and Salway (<xref ref-type="bibr" rid="B77">2005</xref>), where the additional information is new but related information (such as date, country, person&#x00027;s name). Extension is only possible when partial or high objective overlap exists between two modalities. The sub-categories for this relation are described below with their corresponding notations:</p>
<list list-type="bullet">
<list-item><p><italic>No extension</italic>: This is possible in two scenarios, (a) <italic>OO</italic> is <italic>uncorrelated</italic> (<italic>N</italic><sub><italic>O</italic></sub> &#x0003D; &#x02205;), see example (a) in <xref ref-type="fig" rid="F6">Figure 6</xref>, and (b) <italic>OO</italic> is <italic>high overlap</italic> with an one-to-one correspondence between shared concepts and no new related concepts. A typical example for this case is image-caption pairs, with the caption describing the objects and scene as depicted in the image.</p></list-item>
<list-item><p><italic>Text extends image</italic> and <italic>image extends text</italic>: Both classes are possible when <italic>OO</italic> is <italic>partial overlap</italic> or <italic>high overlap</italic> and when one of the modalities has more components. That is |<italic>N</italic><sub><italic>T</italic></sub>|&#x0003E;|<italic>N</italic><sub><italic>I</italic></sub>| for <italic>Text Extends image</italic> and |<italic>N</italic><sub><italic>I</italic></sub>|&#x0003E;|<italic>N</italic><sub><italic>T</italic></sub>| for <italic>Image extends text</italic>. The examples (c) and (d) in <xref ref-type="fig" rid="F6">Figure 6</xref> show these categories. Elaborating further on example (d), text provides details that the appeal is to Bolivians to refrain from violence, while the image depicts the UN chief speaking on a podium.</p></list-item>
<list-item><p><italic>Both extend</italic>: This case is possible in two scenarios, a) <italic>OO</italic> is <italic>partial overlap</italic> and |<italic>N</italic><sub><italic>I</italic></sub>| &#x0003D; |<italic>N</italic><sub><italic>T</italic></sub>|, i.e. that there are still components in both modalities that provide more details than each on its own, and b) <italic>OO</italic> is <italic>high overlap</italic> and there are still additional entities/concepts in both. Example (e) in <xref ref-type="fig" rid="F6">Figure 6</xref> illustrates the latter case, where the image shows a dimly lit kitchen with partially clothed workers working and the text provides details such as Indonesia and burning.</p></list-item>
</list>
</sec>
<sec>
<title>3.2.2. Semantic relations</title>
<p><bold>Semantic relations</bold> cover relationships on the meaning level and overall message of the image-text pair that can go beyond descriptive <italic>Objective relations</italic>. The relations under this category refer to how interpretations and meaning can be constructed from a image-text pair by combining information from both modalities. Typically, understanding the meaning of an image-text pair requires knowledge and a user&#x00027;s ability to identify entities or concepts and how to link them. These relations can be inferred based on publicly accessible knowledge sources (e.g., encyclopedia, knowledge graphs, etc.). The exploitation of such sources would be required for computational approaches to identify semantic relations. <italic>Semantic relations</italic> include three different aspects of capturing relationships on the meaning level: <italic>Semantic Correlation</italic> captures whether image and text belong together, <italic>STATUS</italic> signifies modality importance in terms of context and information, and <italic>Modification</italic> indicates whether one modality changes or strengthens certain aspects provided in the other modality. Please note that there might be a semantic relation between image and text, while there is no <italic>objective relation</italic>. For example, <xref ref-type="fig" rid="F6">Figure 6A</xref> shares no entities or concepts, but 750,000 people in text is abstractly linked to large number of pills in the image.</p>
<p><italic><bold>Semantic correlation</bold></italic> measures the correlation between image and text. It can refer to concrete entities and contextual, interpretative correlations, that is, how much meaning is shared between modalities according to the SC metric from Otto et al. (<xref ref-type="bibr" rid="B93">2019b</xref>). In addition to Otto et al.&#x00027;s (<xref ref-type="bibr" rid="B93">2019b</xref>) SC metric, we replace <italic>Positive correlation</italic> with two sub-classes <italic>Low-positive correlation</italic> and <italic>High-positive correlation</italic> to capture more fine-grained relations. The <italic>Objective relations</italic> only tell us about the objective overlap between image and text, and it remains a possibility that, even with high overlap, the items related can be opposite or contradictory due to context differences at a semantic level. To capture semantic overlap, we include <italic>semantic correlation</italic> in the framework to measure meaning overlap regardless of the shared information.</p>
<p>As a result, we divide <italic>semantic correlation</italic> into four sub-categories:</p>
<list list-type="bullet">
<list-item><p><italic>Zero correlation</italic>: If the two modalities do not have any material, conceptual, abstract or contextual semantic links, then the image-text pairs are considered as unrelated. We provide a sample in <xref ref-type="fig" rid="F7">Figure 7C</xref> where the given image and the text in the news article do not have any implicit or explicit meaning relations.</p></list-item>
<list-item><p><italic>Negative correlation</italic>: This is the case when information in two modalities has different, opposing or contradictory contexts that disturb the overall meaning of the image-text pair. For example, in <xref ref-type="fig" rid="F7">Figures 7A</xref>, <xref ref-type="fig" rid="F7">B</xref>, two samples of negative semantic correlation are shown where the images contradict the news or information provided in the text. On the left, we have a celebratory image of people with a Tokyo 2020 banner, while the news is about alleged labor violations and resulting deaths. In the middle, we again have an image of football players celebrating that resembles the action of playing the match, whereas the text in news article is about an abandoned match.</p></list-item>
<list-item><p><italic>Low-positive correlation</italic>: In this case, image and text are partially linked by sharing one or a few aspects but not the entire message. Typically, images and text are weakly linked in news articles and have low-positive semantic correlation when images in the news article serve no purpose besides being a placeholder or a stock photograph. For example, in <xref ref-type="fig" rid="F7">Figure 7D</xref>, the image of <italic>Emmanuel Macron</italic> with <italic>EU</italic> countries&#x00027; flags beside him is weakly linked since the actual news is about <italic>EU budget</italic>.</p></list-item>
<list-item><p><italic>High-positive correlation</italic>: This is the case when image and text belong together through sharing information and the entire contextual message. However, sharing information, i.e., partial or high <italic>objective overlap</italic>, is optional. Particularly topics such as politics, sports and breaking news have images that can be strongly linked to the news text or the headline. Moreover, satire news can have abstract (aesthetic or artistic) images with zero or very low objective overlap but with high semantic overlap with the text. For example, in <xref ref-type="fig" rid="F7">Figure 7E</xref>, the highlighted text describes the people, image scene, and also provides an additional context of the place and event. Similarly, the example in <xref ref-type="fig" rid="F7">Figure 7F</xref> mentions the effects of climate change on future generations while the image shows the melting ice blocks as a result of climate change.</p></list-item>
</list>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><italic>Semantic Correlation</italic> (SC) and <italic>STATUS</italic> examples. Negative correlation examples have highlighted portion in text (orange) that is opposite in meaning to the image. Low and high-positive correlation examples have highlighted portion in text (yellow) that is aligned in meaning to the image. <bold>(A)</bold> SC: Negative correlation. <bold>(B)</bold> SC: Negative correlation. <bold>(C)</bold> SC: Zero correlation. <bold>(D)</bold> SC: Low-positive correlation. <bold>(E)</bold> SC: High-positive correlation. <bold>(F)</bold> SC: High positive correlation. <bold>(G)</bold> STATUS: Equal. <bold>(H)</bold> STATUS: Image. <bold>(I)</bold> STATUS: Text.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0007.tif"/>
</fig>
<p>The <italic><bold>STATUS</bold></italic> relation indicates the relative importance of image and text with respect to the overall message. Although the text is usually the main modality in a news article [although St&#x000F6;ckl et al. (<xref ref-type="bibr" rid="B116">2020</xref>) describe an emerging genre where this is not the case, identifying the relative importance of a modality (image or text)] can be significant to analyse the importance of images against headlines, captions, short text pairs (social media) or the whole news body text. Therefore, we use the <italic>STATUS</italic> relation as defined by Otto et al. (<xref ref-type="bibr" rid="B92">2019a</xref>) for modeling image-text relative importance. Relative importance is in terms of meaning and the modality central to the overall message of an image-text pair, in contrast to <italic>extension</italic>, which operates in terms of individual concepts. The <italic>STATUS</italic> relation describes the hierarchical relation between image and text to signify whether they are equally important for the overall message or one modality is subordinate (provides additional information while the main message exists in the other modality) to the other.</p>
<p><italic>STATUS</italic> has three sub-categories as follows:</p>
<list list-type="bullet">
<list-item><p><italic>Equal</italic>: Both parts are equally important for the overall message. There is a possibility that image and text are tightly linked here on the level of <italic>objective overlap</italic> (<italic>High Overlap</italic>) and <italic>semantic correlation</italic> (<italic>High-positive Correlation</italic>). Both can also add information to each other that works together for the overall meaning and explanation. A sample in <xref ref-type="fig" rid="F7">Figure 7G</xref> has an <italic>Equal STATUS</italic> relationship. The text mentions the effects of climate change while the image shows how the coral reefs are dying. The information in each modality is unique and equally important with regard to the overall message of <italic>climate change</italic>.</p></list-item>
<list-item><p><italic>Image</italic>: Here, the image is the center of attention and provides additional information or context not present in the text. The text acts as a caption or explicitly references the image, indicating that the image holds vital information regarding the message. It is still possible that text adds context here and fixes the intended meaning of the image-text pair. The sample in <xref ref-type="fig" rid="F7">Figure 7H</xref> is such an example where the table in the image contains more information since the text describes the overall message without any details.</p></list-item>
<list-item><p><italic>Text</italic>: In this case, the image plays a minor role or is exchangeable with another image. The text provides the main context and other unique and important details relevant to the main message. This is the usual case of news articles, where text holds the main information, and one or more images act as illustrations of a few things in the text. The sample in <xref ref-type="fig" rid="F7">Figure 7I</xref> has the main content in its text, and the image serves as a visualization of the message. The image can be replaced with another one, and the image-text pair&#x00027;s overall message would not change. Thus, the given image-text pair sample has <italic>Text</italic> as <italic>STATUS</italic>.</p></list-item>
</list>
<p><italic><bold>Modification</bold></italic> indicates if one modality changes a particular aspect or meaning in the other. Such aspects can be <italic>sentiment</italic>, emotion, quantities, or broad modifiers (adjectives) present in image or text. Unlike <italic>semantic correlation, modification</italic> is directional in nature because it changes or enhances the meaning or interpretation of the other modality&#x00027;s content. It is relevant for news in particular, where the image&#x00027;s role (like amplifying the news) can be explained <italic>via modification</italic> in addition to news values. The relation is motivated from Bateman&#x00027;s (<xref ref-type="bibr" rid="B10">2014</xref>) explanation of Kloepfer&#x00027;s (<xref ref-type="bibr" rid="B62">1976</xref>) taxonomy that covers additive image-text relationships. Although Kloepfer defines modification and amplification as two different relations under additive relations, the taxonomy and terminology are still debatable as amplifying any aspect also means changing it (hence modifying it).</p>
<p>We define the <italic>modification</italic> relation as a parent relation under which several types of modification can be distinctly sub-categorized as follows:</p>
<list list-type="bullet">
<list-item><p><italic>Amplification</italic>: One modality changes the conveyed message by making it, i.e., its meaning or certain aspects in the other modality, stronger (&#x0201C;louder&#x0201D;). Here, the meaning and categories of certain aspects stay the same, but the intensity is increased in the same direction. The bottom image in <xref ref-type="fig" rid="F8">Figure 8A</xref> amplifies the central message in the text, i.e. of exposed landscape due to melting ice contrary to the small iceberg (broken from glacier), which future generations will not get to see and experience. In contrast, the message in the same text is not modified by the top image in <xref ref-type="fig" rid="F8">Figure 8A</xref> since it does not show the drastic changes (drastic melting of snow) as a result of climate change. Similarly, the top image in <xref ref-type="fig" rid="F8">Figure 8B</xref> amplifies the phrase &#x0201C;fans storm&#x0201D; by showing the large number of people gathered for an event mentioned in the text.</p></list-item>
<list-item><p><bold>(A-C)</bold> <italic>Quality modification</italic>: The <italic>Quality modification</italic> changes the conveyed message, its meaning or certain aspects in any way or direction (not the same direction as in Amplify). There is a chance that image and text have a <italic>negative semantic correlation</italic> in such a case. For example, the bottom images in <xref ref-type="fig" rid="F8">Figures 8B</xref>, <xref ref-type="fig" rid="F8">C</xref> are opposite in meaning to &#x0201C;fans storm&#x0201D; and &#x0201C;snowstorms&#x0201D; showing an empty stadium and lush green Texas farms respectively. These examples are specific instances of <italic>Quality modification</italic> due to the stark contrast of focused aspects in the image content when compared to the core theme of the news text.</p></list-item>
</list>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><italic>Modification</italic> examples. Three news headlines each with two different images showing the change in modification relations. &#x0002A;Used ControlNet (Zhang and Agrawala, <xref ref-type="bibr" rid="B147">2023</xref>) to re-create images because of licensing issues.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0008.tif"/>
</fig>
<p>If neither <italic>Amplification</italic> nor <italic>Quality</italic> modification is present between image and text, then there is no modification. The top images in <xref ref-type="fig" rid="F8">Figures 8A</xref>, <xref ref-type="fig" rid="F8">C</xref> show icy mountains and the arctic, with both images possibly representing melting ice; however, they do not modify the message by either amplifying or changing its quality.</p>
</sec>
</sec>
<sec>
<title>3.3. News values</title>
<p>In this section, we modify and expand the news values defined by Caple (<xref ref-type="bibr" rid="B20">2013</xref>), Caple and Bednarek (<xref ref-type="bibr" rid="B21">2016</xref>), and Caple et al. (<xref ref-type="bibr" rid="B23">2020</xref>), as they define news values for both image and text in a discursive and content-specific manner, which is interesting and relevant from a multimodal news analytics perspective. We derive new sub-classes for some of the news values to make them more concrete, discrete or meaningful. In our framework, any news value may have the same or different sub-categories for the image and text, which adds another dimension for comparing both modalities. Having said that, where needed both modalities can be used to estimate a single multimodal news value with respect to the news article (explained in each news value). We discuss each news value and its realizations in each modality, define scope and list visual and verbal cues essential to identify them in image and text. In previous work, news values are indicated by their presence or absence, which we also follow for most news values. We leave out the news value <italic>Aesthetics</italic> for two reasons: first, it is only defined for the image, and second, it is a subjective-only and not concretely defined news value. This makes categorizing an image into non-aesthetic vs. aesthetic from only the perspective content a difficult, and probably inappropriate task. In the following sub-sections, we list the news values (see <xref ref-type="fig" rid="F9">Figure 9</xref>) in our framework and their descriptions from the content-only perspective, which can differ from the user-dependent news values covered in <italic>subjective interpretation</italic> in Section 3.4.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>News values from content and user perspective.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0009.tif"/>
</fig>
<p>We provide five examples of news articles that are categorized with the respective <italic>Proximity, Timeliness</italic>, and <italic>Eliteness</italic> categories in <xref ref-type="fig" rid="F10">Figure 10</xref>; five examples of news articles for which <italic>Sentiment, Superlativeness</italic>, and <italic>Impact</italic> news values are identified in <xref ref-type="fig" rid="F11">Figure 11</xref>; and five examples of news articles that are categorized with <italic>Personalization, Novelty</italic>, and <italic>Consonance</italic> news values in <xref ref-type="fig" rid="F12">Figure 12</xref>. Each sample is matched with a corresponding news value and its sub-category and marked with the modality (I: Image, T: Text) that provides the essential cues.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p><italic>Proximity, Timeliness</italic>, and <italic>Eliteness</italic> news values examples. All news values except domain <italic>proximity</italic> are described as Category, (textual or visual cues), <italic>[Modalities: I (image), T (Text)]</italic>. The segments in text that mention important words or phrases indicating time, persons, locations, etc. are highlighted (yellow). <bold>(A)</bold> Eliteness : Yes, (UN chief, President), [T], Geo. proximity : National, (Bolivia), [I,T], Cul. proximity : Culture-free, [I, T], Dom. proximity : (world affairs, politics, crisis), [I,T], Timeliness : Ongoing, (Thursday, Amidst), [T]. <bold>(B)</bold> Eliteness : No [I, T], Geo. proximity : International, (Venezuela,World), [T], Cul. proximity : Culture-free, [I,T], Dom. proximity : (migration, refugee-crisis) [T], Timeliness : Ongoing, recent, (2020, recent years) [T]. <bold>(C)</bold> Eliteness : No, [I,T], Geo. proximity : National, (UK) [T], Cul. proximity : Culture-bound, (festive, christmas), [T], Dom. proximity : (shopping, christmas, sales), [I,T], Timeliness : Recent, Ongoing, (2019, 2020), [T]. <bold>(D)</bold> Eliteness - Yes, (Queen, David Beckham), [T], Geo. proximity : National, (London, UK), [I,T], Cul. proximity : Culture-free, [I, T], Dom. proximity : (olympic games, sports), [I,T], Timeliness : Old, (Six years, 2012), [T]. <bold>(E)</bold> Eliteness - Yes, (James webb, NASA), [I,T], Geo. proximity : International, (French Guiana, NASA), [I,T], Cul. proximity : Culture-free, [I, T], Dom. proximity : (space, telescopes, science), [I,T], Timeliness : Future, (expected, 18 December), [T].</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0010.tif"/>
</fig>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p><italic>Sentiment, Superlativeness</italic> and <italic>Impact</italic> news values examples. All news values are described as <bold>Category</bold>, (textual or visual cues), <italic>[Modalities: I (image), T (Text)]</italic>. The segments in text that mention important words or phrases that indicate quantifiers, adjectives, mentions of disasters, etc. are highlighted (yellow). *Used ControlNet (Zhang and Agrawala, <xref ref-type="bibr" rid="B147">2023</xref>) to re-create images because of licensing issues. <bold>(A)</bold> Sentiment : Negative, (Angry, unwarranted), [T], Superlativeness : Yes, (Thousands, tractors) [I,T], Impact : Medium, (protest, blocked roads in..) [I,T]. <bold>(B)</bold> Sentiment : Positive, (Win, celebration, happy), [I,T], Superlativeness : Yes, (multiple players, happy), [I], Impact : Medium, (win, celebration), [I,T]. <bold>(C)</bold> Sentiment : Negative, (losses, devastating), [I,T], Superlativeness : Yes, (heaps of trash, billion euros) [I,T], Impact : High, (damage, most devastating), [I,T]. <bold>(D)</bold> Sentiment : Positive (winning a case), [I,T], Superlativeness : No, [I,T], Impact : Low, (wins case, person vs. publisher), [T]. <bold>(E)</bold> Sentiment : Positive, (rises, safest), [T], Superlativeness : Yes, (biggest, safest), [T], Impact : High, (historic achievement), [T].</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0011.tif"/>
</fig>
<fig id="F12" position="float">
<label>Figure 12</label>
<caption><p><italic>Personalisation, Novelty</italic> and <italic>Consonance</italic> news values examples. All news values are described as <bold>Category</bold>, (textual or visual cues), <italic>[Modalities: I (Image), T (Text)]</italic>. The segments in text that mention important words or phrases that indicate new or reoccurring events, mentions of disasters, etc. are highlighted (yellow). *Used ControlNet (Zhang and Agrawala, <xref ref-type="bibr" rid="B147">2023</xref>) to re-create images because of licensing issues. <bold>(A)</bold> Personalisation : Yes, (Personal account), [T], Novelty : No, [I,T], Consonance : No, [I,T]. <bold>(B)</bold> Personalisation : No, [I,T], Novelty : Yes, (new invention) [T], Consonance : No, [I,T]. <bold>(C)</bold> Personalisation : Yes, (Focus on protestor), [I], Novelty : Yes, (contrast b/w police vs protestor) [I], Consonance : No, [I,T]. <bold>(D)</bold> Personalisation : No, [I,T], Novelty : Yes, (record number, police vs immigrants) [I,T], Consonance : Yes, (flooded, drugs, image usage) [I,T]. <bold>(E)</bold> Personalisation : No, [I,T], Novelty : Yes, (precedent), [T], Consonance : Yes, (News is about stereotypes), [T].</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0012.tif"/>
</fig>
<sec>
<title>3.3.1. Proximity</title>
<p><italic>Proximity</italic> is defined by Caple and Bednarek (<xref ref-type="bibr" rid="B21">2016</xref>) as a <italic>geographical or cultural nearness of an event or issue</italic>. Based on this given definition, we divide it into three different components so that each can be analyzed separately. In addition to the geographical and cultural aspect of an event, we propose adding domain <italic>proximity</italic>, which indicates the topics covered in a news article.</p>
<list list-type="bullet">
<list-item><p><italic>Geographical proximity</italic> is the distance between the event and the expected (target) audience. It can be seen as the scope of the event or the news, such that the news piece could directly or indirectly <italic>impact</italic> people who come under this scope. In addition to the individual locations mentioned in the article, we further sub-categorize this into four values which reflect the scope of the news article: <italic>local</italic> (a city), <italic>regional</italic> (multiple cities in the state/province), <italic>national</italic> (country or multiple states/cities in the country) and <italic>global</italic> (multiple countries). While it is relatively easy to identify locations and scope from text, geo-location estimation from images (M&#x000FC;ller-Budack et al., <xref ref-type="bibr" rid="B85">2018</xref>; Theiner et al., <xref ref-type="bibr" rid="B123">2022</xref>) is a harder task that gives us the probable location(s) based on iconic landmarks, cultural symbols, artifacts, various other objects, including prominent people etc. Therefore, geographical <italic>proximity</italic> from images is limited to one (most probable) location unless there is text in the image, which helps to identify the scope. In <xref ref-type="fig" rid="F10">Figure 10D</xref>, the news value of <italic>Proximity</italic> is categorized as <italic>Geographical-National</italic> since the text include entities such as <italic>London, UK</italic> and the image shows a well-known landmark in <italic>London</italic>.</p></list-item>
<list-item><p><italic>Cultural proximity</italic> is the reference to a particular religion, tradition, celebration, festival or ritual. Even though any group of people with some commonalities can constitute a culture, we adhere to the mentioned aspects only. Any other aspects that come under the broad definition of culture can be captured in domain <italic>proximity</italic> as topics in a news article. When words or phrases and visual cues in text and image respectively refer to one of these aspects, they are categorized as <italic>culture-bound</italic>; otherwise, they are <italic>culture-free</italic>. In example <xref ref-type="fig" rid="F10">Figure 10C</xref>, the text includes the word &#x0201C;Christmas&#x0201D;, which makes the post <italic>culture-bound</italic>.</p></list-item>
<list-item><p><italic>Domain proximity</italic> refers to concepts or topics discussed in the news article discernible through image and text. In this case, concepts could be different in image and text, and a combination of both gives the list of all concepts. For example, broad topics such as business, technology, science, and sports are valid domains under which news articles are often organized in conventional newspapers and news websites. In addition, there can be other fine-grained topics based on events (such as Olympics) and themes in the news article. In all the examples in <xref ref-type="fig" rid="F10">Figure 10</xref>, <italic>Domain Proximity</italic> lists topics extracted (manually) from both image and text.</p></list-item>
</list>
</sec>
<sec>
<title>3.3.2. Timeliness</title>
<p><italic>Timeliness</italic> is defined as the relevance of an event in terms of time of occurrence. We divide <italic>timeliness</italic> into four categories based on the date of publication and reference to exact time, date, or season in the news article: <italic>old</italic> (two years or before), <italic>recent</italic> (last two years), <italic>ongoing</italic> and <italic>future</italic>. As explained in Caple and Bednarek (<xref ref-type="bibr" rid="B21">2016</xref>), explicit verbal cues other than a date can be words like <italic>today, yesterday</italic> and verb tenses like <italic>have been trying</italic> in addition to the context around these phrases. For images, in the absence of text on the image, it is very challenging (M&#x000FC;ller et al., <xref ref-type="bibr" rid="B84">2017</xref>) to determine the date from a single image.</p>
</sec>
<sec>
<title>3.3.3. Eliteness</title>
<p><italic>Eliteness</italic> refers to the involvement of known individuals, organizations, or nations in an event. We consider only certain people and organizations under <italic>eliteness</italic> and not nations as defined in Caple (<xref ref-type="bibr" rid="B20">2013</xref>). Although, it can be true that certain developed countries get more coverage and be considered as elite in news value. However, considering certain countries as elite and others as not elite is reinforcing or introducing bias into the models that are (or will be) specifically designed to detect such attributes. As geographical <italic>proximity</italic> provides locations and scope of the event, the aspect of <italic>eliteness</italic> according to countries can be established by publisher-specific analyses based on <italic>proximity</italic>. People like celebrities, politicians, elite professionals (e.g., authors, scientists), and organizations that are well known to operate as for-profit or non-profit are considered of high status. From the text, named entities that have Wikipedia entries can be used to depict <italic>eliteness</italic>, and in the absence of all of these as <italic>no eliteness</italic>. In images, the same can be recognized via faces, context (microphones, cameras, police) and elements (such as a uniform) associated with an elite profession. As discussed in Caple (<xref ref-type="bibr" rid="B20">2013</xref>), <italic>eliteness</italic> can be identified via technical aspects such as camera angle with respect to the individual, where a low angle (viewed as looking up) can be hypothesized as suggesting the high status of the person in the image, although such attributions must always be treated with caution as they generally depend on other supporting accompanying cues. The example in <xref ref-type="fig" rid="F10">Figure 10D</xref> has a news value <italic>Eliteness</italic> based on the following known persons mentioned in text: <italic>David Beckham, James Bond, Mr. Bean</italic> or <italic>the Queen</italic>.</p>
</sec>
<sec>
<title>3.3.4. Sentiment</title>
<p>Sentiment refers to the negative and positive evaluation of events etc. in addition to language and vocabulary used in the news article. Although it is mentioned as <italic>negativity</italic> in journalism studies, we prefer the broader term <italic>Sentiment</italic> as both negative (e.g., disasters) and positive (e.g., winning a World Cup) events are covered as stories in the news. In the text, this can refer to certain events (like disasters, accidents, new invention, a peace treaty), and usage of positive vocabulary (tone/language) (e.g., happy, saved, beautiful) or negative vocabulary (e.g., terrible, bad, worried). In images, people with negative or positive emotions and effects of certain events mentioned before (e.g., football players celebrating a win, the aftermath of natural disasters) correspond to negative or positive <italic>sentiment</italic>. Similar to <italic>eliteness</italic>, technical aspects like high camera angle (viewed as looking down upon) with respect to the person of interest in the image may depict negative <italic>sentiment</italic>, as discussed in Caple (<xref ref-type="bibr" rid="B20">2013</xref>). The news value <italic>Sentiment</italic> is categorized into three classes, i.e., <italic>positive, negative</italic>, and <italic>neutral</italic>. For example, the sample in <xref ref-type="fig" rid="F11">Figure 11E</xref> has a <italic>positive sentiment</italic> as the number of endangered species have increased. Whereas examples (a) and (c) have <italic>negative sentiment</italic>, because of the words like protest, angry, damage and devastating, and image in (c) showing damage and destruction.</p>
</sec>
<sec>
<title>3.3.5. Superlativeness</title>
<p><italic>Superlativeness</italic> refers to intensified or maximized aspects in the news article. The intensified aspects can be both in the positive (e.g., millions of dollars in profit) and negative (e.g., sharpest drop in the stock) direction, such that it is independent of the context. In the text, maximized aspects of the event can be identified via quantifiers (large numbers), intensifiers that emphasize amount, scale or size, and superlatives, all of which can be identified using the grammatical structure of text documents. It can also be established with metaphors and similes (like a raging river) that reflect the intensity or <italic>impact</italic> of the event. For images, <italic>superlativeness</italic> can be identified through repetition of key elements (e.g., soldiers marching, cars stuck in traffic), depiction of extreme emotions (anger, shock, fear) and placement of contrasting elements (e.g., a big new airplane standing next to a smaller one). <italic>Superlativeness</italic> can be classified into two categories, <italic>yes</italic> (presence of multiple identifiers) and <italic>no</italic> (absence of identifiers). For example, sample (e) in <xref ref-type="fig" rid="F11">Figure 11</xref> has <italic>yes</italic> as <italic>superlativess</italic> news value where the news article mentions it with words such as &#x0201C;safest&#x0201D; or &#x0201C;biggest&#x0201D;.</p>
</sec>
<sec>
<title>3.3.6. Impact</title>
<p><italic>Impact</italic> refers to positive or negative effects and consequences of the event covered in the news article. Although the <italic>impact</italic> is sometimes closely related to <italic>superlativeness, impact</italic> focuses more on the consequences of an event. In the text, <italic>impact</italic> can be identified via references to effects on individuals (e.g., unwarranted environment regulations on farmers, see <xref ref-type="fig" rid="F11">Figure 11A</xref>), important and relevant consequences (e.g., the aftermath of a natural disaster, celebrations after a World Cup, see <xref ref-type="fig" rid="F11">Figure 11C</xref>) and significance of the event (e.g., historic, momentous day). For images, similar to <italic>superlativeness</italic>, it can be identified through repetition (depicts the extent of <italic>impact</italic>) of certain elements and negative or positive (e.g., damage to public property, Olympic gold medal win) imagery of effects or the event. Like a few other news values, technical aspects in images like blurring because of excessive camera movement can also depict the event&#x00027;s <italic>impact</italic> (e.g., war, hostage situation) when viewed in combination with the text that describes the event. <italic>Impact</italic> is categorized into three classes, <italic>low</italic> [impacts number of individual(s)], <italic>medium</italic> (that affects a group of people in a community) and <italic>high</italic> (that affects larger regions or countries).</p>
</sec>
<sec>
<title>3.3.7. Personalization</title>
<p><italic>Personalization</italic> refers to personal aspects in the news article. <italic>Personalization</italic> is constructed when an abstract issue is made more personal via references to ordinary people, their emotions, experiences (e.g., eyewitness accounts) and stories in the news article. In the text, it can be identified by the context (e.g., detailed description of person&#x00027;s suffering or jubilation) and presence of a unknown person (if the mentioned person does not have an entry in Wikipedia). In images, visual cues like a close-up shot of an individual, focus on emotional response, facial expression, and singling out a person in a group of people (protester and police). <italic>Personalization</italic> is divided into <italic>yes</italic> and <italic>no</italic> categories. The sample in <xref ref-type="fig" rid="F12">Figure 12C</xref> includes a <italic>Personalization</italic> news value because the image focuses on an ordinary person as a demonstrator. Text in sample (a) describes a personal account and situation of a person affected by high rent prices, realizing <italic>personalization</italic>.</p>
</sec>
<sec>
<title>3.3.8. Novelty</title>
<p><italic>Novelty</italic> refers to the new or unexpected aspects of the event. In the text, the aspect of <italic>new</italic> can be identified with evaluative words like <italic>different, strange</italic>, comparisons (with the past) that indicate unexpectedness like <italic>never seen anything</italic> and references to extreme emotions like shock and surprise. Unusual news or happenings that are categorized as bizarre news in media also construct <italic>novelty</italic>. For images, similar to superlativeness, extreme facial expressions indicating people being shocked or surprised and comparison between two or more elements to create a stark contrast in an image can also depict <italic>novelty</italic>. <italic>Novelty</italic> is divided into <italic>yes</italic> and <italic>no</italic> categories. The sample in <xref ref-type="fig" rid="F12">Figure 12B</xref> has a <italic>Novelty</italic> news value because the text mentions the phrases &#x0201C;new invention&#x0201D;. Similarly the image in example <xref ref-type="fig" rid="F12">Figure 12C</xref> realizes <italic>novelty</italic> by showing a contrast scene with one demonstrator surrounded by riot police.</p>
</sec>
<sec>
<title>3.3.9. Consonance</title>
<p><italic>Consonance</italic> refers to the expectedness (unlike <italic>novelty</italic>) of the event or presence of certain stereotypical aspects, which can be with respect to a person, community, organization or country in the news article. Just like <italic>novelty</italic>, expectedness in text can be identified with evaluative phrases like <italic>famed for, notorious for</italic> and comparisons like <italic>yet another, once again</italic>. Moreover, images can show aspects that fit a particular stereotype and combine with text to construct <italic>consonance</italic>. <italic>Consonance</italic> is divided into <italic>yes</italic> and <italic>no</italic> categories. The sample <xref ref-type="fig" rid="F12">Figure 12D</xref> realizes <italic>consonance</italic> news value by associating certain traits (drugs, flooded the border) to immigrants that reinforce stereotypes.</p>
</sec>
</sec>
<sec>
<title>3.4. Subjective interpretation</title>
<p>Recipients perceives information in the news differently depending on their background, knowledge and experience. This source of variation is far from arbitrary, however. With multiple presentation modalities such as image and text, different kinds of objective and semantic relations can be established between modalities from the available content and information, as explained in detail above. But, in addition, the differing sets of traits, such as <italic>age, gender, education, reader&#x00027;s location, language, domains of interest, religion, click history, etc</italic>., attributable to each recipient (reader) induce further changes in relations when contrasted with the content-based perspective. This change in objective and semantic relations, as well as some news values, is called <italic>Subjective Interpretation</italic>. One use case for this kind of analysis is reader-specific studies based on demographics and how the information composed using presentation modalities varies from the intended purpose. From a media studies perspective, given the cross-modal relations extracted from a news publisher&#x00027;s corpus and its reader base characteristics, one can analyse links between the content used for news vs the reader demographics. For instance, if certain image-text compositions are more prevalent for one type of readers&#x00027; demographics or characteristics than the other.</p>
<p>The analysis can be further refined based on certain news values that are highly reader-centric and exhibit a different value from the reader&#x00027;s perspective. All news values except <italic>timeliness</italic> and <italic>superlativeness</italic> are constructed in the view of the target audience. For example, consider a news article about Germany winning a football match against France. The content would be mostly about how well Germany played, won the game, and construct the story with certain statistics about the game. Even though the article has a positive connotation, a French team supporter living in either Germany or France would probably perceive it differently from German football fans or even perceive it negatively. Here, we describe how news values may be varied from a recipient&#x00027;s perspective:</p>
<list list-type="bullet">
<list-item><p><italic>Proximity</italic>: Based on the locations and places identified from the news article, geographical <italic>proximity</italic> for a recipient is categorized as <italic>low</italic> (near), <italic>medium</italic>, and <italic>high</italic> (far) based on geo-location and nationality of the reader. The sub-category for a reader can be computed using a threshold. For cultural <italic>proximity, culture-bound</italic>, and <italic>culture-free</italic> for a reader is based on its traits and cultural aspects (if any) and context in the news article. Lastly, domain <italic>proximity</italic> is based on the overlap of topics in the article and the reader&#x00027;s interests.</p></list-item>
<list-item><p><italic>Sentiment</italic>: Each recipient perceives the news and content in the article differently that can vary from the intended purpose of the article. In addition to polarity of the <italic>sentiment</italic> (<italic>positive, negative and neutral</italic>), it is further categorized into eight emotions (Mikels et al., <xref ref-type="bibr" rid="B80">2005</xref>) as <italic>fear, sadness, disgust, anger, amusement, contentment, awe</italic>, and <italic>excitement</italic>.</p></list-item>
<list-item><p><italic>Eliteness, impact, personalization, novelty, consonance</italic>: All other news values can be different based on the recipients&#x00027; traits and are sub-categorized as <italic>yes</italic> and <italic>no</italic>. For instance, a prominent individual or an organization mentioned in the article might not be relevant and not considered elite by the reader because of a different location and interests. <italic>Impact</italic>, in this case, refers to whether the reader is affected by the event or its effects and consequences. As <italic>personalization</italic> is the &#x0201C;human&#x0201D; aspect in the article, it refers to whether a recipient relates to the personal aspects or not (e.g., photo of a person crying in front of a destroyed building after Earthquake). Lastly, whether a reader considers the events as expected or unexpected is covered in <italic>consonance</italic> and <italic>novelty</italic>, respectively.</p></list-item>
</list>
<p>From the computational perspective, we have three kinds of resources here, first is cross-modal relations or news values based on content, second are reader characteristics, and lastly the reader-centric news values. Treating two as inputs and one as the output in a predictive task, we can consider these problems under user modeling to predict readers&#x00027; preferences, readers&#x00027; characteristics or improvise information dissemination (language and image usage).</p>
</sec>
<sec>
<title>3.5. Interplay between different aspects</title>
<p>In this section, we discuss some illustrative articles from different news media channels. We particularly pay attention to author intent, cross-modal relations and news values in order to show the necessity of adequately reflecting the sometimes quite complex interplay between these different aspects. In the previous sections, we presented examples of each aspect with pairs of image-short text (headline or caption). Here, we extend the focus to include an article&#x00027;s components, including the headline, image captions and some parts of the main text. Structural analysis of this kind has also been performed and shown specifically for news values by Caple and Bednarek (<xref ref-type="bibr" rid="B22">2017</xref>), enabling the comparison and interaction of news values across an entire article. Now, we show how it is also beneficial to perform analyses that incorporate our broader set of perspectives that go beyond news values alone.</p>
<p>In <xref ref-type="fig" rid="F13">Figure 13</xref>, we present four articles for comparison, all related to COVID-19 and taken from top-read daily news media channels in the United Kingdom. In the top row (<xref ref-type="fig" rid="F13">Figure 13A</xref>), we see two articles from <italic><ext-link ext-link-type="uri" xlink:href="https://dailymail.co.uk/">dailymail.co.uk</ext-link></italic> (a) and (b); bottom left an article from the <italic><ext-link ext-link-type="uri" xlink:href="https://guardian.com/">guardian.com</ext-link></italic> (<xref ref-type="fig" rid="F13">Figure 13C</xref>); and bottom right an article from the <italic><ext-link ext-link-type="uri" xlink:href="https://mirror.co.uk/">mirror.co.uk</ext-link></italic> (<xref ref-type="fig" rid="F13">Figure 13D</xref>). The articles in each column are drawn from the same dates in March and September 2022, respectively. The left column articles are about the &#x0201C;<italic>rise in weekly COVID-19 cases by 50%&#x0201D;</italic>, and the right column articles are about &#x0201C;<italic>WHO&#x00027;s announcement or statement about the end of the COVID-19 pandemic&#x0201D;</italic>. In the following, we discuss differences and similarities across the three media channels drawing on the perspectives defined in our analytic framework.</p>
<fig id="F13" position="float">
<label>Figure 13</label>
<caption><p>Examples of news articles on COVID-19 with multiple images, captions and parts from the main text. Left and right column news are from 10<sup><italic>th</italic></sup> March and 15<sup><italic>th</italic></sup> September 2022, respectively. Yellow highlighted text - refers to entities/concepts and news values clues. Green highlighted text - statements being informative (reporting facts or opinions of experts) and orange highlighted text - statements being persuasive (author&#x00027;s own analysis/opinions). <bold>(A, B)</bold> Daily mail. <bold>(C)</bold> The guardian. <bold>(D)</bold> Daily mirror.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1125533-g0013.tif"/>
</fig>
<p>Starting with the two <italic>Daily Mail</italic> articles, a common observation is the use of images and text to provide both a supporting and a contrasting viewpoint with respect to the headline. In the left article,<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> several graph images (three with black background, one shown in <xref ref-type="fig" rid="F13">Figure 13A</xref>) are used without direct references in the text and no captions (<italic>Uncorrelated OO and no extension</italic>). Here, the text comes to the same conclusion (<italic>Equal STATUS</italic> and there is <italic>no modification</italic>) evident in the image (<italic>high-positive SC</italic>). These three images and the supporting texts followed by the author&#x00027;s own opinions (sometimes referring to theories behind the increase without citing sources) promote (<italic>persuasive author intent</italic>) the main theme of &#x0201C;rise of infections&#x0201D;. Conversely, three more graph images are included in the article to provide a different viewpoint on &#x0201C;covid being less deadly than the flu&#x0201D; and &#x0201C;covid impact on rise in the waitlist of routine treatments&#x0201D;. One of these images (shown in <xref ref-type="fig" rid="F13">Figure 13A</xref>) has a caption which refers (<italic>Image STATUS</italic>) to the image (has a common entity &#x0201C;NHS&#x0201D;, <italic>partial OO</italic>), summarizes the graph (<italic>high-positive SC</italic>), and adds more entities (<italic>Text extends Image</italic>). The image, on the other hand, amplifies (<italic>amplification</italic>) the &#x0201C;one in nine people waiting&#x0201D; messaging by showing the difference in people on the waitlist now vs. pre-covid times. Interestingly, this image caption is not discussed and referenced anywhere in the main article but instead in a subordinated text fragment embedded within the article, which reflects the complex structure used by the <italic>Daily Mail</italic>. In addition, another image (not shown in <xref ref-type="fig" rid="F13">Figure 13A</xref>), which also has a caption, exhibits similar image-text relations, presenting in detail opinions from the author and quoted excerpts from cited experts (<italic>informative, reports facts, and opinions</italic>). In summary, the majority of the article has image-text pairs that exhibit different relations as in the beginning of the article, and promotes (<italic>persuasive author intent</italic>) the &#x0201C;living with covid&#x0201D; perspective, in contrast to the headline (<italic>demotes, persuasive author intent</italic>). This demonstrates clearly how important it is that any analysis is capable both of differentiating finer structural units between which relations can hold and of allowing those relations to be divergent.</p>
<p>The article also constructs distinct news values of <italic>superlativeness</italic> (large numbers, rise of cases and waitlist), <italic>impact</italic> (infection rise, deaths, routine treatments), <italic>proximity</italic> (national-UK and international-European), <italic>timeliness</italic> (recent, ongoing, and old), <italic>sentiment</italic> (positive and negative), and <italic>eliteness</italic> (NHS, other organizations, experts). As with the image-text relations, here there is variation over the extent of the article of news values as well as between the image and text. For instance, the main article text constructs the news value of <italic>timeliness</italic> as <italic>ongoing</italic>, whereas images explicitly construct <italic>recent and old</italic> by showing years as far back as 2007. Similarly, <italic>superlativeness</italic> (from thousands of cases to millions of people on the waiting list), <italic>impact</italic> and <italic>sentiment</italic> (from <italic>negative</italic>-rise in cases to <italic>positive</italic>-covid less deadly than flu to again <italic>negative</italic>-rise in wait list) are constructed in different ways across the extent of the article.</p>
<p>In a similar manner, the article<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> on the right, first, primarily reports (<italic>informative, facts and opinions</italic>) regarding the statement given by WHO and, later, provides a contrasting viewpoint on the long period lockdowns in China and their impact on residents. Two out of three images (one shown on the left of <xref ref-type="fig" rid="F13">Figure 13B</xref>) are stock photographs (<italic>partial OO</italic> and <italic>low-positive SC</italic> and <italic>no modification</italic>) with captions providing the main message (<italic>Text STATUS</italic>) and additional entities (<italic>Text extends Image</italic>). A notable difference from the other <italic>Daily Mail</italic> article is the absence of the connection to the rest of the main article text, except for the news value of <italic>eliteness</italic> via images. This could be because the first article is written by a health and science editor for an in-depth analysis. The third image (shown on right in <xref ref-type="fig" rid="F13">Figure 13B</xref>) depicts the situation in the caption (<italic>high OO</italic> and <italic>high-positive SC</italic>) and exhibits the same relations as the previous two image-caption pairs. In addition, the caption here exhibits the different news value of <italic>timeliness</italic> (June) in comparison to the lockdown news (from September) from China. Furthermore, similar to the other article, the later text on the lockdown in China consists of the author&#x00027;s claims and opinions without concrete source references or quotes from experts. In summary, several distinct news values are constructed &#x0201C;simultaneously&#x0201D; in the article: <italic>eliteness</italic> via both image and text (Director-General and WHO), <italic>novelty</italic> (lowest figures) via text, <italic>international proximity</italic> (U.S., England, China, World) in text, positive <italic>high impact</italic> (lowest figures, end of pandemic) and negative <italic>medium impact</italic> (resident&#x00027;s condition in a Chinese city), and lastly <italic>sentiment</italic> (from <italic>positive</italic>-lowest cases worldwide to <italic>negative</italic>-hunger, forced quarantines in China).</p>
<p>In contrast to the <italic>Daily Mail</italic>&#x00027;s article on the left, <italic>The Guardian</italic><xref ref-type="fn" rid="fn0004"><sup>4</sup></xref> article (also written by a science editor) covers the news on the rise in cases and its impact on people aged 55 and above. The article focuses on facts from government reports and figures, scientific reports and opinions of five different experts (<italic>informative, facts, and opinions</italic>). Moreover, the structure of the text (paragraphs instead of sentences) and the use of images is quite different from the <italic>Daily Mail</italic>&#x00027;s article. For example, the left image is not referenced in the main text, but the image-caption pair is relevant for the construction of news values of <italic>personalization</italic> and <italic>impact</italic>, resulting in <italic>high-positive SC but no modification</italic> with the caption and some of the text snippets in the main text. In addition, the image-caption pair has a <italic>partial OO</italic> and <italic>both extend</italic> in terms of new information, with the text as the main message <italic>(Text STATUS)</italic> and constructing news value of <italic>local proximity</italic> (Dedworth, Berkshire), <italic>timeliness</italic> as ongoing and future, and <italic>eliteness</italic> (NHS, NHS personnel). Conversely, the second graph image from <italic>gov.uk</italic> acts as a placeholder image (in this case, of COVID-19 patients in hospitals) and is not referenced in the main text; it is then only weakly linked (<italic>low-positive SC</italic>) to show the rise in cases and hospitalizations (<italic>quality modification</italic>, not matching the figures). Such image-text relations consequently suggest a sub-optimal combination of image and text in terms of information dissemination. In addition, several other news values are constructed in the rest of the text: <italic>sentiment</italic> as mostly <italic>negative, superlativeness</italic> with large numbers, <italic>high negative impact</italic> for the nation&#x00027;s elderly, <italic>local and national proximity</italic> (Berkshire and England), <italic>timeliness</italic> as <italic>recent, ongoing, and future</italic>, and <italic>eliteness</italic> (institutes, experts, government, and NHS).</p>
<p>Lastly, the <italic>Daily Mirror</italic><xref ref-type="fn" rid="fn0005"><sup>5</sup></xref> article shows the most contrasting use of images with respect to the article and the main text. In three (two shown in <xref ref-type="fig" rid="F13">Figure 13D</xref>) out of five images in the article, the news value of <italic>personalization</italic> (focus on laypersons) is constructed. The article&#x00027;s main text stands in contrast to the <italic>Daily Mail</italic> article in that most of the focus is on disseminating the WHO&#x00027;s announcement and information from documents shared by the WHO (<italic>informative, facts, and opinions</italic>). Moreover, the rest of the text only points to vaccine strength, decreasing the number of cases, and the alert levels in the United Kingdom. In contrast, two image-caption pairs (one shown on the left in <xref ref-type="fig" rid="F13">Figure 13D</xref>) exhibit <italic>negative SC</italic> and <italic>quality modification</italic> (image showing people wearing masks vs. mask mandate ending in February in text). In terms of objective relations, there is <italic>partial OO</italic> and <italic>Image extends Text</italic>. This could be because of either the use of older dated image or to signify that people are still cautious in public places, and so serves to persuade the reader to be cautious (<italic>persuasive</italic>, promote the &#x0201C;cautious perspective&#x0201D; or demote the &#x0201C;end of covid&#x0201D; message). In addition, the news value of <italic>superlativeness</italic> is constructed with multiple people in the image to emphasize this cautiousness aspect. This is a notable difference between the <italic>Daily Mail</italic> and the <italic>Daily Mirror</italic> articles: both are <italic>informative</italic> and <italic>persuasive</italic>, but compose their articles via different usages of text, image-caption pairs and news values. While the <italic>Daily Mail</italic> relies on expanding proximity (to China&#x00027;s condition) and <italic>high-positive SC</italic> image-text pairs, the <italic>Daily Mirror</italic> shows <italic>negative SC</italic> image-text pairs and news value of <italic>personalization</italic> to contrast with the central theme of the article.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Discussion and conclusions</title>
<p>In this final section, we briefly discuss the computational feasibility of different parts of the framework. Moreover, we outline use cases and applications of the framework that can be interesting directions of research in computational modeling, multimodal analytics and social science and then briefly conclude.</p>
<sec>
<title>4.1. Feasibility of computational approaches</title>
<p><italic><bold>Author intent</bold></italic>: Previous work related to author intent tackles the problem of identifying news (reporting of factual information) vs. opinion (Kr&#x000FC;ger et al., <xref ref-type="bibr" rid="B63">2017</xref>), the persuasive effect of news editorials (Baff et al., <xref ref-type="bibr" rid="B7">2020</xref>) and argumentation strategies in news (Khatib et al., <xref ref-type="bibr" rid="B60">2017</xref>). These earlier approaches rely on engineered lexical features which are not robust against changes in topics and not generalisable to unseen (during training) news publishers. Recently, Alhindi et al. (<xref ref-type="bibr" rid="B3">2020</xref>) combine argumentation features and computational language models like BERT<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref> (Devlin et al., <xref ref-type="bibr" rid="B32">2019</xref>) to identify news vs. opinion. They show promising performance in comparison to previous approaches across datasets and publishers. For multimodal author intent detection, Kruk et al. (<xref ref-type="bibr" rid="B64">2019</xref>) benchmark several large multimodal model features to classify author intent from Instagram posts. The recent work suggests that a combination of unique lexical and multimodal features can be used to classify author intent classes as described in our framework.</p>
<p><italic><bold>Cross-modal relations</bold></italic>: There is a wide body of work on vision-language modeling (Gan et al., <xref ref-type="bibr" rid="B38">2022</xref>), where the goal is to learn visual and multimodal features from a large-scale corpus of image-text pairs from the Web. These models often have deep network backbones like BERT (Devlin et al., <xref ref-type="bibr" rid="B32">2019</xref>) and are trained with loss functions such as masked language modeling, masked image region prediction and image-text matching. Recent work on image-text matching and retrieval (Cao et al., <xref ref-type="bibr" rid="B19">2022</xref>) has explicitly focused on refining cross-attention mechanism and local alignment to get better retrieval performance. Consequently, these models focus more toward aligning literal (objective relations) relations than abstract and high-level semantic relations (semantic correlation STATUS, modification). Recent work (Xu et al., <xref ref-type="bibr" rid="B140">2022</xref>) suggests that weakly aligned (e.g., low-positive SC) and unaligned (e.g., zero, negative SC) pre-training can be beneficial and is still understudied. On the contrary, recent work such as Henning and Ewerth (<xref ref-type="bibr" rid="B51">2018</xref>) and Otto et al. (<xref ref-type="bibr" rid="B94">2020</xref>) have tailored multimodal neural networks in order to predict relationship metrics like CMI, SC, and STATUS. M&#x000FC;ller-Budack et al. (<xref ref-type="bibr" rid="B87">2021b</xref>) have taken a different approach of predicting image-text consistency in news over named entities like persons, locations and events, using several supervised pre-trained model features for face recognition, geo-location estimation and event detection.</p>
<p><italic><bold>News values</bold></italic>: For all the news values except <italic>sentiment</italic>, parts of speech tagging (Chiche and Yitagesu, <xref ref-type="bibr" rid="B27">2022</xref>) and named entity recognition and linking (Luo et al., <xref ref-type="bibr" rid="B71">2015</xref>) can identify words and phrases for each news value. With the identified words, one can either create a rule-based classifier to classify the news values or train a data-driven classifier. In addition, temporal reasoning and extracting temporal and causal relations can be beneficial to identify the impact and novelty of the event in news (Caselli and Vossen, <xref ref-type="bibr" rid="B24">2017</xref>). Both lexical and neural network-based classifiers are well explored for sentiment and can be used to detect sentiment in news (Mello et al., <xref ref-type="bibr" rid="B79">2022</xref>). Moreover, convolutional neural networks (CNN) and multimodal neural networks can also be applied for visual (Ortis et al., <xref ref-type="bibr" rid="B91">2020</xref>) and multimodal (Soleymani et al., <xref ref-type="bibr" rid="B112">2017</xref>) sentiment analysis. Similarly, topic modeling can be performed using recent state-of-the-art approaches like BERTopic (Grootendorst, <xref ref-type="bibr" rid="B42">2022</xref>), which combines word embeddings with TF-IDF (Term Frequency-Inverse Document Frequency) to create topic representations. Recently, some works have explored deep networks for image understanding in combination with text-based topic detection for news videos (Li et al., <xref ref-type="bibr" rid="B68">2017</xref>) and Twitter posts (Zhang et al., <xref ref-type="bibr" rid="B144">2019</xref>). For images, while there is ongoing research for tasks related to some news values, others, like stereotype identification and novelty detection, are still yet to be explored. For other news values, such as identifying location, visual cultural symbols and identifying objects for more context, open vocabulary multimodal models like CLIP (Contrastive Language Image Pre-training) (Radford et al., <xref ref-type="bibr" rid="B102">2021</xref>) can be used to detect the presence of things of interest. Some news values require face recognition (Zhu et al., <xref ref-type="bibr" rid="B151">2021</xref>) and facial analysis (Li and Deng, <xref ref-type="bibr" rid="B67">2022</xref>), which can be efficiently done with the use of CNN-based face detection and facial expression prediction. For identifying the impact of events in images, deep networks can be applied for a disaster impact assessment on aerial imagery (Gupta et al., <xref ref-type="bibr" rid="B44">2021</xref>) and image-text pairs of tweets (Rizk et al., <xref ref-type="bibr" rid="B105">2019</xref>). Lastly, there are a few works where a BERT-based model has been used for detecting racist (Fokkens et al., <xref ref-type="bibr" rid="B36">2018</xref>), gender (Chiril et al., <xref ref-type="bibr" rid="B29">2021</xref>) and immigrant (S&#x000E1;nchez-Junquera et al., <xref ref-type="bibr" rid="B107">2021</xref>) stereotypes in news, tweets and political debates respectively. As each news value is a separate task, an ensemble or a rule-based model using different features and labels extracted from the discussed models could also be explored in a multi-task and multimodal learning setting.</p>
</sec>
<sec>
<title>4.2. Applications and use cases</title>
<p>As the computational modeling of cross-modal relations is an active research area (see Section 2.2), both the objective and the semantic relations are interesting. For example, they can serve as the basis for creating novel datasets for news analysis for multimodal learning tasks to foster further research in the area. Recently, there have been attempts to create datasets for non-literal relations between image and text in advertisements [e.g., Zhang et al. (<xref ref-type="bibr" rid="B145">2018</xref>)], tweets [e.g., Vempala and Preotiuc-Pietro (<xref ref-type="bibr" rid="B130">2019</xref>)], Instagram posts (Kruk et al., <xref ref-type="bibr" rid="B64">2019</xref>), and across domains (Otto et al., <xref ref-type="bibr" rid="B93">2019b</xref>). However, only Otto et al. (<xref ref-type="bibr" rid="B93">2019b</xref>) decouple relations at conceptual and meaning level motivated from semiotics and none of them are in the news domain. Recent work by Alikhani et al. (<xref ref-type="bibr" rid="B4">2020</xref>), Sosea et al. (<xref ref-type="bibr" rid="B113">2021</xref>), and Utescher and Zarrie&#x000DF; (<xref ref-type="bibr" rid="B127">2021</xref>) have posed new research directions and interesting applications of modeling a variety of image-text relations. From the perspective of news values, recent work on automatic extraction of news values from news text (di Buono et al., <xref ref-type="bibr" rid="B33">2017</xref>; Piotrkowicz et al., <xref ref-type="bibr" rid="B97">2017</xref>; Belyaeva et al., <xref ref-type="bibr" rid="B16">2018</xref>) can be extended to both images and text with novel datasets and multimodal modeling contributions. The challenge of detecting news values in both image and text can be at the level of just predicting the type and class of news value to detecting spans of text and regions in the image signifying the news value. From the user modeling perspective, there has been much work on news recommendation based on Twitter user profiles (Abel et al., <xref ref-type="bibr" rid="B1">2011</xref>, <xref ref-type="bibr" rid="B2">2013</xref>), user graphs (Wu et al., <xref ref-type="bibr" rid="B132">2021</xref>), click patterns and behaviors (Wu et al., <xref ref-type="bibr" rid="B131">2019</xref>, <xref ref-type="bibr" rid="B134">2020</xref>) and multimodality (Wu et al., <xref ref-type="bibr" rid="B135">2022b</xref>). Wu et al. (<xref ref-type="bibr" rid="B133">2022a</xref>) present a detailed survey of personalized news recommendation and organize relevant literature by the type of news content, properties (like a publisher), dynamic information (like popularity, click-through rate) and user features (like user networks). Our framework includes the recipients&#x00027; perspectives in the news process and provides a combined view with news values where contributions can be made from both the content and user modeling for further research in the area.</p>
<p>Modeling semiotic relations at the conceptual and meaning level for news can be important for news-specific downstream tasks like the detection of out-of-context images (Aneja et al., <xref ref-type="bibr" rid="B5">2021</xref>; Luo et al., <xref ref-type="bibr" rid="B70">2021</xref>) or fake news (Giachanou et al., <xref ref-type="bibr" rid="B39">2020</xref>; Singh and Sharma, <xref ref-type="bibr" rid="B108">2021</xref>). Both applications can benefit from predicting additional attributes relevant to news dissemination when coupled with news values from a multimodal perspective. For instance, Springstein et al. (<xref ref-type="bibr" rid="B114">2021</xref>) present an application where cross-modal consistency between text and photo is checked for entity types persons, locations and events. Additional attributes like news values can be added as consistency parameters to enrich existing models. Accurate identification of news values in text and images can also help estimate the newsworthiness of informal news on social media. As mentioned before, the news recommendation system is an interesting application that can be enriched with predictable news values from the content combined with available user traits and click patterns. News retrieval (Tahmasebzadeh et al., <xref ref-type="bibr" rid="B118">2021</xref>) can be enriched based on image-text relations and news values. For instance, news can be retrieved and ranked higher of a particular event where images are relevant (<italic>high objective overlap and semantic correlation</italic>) to the news text and includes personal stories or eyewitness accounts (<italic>personalization</italic>).</p>
<p>From the social science perspective, some of the models and applications discussed above enable news analysis, user studies and news comparisons. For instance, the automatic identification of news values in text and image can aid in studying narrative differences between mainstream news and news on social/alternative media. In terms of news comparison and supporting discourse analysis, Pollak et al. (<xref ref-type="bibr" rid="B98">2011</xref>) use text mining to find contrasting patterns between coverage of &#x0201C;2007 Kenyan elections and post-election crisis&#x0201D; in local (Kenyan) and Western (British and US) newspapers. As presented in the framework, fine-grained relations between image and text coupled with news values can provide a multimodal perspective of image and text usage in important events (like an election crisis) across news publishing platforms. Recently, O&#x00027;Halloran et al. (<xref ref-type="bibr" rid="B90">2021</xref>) have introduced a platform called Multimodal Analysis Platform (MAP) for searching, storing and analysing multimodal content (text, images, and videos) in online social and news media. However, the tool focuses on extracting information from individual modalities and so far does not integrate this as would be necessary to operationalize meaningful search and descriptive analysis of meanings emerging from combinations of modalities. The integration of analytical capabilities of image-text relations and news values would therefore be an important additions for such a tool in order to perform analyses going beyond the descriptive aspects.</p>
</sec>
<sec>
<title>4.3. Conclusions</title>
<p>In this paper, we have reviewed multimodal semiotics and computational science literature to derive a framework of image-text relations, news values, and author intent for the support of richer multimodal news analysis. The presented framework covers the entire spectrum of the news process, from news production to news consumption. Taking inspiration from semiotics as well as recent work in the fields of multimedia and multimodal machine learning on relations between image and text, we outlined a framework based on objective information (<italic>objective relations</italic>), meaning formation (<italic>semantic relations</italic>) and relative relevance (<italic>STATUS</italic>) between the two presentation modalities. We also introduced a new relation called <italic>modification</italic> as a sub-category of semantic relations better to capture the role of images in news dissemination. Further, for news-specific analysis, we introduced the aspect of <italic>author intent</italic> to identify the author&#x00027;s primary purpose behind an article and <italic>news values</italic>. While there have been a few attempts at automatically detecting some news values from text and particularly news headlines, image and text together have yet to be considered in computational research, so this is an upcoming specific research challenge. Finally, we included <italic>subjective interpretation</italic>, which captures the end user&#x00027;s and reader&#x00027;s subjectivity toward presentation modalities and news value aspects. Relevant examples for each part of the framework and comprehensive comparisons with existing work were also presented. Finally, we concluded the paper with possible research directions, applications and relevance of the work from a social science perspective.</p>
</sec>
</sec>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s6">
<title>Author contributions</title>
<p>GC devised the project, developed main ideas, prepared examples, theory and review, and took lead in writing the manuscript. SH, EM-B, and CO provided feedback and fruitful discussions, edited parts of the manuscript, and helped shaping the research and ideas. JB and RE provided critical feedback, helped in shaping the manuscript and bring out contributions, and edited the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>This work was funded by European Union&#x00027;s Horizon 2020 research and innovation programme under the Marie Sk&#x00142;odowska-Curie grant agreement no 812997 (CLEOPATRA project) and by the German Federal Ministry of Education and Research (BMBF, FakeNarratives project, no. 16KIS1517). The publication of this article was funded by the Open Access Fund of Technische Informationsbibliothek (TIB).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://theconversation.com/3-2-billion-images-and-720-000-hours-of-video-are-shared-online-daily-can-you-sort-real-from-fake-148630">https://theconversation.com/3-2-billion-images-and-720-000-hours-of-video-are-shared-online-daily-can-you-sort-real-from-fake-148630</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://www.dailymail.co.uk/news/article-10599205/Daily-Covid-admissions-rise-7TH-day-row-cases-jump-56-week.html">https://www.dailymail.co.uk/news/article-10599205/Daily-Covid-admissions-rise-7TH-day-row-cases-jump-56-week.html</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup><ext-link ext-link-type="uri" xlink:href="https://www.dailymail.co.uk/news/article-11213749/WHO-boss-says-end-Covid-sight-says-means-worst-time-up.html">https://www.dailymail.co.uk/news/article-11213749/WHO-boss-says-end-Covid-sight-says-means-worst-time-up.html</ext-link></p></fn>
<fn id="fn0004"><p><sup>4</sup><ext-link ext-link-type="uri" xlink:href="https://www.theguardian.com/world/2022/mar/10/uk-covid-cases-rising-among-those-aged-55-and-over">https://www.theguardian.com/world/2022/mar/10/uk-covid-cases-rising-among-those-aged-55-and-over</ext-link></p></fn>
<fn id="fn0005"><p><sup>5</sup><ext-link ext-link-type="uri" xlink:href="https://www.mirror.co.uk/news/world-news/end-covid-in-sight-says-27996521">https://www.mirror.co.uk/news/world-news/end-covid-in-sight-says-27996521</ext-link></p></fn>
<fn id="fn0006"><p><sup>6</sup>Bidirectional encoder representations from transformers.</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abel</surname> <given-names>F.</given-names></name> <name><surname>Gao</surname> <given-names>Q.</given-names></name> <name><surname>Houben</surname> <given-names>G.</given-names></name> <name><surname>Tao</surname> <given-names>K.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Analyzing user modeling on twitter for personalized news recommendations,&#x0201D;</article-title> in <source>User Modeling, Adaption and Personalization - 19th International Conference, UMAP 2011</source> (<publisher-loc>Girona, Spain</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-22362-4_1</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abel</surname> <given-names>F.</given-names></name> <name><surname>Gao</surname> <given-names>Q.</given-names></name> <name><surname>Houben</surname> <given-names>G.</given-names></name> <name><surname>Tao</surname> <given-names>K.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Twitter-based user modeling for news recommendations,&#x0201D;</article-title> in <source>IJCAI 2013, Proceedings of the 23rd International Joint Conference on Artificial Intelligence</source> (<publisher-loc>Beijing, China</publisher-loc>) <fpage>2962</fpage>&#x02013;<lpage>2966</lpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Alhindi</surname> <given-names>T.</given-names></name> <name><surname>Muresan</surname> <given-names>S.</given-names></name> <name><surname>Preotiuc-Pietro</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Fact vs. opinion: the role of argumentation features in news classification,&#x0201D;</article-title> in <source>Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020</source> (<publisher-loc>Barcelona, Spain</publisher-loc>: <publisher-name>International Committee on Computational Linguistics</publisher-name>) <fpage>6139</fpage>&#x02013;<lpage>6149</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2020.coling-main.540</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alikhani</surname> <given-names>M.</given-names></name> <name><surname>Sharma</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Soricut</surname> <given-names>R.</given-names></name> <name><surname>Stone</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Cross-modal coherence modeling for caption generation,&#x0201D;</article-title> in <source>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source> <fpage>6525</fpage>&#x02013;<lpage>6535</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2020.acl-main.583</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aneja</surname> <given-names>S.</given-names></name> <name><surname>Bregler</surname> <given-names>C.</given-names></name> <name><surname>Nie&#x000DF;ner</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <publisher-name>Catching out-of-context misinformation with self-supervised learning. CoRR, abs/2101.06278</publisher-name>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Araujo</surname> <given-names>T.</given-names></name> <name><surname>van der Meer</surname> <given-names>T. G.</given-names></name></person-group> (<year>2020</year>). <article-title>News values on social media: Exploring what drives peaks in user activity about organizations on twitter</article-title>. <source>Journalism</source> <volume>21</volume>, <fpage>633</fpage>&#x02013;<lpage>651</lpage>. <pub-id pub-id-type="doi">10.1177/1464884918809299</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baff</surname> <given-names>R. E.</given-names></name> <name><surname>Wachsmuth</surname> <given-names>H.</given-names></name> <name><surname>Khatib</surname> <given-names>K. A.</given-names></name> <name><surname>Stein</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Analyzing the persuasive effect of style in news editorial argumentation,&#x0201D;</article-title> in <source>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020</source> (<publisher-loc>Association for Computational Linguistics</publisher-loc>) <fpage>3154</fpage>&#x02013;<lpage>3160</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2020.acl-main.287</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baltrusaitis</surname> <given-names>T.</given-names></name> <name><surname>Ahuja</surname> <given-names>C.</given-names></name> <name><surname>Morency</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>Multimodal machine learning: A survey and taxonomy</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>41</volume>, <fpage>423</fpage>&#x02013;<lpage>443</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2018.2798607</pub-id><pub-id pub-id-type="pmid">29994351</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Barthes</surname> <given-names>R.</given-names></name></person-group> (<year>1977</year>). <source>Image-Music-Text</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Fontana Press</publisher-name>.</citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bateman</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <source>Text and Image: A Critical Introduction to the Visual/Verbal Divide</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>. <pub-id pub-id-type="doi">10.4324/9781315773971</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bednarek</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Investigating evaluation and news values in news items that are shared through social media</article-title>. <source>Corpora</source> <volume>11</volume>, <fpage>227</fpage>&#x02013;<lpage>257</lpage>. <pub-id pub-id-type="doi">10.3366/cor.2016.0093</pub-id><pub-id pub-id-type="pmid">33914809</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bednarek</surname> <given-names>M.</given-names></name> <name><surname>Caple</surname> <given-names>H.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;value added&#x0201D;: Language, image and news values</article-title>. <source>Discour. Context Media</source> <volume>1</volume>, <fpage>103</fpage>&#x02013;<lpage>113</lpage>. <pub-id pub-id-type="doi">10.1016/j.dcm.2012.05.006</pub-id><pub-id pub-id-type="pmid">27885969</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bednarek</surname> <given-names>M.</given-names></name> <name><surname>Caple</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <source>The Discourse of News Values: How News Organizations Create Newsworthiness</source>. <publisher-loc>Oxford and New York</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>. <pub-id pub-id-type="doi">10.1093/acprof:oso/9780190653934.001.0001</pub-id><pub-id pub-id-type="pmid">36389024</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bednarek</surname> <given-names>M.</given-names></name> <name><surname>Caple</surname> <given-names>H.</given-names></name> <name><surname>Huan</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>Computer-based analysis of news values: A case study on national day reporting</article-title>. <source>Journal. Stud</source>. <volume>22</volume>, <fpage>702</fpage>&#x02013;<lpage>722</lpage>. <pub-id pub-id-type="doi">10.1080/1461670X.2020.1807393</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bell</surname> <given-names>A.</given-names></name></person-group> (<year>1991</year>). <source>The Language of News Media</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Blackwell</publisher-name>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Belyaeva</surname> <given-names>E.</given-names></name> <name><surname>Kosmerlj</surname> <given-names>A.</given-names></name> <name><surname>Mladenic</surname> <given-names>D.</given-names></name> <name><surname>Leban</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>Automatic estimation of news values reflecting importance and closeness of news events</article-title>. <source>Informatica</source> <volume>42</volume>, <fpage>1132</fpage>. <pub-id pub-id-type="doi">10.31449/inf.v42i4.1132</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Biber</surname> <given-names>D.</given-names></name></person-group> (<year>1988</year>). <source>Variation Across Speech and Writing</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/CBO9780511621024</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brighton</surname> <given-names>P.</given-names></name> <name><surname>Foy</surname> <given-names>D.</given-names></name></person-group> (<year>2007</year>). <source>News Values</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Sage</publisher-name>. <pub-id pub-id-type="doi">10.4135/9781446216026</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cao</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Nie</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>Image-text retrieval: A survey on recent research and development,&#x0201D;</article-title> in <source>Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022</source> (<publisher-loc>Vienna, Austria</publisher-loc>) <fpage>5410</fpage>&#x02013;<lpage>5417</lpage>. <pub-id pub-id-type="doi">10.24963/ijcai.2022/759</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Caple</surname> <given-names>H.</given-names></name></person-group> (<year>2013</year>). <source>Photojournalism: A Social Semiotic Approach</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Springer</publisher-name>. <pub-id pub-id-type="doi">10.1057/9781137314901</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caple</surname> <given-names>H.</given-names></name> <name><surname>Bednarek</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Rethinking news values: What a discursive approach can tell us about the construction of news discourse and news photography</article-title>. <source>Journalism</source> <volume>17</volume>, <fpage>435</fpage>&#x02013;<lpage>455</lpage>. <pub-id pub-id-type="doi">10.1177/1464884914568078</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Caple</surname> <given-names>H.</given-names></name> <name><surname>Bednarek</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <source>DNVA and Intratextual Analysis</source>. <publisher-name>Discursive News Values Analysis (DNVA)</publisher-name>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caple</surname> <given-names>H.</given-names></name> <name><surname>Huan</surname> <given-names>C.</given-names></name> <name><surname>Bednarek</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <source>Multimodal News Analysis across Cultures</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/9781108886048</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Caselli</surname> <given-names>T.</given-names></name> <name><surname>Vossen</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;The event storyline corpus: A new benchmark for causal and temporal relation extraction,&#x0201D;</article-title> in <source>Proceedings of the Events and Stories in the News Workshop&#x00040;ACL 2017</source> (<publisher-loc>Vancouver, Canada</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>77</fpage>&#x02013;<lpage>86</lpage>. <pub-id pub-id-type="doi">10.18653/v1/W17-2711</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>T.</given-names></name> <name><surname>Lu</surname> <given-names>D.</given-names></name> <name><surname>Kan</surname> <given-names>M.</given-names></name> <name><surname>Cui</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Understanding and classifying image tweets,&#x0201D;</article-title> in <source>ACM Multimedia Conference, MM &#x00027;13</source> (<publisher-loc>Barcelona, Spain</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>781</fpage>&#x02013;<lpage>784</lpage>. <pub-id pub-id-type="doi">10.1145/2502081.2502203</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Yu</surname> <given-names>L.</given-names></name> <name><surname>Kholy</surname> <given-names>A. E.</given-names></name> <name><surname>Ahmed</surname> <given-names>F.</given-names></name> <name><surname>Gan</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;UNITER: universal image-text representation learning,&#x0201D;</article-title> in <source>Computer Vision - ECCV 2020 - 16th European Conference</source> (<publisher-loc>Glasgow, UK</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>104</fpage>&#x02013;<lpage>120</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-58577-8_7</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chiche</surname> <given-names>A.</given-names></name> <name><surname>Yitagesu</surname> <given-names>B.</given-names></name></person-group> (<year>2022</year>). <article-title>Part of speech tagging: a systematic review of deep learning and machine learning approaches</article-title>. <source>J. Big Data</source> <volume>9</volume>, <fpage>10</fpage>. <pub-id pub-id-type="doi">10.1186/s40537-022-00561-y</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chinnappa</surname> <given-names>D.</given-names></name> <name><surname>Murugan</surname> <given-names>S.</given-names></name> <name><surname>Blanco</surname> <given-names>E.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Extracting possessions from social media: Images complement language,&#x0201D;</article-title> in <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019</source> (<publisher-loc>Hong Kong, China</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>663</fpage>&#x02013;<lpage>672</lpage>. <pub-id pub-id-type="doi">10.18653/v1/D19-1061</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chiril</surname> <given-names>P.</given-names></name> <name><surname>Benamara</surname> <given-names>F.</given-names></name> <name><surname>Moriceau</surname> <given-names>V.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Be nice to your wife! the restaurants are closed&#x0201D;: Can gender stereotype detection improve sexism classification?,&#x0201D;</article-title> in <source>Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event</source> (<publisher-loc>Punta Cana, Dominican Republic</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>2833</fpage>&#x02013;<lpage>2844</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2021.findings-emnlp.242</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cortes</surname> <given-names>C.</given-names></name> <name><surname>Vapnik</surname> <given-names>V.</given-names></name></person-group> (<year>1995</year>). <article-title>Support-vector networks</article-title>. <source>Mach. Learn</source>. <volume>20</volume>, <fpage>273</fpage>&#x02013;<lpage>297</lpage>. <pub-id pub-id-type="doi">10.1007/BF00994018</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;Imagenet: A large-scale hierarchical image database,&#x0201D;</article-title> in <source>2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009)</source> (<publisher-loc>Miami, Florida, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>) <fpage>248</fpage>&#x02013;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206848</pub-id><pub-id pub-id-type="pmid">26886976</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;BERT: pre-training of deep bidirectional transformers for language understanding,&#x0201D;</article-title> in <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019</source> (<publisher-loc>Minneapolis, MN, USA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>4171</fpage>&#x02013;<lpage>4186</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>di Buono</surname> <given-names>M. P.</given-names></name> <name><surname>Snajder</surname> <given-names>J.</given-names></name> <name><surname>Basic</surname> <given-names>B. D.</given-names></name> <name><surname>Glavas</surname> <given-names>G.</given-names></name> <name><surname>Tutek</surname> <given-names>M.</given-names></name> <name><surname>Milic-Frayling</surname> <given-names>N.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Predicting news values from headline text and emotions,&#x0201D;</article-title> in <source>Proceedings of the 2017 Workshop: Natural Language Processing meets Journalism, NLPmJ&#x00040;EMNLP</source> (<publisher-loc>Copenhagen, Denmark</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.18653/v1/W17-4201</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Diakopoulos</surname> <given-names>N.</given-names></name> <name><surname>Trielli</surname> <given-names>D.</given-names></name> <name><surname>Lee</surname> <given-names>G.</given-names></name></person-group> (<year>2021</year>). <article-title>Towards understanding and supporting journalistic practices using semi-automated news discovery tools</article-title>. <source>Proc. ACM Human-Comput. Inter</source>. <volume>5</volume>, <fpage>1</fpage>&#x02013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1145/3479550</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>D&#x00027;Ignazio</surname> <given-names>C.</given-names></name> <name><surname>Bhargava</surname> <given-names>R.</given-names></name> <name><surname>Zuckerman</surname> <given-names>E.</given-names></name> <name><surname>Beck</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <source>Cliff-clavin: Determining geographic focus for news articles</source>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fokkens</surname> <given-names>A.</given-names></name> <name><surname>Ruigrok</surname> <given-names>N.</given-names></name> <name><surname>Beukeboom</surname> <given-names>C. J.</given-names></name> <name><surname>Sarah</surname> <given-names>G.</given-names></name> <name><surname>Attveldt</surname> <given-names>W. V.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Studying muslim stereotyping through microportrait extraction,&#x0201D;</article-title> in <source>Proceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018</source> (<publisher-loc>Miyazaki, Japan</publisher-loc>: <publisher-name>European Language Resources Association (ELRA)</publisher-name>).</citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galtung</surname> <given-names>J.</given-names></name> <name><surname>Ruge</surname> <given-names>M. H.</given-names></name></person-group> (<year>1965</year>). <article-title>The structure of foreign news: The presentation of the congo, cuba and cyprus crises in four norwegian newspapers</article-title>. <source>J. Peace Res</source>. <volume>2</volume>, <fpage>64</fpage>&#x02013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1177/002234336500200104</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gan</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Gao</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Vision-language pre-training: Basics, recent advances, and future trends</article-title>. <source>Found. Trends Comput. Graph. Vis</source>. <volume>14</volume>, <fpage>163</fpage>&#x02013;<lpage>352</lpage>. <pub-id pub-id-type="doi">10.1561/0600000105</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Giachanou</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>G.</given-names></name> <name><surname>Rosso</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Multimodal fake news detection with textual, visual and semantic information,&#x0201D;</article-title> in <source>Text, Speech, and Dialogue - 23rd International Conference, TSD 2020</source> (<publisher-loc>Brno, Czech Republic</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>30</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-58323-1_3</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Godbole</surname> <given-names>N.</given-names></name> <name><surname>Srinivasaiah</surname> <given-names>M.</given-names></name> <name><surname>Skiena</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Large-scale sentiment analysis for news and blogs,&#x0201D;</article-title> in <source>Proceedings of the First International Conference on Weblogs and Social Media, ICWSM 2007</source> (<publisher-loc>Boulder, Colorado, USA</publisher-loc>).</citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Golbeck</surname> <given-names>J.</given-names></name> <name><surname>Mauriello</surname> <given-names>M. L.</given-names></name> <name><surname>Auxier</surname> <given-names>B.</given-names></name> <name><surname>Bhanushali</surname> <given-names>K. H.</given-names></name> <name><surname>Bonk</surname> <given-names>C.</given-names></name> <name><surname>Bouzaghrane</surname> <given-names>M. A.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Fake news vs satire: A dataset and analysis,&#x0201D;</article-title> in <source>Proceedings of the 10th ACM Conference on Web Science, WebSci 2018</source> (<publisher-loc>Amsterdam, The Netherlands</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>17</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1145/3201064.3201100</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grootendorst</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>Bertopic: Neural topic modeling with a class-based TF-IDF procedure. CoRR, abs/2203.05794</article-title>.</citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>C.</given-names></name> <name><surname>Sun</surname> <given-names>C.</given-names></name> <name><surname>Ross</surname> <given-names>D. A.</given-names></name> <name><surname>Vondrick</surname> <given-names>C.</given-names></name> <name><surname>Pantofaru</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>AVA: A video dataset of spatio-temporally localized atomic visual actions,&#x0201D;</article-title> in <source>2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018</source> (<publisher-loc>Salt Lake City, UT, USA</publisher-loc>: <publisher-name>Computer Vision Foundation/IEEE Computer Society</publisher-name>) <fpage>6047</fpage>&#x02013;<lpage>6056</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00633</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gupta</surname> <given-names>A.</given-names></name> <name><surname>Watson</surname> <given-names>S.</given-names></name> <name><surname>Yin</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep learning-based aerial image segmentation with open data for disaster impact assessment</article-title>. <source>Neurocomputing</source> <volume>439</volume>, <fpage>22</fpage>&#x02013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.02.139</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Halliday</surname> <given-names>M. A.</given-names></name></person-group> (<year>1985</year>). <source>An Introduction to Functional Grammar</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>.</citation>
</ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Halliday</surname> <given-names>M. A. K.</given-names></name> <name><surname>Matthiessen</surname> <given-names>C.</given-names></name> <name><surname>Halliday</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <source>An Introduction to Functional Grammar</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>. <pub-id pub-id-type="doi">10.4324/9780203783771</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hanselowski</surname> <given-names>A. S A. P. V</given-names></name> <name><surname>Schiller</surname> <given-names>B.</given-names></name> <name><surname>Caspelherr</surname> <given-names>F.</given-names></name> <name><surname>Chaudhuri</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;A retrospective analysis of the fake news challenge stance-detection task,&#x0201D;</article-title> in <source>Proceedings of the 27th International Conference on Computational Linguistics, COLING 2018</source> (<publisher-loc>Santa Fe, New Mexico, USA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>1859</fpage>&#x02013;<lpage>1874</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harcup</surname> <given-names>T.</given-names></name> <name><surname>O&#x00027;Neill</surname> <given-names>D.</given-names></name></person-group> (<year>2001</year>). <article-title>What is news? Galtung and ruge revisited</article-title>. <source>Journal. Stud</source>. <volume>2</volume>, <fpage>261</fpage>&#x02013;<lpage>280</lpage>. <pub-id pub-id-type="doi">10.1080/14616700118449</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harcup</surname> <given-names>T.</given-names></name> <name><surname>O&#x00027;Neill</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>What is news? News values revisited (again)</article-title>. <source>Journal. Stud</source>. <volume>18</volume>, <fpage>1470</fpage>&#x02013;<lpage>1488</lpage>. <pub-id pub-id-type="doi">10.1080/1461670X.2016.1150193</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Heilbron</surname> <given-names>F. C.</given-names></name> <name><surname>Escorcia</surname> <given-names>V.</given-names></name> <name><surname>Ghanem</surname> <given-names>B.</given-names></name> <name><surname>Niebles</surname> <given-names>J. C.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Activitynet: A large-scale video benchmark for human activity understanding,&#x0201D;</article-title> in <source>IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015</source> (<publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>) <fpage>961</fpage>&#x02013;<lpage>970</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298698</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henning</surname> <given-names>C. A.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>Estimating the information gap between textual and visual representations</article-title>. <source>Int. J. Multim. Inf. Retr</source>. <volume>7</volume>, <fpage>43</fpage>&#x02013;<lpage>56</lpage>. <pub-id pub-id-type="doi">10.1007/s13735-017-0142-y</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hogan</surname> <given-names>B.</given-names></name></person-group> (<year>2010</year>). <article-title>The presentation of self in the age of social media: Distinguishing performances and exhibitions online</article-title>. <source>Bull. Sci. Technol. Soc</source>. <volume>30</volume>, <fpage>377</fpage>&#x02013;<lpage>386</lpage>. <pub-id pub-id-type="doi">10.1177/0270467610385893</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hossain</surname> <given-names>M. Z.</given-names></name> <name><surname>Sohel</surname> <given-names>F.</given-names></name> <name><surname>Shiratuddin</surname> <given-names>M. F.</given-names></name> <name><surname>Laga</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>A comprehensive survey of deep learning for image captioning</article-title>. <source>ACM Comput. Surv</source>. <volume>51</volume>, <fpage>1</fpage>&#x02013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1145/3295748</pub-id><pub-id pub-id-type="pmid">35130142</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Imani</surname> <given-names>M. B.</given-names></name> <name><surname>Chandra</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <name><surname>Khan</surname> <given-names>L.</given-names></name> <name><surname>Thuraisingham</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Focus location extraction from political news reports with bias correction,&#x0201D;</article-title> in <source>2017 IEEE International Conference on Big Data (IEEE BigData 2017)</source> (<publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>) <fpage>1956</fpage>&#x02013;<lpage>1964</lpage>. <pub-id pub-id-type="doi">10.1109/BigData.2017.8258141</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jia</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Xia</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Parekh</surname> <given-names>Z.</given-names></name> <name><surname>Pham</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Scaling up visual and vision-language representation learning with noisy text supervision,&#x0201D;</article-title> in <source>Proceedings of the 38th International Conference on Machine Learning, ICML 2021</source> (<publisher-loc>PMLR</publisher-loc>) <fpage>4904</fpage>&#x02013;<lpage>4916</lpage>.</citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Judina</surname> <given-names>D.</given-names></name> <name><surname>Platonov</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>Newsworthiness and the public&#x00027;s response in russian social media: A comparison of state and private news organizations</article-title>. <source>Media Communic</source>. <volume>7</volume>, <fpage>157</fpage>&#x02013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.17645/mac.v7i3.1910</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Karlsson</surname> <given-names>M.</given-names></name> <name><surname>Sj&#x000F8;vaag</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>Content analysis and online news: epistemologies of analysing the ephemeral web</article-title>. <source>Digital Journal</source>. <volume>4</volume>, <fpage>177</fpage>&#x02013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1080/21670811.2015.1096619</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Karpathy</surname> <given-names>A.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Deep visual-semantic alignments for generating image descriptions,&#x0201D;</article-title> in <source>IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015</source> (<publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>) <fpage>3128</fpage>&#x02013;<lpage>3137</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298932</pub-id><pub-id pub-id-type="pmid">27514036</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kastner</surname> <given-names>M. A.</given-names></name> <name><surname>Ide</surname> <given-names>I.</given-names></name> <name><surname>Nack</surname> <given-names>F.</given-names></name> <name><surname>Kawanishi</surname> <given-names>Y.</given-names></name> <name><surname>Hirayama</surname> <given-names>T.</given-names></name> <name><surname>Deguchi</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Estimating the imageability of words by mining visual characteristics from crawled image data</article-title>. <source>Multim. Tools Appl</source>. <volume>79</volume>, <fpage>18167</fpage>&#x02013;<lpage>18199</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-019-08571-4</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Khatib</surname> <given-names>K. A.</given-names></name> <name><surname>Wachsmuth</surname> <given-names>H.</given-names></name> <name><surname>Hagen</surname> <given-names>M.</given-names></name> <name><surname>Stein</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Patterns of argumentation strategies across topics,&#x0201D;</article-title> in <source>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017</source> (<publisher-loc>Copenhagen, Denmark</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>1351</fpage>&#x02013;<lpage>1357</lpage>. <pub-id pub-id-type="doi">10.18653/v1/D17-1141</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kiros</surname> <given-names>R.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name> <name><surname>Zemel</surname> <given-names>R. S.</given-names></name></person-group> (<year>2014</year>). <article-title>Unifying visual-semantic embeddings with multimodal neural language models. CoRR, abs/1411.2539</article-title>.</citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kloepfer</surname> <given-names>R.</given-names></name></person-group> (<year>1976</year>). <article-title>Komplementarit&#x000E4;t von sprache und bild am beispiel von comic, karikatur und reklame.(la compl&#x000E9;mentarit&#x000E9; de la langue et de l&#x00027;image. l&#x00027;exemple des bandes dessin&#x000E9;es, des caricatures et des r&#x000E9;clames)</article-title>. <source>Sprache Techn. Zeitalter Stuttgart</source>. <volume>57</volume>, <fpage>42</fpage>&#x02013;<lpage>56</lpage>.</citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kr&#x000FC;ger</surname> <given-names>K. R.</given-names></name> <name><surname>Lukowiak</surname> <given-names>A.</given-names></name> <name><surname>Sonntag</surname> <given-names>J.</given-names></name> <name><surname>Warzecha</surname> <given-names>S.</given-names></name> <name><surname>Stede</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Classifying news versus opinions in newspapers: Linguistic features for domain independence</article-title>. <source>Nat. Lang. Eng</source>. <volume>23</volume>, <fpage>687</fpage>&#x02013;<lpage>707</lpage>. <pub-id pub-id-type="doi">10.1017/S1351324917000043</pub-id></citation>
</ref>
<ref id="B64">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kruk</surname> <given-names>J.</given-names></name> <name><surname>Lubin</surname> <given-names>J.</given-names></name> <name><surname>Sikka</surname> <given-names>K.</given-names></name> <name><surname>Lin</surname> <given-names>X.</given-names></name> <name><surname>Jurafsky</surname> <given-names>D.</given-names></name> <name><surname>Divakaran</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Integrating text and image: Determining multimodal document intent in instagram posts,&#x0201D;</article-title> in <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019</source> (<publisher-loc>Hong Kong, China</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>4621</fpage>&#x02013;<lpage>4631</lpage>. <pub-id pub-id-type="doi">10.18653/v1/D19-1469</pub-id></citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Haffner</surname> <given-names>P.</given-names></name></person-group> (<year>1998</year>). <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc. IEEE</source> <volume>86</volume>, <fpage>2278</fpage>&#x02013;<lpage>2324</lpage>. <pub-id pub-id-type="doi">10.1109/5.726791</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lemke</surname> <given-names>J. L.</given-names></name></person-group> (<year>1998</year>). <article-title>Multiplying meaning: visual and verbal semiotics in scientific text,&#x0201D;</article-title> in <source>Reading science: critical and functional perspectives on discourses of science</source> (<publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>) <fpage>87</fpage>&#x02013;<lpage>113</lpage>.</citation>
</ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Deng</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>Deep facial expression recognition: A survey</article-title>. <source>IEEE Trans. Affect. Comput</source>. <volume>13</volume>, <fpage>1195</fpage>&#x02013;<lpage>1215</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2020.2981446</pub-id></citation>
</ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Joo</surname> <given-names>J.</given-names></name> <name><surname>Qi</surname> <given-names>H.</given-names></name> <name><surname>Zhu</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Joint image-text news topic detection and tracking by multimodal topic and-or graph</article-title>. <source>IEEE Trans. Multim</source>. <volume>19</volume>, <fpage>367</fpage>&#x02013;<lpage>381</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2016.2616279</pub-id></citation>
</ref>
<ref id="B69">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>T.</given-names></name> <name><surname>Ordonez</surname> <given-names>V.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Visual news: Benchmark and challenges in news image captioning,&#x0201D;</article-title> in <source>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021</source> (<publisher-loc>Punta Cana, Dominican Republic</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>6761</fpage>&#x02013;<lpage>6771</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2021.emnlp-main.542</pub-id><pub-id pub-id-type="pmid">27875221</pub-id></citation></ref>
<ref id="B70">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>G.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Rohrbach</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Newsclippings: Automatic generation of out-of-context multimodal media,&#x0201D;</article-title> in <source>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021</source> (<publisher-loc>Virtual Event / Punta Cana, Dominican Republic</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>6801</fpage>&#x02013;<lpage>6817</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2021.emnlp-main.545</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B71">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>G.</given-names></name> <name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Lin</surname> <given-names>C.</given-names></name> <name><surname>Nie</surname> <given-names>Z.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Joint entity recognition and disambiguation,&#x0201D;</article-title> in <source>Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015</source> (<publisher-loc>Lisbon, Portugal</publisher-loc>: <publisher-name>The Association for Computational Linguistics</publisher-name>) <fpage>879</fpage>&#x02013;<lpage>888</lpage>. <pub-id pub-id-type="doi">10.18653/v1/D15-1104</pub-id></citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mahoney</surname> <given-names>J.</given-names></name> <name><surname>Feltwell</surname> <given-names>T.</given-names></name> <name><surname>Ajuruchi</surname> <given-names>O.</given-names></name> <name><surname>Lawson</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Constructing the visual online political self: an analysis of instagram use by the scottish electorate,&#x0201D;</article-title> in <source>Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems</source> <fpage>3339</fpage>&#x02013;<lpage>3351</lpage>. <pub-id pub-id-type="doi">10.1145/2858036.2858160</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mansimov</surname> <given-names>E.</given-names></name> <name><surname>Parisotto</surname> <given-names>E.</given-names></name> <name><surname>Ba</surname> <given-names>L. J.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Generating images from captions with attention,&#x0201D;</article-title> in <source>4th International Conference on Learning Representations, ICLR 2016</source> (<publisher-loc>San Juan, Puerto Rico</publisher-loc>).</citation>
</ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marsh</surname> <given-names>E. E.</given-names></name> <name><surname>White</surname> <given-names>M. D.</given-names></name></person-group> (<year>2003</year>). <article-title>A taxonomy of relationships between images and text</article-title>. <source>J. Document</source>. <volume>59</volume>, <fpage>647</fpage>&#x02013;<lpage>672</lpage>. <pub-id pub-id-type="doi">10.1108/00220410310506303</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martin</surname> <given-names>J. R.</given-names></name></person-group> (<year>1994</year>). <article-title>Macro-genres: the ecology of the page</article-title>. <source>Network</source> <volume>21</volume>, <fpage>29</fpage>&#x02013;<lpage>52</lpage>.</citation>
</ref>
<ref id="B76">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Martin</surname> <given-names>J. R.</given-names></name> <name><surname>Rose</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <source>Genre Relations: Mapping Culture</source>. <publisher-loc>London and New York</publisher-loc>: <publisher-name>Equinox</publisher-name>.<pub-id pub-id-type="pmid">33465413</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martinec</surname> <given-names>R.</given-names></name> <name><surname>Salway</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>A system for image-text relations in new (and old) media</article-title>. <source>Visual Communic</source>. <volume>4</volume>, <fpage>337</fpage>&#x02013;<lpage>371</lpage>. <pub-id pub-id-type="doi">10.1177/1470357205055928</pub-id></citation>
</ref>
<ref id="B78">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mehmet</surname> <given-names>M.</given-names></name> <name><surname>Clarke</surname> <given-names>R. J.</given-names></name> <name><surname>Kautz</surname> <given-names>K.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Social media semantics: Analysing meanings in multimodal online conversations,&#x0201D;</article-title> in <source>Proceedings of the International Conference on Information Systems - Building a Better World through Information Systems, ICIS 2014</source> (<publisher-loc>Auckland, New Zealand</publisher-loc>: <publisher-name>Association for Information Systems</publisher-name>).</citation>
</ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mello</surname> <given-names>C.</given-names></name> <name><surname>Cheema</surname> <given-names>G. S.</given-names></name> <name><surname>Thakkar</surname> <given-names>G.</given-names></name></person-group> (<year>2022</year>). <article-title>Combining sentiment analysis classifiers to explore multilingual news articles covering london 2012 and rio 2016 olympics</article-title>. <source>Int. J. Digital Human</source>. <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1007/s42803-022-00052-9</pub-id><pub-id pub-id-type="pmid">36407478</pub-id></citation></ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikels</surname> <given-names>J. A.</given-names></name> <name><surname>Fredrickson</surname> <given-names>B. L.</given-names></name> <name><surname>Larkin</surname> <given-names>G. R.</given-names></name> <name><surname>Lindberg</surname> <given-names>C. M.</given-names></name> <name><surname>Maglio</surname> <given-names>S. J.</given-names></name> <name><surname>Reuter-Lorenz</surname> <given-names>P. A.</given-names></name></person-group> (<year>2005</year>). <article-title>Emotional category data on images from the international affective picture system</article-title>. <source>Behav. Res. Methods</source> <volume>37</volume>, <fpage>626</fpage>&#x02013;<lpage>630</lpage>. <pub-id pub-id-type="doi">10.3758/BF03192732</pub-id><pub-id pub-id-type="pmid">16629294</pub-id></citation></ref>
<ref id="B81">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>C. R.</given-names></name></person-group> (<year>1994</year>). <article-title>&#x0201C;Genre as social action,&#x0201D;</article-title> in <source>Genre and the New Rhetoric, Chapter 2</source> (<publisher-loc>London</publisher-loc>: <publisher-name>Taylor and Francis</publisher-name>) <fpage>23</fpage>&#x02013;<lpage>42</lpage>.</citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Motta</surname> <given-names>E.</given-names></name> <name><surname>Daga</surname> <given-names>E.</given-names></name> <name><surname>Opdahl</surname> <given-names>A. L.</given-names></name> <name><surname>Tessem</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Analysis and design of computational news angles</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>120613</fpage>&#x02013;<lpage>120626</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3005513</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Moya Guijarro</surname> <given-names>A. J.</given-names></name></person-group> (<year>2014</year>). <source>A Multimodal Analysis of Picture Books for Children: A Systemic Functional Approach</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Equinox</publisher-name>.</citation>
</ref>
<ref id="B84">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>E.</given-names></name> <name><surname>Springstein</surname> <given-names>M.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;When was this picture taken? Image date estimation in the wild,&#x0201D;</article-title> in <source>Advances in Information Retrieval - 39th European Conference on IR Research, ECIR 2017</source> (<publisher-loc>Aberdeen, UK</publisher-loc>) <fpage>619</fpage>&#x02013;<lpage>625</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-56608-5_57</pub-id></citation>
</ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>M&#x000FC;ller-Budack</surname> <given-names>E.</given-names></name> <name><surname>Pustu-Iren</surname> <given-names>K.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Geolocation estimation of photos using a hierarchical model and scene classification,&#x0201D;</article-title> in <source>Computer Vision - ECCV 2018 - 15th European Conference</source> (<publisher-loc>Munich, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>575</fpage>&#x02013;<lpage>592</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01258-8_35</pub-id></citation>
</ref>
<ref id="B86">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>M&#x000FC;ller-Budack</surname> <given-names>E.</given-names></name> <name><surname>Springstein</surname> <given-names>M.</given-names></name> <name><surname>Hakimov</surname> <given-names>S.</given-names></name> <name><surname>Mrutzek</surname> <given-names>K.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2021a</year>). <article-title>Ontology-driven event type classification in images,&#x0201D;</article-title> in <source>IEEE Winter Conference on Applications of Computer Vision, WACV 2021</source> (<publisher-loc>Waikoloa, HI, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>) <fpage>2927</fpage>&#x02013;<lpage>2937</lpage>. <pub-id pub-id-type="doi">10.1109/WACV48630.2021.00297</pub-id></citation>
</ref>
<ref id="B87">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>M&#x000FC;ller-Budack</surname> <given-names>E.</given-names></name> <name><surname>Theiner</surname> <given-names>J.</given-names></name> <name><surname>Diering</surname> <given-names>S.</given-names></name> <name><surname>Idahl</surname> <given-names>M.</given-names></name> <name><surname>Hakimov</surname> <given-names>S.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2021b</year>). <article-title>Multimodal news analytics using measures of cross-modal entity and context consistency</article-title>. <source>Int. J. Multim. Inf. Retr</source>. <volume>10</volume>, <fpage>111</fpage>&#x02013;<lpage>125</lpage>. <pub-id pub-id-type="doi">10.1007/s13735-021-00207-4</pub-id></citation>
</ref>
<ref id="B88">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ngiam</surname> <given-names>J.</given-names></name> <name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>M.</given-names></name> <name><surname>Nam</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Multimodal deep learning,&#x0201D;</article-title> in <source>Proceedings of the 28th International Conference on Machine Learning, ICML 2011</source> (<publisher-loc>Bellevue, Washington, USA</publisher-loc>: <publisher-name>Omnipress</publisher-name>) <fpage>689</fpage>&#x02013;<lpage>696</lpage>.</citation>
</ref>
<ref id="B89">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nhat</surname> <given-names>T. N. M.</given-names></name> <name><surname>Pha</surname> <given-names>N. T. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Exploring text-image relations in english comics for children: The case of &#x0201C;little red riding hood&#x0201D;</article-title>. <source>VNU J. Foreign Stud</source>. <volume>35</volume>, <fpage>4372</fpage>. <pub-id pub-id-type="doi">10.25073/2525-2445/vnufs.4372</pub-id></citation>
</ref>
<ref id="B90">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Halloran</surname> <given-names>K. L.</given-names></name> <name><surname>Pal</surname> <given-names>G.</given-names></name> <name><surname>Jin</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Multimodal approach to analysing big social and news media data</article-title>. <source>Discourse, Context Media</source> <volume>40</volume>, <fpage>100467</fpage>. <pub-id pub-id-type="doi">10.1016/j.dcm.2021.100467</pub-id></citation>
</ref>
<ref id="B91">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ortis</surname> <given-names>A.</given-names></name> <name><surname>Farinella</surname> <given-names>G. M.</given-names></name> <name><surname>Battiato</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Survey on visual sentiment analysis</article-title>. <source>IET Image Process</source>. <volume>14</volume>, <fpage>1440</fpage>&#x02013;<lpage>1456</lpage>. <pub-id pub-id-type="doi">10.1049/iet-ipr.2019.1270</pub-id></citation>
</ref>
<ref id="B92">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Otto</surname> <given-names>C.</given-names></name> <name><surname>Holzki</surname> <given-names>S.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2019a</year>). <article-title>&#x0201C;Is this an example image?&#x0201D; Predicting the relative abstractness level of image and text,</article-title> in <source>Advances in Information Retrieval - 41st European Conference on IR Research, ECIR 2019</source> (<publisher-loc>Cologne, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>711</fpage>&#x02013;<lpage>725</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-15712-8_46</pub-id><pub-id pub-id-type="pmid">24168217</pub-id></citation></ref>
<ref id="B93">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Otto</surname> <given-names>C.</given-names></name> <name><surname>Springstein</surname> <given-names>M.</given-names></name> <name><surname>Anand</surname> <given-names>A.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2019b</year>). <article-title>Understanding, categorizing and predicting semantic image-text relations,&#x0201D;</article-title> in <source>Proceedings of the 2019 on International Conference on Multimedia Retrieval, ICMR 2019</source> (<publisher-loc>Ottawa, ON, Canada</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>168</fpage>&#x02013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.1145/3323873.3325049</pub-id></citation>
</ref>
<ref id="B94">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Otto</surname> <given-names>C.</given-names></name> <name><surname>Springstein</surname> <given-names>M.</given-names></name> <name><surname>Anand</surname> <given-names>A.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Characterization and classification of semantic image-text relations</article-title>. <source>Int. J. Multim. Inf. Retr</source>. <volume>9</volume>, <fpage>31</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1007/s13735-019-00187-6</pub-id></citation>
</ref>
<ref id="B95">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Parekh</surname> <given-names>Z.</given-names></name> <name><surname>Baldridge</surname> <given-names>J.</given-names></name> <name><surname>Cer</surname> <given-names>D.</given-names></name> <name><surname>Waters</surname> <given-names>A.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO,&#x0201D;</article-title> in <source>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021</source> (<publisher-loc>Association for Computational Linguistics</publisher-loc>) <fpage>2855</fpage>&#x02013;<lpage>2870</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2021.eacl-main.249</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B96">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>C. S.</given-names></name> <name><surname>Kaye</surname> <given-names>B. K.</given-names></name></person-group> (<year>2021</year>). <article-title>Applying news values theory to liking, commenting and sharing mainstream news articles on facebook</article-title>. <source>Journalism</source> <volume>24</volume>, <fpage>14648849211019895</fpage>. <pub-id pub-id-type="doi">10.1177/14648849211019895</pub-id></citation>
</ref>
<ref id="B97">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Piotrkowicz</surname> <given-names>A.</given-names></name> <name><surname>Dimitrova</surname> <given-names>V.</given-names></name> <name><surname>Markert</surname> <given-names>K.</given-names></name></person-group> (<year>2017</year>). <article-title>Automatic extraction of news values from headline text,&#x0201D;</article-title> in <source>Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017</source> (<publisher-loc>Valencia, Spain</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>64</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.18653/v1/E17-4007</pub-id></citation>
</ref>
<ref id="B98">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pollak</surname> <given-names>S.</given-names></name> <name><surname>Coesemans</surname> <given-names>R.</given-names></name> <name><surname>Daelemans</surname> <given-names>W.</given-names></name> <name><surname>Lavra&#x0010D;</surname> <given-names>N.</given-names></name></person-group> (<year>2011</year>). <article-title>Detecting contrast patterns in newspaper articles by combining discourse analysis and text mining</article-title>. <source>Pragmatics</source> <volume>21</volume>, <fpage>647</fpage>&#x02013;<lpage>683</lpage>. <pub-id pub-id-type="doi">10.1075/prag.21.4.07pol</pub-id></citation>
</ref>
<ref id="B99">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Poria</surname> <given-names>S.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name> <name><surname>Hazarika</surname> <given-names>D.</given-names></name> <name><surname>Majumder</surname> <given-names>N.</given-names></name> <name><surname>Zadeh</surname> <given-names>A.</given-names></name> <name><surname>Morency</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Context-dependent sentiment analysis in user-generated videos,&#x0201D;</article-title> in <source>Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017</source> (<publisher-loc>Vancouver, Canada</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>). <pub-id pub-id-type="doi">10.18653/v1/P17-1081</pub-id></citation>
</ref>
<ref id="B100">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Potts</surname> <given-names>A.</given-names></name> <name><surname>Bednarek</surname> <given-names>M.</given-names></name> <name><surname>Caple</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>How can computer-based methods help researchers to investigate news values in large datasets? A corpus linguistic study of the construction of newsworthiness in the reporting on hurricane katrina</article-title>. <source>Discour. Commun</source>. <volume>9</volume>, <fpage>149</fpage>&#x02013;<lpage>172</lpage>. <pub-id pub-id-type="doi">10.1177/1750481314568548</pub-id></citation>
</ref>
<ref id="B101">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Qiao</surname> <given-names>T.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Tao</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Mirrorgan: Learning text-to-image generation by redescription,&#x0201D;</article-title> in <source>IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019</source> (<publisher-loc>Long Beach, CA, USA</publisher-loc>: <publisher-name>Computer Vision Foundation/IEEE</publisher-name>) <fpage>1505</fpage>&#x02013;<lpage>1514</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00160</pub-id></citation>
</ref>
<ref id="B102">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>J. W.</given-names></name> <name><surname>Hallacy</surname> <given-names>C.</given-names></name> <name><surname>Ramesh</surname> <given-names>A.</given-names></name> <name><surname>Goh</surname> <given-names>G.</given-names></name> <name><surname>Agarwal</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Learning transferable visual models from natural language supervision,&#x0201D;</article-title> in <source>Proceedings of the 38th International Conference on Machine Learning, ICML 2021</source> (<publisher-loc>PMLR</publisher-loc>) <fpage>8748</fpage>&#x02013;<lpage>8763</lpage>.</citation>
</ref>
<ref id="B103">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ramesh</surname> <given-names>A.</given-names></name> <name><surname>Pavlov</surname> <given-names>M.</given-names></name> <name><surname>Goh</surname> <given-names>G.</given-names></name> <name><surname>Gray</surname> <given-names>S.</given-names></name> <name><surname>Voss</surname> <given-names>C.</given-names></name> <name><surname>Radford</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Zero-shot text-to-image generation,&#x0201D;</article-title> in <source>Proceedings of the 38th International Conference on Machine Learning, ICML 2021</source> (<publisher-loc>PMLR</publisher-loc>) <fpage>8821</fpage>&#x02013;<lpage>8831</lpage>.<pub-id pub-id-type="pmid">36927634</pub-id></citation></ref>
<ref id="B104">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ramisa</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Multimodal news article analysis,&#x0201D;</article-title> in <source>Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017</source> (<publisher-loc>Melbourne, Australia</publisher-loc>) <fpage>5136</fpage>&#x02013;<lpage>5140</lpage>. <pub-id pub-id-type="doi">10.24963/ijcai.2017/737</pub-id></citation>
</ref>
<ref id="B105">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rizk</surname> <given-names>Y.</given-names></name> <name><surname>Jomaa</surname> <given-names>H. S.</given-names></name> <name><surname>Awad</surname> <given-names>M.</given-names></name> <name><surname>Castillo</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;A computationally efficient multi-modal classification approach of disaster-related twitter images,&#x0201D;</article-title> in <source>Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing, SAC 2019</source> (<publisher-loc>Limassol, Cyprus</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>2050</fpage>&#x02013;<lpage>2059</lpage>. <pub-id pub-id-type="doi">10.1145/3297280.3297481</pub-id></citation>
</ref>
<ref id="B106">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Royce</surname> <given-names>T.</given-names></name></person-group> (<year>1998</year>). <article-title>Synergy on the page: Exploring intersemiotic complementarity in page-based multimodal text</article-title>. <source>JASFL Occas</source>. <volume>1</volume>, <fpage>25</fpage>&#x02013;<lpage>49</lpage>.</citation>
</ref>
<ref id="B107">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>S&#x000E1;nchez-Junquera</surname> <given-names>J.</given-names></name> <name><surname>Chulvi</surname> <given-names>B.</given-names></name> <name><surname>Rosso</surname> <given-names>P.</given-names></name> <name><surname>Ponzetto</surname> <given-names>S. P.</given-names></name></person-group> (<year>2021</year>). <article-title>How do you speak about immigrants? Taxonomy and stereoimmigrants dataset for identifying stereotypes about immigrants</article-title>. <source>Appl. Sci</source>. <volume>11</volume>, <fpage>3610</fpage>. <pub-id pub-id-type="doi">10.3390/app11083610</pub-id></citation>
</ref>
<ref id="B108">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>B.</given-names></name> <name><surname>Sharma</surname> <given-names>D. K.</given-names></name></person-group> (<year>2021</year>). <article-title>Predicting image credibility in fake news over social media using multi-modal approach</article-title>. <source>Neural Comput. Applic</source>. <volume>34</volume>, <fpage>21503</fpage>&#x02013;<lpage>21517</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-021-06086-4</pub-id><pub-id pub-id-type="pmid">34054227</pub-id></citation></ref>
<ref id="B109">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>V. K.</given-names></name> <name><surname>Ghosh</surname> <given-names>I.</given-names></name> <name><surname>Sonagara</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <article-title>Detecting fake news stories via multimodal analysis</article-title>. <source>J. Assoc. Inf. Sci. Technol</source>. <volume>72</volume>, <fpage>3</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1002/asi.24359</pub-id></citation>
</ref>
<ref id="B110">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smeulders</surname> <given-names>A. W. M.</given-names></name> <name><surname>Worring</surname> <given-names>M.</given-names></name> <name><surname>Santini</surname> <given-names>S.</given-names></name> <name><surname>Gupta</surname> <given-names>A.</given-names></name> <name><surname>Jain</surname> <given-names>R. C.</given-names></name></person-group> (<year>2000</year>). <article-title>Content-based image retrieval at the end of the early years</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>22</volume>, <fpage>1349</fpage>&#x02013;<lpage>1380</lpage>. <pub-id pub-id-type="doi">10.1109/34.895972</pub-id></citation>
</ref>
<ref id="B111">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Karpathy</surname> <given-names>A.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name> <name><surname>Manning</surname> <given-names>C. D.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2014</year>). <article-title>Grounded compositional semantics for finding and describing images with sentences</article-title>. <source>Trans. Assoc. Comput. Linguist</source>. <volume>2</volume>, <fpage>207</fpage>&#x02013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1162/tacl_a_00177</pub-id></citation>
</ref>
<ref id="B112">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soleymani</surname> <given-names>M.</given-names></name> <name><surname>Garc&#x000ED;a</surname> <given-names>D.</given-names></name> <name><surname>Jou</surname> <given-names>B.</given-names></name> <name><surname>Schuller</surname> <given-names>B. W.</given-names></name> <name><surname>Chang</surname> <given-names>S.</given-names></name> <name><surname>Pantic</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>A survey of multimodal sentiment analysis</article-title>. <source>Image Vis. Comput</source>. <volume>65</volume>, <fpage>3</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1016/j.imavis.2017.08.003</pub-id></citation>
</ref>
<ref id="B113">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sosea</surname> <given-names>T.</given-names></name> <name><surname>Sirbu</surname> <given-names>I.</given-names></name> <name><surname>Caragea</surname> <given-names>C.</given-names></name> <name><surname>Caragea</surname> <given-names>D.</given-names></name> <name><surname>Rebedea</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Using the image-text relationship to improve multimodal disaster tweet classification,&#x0201D;</article-title> in <source>The 18th International Conference on Information Systems for Crisis Response and Management (ISCRAM 2021)</source>.</citation>
</ref>
<ref id="B114">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Springstein</surname> <given-names>M.</given-names></name> <name><surname>M&#x000FC;ller-Budack</surname> <given-names>E.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Quti! quantifying text-image consistency in multimodal documents,&#x0201D;</article-title> in <source>SIGIR &#x00027;21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval</source> (<publisher-loc>Canada</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>2575</fpage>&#x02013;<lpage>2579</lpage>. <pub-id pub-id-type="doi">10.1145/3404835.3462796</pub-id></citation>
</ref>
<ref id="B115">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>St&#x000F6;ckl</surname> <given-names>H.</given-names></name></person-group> (<year>1997</year>). <source>Textstil und Semiotik englischsprachiger Anzeigenwerbung</source>. <publisher-loc>Frankfurt am Main</publisher-loc>: <publisher-name>Peter Lang</publisher-name>.</citation>
</ref>
<ref id="B116">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>St&#x000F6;ckl</surname> <given-names>H.</given-names></name> <name><surname>Caple</surname> <given-names>H.</given-names></name> <name><surname>Pflaeging</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <source>Shifts Towards Image-Centricity in Contemporary Multimodal Practices</source>. <publisher-loc>London and New York</publisher-loc>: <publisher-name>Routledge</publisher-name>. <pub-id pub-id-type="doi">10.4324/9780429487965</pub-id><pub-id pub-id-type="pmid">32003690</pub-id></citation></ref>
<ref id="B117">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Swales</surname> <given-names>J. M.</given-names></name></person-group> (<year>1990</year>). <source>Genre Analysis: English in Academic and Research Settings</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B118">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tahmasebzadeh</surname> <given-names>G.</given-names></name> <name><surname>Kacupaj</surname> <given-names>E.</given-names></name> <name><surname>M&#x000FC;ller-Budack</surname> <given-names>E.</given-names></name> <name><surname>Hakimov</surname> <given-names>S.</given-names></name> <name><surname>Lehmann</surname> <given-names>J.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>Geowine: Geolocation based wiki, image, news and event retrieval,&#x0201D;</article-title> in <source>SIGIR &#x00027;21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval</source> (<publisher-loc>Canada</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>2565</fpage>&#x02013;<lpage>2569</lpage>. <pub-id pub-id-type="doi">10.1145/3404835.3462786</pub-id></citation>
</ref>
<ref id="B119">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tahmasebzadeh</surname> <given-names>G.</given-names></name> <name><surname>M&#x000FC;ller-Budack</surname> <given-names>E.</given-names></name> <name><surname>Hakimov</surname> <given-names>S.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>Mm-locate-news: Multimodal focus location estimation in news</article-title>. arXiv preprint arXiv:2211.08042. <pub-id pub-id-type="doi">10.1007/978-3-031-28238-6_14</pub-id></citation>
</ref>
<ref id="B120">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taj</surname> <given-names>S.</given-names></name> <name><surname>Shaikh</surname> <given-names>B. B.</given-names></name> <name><surname>Meghji</surname> <given-names>A. F.</given-names></name></person-group> (<year>2019</year>). Sentiment analysis of news articles: a lexicon based approach,&#x0201D; in <italic>2019 2nd International Conference on Computing, Mathematics and Engineering Technologies (iCoMET)</italic> (IEEE) <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/ICOMET.2019.8673428</pub-id></citation>
</ref>
<ref id="B121">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tandoc</surname> <given-names>E. C.</given-names> <suffix>Jr</suffix></name> <name><surname>Thomas</surname> <given-names>R. J.</given-names></name> <name><surname>Bishop</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>What is (fake) news? Analyzing news values (and more) in fake stories</article-title>. <source>Media Communic</source>. <volume>9</volume>, <fpage>110</fpage>&#x02013;<lpage>119</lpage>. <pub-id pub-id-type="doi">10.17645/mac.v9i1.3331</pub-id></citation>
</ref>
<ref id="B122">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tessem</surname> <given-names>B.</given-names></name> <name><surname>Nyre</surname> <given-names>L.</given-names></name> <name><surname>Mesquita</surname> <given-names>M. d,. S</given-names></name> <name><surname>Mulholland</surname> <given-names>P.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Deep learning to encourage citizen involvement in local journalism,&#x0201D;</article-title> in <source>Futures of Journalism: Technology-stimulated Evolution in the Audience-News Media Relationship</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>211</fpage>&#x02013;<lpage>226</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-95073-6_14</pub-id></citation>
</ref>
<ref id="B123">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Theiner</surname> <given-names>J.</given-names></name> <name><surname>M&#x000FC;ller-Budack</surname> <given-names>E.</given-names></name> <name><surname>Ewerth</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Interpretable semantic photo geolocation,&#x0201D;</article-title> in <source>IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2022</source> (<publisher-loc>Waikoloa, HI, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>) <fpage>1474</fpage>&#x02013;<lpage>1484</lpage>. <pub-id pub-id-type="doi">10.1109/WACV51458.2022.00154</pub-id></citation>
</ref>
<ref id="B124">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thomee</surname> <given-names>B.</given-names></name> <name><surname>Shamma</surname> <given-names>D. A.</given-names></name> <name><surname>Friedland</surname> <given-names>G.</given-names></name> <name><surname>Elizalde</surname> <given-names>B.</given-names></name> <name><surname>Ni</surname> <given-names>K.</given-names></name> <name><surname>Poland</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>YFCC100M: the new data in multimedia research</article-title>. <source>Commun. ACM</source> <volume>59</volume>, <fpage>64</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1145/2812802</pub-id></citation>
</ref>
<ref id="B125">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Trattner</surname> <given-names>C.</given-names></name> <name><surname>Jannach</surname> <given-names>D.</given-names></name> <name><surname>Motta</surname> <given-names>E.</given-names></name> <name><surname>Costera Meijer</surname> <given-names>I.</given-names></name> <name><surname>Diakopoulos</surname> <given-names>N.</given-names></name> <name><surname>Elahi</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Responsible media technology and ai: challenges and research directions</article-title>. <source>AI Ethics</source>. <volume>2</volume>, <fpage>585</fpage>&#x02013;<lpage>594</lpage>. <pub-id pub-id-type="doi">10.1007/s43681-021-00126-4</pub-id></citation>
</ref>
<ref id="B126">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Unsworth</surname> <given-names>L.</given-names></name></person-group> (<year>2007</year>). <article-title>Image/text relations and intersemiosis: Towards multimodal text description for multiliteracies education,&#x0201D;</article-title> in <source>Proceedings of the 33rd IFSC: International Systemic Functional Congress</source>. <publisher-name>Pontif&#x000ED;cia Universidade Cat&#x000F3;lica de S&#x000E3;o Paulo</publisher-name>.</citation>
</ref>
<ref id="B127">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Utescher</surname> <given-names>R.</given-names></name> <name><surname>Zarrie&#x000DF;</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>What did this castle look like before? exploring referential relations in naturally occurring multimodal texts,&#x0201D;</article-title> in <source>Proceedings of the Third Workshop on Beyond Vision and LANguage: inTEgrating Real-world kNowledge (LANTERN)</source> (<publisher-loc>Kyiv, Ukraine</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>53</fpage>&#x02013;<lpage>60</lpage>.</citation>
</ref>
<ref id="B128">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Leeuwen</surname> <given-names>T.</given-names></name></person-group> (<year>1991</year>). <article-title>Conjunctive structure in documentary film and television</article-title>. <source>Continuum J. Media Cult. Stud</source>. <volume>5</volume>, <fpage>76</fpage>&#x02013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1080/10304319109388216</pub-id></citation>
</ref>
<ref id="B129">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>van Leeuwen</surname> <given-names>T.</given-names></name></person-group> (<year>2005</year>). <source>Introducing Social Semiotics</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Psychology Press</publisher-name>. <pub-id pub-id-type="doi">10.4324/9780203647028</pub-id></citation>
</ref>
<ref id="B130">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vempala</surname> <given-names>A.</given-names></name> <name><surname>Preotiuc-Pietro</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Categorizing and inferring the relationship between the text and image of twitter posts,&#x0201D;</article-title> in <source>Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019</source> (<publisher-loc>Florence, Italy</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>) <fpage>2830</fpage>&#x02013;<lpage>2840</lpage>. <pub-id pub-id-type="doi">10.18653/v1/P19-1272</pub-id></citation>
</ref>
<ref id="B131">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>An</surname> <given-names>M.</given-names></name> <name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;NPA: neural news recommendation with personalized attention,&#x0201D;</article-title> in <source>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &#x00026;Data Mining, KDD 2019</source> (<publisher-loc>Anchorage, AK, USA</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>2576</fpage>&#x02013;<lpage>2584</lpage>.</citation>
</ref>
<ref id="B132">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;User-as-graph: User modeling with heterogeneous graph pooling for news recommendation,&#x0201D;</article-title> in <source>Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021</source> (<publisher-loc>Montreal, Canada</publisher-loc>) <fpage>1624</fpage>&#x02013;<lpage>1630</lpage>.</citation>
</ref>
<ref id="B133">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name></person-group> (<year>2022a</year>). <article-title>Personalized news recommendation: Methods and challenges</article-title>. <source>ACM Trans. Inf. Syst</source>. <volume>41</volume>, <fpage>1</fpage>&#x02013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1145/3530257</pub-id></citation>
</ref>
<ref id="B134">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>Qi</surname> <given-names>T.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;User modeling with click preference and reading satisfaction for news recommendation,&#x0201D;</article-title> in <source>Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020</source> <fpage>3023</fpage>&#x02013;<lpage>3029</lpage>. <pub-id pub-id-type="doi">10.24963/ijcai.2020/418</pub-id></citation>
</ref>
<ref id="B135">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>Qi</surname> <given-names>T.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>T.</given-names></name></person-group> (<year>2022b</year>). <article-title>&#x0201C;Mm-rec: Visiolinguistic model empowered multimodal news recommendation,&#x0201D;</article-title> in <source>SIGIR &#x00027;22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval</source> (<publisher-loc>Madrid, Spain</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>2560</fpage>&#x02013;<lpage>2564</lpage>.</citation>
</ref>
<ref id="B136">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>A multimodal analysis of image-text relations in picture books</article-title>. <source>Theory Pract. Langu. Stud</source>.<volume>4</volume>, <fpage>1415</fpage>. <pub-id pub-id-type="doi">10.4304/tpls.4.7.1415-1420</pub-id></citation>
</ref>
<ref id="B137">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wunderli</surname> <given-names>P. S.</given-names></name></person-group> (<year>1995</year>). <article-title>Winfried n&#x000F6;th, handbook of semiotics</article-title>. <source>Zeitschrift Romanische Philol</source>. <volume>111</volume>, <fpage>59</fpage>&#x02013;<lpage>60</lpage>.</citation>
</ref>
<ref id="B138">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiao</surname> <given-names>J.</given-names></name> <name><surname>Hays</surname> <given-names>J.</given-names></name> <name><surname>Ehinger</surname> <given-names>K. A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;SUN database: Large-scale scene recognition from abbey to zoo,&#x0201D;</article-title> in <source>The Twenty-Third IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2010</source> (<publisher-loc>San Francisco, CA, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>) <fpage>3485</fpage>&#x02013;<lpage>3492</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2010.5539970</pub-id></citation>
</ref>
<ref id="B139">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiong</surname> <given-names>Y.</given-names></name> <name><surname>Zhu</surname> <given-names>K.</given-names></name> <name><surname>Lin</surname> <given-names>D.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Recognize complex events from static images by fusing deep channels,&#x0201D;</article-title> in <source>IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015</source> (<publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>) <fpage>1600</fpage>&#x02013;<lpage>1609</lpage>.</citation>
</ref>
<ref id="B140">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>P.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Clifton</surname> <given-names>D. A.</given-names></name></person-group> (<year>2022</year>). <article-title>Multimodal learning with transformers: A survey. CoRR, abs/2206.06488</article-title>.</citation>
</ref>
<ref id="B141">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>R.</given-names></name> <name><surname>Xiong</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Corso</surname> <given-names>J. J.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Jointly modeling deep video and compositional text to bridge vision and language in a unified framework,&#x0201D;</article-title> in <source>Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence</source> (<publisher-loc>Austin, Texas, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>) <fpage>2346</fpage>&#x02013;<lpage>2352</lpage>.</citation>
</ref>
<ref id="B142">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xue</surname> <given-names>J.</given-names></name> <name><surname>Du</surname> <given-names>Y.</given-names></name> <name><surname>Shui</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Semantic correlation mining between images and texts with global semantics and local mapping,&#x0201D;</article-title> in <source>MultiMedia Modeling - 21st International Conference, MMM 2015</source> (<publisher-loc>Sydney, NSW, Australia</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>427</fpage>&#x02013;<lpage>435</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-14442-9_48</pub-id></citation>
</ref>
<ref id="B143">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yanai</surname> <given-names>K.</given-names></name> <name><surname>Barnard</surname> <given-names>K.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x0201C;Image region entropy: a measure of &#x0201C;visualness&#x0201D; of web images associated with one concept,&#x0201D;</article-title> in <source>Proceedings of the 13th ACM International Conference on Multimedia</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>419</fpage>&#x02013;<lpage>422</lpage>. <pub-id pub-id-type="doi">10.1145/1101149.1101241</pub-id></citation>
</ref>
<ref id="B144">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Lu</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Xiao</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Chen</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>A novel hot topic detection framework with integration of image and short text information from twitter</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>9225</fpage>&#x02013;<lpage>9231</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2018.2886366</pub-id></citation>
</ref>
<ref id="B145">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>M.</given-names></name> <name><surname>Hwa</surname> <given-names>R.</given-names></name> <name><surname>Kovashka</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Equal but not the same: Understanding the implicit relationship between persuasive images and text,&#x0201D;</article-title> in <source>British Machine Vision Conference 2018, BMVC 2018</source> (<publisher-loc>Newcastle, UK</publisher-loc>: <publisher-name>BMVA Press</publisher-name>) <volume>8</volume>.</citation>
</ref>
<ref id="B146">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Schneider</surname> <given-names>J. G.</given-names></name> <name><surname>Dubrawski</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Learning the semantic correlation: An alternative way to gain from unlabeled text,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems 21, Proceedings of the Twenty-Second Annual Conference on Neural Information Processing Systems</source> (<publisher-loc>Vancouver, British Columbia, Canada</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>) <fpage>1945</fpage>&#x02013;<lpage>1952</lpage>.</citation>
</ref>
<ref id="B147">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Agrawala</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>Adding conditional control to text-to-image diffusion models</article-title>. <source>arXiv [Preprint].arXiv: 2302.05543</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2302.05543</pub-id></citation>
</ref>
<ref id="B148">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhen</surname> <given-names>L.</given-names></name> <name><surname>Hu</surname> <given-names>P.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Peng</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Deep supervised cross-modal retrieval,&#x0201D;</article-title> in <source>IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019</source> (<publisher-loc>Long Beach, CA, USA</publisher-loc>: <publisher-name>Computer Vision Foundation/IEEE</publisher-name>) <fpage>10394</fpage>&#x02013;<lpage>10403</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01064</pub-id></citation>
</ref>
<ref id="B149">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Lapedriza</surname> <given-names>&#x000C0;.</given-names></name> <name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Places: A 10 million image database for scene recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>40</volume>, <fpage>1452</fpage>&#x02013;<lpage>1464</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2723009</pub-id><pub-id pub-id-type="pmid">28692961</pub-id></citation></ref>
<ref id="B150">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Geo-location inference on news articles via multimodal plsa,&#x0201D;</article-title> in <source>Proceedings of the 20th ACM Multimedia Conference, MM&#x00027;12</source> (<publisher-loc>Nara, Japan</publisher-loc>: <publisher-name>ACM</publisher-name>) <fpage>741</fpage>&#x02013;<lpage>744</lpage>. <pub-id pub-id-type="doi">10.1145/2393347.2396301</pub-id></citation>
</ref>
<ref id="B151">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Ye</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Webface260m: A benchmark unveiling the power of million-scale deep face recognition,&#x0201D;</article-title> in <source>IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021</source> (Computer Vision Foundation/IEEE) <fpage>10492</fpage>&#x02013;<lpage>10502</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01035</pub-id></citation>
</ref>
</ref-list>
</back>
</article>
