<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="methods-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2024.1365508</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A step-by-step method for cultural annotation by LLMs</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Dubourg</surname> <given-names>Edgar</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1357218/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Thouzeau</surname> <given-names>Valentin</given-names></name>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Baumard</surname> <given-names>Nicolas</given-names></name>
<uri xlink:href="https://loop.frontiersin.org/people/67085/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>D&#x00E9;partement d&#x2019;&#x00E9;tudes cognitives, Institut Jean Nicod, &#x00C9;cole normale sup&#x00E9;rieure, Universit&#x00E9; PSL</institution>, <addr-line>Paris</addr-line>, <country>France</country></aff>
<author-notes>
<fn fn-type="edited-by" id="fn0001">
<p>Edited by: Yunhyong Kim, University of Glasgow, United Kingdom</p>
</fn>
<fn fn-type="edited-by" id="fn0002">
<p>Reviewed by: Kate Simpson, The University of Sheffield, United Kingdom</p>
<p>Vladimir Adamovich Golovko, John Paul II University in Bia&#x0142;a Podlaska, Poland</p>
</fn>
<corresp id="c001">&#x002A;Correspondence: Edgar Dubourg, <email>edgar.dubourg@gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>05</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>7</volume>
<elocation-id>1365508</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>01</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>05</day>
<month>04</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2024 Dubourg, Thouzeau and Baumard.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Dubourg, Thouzeau and Baumard</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Building on the growing body of research highlighting the capabilities of Large Language Models (LLMs) like Generative Pre-trained Transformers (GPT), this paper presents a structured pipeline for the annotation of cultural (big) data through such LLMs, offering a detailed methodology for leveraging GPT&#x2019;s computational abilities. Our approach provides researchers across various fields with a method for efficient and scalable analysis of cultural phenomena, showcasing the potential of LLMs in the empirical study of human cultures. LLMs proficiency in processing and interpreting complex data finds relevance in tasks such as annotating descriptions of non-industrial societies, measuring the importance of specific themes in stories, or evaluating psychological constructs in texts across societies or historical periods. These applications demonstrate the model&#x2019;s versatility in serving disciplines like cultural anthropology, cultural psychology, cultural history, and cultural sciences at large.</p>
</abstract>
<kwd-group>
<kwd>automatic annotation</kwd>
<kwd>human cultures</kwd>
<kwd>large language models</kwd>
<kwd>annotation loop</kwd>
<kwd>tutorial</kwd>
</kwd-group>
<counts>
<fig-count count="1"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="57"/>
<page-count count="10"/>
<word-count count="8109"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Machine Learning and Artificial Intelligence</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1">
<label>1</label>
<title>Introduction</title>
<p>The study of human cultures has always presented a formidable challenge to researchers aiming for a scientific and empirical approach (e.g., <xref ref-type="bibr" rid="ref28">Gottschall, 2008</xref>; <xref ref-type="bibr" rid="ref37">Moretti, 2014</xref>). This challenge arises from the volume and diversity of data that needs to be handled, processed, and analyzed consistently. This has become even more evident since the emergence of the digital age (<xref ref-type="bibr" rid="ref2">Acerbi, 2020</xref>). Cultural data is often vast, but also scattered and heterogeneous, making it a daunting task to gather and interpret it meaningfully.</p>
<p>Historically, the computational humanities have employed several methods for large-scale data collection and analysis. However, these methodologies have been shown to be inherently limited. For instance, participatory manual collection involves direct data gathering from individuals through surveys and observations, but it is considered costly, time-consuming, and limited in scale (<xref ref-type="bibr" rid="ref53">Wang et al., 2021</xref>). Another approach is interrogating pre-existing databases which may contain historical records and artifact descriptions, such as IMDb for movies (e.g., <xref ref-type="bibr" rid="ref49">Sreenivasan, 2013</xref>; <xref ref-type="bibr" rid="ref11">Canet, 2016</xref>; <xref ref-type="bibr" rid="ref9001">Dubourg et al., 2023</xref>) or HRAF for anthropological texts (e.g., <xref ref-type="bibr" rid="ref25">Garfield et al., 2016</xref>; <xref ref-type="bibr" rid="ref8">Boyer, 2020</xref>; <xref ref-type="bibr" rid="ref48">Singh, 2021</xref>). While these databases offer a wealth of information, they often lack uniformity and consistency, making it difficult to use them to compare different cultures and time periods. For instance, we cannot straightforwardly use the Science Fiction tag of Wikipedia to track the evolution of the number of Science Fiction works over time, as this category emerged quite late in history, even though it may be applicable to earlier works. Additionally, methods like embedding and bag-of-words have been used to analyze textual descriptions, converting text into numerical vectors for large dataset analysis (e.g., <xref ref-type="bibr" rid="ref9003">Martins and Baumard, 2020</xref>). Despite their utility, these computational techniques are hard to implement and sometimes fall short in capturing the contextual understanding necessary for classifying or rating cultural items.</p>
<p>In line with many other researchers, we believe that recent advancements in language models pre-trained on massive amounts of textual information have heralded a new era in the scientific study of culture (see <xref ref-type="bibr" rid="ref3">Bail, 2023</xref>, for a review; see <xref ref-type="bibr" rid="ref6">Binz and Schulz, 2023</xref>). These LLMs, like Generative Pre-trained Transformers (GPT), appear for the first time to be capable of annotating and parsing cultural data on a vast scale, and at an abstract level, by leveraging the contextual understanding inherent in these models (<xref ref-type="bibr" rid="ref10">Brown et al., 2020</xref>; <xref ref-type="bibr" rid="ref29">Grossmann et al., 2023</xref>). Unlike static online databases, GPT cross-references and synthesizes a wide range of texts, creating uniform annotations by linking information not explicitly connected in its training data. While LLMs definitively cannot replace human expertise in all aspects of scientific fields interested in human culture, they can replace cultural data annotation tasks, with multiple advantages for research, including: (1) Cost-effectiveness&#x2014;GPT can annotate tens of thousands of cultural descriptions within a few hours. (2) Uniformity&#x2014;GPT facilitates a consistent annotation process capable of handling descriptions or titles across various historical periods, societies, and media types. (3) Objectivity&#x2014;compared to human annotations for similar tasks, GPT&#x2019;s approach is less dependent upon idiosyncrasies (see <xref ref-type="bibr" rid="ref33">Kjeldgaard-Christiansen, 2024</xref>; see Section 3).</p>
<p>This paper primarily serves as a practical guide for their application. We introduce a detailed, step-by-step pipeline for utilizing LLMs in empirical studies. This guide is designed to provide practical insights for various applications of such automatic annotation, including annotating Human Relations Area Files (HRAF) descriptive accounts of non-industrial societies, generating or annotating descriptions of cultural items such as novels, video games or technological patents, analyzing folklore narratives, or extracting thematic elements from human-generated texts. Note that we strongly advocate for pairing LLM methods with other more established research techniques in all studies where it is possible, enabling case-by-case convergence testing and facilitating future meta-analyses. In all, this paper does not seek to debate whether LLMs can (<xref ref-type="bibr" rid="ref1">Abdurahman et al., 2023</xref>) or should (<xref ref-type="bibr" rid="ref16">Crockett and Messeri, 2023</xref>) be used in science (see also <xref ref-type="bibr" rid="ref9">Brinkmann et al., 2023</xref>); instead, it provides a concrete how-to guide for applying these models to large cultural datasets.</p>
</sec>
<sec id="sec2">
<label>2</label>
<title>Method for automatic annotation</title>
<p>Throughout this guide, specific R code snippets are included in the text. These can be directly copied and pasted into your R programming console. This script can be downloaded in an R format here: <ext-link xlink:href="https://osf.io/3q6zb/" ext-link-type="uri">https://osf.io/3q6zb/</ext-link>. Alongside the code, we provide explanations for each step and decision in the methodology. For those who prefer to see an application of the methodology before diving into the details, refer to Section 2.7. This section presents a practical example that demonstrates how the methodology can be applied to a real-world research project. It serves as a reference point to contextualize the steps discussed throughout the tutorial (see <xref ref-type="fig" rid="fig1">Figure 1</xref>).</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Schematic representation of all the steps for the cultural data annotation process.</p>
</caption>
<graphic xlink:href="frai-07-1365508-g001.tif"/>
</fig>
<sec id="sec3">
<label>2.1</label>
<title>Preliminary setup</title>
<p>In this methodology, we have selected R and GPT-4 based on their widespread use in relevant fields (see <ext-link xlink:href="https://rstudio-education.github.io/hopr/starting.html" ext-link-type="uri">https://rstudio-education.github.io/hopr/starting.html</ext-link> to install R and RStudio). However, note that this approach can be adapted to other programming languages and LLMs. Python is a suitable alternative for R, and models like Bard or LLaMA can be used instead of GPT-4 (with some adjustments in the code). These alternatives can be chosen based on the researcher&#x2019;s familiarity and the specific requirements of their project. To start, the &#x2018;httr&#x2019; and &#x2018;tidyverse&#x2019; packages in R are essential for interfacing with web APIs, particularly for accessing the OpenAI platform and GPT-4. Install this package using the following code:<preformat>
             install.packages(&#x2018;httr&#x2019;)
             library(httr)
             install.packages(&#x2018;tidyverse&#x2019;)
             library(tidyverse)</preformat>
</p>
<p>In addition to the programming setup, a premium OpenAI account is required to access GPT&#x2019;s capabilities via its API. This account provides an API key, a unique identifier necessary to authenticate and make requests to the OpenAI services. To obtain this key, log in to your OpenAI account, navigate to the &#x2018;Account&#x2019; section, and then to the &#x2018;Keys&#x2019; subsection. Here, you can find or generate your API key. This key will be crucial in the subsequent steps of the methodology, as it allows your R scripts to communicate with the GPT-4 model and send data for annotation. To effectively manage your usage and track expenses associated with using the OpenAI API, you can monitor the &#x2018;Usage&#x2019; section in your OpenAI account. As of the latest information (December 2023), using GPT-3.5-turbo is priced at $0.0080 per 1,000 tokens (a token being roughly equivalent to a word). To annotate 50 book summaries of 500 tokens, it would therefore cost $0.20. Using the updated pricing for GPT-4 Turbo, which offers an 8&#x2009;k context length, annotating 50 book summaries, each 500 tokens long, totals $1.50 at $0.06 per 1,000 output tokens.</p>
</sec>
<sec id="sec4">
<label>2.2</label>
<title>Data preparation</title>
<p>The dataset that is going to be annotated could encompass anything from novel titles or video game titles to more detailed texts like summaries of books or descriptions of social practices. Let us note that there are two distinct approaches to using GPT for data annotation, depending on the nature of your dataset:</p>
<list list-type="simple">
<list-item>
<p><italic>Title-based knowledge retrieval</italic>: In cases where the dataset consists of brief information like titles of well-known cultural artifacts, GPT can leverage its vast pre-existing knowledge. For example, when given just the title of a video game, GPT can use its comprehensive database to provide accurate ratings or categorizations (e.g., <xref ref-type="bibr" rid="ref9004">Dubourg and Chambon, 2023</xref>). This approach relies on the model&#x2019;s ability to tap into a wealth of accumulated information about widely recognized items (see <xref ref-type="bibr" rid="ref13">Chang et al., 2023</xref>, for a study about books known by GPT-4).</p>
</list-item>
<list-item>
<p><italic>Textual annotation</italic>: This approach applies when the dataset includes more detailed textual content, such as plot summaries or descriptions of artifacts. For example, if the dataset contains descriptive accounts of hunter-gatherer social behaviors, you can use GPT to assess specific cultural aspects, like the presence of third-party punishment. This method relies on GPT&#x2019;s ability to interpret and analyze the given text.</p>
</list-item>
</list>
<p>It is crucial to understand that even in the Textual Annotation approach, GPT&#x2019;s vast knowledge plays a significant role. For instance, if you prompt GPT with a movie plot summary asking for an evaluation of a certain aspect (like the importance of the theme of love), GPT is likely to identify the movie from the summary and assess the requested feature based on its extensive understanding of the movie, beyond just the plot details provided. Therefore, the distinction between title-based knowledge retrieval and textual annotation is only significant in the extent to which it guides how you structure your dataset, how you design your prompts, and how you test the validity of the annotation (e.g., title-based knowledge retrieval relies entirely on GPT&#x2019;s internal knowledge, necessitating careful verification). However, the overall process of interfacing with GPT remains consistent regardless of the approach.</p>
<p>Your final dataset (let us call it <monospace>data</monospace>) should have a column called          <monospace>to_annotate</monospace> which will be taken as input in the GPT prompt. It could therefore be a string of titles or of detailed textual information.</p>
</sec>
<sec id="sec5">
<label>2.3</label>
<title>Annotation loop</title>
<p>The annotation process using GPT-4 is implemented through a script. Here&#x2019;s a breakdown of the script&#x2019;s steps with portions of the code included:</p>
<p>Start by setting up your OpenAI API key for authentication. Place your API key in the script:<preformat>
             my_API &#x003C;- &#x2018;PUT YOUR API HERE&#x2019;</preformat></p>
<p>Then, create a function named <monospace>hey_chatGPT</monospace> to handle sending prompts to GPT and receiving responses (adapted from: <ext-link xlink:href="https://rpubs.com/nirmal/setting_chat_gpt_R" ext-link-type="uri">https://rpubs.com/nirmal/setting_chat_gpt_R</ext-link>). The function is structured to make POST requests to the OpenAI API, using your API key for authentication. The <monospace>hey_chatGPT</monospace> function includes error handling and retry logic to ensure reliable communication with the API. If a request fails or an error occurs, the function retries up to a maximum number of times, defined in the script (here, 3).
<preformat>
            hey_chatGPT &#x003C;- function(prompt) {
              retries &#x003C;- 0
              max_retries &#x003C;- 3  # Set a maximum number of retries
              while (retries &#x003C; max_retries) {
                tryCatch({
                  chat_GPT_answer &#x003C;- POST(
                    url = "
<ext-link xlink:href="https://api.openai.com/v1/chat/completions" ext-link-type="uri">https://api.openai.com/v1/chat/completions</ext-link>",
                    add_headers(Authorization = paste("Bearer", my_API)),
                    content_type_json(),
                    encode = "json",
                    body = list(
                      model = "PUT A GPT MODEL HERE",
                      temperature = 0,
                      messages = list(
                        list(role = "system", content = "PUT THE ROLE OF GPT HERE"),
                        list(role = "user", content = prompt)
                      )
                    )
                  )

                  if (status_code(chat_GPT_answer) != 200) {
                    print(paste("API request failed with status", status_code(chat_GPT_answer)))
                    retries &#x003C;- retries + 1
                    Sys.sleep(1)  # Wait a second before retrying
                  } else {
                    result &#x003C;- content(chat_GPT_answer)$choices[[1]]$message$content
                    if (nchar(result) &#x003E; 0) {
                      return(str_trim(result))
                    } else {
                      print("Received empty result, retrying...")
                      retries &#x003C;- retries + 1
                      Sys.sleep(1)  # Wait a second before retrying
                    }
                  }
                }, error = function(e) {
                  print(paste("Error occurred:", e))
                  retries &#x003C;- retries + 1
                  Sys.sleep(1)  # Wait a second before retrying
                })
              }
              return(NA)  # Return NA if all retries failed
            }
</preformat>
</p>
<p>Note that in the function, two important parameters need to be customized according to your specific research needs. First, within the function, there is a placeholder for specifying which GPT model you intend to use, indicated by &#x201C;<monospace>PUT A GPT MODEL HERE</monospace>.&#x201D; GPT-4, being the most advanced and efficient model available, is typically the preferred choice for difficult tasks. However, it is important to note that GPT-4 might also incur higher costs. Depending on your project&#x2019;s requirements and budget, you might opt for other models like GPT-3.5 or earlier versions. Replace the placeholder with the model ID of your chosen GPT version (models&#x2019; name to find here: <ext-link xlink:href="https://platform.openai.com/docs/models" ext-link-type="uri">https://platform.openai.com/docs/models</ext-link>). Second, the role parameter, indicated by &#x201C;<monospace>PUT THE ROLE OF GPT HERE</monospace>,&#x201D; is crucial in guiding the kind of responses you expect from GPT. This role defines the nature of GPT&#x2019;s interaction in the conversation. For instance, if your project involves analyzing films, you might define GPT&#x2019;s role as a &#x201C;film expert.&#x201D; This role setting helps in aligning GPT&#x2019;s responses with the specific perspective or context required for your research. Modify this parameter to reflect the role that best fits the context of your project.</p>
<p>After customizing these parameters, prepare the dataset for annotation by creating a new column to hold GPT&#x2019;s annotations. Assign a placeholder (NA) to each entry in this new column:
<preformat>
             data$annotation &#x003C;- NA</preformat></p>
<p>Then, the following script runs a loop over each row of the <monospace>to_annotate</monospace> column in your dataset. In each iteration, it constructs a prompt by combining a predefined question or statement with the specific data point and then sends this concatenated prompt to GPT:
<preformat>
             for (i in 1:nrow(data)) {
               prompt &#x003C;- "PUT YOUR PROMPT HERE"
               to_annotate &#x003C;- data$to_annotate[i]
               concat &#x003C;- paste(prompt, to_annotate)
               result &#x003C;- hey_chatGPT(concat)
               data$annotation[i] &#x003C;- result
               }
</preformat>
</p>
<p>As each prompt is processed by GPT, the response is captured and stored in the designated annotation column. The loop prints the response for monitoring.</p>
<p>Note that, in the script, the &#x201C;<monospace>PUT YOUR PROMPT HERE</monospace>&#x201D; placeholder within the loop is where you need to insert the specific prompt that will guide GPT&#x2019;s annotation process. This prompt should be crafted to elicit the desired information from GPT (see Section 2.7).</p>
<p>When constructing the prompt for each data point, it is important to frame it based on the type of annotation approach you are using. For title-based knowledge retrieval, the prompt should end with &#x201C;<monospace>The title is:</monospace>.&#x201D; This prefix indicates to GPT-4 that what follows is a title, and it should tap into its vast knowledge base for annotation. For textual annotation, the prompt should end with &#x201C;<monospace>The description is:</monospace>.&#x201D; This tells GPT-4 that it will be analyzing a more detailed textual description. The content of the <monospace>to_annotate</monospace> column, either titles or textual descriptions, follows this initial prompt, in an iterative manner.</p>
<p>In your prompt construction, also consider the type of output you require from GPT, which can be broadly categorized into two approaches:</p>
<list list-type="simple">
<list-item>
<p><italic>Categorical approach</italic>: Here, you ask GPT to classify each item into one of several predefined categories. For example, in speculative genre labeling of movies, the prompt could be: &#x201C;<monospace>Assign each movie to either Fantasy or Science Fiction.</monospace>&#x201D; This approach is useful for sorting items into distinct groups or labels.</p>
</list-item>
<list-item>
<p><italic>Dimensional approach</italic>: Alternatively, you might ask GPT to rate an item on a particular dimension, providing a numerical score. For instance, you could ask: &#x201C;<monospace>On a scale from 0 to 10, rate the importance of the imaginary world in the story.</monospace>&#x201D; If a numerical score is requested, it is helpful to specify in the prompt that the response should end with a digit, such as: &#x201C;<monospace>The end of your response should be \Score&#x2009;=&#x2009;\, with a digit between 0 and 10, with no text, letters, or symbols after.</monospace>&#x201D; This specification aids in the extraction of numerical data for analysis (see Section 2.3).</p>
</list-item>
</list>
<p>To enhance the usefulness and interpretability of GPT&#x2019;s responses, especially in the dimensional approach, you can prompt GPT to provide a brief justification before the numerical value. This not only gives context to GPT (leading to more accurate ratings) but also allows for a qualitative assessment of GPT&#x2019;s reasoning process. For example, &#x201C;<monospace>Explain briefly why you assign this score, followed by the score itself, which should be a single digit between 0 and 10.</monospace>&#x201D; Prompting GPT to first analyze and then rate encourages its neural network to deeply process the context before quantifying, utilizing the layered design for more context-informed ratings.</p>
<p>Finally, to prevent data inaccuracies and avoid GPT &#x2018;hallucinating&#x2019; responses when lacking information (in the textual annotation approach) or knowledge (in the knowledge retrieval approach), include in the prompt: &#x201C;<monospace>If unable to annotate due to insufficient data or unrecognized titles, respond with &#x2018;NA&#x2019;.</monospace>&#x201D; This ensures more reliable annotations.</p>
<p>Incorporating these considerations in prompt engineering will help tailor GPT&#x2019;s output to your specific research needs. This whole script is therefore designed to facilitate the efficient use of GPT-4 for annotating a wide range of textual data. It is particularly suitable for researchers in cultural studies looking to leverage AI for large-scale data analysis.</p>
</sec>
<sec id="sec6">
<label>2.4</label>
<title>Cleaning output</title>
<p>After running the annotation loop, it is often necessary to clean the output from GPT, especially if it includes numerical annotations like scores or binary classifications (i.e., with the dimensional approach). This step ensures that the data is in a usable format for analysis. The script uses a regular expression pattern to identify and extract numerical values from GPT&#x2019;s output. The pattern is designed to recognize numbers (single or double digits) that appear at the end of a sentence or before a space. This is crucial when GPT outputs a mix of text and numbers, and you only need the numerical value. After identifying the numerical patterns in GPT&#x2019;s responses, the script extracts these numbers and converts them into a numeric format.<preformat>
             pattern &#x003C;- "(\\b\\d{1,2}\\b)(\\.\\s&#x002A;|$)"
          data$annotation_cleaned &#x003C;- as.numeric(
                as.integer(str_extract(data$annotation, pattern)))
</preformat>
</p>
<p>This cleaning process is especially important in studies where quantitative measures are derived from qualitative data, such as rating scales or classifications.</p>
</sec>
<sec id="sec7">
<label>2.5</label>
<title>Validity check</title>
<sec id="sec8">
<label>2.5.1</label>
<title>Internal validity check</title>
<p>Internal validity, which refers to the consistency of the results within the scope of the study, can be checked to establish the reliability of the study&#x2019;s outcomes. To assess the internal validity, the main method involves using multiple annotations with different prompts. This technique allows for checking the inter-rater agreement between various iterations of GPT annotations on the same data. For instance, annotating the same set of cultural artifacts with slightly varied prompts and comparing the consistency of GPT-4&#x2019;s responses. Let us note that, in the <monospace>hey_chatGPT</monospace> function, the temperature parameter is set to 0. A temperature of 0 means that GPT will generate the most likely response, reducing randomness, and therefore enhancing the reproducibility of the output. This setting is especially important when aiming for consistent and precise annotations across multiple runs. Note that increasing the temperature can be useful in tasks where creativity or diversity in responses is desired.</p>
</sec>
<sec id="sec9">
<label>2.5.2</label>
<title>External validity check</title>
<p>In the context of Automatic Annotation, external validity assesses whether the LLM&#x2019;s annotations accurately reflect real-world phenomena. This is important because we aim to ensure that GPT&#x2019;s interpretations or classifications are not just internally consistent but also truly representative of the cultural artifacts or behaviors they are annotating. The methods to check the external validity include:</p>
<list list-type="simple">
<list-item>
<p><italic>Random sampling for qualitative evaluation</italic>: A straightforward method is to manually review a random sample of GPT&#x2019;s annotations. This review can confirm whether the annotations align with the actual content or nature of the data points.</p>
</list-item>
<list-item>
<p><italic>Statistical comparison with manual annotation</italic>: For a random subsample, compare GPT&#x2019;s annotations with those made by human annotators. This involves statistical analysis to see how closely GPT&#x2019;s ratings or categorizations match with those done manually. A high degree of correlation would indicate good external validity.</p>
</list-item>
<list-item>
<p><italic>Statistical comparison with expected metadata</italic>: This involves checking GPT&#x2019;s annotations against existing relevant metadata. For example, in annotating movies for the presence of love, one would expect these annotations to correlate with the &#x2018;Romance&#x2019; tag in IMDb metadata. However, perfect correlation is not always expected. Here, genres are broad categories and may not precisely capture all content nuances (e.g., a movie can include a love story without being tagged as a Romance). A statistically significant association, nonetheless, would suggest that GPT&#x2019;s annotations are valid in reflecting real-world characteristics.</p>
</list-item>
<list-item>
<p><italic>Statistical comparison with other methods</italic>: Other computational linguistic techniques can be used, such as word frequency analysis, semantic embeddings, and topic modeling, to independently assess cultural data.</p>
</list-item>
<list-item>
<p><italic>Expert consultation</italic>: Engaging experts in the annotation process can provide insights into the nuances that LLMs might overlook or misinterpret. For instance, when annotating cultural artifacts or practices from various societies, experts can help ensure that the annotations respect the subtleties of those cultures.</p>
</list-item>
</list>
<p>To enhance the credibility and reproducibility of research, it is highly recommended to make the full prompts and outputs used in the study transparent. This can be achieved by including detailed appendices in publications or making the data available in public repositories.</p>
</sec>
</sec>
<sec id="sec10">
<label>2.6</label>
<title>Prompt engineering</title>
<p>Prompt engineering is a crucial preliminary step in the process of using LLMs like GPT for automatic annotation. It involves crafting the queries or instructions (prompts) that are fed to the model to elicit the most accurate and relevant responses. This step is essential because the quality and specificity of the prompts significantly influence the model&#x2019;s output. Notably, <xref ref-type="bibr" rid="ref36">Liu et al. (2022)</xref> revealed that through optimized prompt-tuning results akin to fine-tuning can be achieved. Unlike fine-tuning, which requires retraining the model on a specific dataset to adjust its parameters (and therefore demands high levels of computational resources and programming skills), prompt engineering simply involves formulating effective prompts that direct the existing model&#x2019;s capabilities. This emphasizes the potential of prompt engineering as a viable alternative to more resource-intensive model training methods. This process should be undertaken before initiating the annotation loop (Section 2.3) to ensure the model is properly guided to provide the desired output.</p>
<p>The method outlined below for prompt engineering is not the only or definitive approach; rather, it serves as a guiding framework. Researchers should feel encouraged to adapt and evolve this process based on their specific project needs.</p>
<list list-type="simple">
<list-item>
<p><italic>Trial and error with multiple prompts</italic>: Experiment with various prompts on GPT&#x2019;s playground (<ext-link xlink:href="https://platform.openai.com/playground" ext-link-type="uri">https://platform.openai.com/playground</ext-link>) or within ChatGPT (noting that ChatGPT does not allow for temperature adjustments). The goal is to find prompts that consistently yield coherent and relevant outputs.</p>
</list-item>
<list-item>
<p><italic>Adjustment of prompts</italic>: If a response from GPT seems off, ask the model to explain its rating or annotation, providing insights into its reasoning process. For further refinement, initiate a different discussion where you present GPT with the prompt, the content to be annotated, and the model&#x2019;s initial response. Explicitly point out inaccuracies or issues in the response and then ask GPT how it would reconstruct the prompt to avoid such errors. This iterative process helps in fine-tuning the prompts.</p>
</list-item>
<list-item>
<p><italic>Iterative testing and refinement</italic>: Repeat this process for about 20 cultural artifacts or descriptions that are well-known to the researchers. This familiarity allows for a better assessment of GPT&#x2019;s responses. This iterative approach helps in identifying the most effective prompt structure for your specific annotation task.</p>
</list-item>
</list>
<p>After testing multiple prompts, you can conduct a qualitative analysis of the outputs. Review the responses to determine which prompts consistently produced the most accurate and relevant annotations. Select the prompt or prompts that work best for your dataset and research objectives. Prompt engineering is a dynamic and iterative process that requires experimentation. The time invested in this stage can significantly enhance the quality of the data annotation process, leading to more reliable research findings.</p>
</sec>
<sec id="sec11">
<label>2.7</label>
<title>Example</title>
<p>To illustrate a practical application of the methodology outlined in the previous sections, we present an example from a recent study that utilized GPT for automatically annotating video game titles based on specific dimensions. This study, conducted by <xref ref-type="bibr" rid="ref9004">Dubourg and Chambon (2023)</xref>, aimed to score video games on the dimensions of the DEEP model&#x2013;Discovering, Experimenting, Expanding, and Performing.</p>
<p>The dataset comprised titles of 16,000 video games, which were to be annotated along the DEEP dimensions, with a title-based knowledge retrieval approach. The prompts were theory-driven, reflecting each of the DEEP dimensions. For instance, for the Discovering dimension, the prompt was:</p>
<disp-quote>
<p>Discovering is about using novel and innovative actions or strategies to achieve abstract goals. Discovering involves actively exploring the game world, uncovering hidden secrets, and engaging in non-linear gameplay elements. It includes the ability to undertake side quests or optional objectives that offer new insights, items, or areas to explore. You will rate a video game on a scale of 0 to 100 on this dimension. This score will reflect how well the game aligns with the characteristics and potential of this dimension. Give a single number, without text. The video game is:</p>
</disp-quote>
<p>This prompt encapsulates the theoretical underpinnings of the dimension and is used to instruct GPT to tap into its extensive knowledge base for rating.</p>
<p>The same R script and annotation loop described in Section 2.3 were employed. The script was designed to prompt GPT-3.5 with each video game title and the relevant DEEP dimension, receive the model&#x2019;s rating, and then store this rating in the dataset (for a total cost of less than $200 for 16,000 video games annotated along 4 dimensions). The output from GPT-3.5 was cleaned and processed as described in Section 2.4. This step ensured that the ratings were in a numerical format suitable for analysis.</p>
<p>Here is, as a second example, a potential prompt to measure love in movies:</p>
<disp-quote>
<p>Rate the significance of love in each movie on a scale from 0 to 10, where 0 represents a complete absence or irrelevance of love, and 10 indicates that love is central and extremely relevant. Focus exclusively on the presence and impact of love between partners, setting aside other forms of love or relationships such as familial bonds or friendships. Provide a very brief justification. You must write the score after / SCORE = / at the end, with no text nor symbol after. If you don&#x2019;t know the movie, score NA. The movie is:</p>
</disp-quote>
</sec>
<sec id="sec12">
<label>2.8</label>
<title>Further developments</title>
<sec id="sec13">
<label>2.8.1</label>
<title>Confidence intervals</title>
<p>Evaluating GPT&#x2019;s confidence in its annotations can provide additional insight into the reliability of the output. Understanding the model&#x2019;s self-assessed certainty can help gauge the robustness of the annotations and identify areas where GPT might be less reliable. Here are some methods for assessing GPT&#x2019;s confidence:</p>
<list list-type="simple">
<list-item>
<p><italic>Evaluating prior knowledge in title-based knowledge retrieval</italic>: In cases where titles are used, ask GPT to rate its knowledge about the title on a scale from 0 to 10 before providing the annotation. This can be done using a prompt like: &#x201C;<monospace>On a scale from 0 to 10, how familiar are you with this work. Please provide a number representing your level of knowledge. The work is:</monospace>&#x201D; Consider using only responses where GPT rates its knowledge level as 7 or higher for more reliable annotations, for instance.</p>
</list-item>
<list-item>
<p><italic>Rating confidence in annotations</italic>: Regardless of the approach (title-based retrieval or textual approach), after receiving an annotation, you can prompt GPT to rate its confidence in that annotation. For instance: &#x201C;<monospace>Rate your confidence in your previous answer on a scale from 0 to 10.</monospace>&#x201D; Select annotations where GPT&#x2019;s self-rated confidence is high for increased reliability.</p>
</list-item>
<list-item>
<p><italic>Direct confidence interval in dimensional approach</italic>: In scenarios where GPT is asked to provide a score on a dimension, you can also ask for a confidence interval along with the score. A prompt might include: &#x201C;<monospace>Provide a score between 0 and 10 for the following aspect, and also give a confidence interval for your score.</monospace>&#x201D; The confidence interval can be a range (e.g., 5&#x2013;7), indicating the range within which GPT believes the true score lies. This method adds a layer of probabilistic assessment to the annotations, giving a sense of the range of uncertainty in GPT&#x2019;s response.</p>
</list-item>
</list>
<p>Incorporating confidence evaluations in the annotation process adds reliability to the analysis, allowing researchers to distinguish between more and less certain annotations. This approach can refine the data selection process, leading to a more nuanced understanding of the model&#x2019;s capabilities and limitations.</p>
</sec>
<sec id="sec14">
<label>2.8.2</label>
<title>Chains of thoughts</title>
<p>Chain-of-thought prompting is a technique that has the potential to enhance the precision of annotations in complex tasks (<xref ref-type="bibr" rid="ref55">Wei et al., 2023</xref>). This method involves prompting language models like GPT to decompose tasks into intermediate steps, providing a detailed breakdown of the reasoning process. This approach not only improves the accuracy of the model&#x2019;s responses in complex reasoning scenarios but also offers interpretability by revealing how conclusions are reached. It is particularly beneficial in cultural or social science research where the data often involves multiple layers of interpretation. For instance, in a study analyzing the cultural significance of ritual practices in various societies, chain-of-thought prompting could be used to ask GPT to first identify key elements of a ritual described in a text, then analyze their symbolic meanings, and finally assign them a numerical rating on some relevant dimensions. By using chain-of-thought prompting, researchers can more reliably interpret the model&#x2019;s analysis, making it a valuable tool for analysis in cultural studies.</p>
</sec>
</sec>
</sec>
<sec id="sec15">
<label>3</label>
<title>Advantages and limits of automatic cultural annotation</title>
<sec id="sec16">
<label>3.1</label>
<title>Cost-effectiveness: the capabilities of LLMs at zero-shot knowledge intensive tasks</title>
<p>It has been argued that the ability to generate domain-specific and structured data rapidly, as well as convert structured knowledge into natural sentences, opens new possibilities for data collection (<xref ref-type="bibr" rid="ref19">Ding et al., 2023</xref>; see <xref ref-type="bibr" rid="ref1">Abdurahman et al., 2023</xref>, for a review). The use of GPT-3 for data annotation has shown promising results, with its accuracy and intercoder agreement surpassing those of human annotators in many Natural-Language Processing (NLP) tasks (<xref ref-type="bibr" rid="ref53">Wang et al., 2021</xref>; <xref ref-type="bibr" rid="ref26">Gilardi et al., 2023</xref>; <xref ref-type="bibr" rid="ref44">Rathje et al., 2023</xref>; <xref ref-type="bibr" rid="ref54">Webb et al., 2023</xref>). Even more crucially, GPT is now adept at <italic>zero-shot learning</italic> applications: it performs well without any re-training of the model (see <xref ref-type="bibr" rid="ref43">Qin et al., 2023</xref>, for an evaluation of the capacity of ChatGPT to zero-shot learning in 20 different NLP tasks; see <xref ref-type="bibr" rid="ref35">Kuzman et al., 2023</xref>, on a genre identification task; see also <xref ref-type="bibr" rid="ref41">Pei et al., 2023</xref>, for a one-shot tuning phase approach).</p>
<p>But what about tasks that require extensive background knowledge? The quality of cultural data annotation is capital, and it often depends on the expertise of the annotators. Pre-trained transformers seem capable of handling even that. For instance, GPT-3 has been employed to automatically generate textual information sheets for artworks, displaying excellent knowledge of art concepts and specific paintings (<xref ref-type="bibr" rid="ref7">Bongini et al., 2023</xref>). Crucially, this study concludes that there is no need to retrain the model to incorporate new knowledge about the cultural artifacts: &#x201C;this is possible thanks to the memorization capabilities of GPT-3, which at training time has observed millions of tokens regarding domain-specific knowledge&#x201D; (<xref ref-type="bibr" rid="ref7">Bongini et al., 2023</xref>). Let us take an example from another domain of human cultures: GPT-4 performed comparably to well-trained law student annotators in analyzing legal texts, demonstrating again its potential in tasks requiring highly specialized domain expertise (<xref ref-type="bibr" rid="ref46">Savelka, 2023</xref>; <xref ref-type="bibr" rid="ref47">Savelka et al., 2023</xref>; see also, <xref ref-type="bibr" rid="ref24">Fink et al., 2023</xref>, for knowledge extraction from lung cancer report; see <xref ref-type="bibr" rid="ref31">Hou and Ji, 2023</xref>, for cell type annotation for RNA-seq analysis; see <xref ref-type="bibr" rid="ref34">Kjell et al., 2023</xref>, for mental health assessment; see <xref ref-type="bibr" rid="ref18">Dillion et al., 2023</xref>, for moral judgements). In all, GPT appears excellent at zero-shot <italic>knowledge intensive</italic> tasks (see <xref ref-type="bibr" rid="ref56">Yang et al., 2023</xref>, for a review). These models therefore offer a highly efficient and cost-effective method for cultural data annotation.</p>
</sec>
<sec id="sec17">
<label>3.2</label>
<title>Uniformity in annotation: facilitating comparative approaches</title>
<p>There are specific tasks within related fields where LLMs can be particularly beneficial, especially when objective properties need to be extracted from large volumes of data. For instance, quantifying the presence of a theme such as love in a story is a task where objectivity is key. Humans making such assessments may be influenced by their personal interests in romantic love, leading to variability in interpretations. In contrast, LLMs can offer more objective measurements. For scientific studies of human cultures, another crucial aspect is the consistent handling of data from diverse societies, languages, or historical periods. This uniform approach is essential in ensuring analysis of various cultural practices and artifacts, enabling cross-cultural comparisons. In their extensive study across 15 datasets, including 31,789 manually annotated tweets and news headlines, <xref ref-type="bibr" rid="ref44">Rathje et al. (2023)</xref> tested GPT-3.5 and GPT-4&#x2019;s ability to accurately detect psychological constructs like sentiment, discrete emotions, and offensiveness across 12 languages, including Turkish, Indonesian, and eight African languages such as Swahili and Amharic. They found that GPT outperformed traditional dictionary-based text analysis methods in these tasks. It highlights GPT&#x2019;s capability to provide more consistent and accurate cross-cultural comparisons.</p>
<p>The advent of advanced language models like GPT-4 also presents opportunities for comparative analysis across various media forms. Traditional approaches have struggled to directly compare cultural artifacts like movies, video games, TV series, and novels, due to their distinct genres and sensory modalities&#x2014;visual and auditory in films and games, versus imaginative in literature. However, LLMs enable a more abstract level of analysis, focusing on underlying themes and psychological constructs. By processing and interpreting data at this higher level of abstraction, LLMs can identify and compare core thematic elements, such as love or imaginary worlds, regardless of the medium. It should open new avenues in psychological research studying cultural trends. For instance, an uptick in horror fiction across movies, video games, and novels may not occur uniformly due to each medium&#x2019;s ability to elicit fear. By analyzing these diverse expressions of a genre like horror throughout all cultural productions, LLMs can provide a comprehensive view of the overall appeal and evolution of recreational horror (see <xref ref-type="bibr" rid="ref14">Clasen, 2017</xref>; <xref ref-type="bibr" rid="ref57">Yang and Zhang, 2022</xref>).</p>
</sec>
<sec id="sec18">
<label>3.3</label>
<title>Annotating fuzzy concepts and loose categories: facilitating the dimensional approach</title>
<p>One difficulty with cultural annotation is that many concepts used in social and cultural sciences are often difficult to express clearly: What is &#x201C;agency&#x201D; exactly (<xref ref-type="bibr" rid="ref30">Haggard and Chambon, 2012</xref>; <xref ref-type="bibr" rid="ref12">Chambon et al., 2014</xref>)? How do you define &#x201C;puritanism&#x201D; (<xref ref-type="bibr" rid="ref9002">Fitouchi et al., 2023</xref>)? What about &#x201C;fictionality&#x201D; (<xref ref-type="bibr" rid="ref38">Nielsen et al., 2015</xref>; <xref ref-type="bibr" rid="ref39">Paige, 2020</xref>)? These concepts are fuzzy, and the categories are loose. Characters can be more or less agentic, stories can be more or less fictional, and there are different ways to define puritanism. In fact, scientists often use scales, based on multiple questions. This means that it would be hard to ask a human annotator to rate cultural items on fictionality for instance using a single definition. In these cases, LLM can offer invaluable help. For instance, one can give a LLM the questionnaire used by psychologists to measure agency (<xref ref-type="bibr" rid="ref52">Vallacher and Wegner, 1989</xref>), and prompt the LLM to rate the level of agency of characters using this questionnaire. GPT will estimate the intensity of the semantic association between the questions and the words associated with the character in a story. LLMs thus enable systematic data collection on complex and hard-to-define aspects of cultural artifacts, fostering the extraction of novel insights (e.g., <xref ref-type="bibr" rid="ref42">Piper and Toubia, 2023</xref>).</p>
<p>We posit that scientific approaches to human cultures can gain from leveraging LLMs for one last key reason: the facilitation of a dimensional approach (see Section 2.3). The standard approach in cultural sciences has often been categorical. For instance, historians have long debated whether romantic love is a Western medieval invention (<xref ref-type="bibr" rid="ref17">De Rougemont, 2016 [1939]</xref>), and whether a similar phenomenon is present or not outside the West (<xref ref-type="bibr" rid="ref27">Goody, 1998</xref>; <xref ref-type="bibr" rid="ref40">Pan, 2015</xref>). Yet, the aspects we seek to analyze, such as psychological constructs in texts or thematic elements in stories, often are not binary but exist on a continuum. These elements vary in degrees, be it in sensitivity, intensity, prevalence, or importance. Adopting a dimensional approach is more fruitful (<xref ref-type="bibr" rid="ref51">Trull and Durrett, 2005</xref>). Unlike the categorical approach, which determines whether an individual possesses a trait or not, the dimensional approach used in personality psychology, psychiatry or behavioral ecology uses a continuum to determine the various levels of a trait that an individual may possess. In this perspective, the right question is not whether people in the medieval period fell in love or not but rather quantifying and comparing the intensity of this feeling to other periods (<xref ref-type="bibr" rid="ref4">Baumard et al., 2022</xref>, <xref ref-type="bibr" rid="ref5">2023</xref>).</p>
<p>The move from categorical to dimensional analysis represents a shift in accurately measuring and understanding cultural phenomena. Recognizing this continuous nature is crucial, as it aligns more closely with the real-world diversity of cultural artifacts or social behaviors. In personality psychology, a similar transition occurred: from rigid categorical classifications to a dimensional approach with traits, which better captured the variability of people&#x2019;s patterns of thoughts and behaviors (e.g., <xref ref-type="bibr" rid="ref15">Costa and McCrae, 1992</xref>; <xref ref-type="bibr" rid="ref32">Kashdan et al., 2018</xref>; <xref ref-type="bibr" rid="ref20">Dubois et al., 2020</xref>). This change was largely enabled by methodological advancements like the emergence of Likert scales and the elaboration of factor analysis. Similarly, in cultural and social sciences, technologies like GPT arguably offer the potential to move beyond categorical thinking. For instance, elements such as love, conflict, or adventure, typically associated with specific genres in movies or literary works, can actually be present to varying degrees across a broad spectrum of works. GPT&#x2019;s capacity to quantitatively analyze these elements on a scale allows for capturing their presence in a more nuanced manner.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="sec19">
<label>4</label>
<title>Conclusion</title>
<p>The ability of GPT to process vast amounts of data with nuanced understanding can show potential in tasks ranging from annotating descriptions of non-industrial societies to extracting psychological constructs from texts, thereby serving a wide array of disciplines including anthropology, psychology, and history. This methodology can also foster interdisciplinary connections. For example, understanding cultural artifacts as cognitive fossils&#x2014;physical imprints of the psychological traits of their creators or consumers&#x2014;can bridge gaps between cultural, historical, and psychological sciences (<xref ref-type="bibr" rid="ref5">Baumard et al., 2023</xref>). The use of LLMs in Automatic Annotation of cultural data can contribute to this interdisciplinary bridge. By enabling the homogenization and comparison of diverse cultural data, LLMs provide a unified methodological ground for multiple research areas interested in human cultures.</p>
<p>Before concluding, it is important to stress again that, for now, we believe that the use of LLMs for cultural data annotation should not be seen as a standalone solution, especially in fields where operationalization and quantification is hard. Rather, it should be used with other traditional research methods or subjected to rigorous validity checks, both internal and external. Multiple studies have highlighted that LLMs can exhibit biases, often reflecting limitations in representing the diversity of personalities, opinions, beliefs, etc. (<xref ref-type="bibr" rid="ref45">Santurkar et al., 2023</xref>; see <xref ref-type="bibr" rid="ref1">Abdurahman et al., 2023</xref>, for a review). These biases stem from the data on which these models are trained, potentially impacting their ability to fully grasp and represent the vast spectrum of human experiences and perspectives. Thus, Automatic Cultural Annotation does not advocate for the complete replacement of human judgment in psychological and cultural studies with LLMs.</p>
</sec>
<sec sec-type="data-availability" id="sec20">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="sec21">
<title>Author contributions</title>
<p>ED: Methodology, Project administration, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. VT: Methodology, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. NB: Conceptualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="sec22">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This work was supported by a FrontCog funding (ANR-17-EURE-0017).</p>
</sec>
<ack>
<p>We thank Quentin Borredon, Robin Fraynet, Marius Mercier, Charlotte Touzeau, Jeanne Boll&#x00E9;e, Thomas Monnier for their valuable feedback on this manuscript, and all the participants of the &#x201C;Automatic Cultural Annotation&#x201D; workshop of the FRESH Conference, Paris, 2023.</p>
</ack>
<sec sec-type="COI-statement" id="sec23">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abdurahman</surname> <given-names>S.</given-names></name> <name><surname>Atari</surname> <given-names>M.</given-names></name> <name><surname>Karimi-Malekabadi</surname> <given-names>F.</given-names></name> <name><surname>Xue</surname> <given-names>M. J.</given-names></name> <name><surname>Trager</surname> <given-names>J.</given-names></name> <name><surname>Park</surname> <given-names>P. S.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Perils and opportunities in using large language models in psychological research</article-title>. <source>OSF Preprints</source>. doi: <pub-id pub-id-type="doi">10.31219/osf.io/tg79n</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Acerbi</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <source>Cultural evolution in the digital age (first edition)</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bail</surname> <given-names>C. A.</given-names></name></person-group> (<year>2023</year>). <article-title>Can generative AI improve social science?</article-title> <source>Soc ArXiv</source>. doi: <pub-id pub-id-type="doi">10.31235/osf.io/rwtzs</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baumard</surname> <given-names>N.</given-names></name> <name><surname>Huillery</surname> <given-names>E.</given-names></name> <name><surname>Hyafil</surname> <given-names>A.</given-names></name> <name><surname>Safra</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <article-title>The cultural evolution of love in literary history</article-title>. <source>Nat. Hum. Behav.</source> <volume>6</volume>, <fpage>506</fpage>&#x2013;<lpage>522</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41562-022-01292-z</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baumard</surname> <given-names>N.</given-names></name> <name><surname>Safra</surname> <given-names>L.</given-names></name> <name><surname>Martins</surname> <given-names>M. D. J. D.</given-names></name> <name><surname>Chevallier</surname> <given-names>C.</given-names></name></person-group> (<year>2023</year>). <article-title>Cognitive fossils: using cultural artifacts to reconstruct psychological changes throughout history</article-title>. <source>Trends Cogn. Sci.</source> <volume>28</volume>, <fpage>172</fpage>&#x2013;<lpage>186</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tics.2023.10.001</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Binz</surname> <given-names>M.</given-names></name> <name><surname>Schulz</surname> <given-names>E.</given-names></name></person-group> (<year>2023</year>). <article-title>Using cognitive psychology to understand GPT-3</article-title>. <source>Proc. Natl. Acad. Sci. USA</source> <volume>120</volume> Scopus:<fpage>e2218523120</fpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.2218523120</pub-id>, PMID: <pub-id pub-id-type="pmid">36730192</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Bongini</surname> <given-names>P.</given-names></name> <name><surname>Becattini</surname> <given-names>F.</given-names></name> <name><surname>Del Bimbo</surname> <given-names>A.</given-names></name></person-group> (<year>2023</year>). &#x201C;<article-title>Is GPT-3 all you need for visual question answering in cultural heritage?</article-title>&#x201D; in <source>Computer vision &#x2013; ECCV 2022 workshops</source>. eds. <person-group person-group-type="editor"><name><surname>Karlinsky</surname> <given-names>L.</given-names></name> <name><surname>Michaeli</surname> <given-names>T.</given-names></name> <name><surname>Nishino</surname> <given-names>K.</given-names></name></person-group>, vol. <volume>13801</volume> (<publisher-loc>Switzerland</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>), <fpage>268</fpage>&#x2013;<lpage>281</lpage>.</citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boyer</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Informal religious activity outside hegemonic religions: wild traditions and their relevance to evolutionary models</article-title>. <source>Relig. Brain Behav.</source> <volume>10</volume>, <fpage>459</fpage>&#x2013;<lpage>472</lpage>. doi: <pub-id pub-id-type="doi">10.1080/2153599X.2019.1678518</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brinkmann</surname> <given-names>L.</given-names></name> <name><surname>Baumann</surname> <given-names>F.</given-names></name> <name><surname>Bonnefon</surname> <given-names>J.-F.</given-names></name> <name><surname>Derex</surname> <given-names>M.</given-names></name> <name><surname>M&#x00FC;ller</surname> <given-names>T. F.</given-names></name> <name><surname>Nussberger</surname> <given-names>A.-M.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Machine culture</article-title>. <source>Nat. Hum. Behav.</source> <volume>7</volume>, <fpage>1855</fpage>&#x2013;<lpage>1868</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41562-023-01742-2</pub-id>, PMID: <pub-id pub-id-type="pmid">37985914</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>T. B.</given-names></name> <name><surname>Mann</surname> <given-names>B.</given-names></name> <name><surname>Ryder</surname> <given-names>N.</given-names></name> <name><surname>Subbiah</surname> <given-names>M.</given-names></name> <name><surname>Kaplan</surname> <given-names>J.</given-names></name> <name><surname>Dhariwal</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Language models are few-shot learners (arXiv:2005.14165)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2005.14165</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Canet</surname> <given-names>F.</given-names></name></person-group> (<year>2016</year>). <article-title>Quantitative approaches for evaluating the influence of films using the IMDb database</article-title>. <source>Commun. Soc.</source> <volume>29</volume>, <fpage>151</fpage>&#x2013;<lpage>172</lpage>. doi: <pub-id pub-id-type="doi">10.15581/003.29.2.151-172</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chambon</surname> <given-names>V.</given-names></name> <name><surname>Sidarus</surname> <given-names>N.</given-names></name> <name><surname>Haggard</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <article-title>From action intentions to action effects: how does the sense of agency come about?</article-title> <source>Front. Hum. Neurosci.</source> <volume>8</volume>. doi: <pub-id pub-id-type="doi">10.3389/fnhum.2014.00320</pub-id>, PMID: <pub-id pub-id-type="pmid">24860486</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>K. K.</given-names></name> <name><surname>Cramer</surname> <given-names>M.</given-names></name> <name><surname>Soni</surname> <given-names>S.</given-names></name> <name><surname>Bamman</surname> <given-names>D.</given-names></name></person-group> (<year>2023</year>). <article-title>Speak, memory: an archaeology of books known to ChatGPT/GPT-4 (arXiv:2305.00118)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2305.00118</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Clasen</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <source>Why horror seduces</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Costa</surname> <given-names>P. T.</given-names></name> <name><surname>McCrae</surname> <given-names>R. R.</given-names></name></person-group> (<year>1992</year>). <article-title>Four ways five factors are basic</article-title>. <source>Personal. Individ. Differ.</source> <volume>13</volume>, <fpage>653</fpage>&#x2013;<lpage>665</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0191-8869(92)90236-I</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crockett</surname> <given-names>M.</given-names></name> <name><surname>Messeri</surname> <given-names>L.</given-names></name></person-group> (<year>2023</year>). <article-title>Should large language models replace human participants? [preprint]</article-title>. <source>PsyArXiv</source>. doi: <pub-id pub-id-type="doi">10.31234/osf.io/4zdx9</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="other"><person-group person-group-type="author"><name><surname>De Rougemont</surname> <given-names>D.</given-names></name></person-group> <year>2016</year> (1939). <source>L&#x2019;amour et l&#x2019;Occident</source>, <fpage>10</fpage>&#x2013;<lpage>18</lpage>.</citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dillion</surname> <given-names>D.</given-names></name> <name><surname>Tandon</surname> <given-names>N.</given-names></name> <name><surname>Gu</surname> <given-names>Y.</given-names></name> <name><surname>Gray</surname> <given-names>K.</given-names></name></person-group> (<year>2023</year>). <article-title>Can AI language models replace human participants?</article-title> <source>Trends Cogn. Sci.</source> <volume>27</volume>, <fpage>597</fpage>&#x2013;<lpage>600</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tics.2023.04.008</pub-id>, PMID: <pub-id pub-id-type="pmid">37173156</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ding</surname> <given-names>B.</given-names></name> <name><surname>Qin</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Chia</surname> <given-names>Y. K.</given-names></name> <name><surname>Joty</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Is GPT-3 a good data annotator? (arXiv:2212.10450)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2212.10450</pub-id></citation></ref>
<ref id="ref9004"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dubourg</surname> <given-names>E.</given-names></name> <name><surname>Chambon</surname> <given-names>V</given-names></name></person-group>. (<year>2023</year>). <article-title>DEEP: A model of gaming preferences informed by the hierarchical nature of goal-oriented cognition</article-title>. <source>In Review</source>.</citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dubois</surname> <given-names>J.</given-names></name> <name><surname>Eberhardt</surname> <given-names>F.</given-names></name> <name><surname>Paul</surname> <given-names>L. K.</given-names></name> <name><surname>Adolphs</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Personality beyond taxonomy</article-title>. <source>Nat. Hum. Behav.</source> <volume>4</volume>, <fpage>1110</fpage>&#x2013;<lpage>1117</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41562-020-00989-3</pub-id></citation></ref>
<ref id="ref9001"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dubourg</surname> <given-names>E.</given-names></name> <name><surname>Thouzeau</surname> <given-names>V.</given-names></name> <name><surname>de Dampierre</surname> <given-names>C.</given-names></name> <name><surname>Mogoutov</surname> <given-names>A.</given-names></name> <name><surname>Baumard</surname> <given-names>N.</given-names></name></person-group> (<year>2023</year>). <article-title>Exploratory preferences explain the human fascination for imaginary worlds</article-title>. <source>Scientific Reports</source>, <volume>13</volume>. doi: <pub-id pub-id-type="doi">10.31234/osf.io/d9uqs</pub-id></citation></ref>
<ref id="ref9002"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Fitouchi</surname> <given-names>L.</given-names></name> <name><surname>Andr&#x00E9;</surname> <given-names>J.-B.</given-names></name> <name><surname>Baumard</surname> <given-names>N.</given-names></name></person-group> (<year>2023</year>). <source>Moral disciplining: The cognitive and evolutionary foundations of puritanical morality</source>. <publisher-name>Behavioral and Brain Sciences</publisher-name>. Available at: <ext-link xlink:href="https://osf.io/2stcv" ext-link-type="uri">https://osf.io/2stcv</ext-link>.</citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fink</surname> <given-names>M. A.</given-names></name> <name><surname>Bischoff</surname> <given-names>A.</given-names></name> <name><surname>Fink</surname> <given-names>C. A.</given-names></name> <name><surname>Moll</surname> <given-names>M.</given-names></name> <name><surname>Kroschke</surname> <given-names>J.</given-names></name> <name><surname>Dulz</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Potential of ChatGPT and GPT-4 for data mining of free-text CT reports on lung cancer</article-title>. <source>Radiology</source> <volume>308</volume>:<fpage>e231362</fpage>. doi: <pub-id pub-id-type="doi">10.1148/radiol.231362</pub-id>, PMID: <pub-id pub-id-type="pmid">37724963</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Garfield</surname> <given-names>Z. H.</given-names></name> <name><surname>Garfield</surname> <given-names>M. J.</given-names></name> <name><surname>Hewlett</surname> <given-names>B. S.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>A cross-cultural analysis of hunter-gatherer social learning</article-title>&#x201D; in <source>Social learning and innovation in contemporary hunter-gatherers: Evolutionary and ethnographic perspectives</source>. eds. <person-group person-group-type="editor"><name><surname>Terashima</surname> <given-names>H.</given-names></name> <name><surname>Hewlett</surname> <given-names>B. S.</given-names></name></person-group> (<publisher-loc>Japan</publisher-loc>: <publisher-name>Springer Japan</publisher-name>), <fpage>19</fpage>&#x2013;<lpage>34</lpage>.</citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gilardi</surname> <given-names>F.</given-names></name> <name><surname>Alizadeh</surname> <given-names>M.</given-names></name> <name><surname>Kubli</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>ChatGPT outperforms crowd workers for text-annotation tasks</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>120</volume>:<fpage>e2305016120</fpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.2305016120</pub-id>, PMID: <pub-id pub-id-type="pmid">37463210</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Goody</surname> <given-names>J.</given-names></name></person-group> (<year>1998</year>). <source>Food and love: A cultural history of east and west</source> <publisher-name>Verso</publisher-name>.</citation></ref>
<ref id="ref28"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Gottschall</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). &#x201C;<article-title>On method</article-title>&#x201D; in <source>Literature, science, and a new humanities</source> (<publisher-loc>Springer, New York</publisher-loc>: <publisher-name>Springer</publisher-name>).</citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grossmann</surname> <given-names>I.</given-names></name> <name><surname>Feinberg</surname> <given-names>M.</given-names></name> <name><surname>Parker</surname> <given-names>D. C.</given-names></name> <name><surname>Christakis</surname> <given-names>N. A.</given-names></name> <name><surname>Tetlock</surname> <given-names>P. E.</given-names></name> <name><surname>Cunningham</surname> <given-names>W. A.</given-names></name></person-group> (<year>2023</year>). <article-title>AI and the transformation of social science research</article-title>. <source>Science</source> <volume>380</volume>, <fpage>1108</fpage>&#x2013;<lpage>1109</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.adi1778</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haggard</surname> <given-names>P.</given-names></name> <name><surname>Chambon</surname> <given-names>V.</given-names></name></person-group> (<year>2012</year>). <article-title>Sense of agency</article-title>. <source>Curr. Biol.</source> <volume>22</volume>, <fpage>R390</fpage>&#x2013;<lpage>R392</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cub.2012.02.040</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>W.</given-names></name> <name><surname>Ji</surname> <given-names>Z.</given-names></name></person-group> (<year>2023</year>). Reference-free and cost-effective automated cell type annotation with GPT-4 in single-cell RNA-seq analysis.</citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kashdan</surname> <given-names>T. B.</given-names></name> <name><surname>Stiksma</surname> <given-names>M. C.</given-names></name> <name><surname>Disabato</surname> <given-names>D. J.</given-names></name> <name><surname>McKnight</surname> <given-names>P. E.</given-names></name> <name><surname>Bekier</surname> <given-names>J.</given-names></name> <name><surname>Kaji</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>The five-dimensional curiosity scale: capturing the bandwidth of curiosity and identifying four unique subgroups of curious people</article-title>. <source>J. Res. Pers.</source> <volume>73</volume>, <fpage>130</fpage>&#x2013;<lpage>149</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jrp.2017.11.011</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kjeldgaard-Christiansen</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>What science can&#x2019;t know: on scientific objectivity and the human subject</article-title>. <source>Poetics Today</source> <volume>45</volume>, <fpage>1</fpage>&#x2013;<lpage>16</lpage>. doi: <pub-id pub-id-type="doi">10.1215/03335372-10938579</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kjell</surname> <given-names>O. N. E.</given-names></name> <name><surname>Kjell</surname> <given-names>K.</given-names></name> <name><surname>Schwartz</surname> <given-names>H. A.</given-names></name></person-group> (<year>2023</year>). <article-title>Beyond rating scales: with care for targeted validation large language models are poised for psychological assessment [preprint]</article-title>. <source>PsyArXiv</source>. doi: <pub-id pub-id-type="doi">10.31234/osf.io/yfd8g</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuzman</surname> <given-names>T.</given-names></name> <name><surname>Mozeti&#x010D;</surname> <given-names>I.</given-names></name> <name><surname>Ljube&#x0161;i&#x0107;</surname> <given-names>N.</given-names></name></person-group> (<year>2023</year>). <article-title>ChatGPT: beginning of an end of manual linguistic data annotation? Use case of automatic genre identification (arXiv:2303.03953)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2303.03953</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Ji</surname> <given-names>K.</given-names></name> <name><surname>Fu</surname> <given-names>Y.</given-names></name> <name><surname>Tam</surname> <given-names>W. L.</given-names></name> <name><surname>Du</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>P-tuning v2: prompt tuning can be comparable to fine-tuning universally across scales and tasks (arXiv:2110.07602)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2110.07602</pub-id></citation></ref>
<ref id="ref9003"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martins</surname> <given-names>M. de J. D.</given-names></name> <name><surname>Baumard</surname> <given-names>N</given-names></name></person-group>. (<year>2020</year>). <article-title>The rise of prosociality in fiction preceded democratic revolutions in Early Modern Europe</article-title>. <source>Proceedings of the National Academy of Sciences</source>. <volume>117</volume>:<fpage>202009571</fpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.2009571117</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moretti</surname> <given-names>F.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x201C;Operationalizing&#x201D;: Or, the function of measurement in modern literary theory</article-title>. <source>J. Engl. Lang. Lit</source>. <volume>60</volume>, <fpage>3</fpage>&#x2013;<lpage>19</lpage>. doi: <pub-id pub-id-type="doi">10.15794/JELL.2014.60.1.001</pub-id></citation></ref>
<ref id="ref38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nielsen</surname> <given-names>H. S.</given-names></name> <name><surname>Phelan</surname> <given-names>J.</given-names></name> <name><surname>Walsh</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>Ten theses about fictionality</article-title>. <source>Narrative</source> <volume>23</volume>, <fpage>61</fpage>&#x2013;<lpage>73</lpage>. doi: <pub-id pub-id-type="doi">10.1353/nar.2015.0005</pub-id></citation></ref>
<ref id="ref39"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Paige</surname> <given-names>N. D.</given-names></name></person-group> (<year>2020</year>). <source>Technologies of the novel</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="ref40"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <source>When true love came to China</source>. <publisher-loc>Hong Kong</publisher-loc>: <publisher-name>Hong Kong Univ. Press</publisher-name>.</citation></ref>
<ref id="ref41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pei</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name></person-group> (<year>2023</year>). <article-title>GPT self-supervision for a better data annotator (arXiv:2306.04349)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2306.04349</pub-id></citation></ref>
<ref id="ref42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Piper</surname> <given-names>A.</given-names></name> <name><surname>Toubia</surname> <given-names>O.</given-names></name></person-group> (<year>2023</year>). <article-title>A quantitative study of non-linearity in storytelling</article-title>. <source>Poetics</source> <volume>98</volume>:<fpage>101793</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.poetic.2023.101793</pub-id></citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qin</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Yasunaga</surname> <given-names>M.</given-names></name> <name><surname>Yang</surname> <given-names>D.</given-names></name></person-group> (<year>2023</year>). <article-title>Is ChatGPT a general-purpose natural language processing task solver? (arXiv:2302.06476)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2302.06476</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rathje</surname> <given-names>S.</given-names></name> <name><surname>Mirea</surname> <given-names>D.-M.</given-names></name> <name><surname>Sucholutsky</surname> <given-names>I.</given-names></name> <name><surname>Marjieh</surname> <given-names>R.</given-names></name> <name><surname>Robertson</surname> <given-names>C.</given-names></name> <name><surname>Van Bavel</surname> <given-names>J. J.</given-names></name></person-group> (<year>2023</year>). <article-title>GPT is an effective tool for multilingual psychological text analysis [preprint]</article-title>. <source>PsyArXiv</source>. doi: <pub-id pub-id-type="doi">10.31234/osf.io/sekf5</pub-id></citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Santurkar</surname> <given-names>S.</given-names></name> <name><surname>Durmus</surname> <given-names>E.</given-names></name> <name><surname>Ladhak</surname> <given-names>F.</given-names></name> <name><surname>Lee</surname> <given-names>C.</given-names></name> <name><surname>Liang</surname> <given-names>P.</given-names></name> <name><surname>Hashimoto</surname> <given-names>T.</given-names></name></person-group> (<year>2023</year>). <article-title>Whose opinions do language models reflect? (arXiv:2303.17548) Proceedings of the 40th International Conference on Machine Learning, 202, 29971-30004</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2303.17548</pub-id></citation></ref>
<ref id="ref46"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Savelka</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Unlocking practical applications in legal domain: evaluation of GPT for zero-shot semantic annotation of legal texts</article-title>. In: <source>Proceedings of the Nineteenth International Conference on Artificial Intelligence and Law</source>, pp. <fpage>447</fpage>&#x2013;<lpage>451</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3594536.3595161</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Savelka</surname> <given-names>J.</given-names></name> <name><surname>Ashley</surname> <given-names>K. D.</given-names></name> <name><surname>Gray</surname> <given-names>M. A.</given-names></name> <name><surname>Westermann</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name></person-group> (<year>2023</year>). <article-title>Can GPT-4 support analysis of textual data in tasks requiring highly specialized domain expertise?</article-title> In: <source>Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education, Vol. 1</source>, pp. <fpage>117</fpage>&#x2013;<lpage>123</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3587102.3588792</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Magic, explanations, and evil: on the origins and design of witches and sorcerers</article-title>. <source>Curr. Anthropol.</source> <volume>62</volume>, <fpage>2</fpage>&#x2013;<lpage>29</lpage>. doi: <pub-id pub-id-type="doi">10.31235/osf.io/pbwc7</pub-id></citation></ref>
<ref id="ref49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sreenivasan</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Quantitative analysis of the evolution of novelty in cinema through crowdsourced keywords</article-title>. <source>Sci. Rep.</source> <volume>3</volume>:<fpage>2758</fpage>. doi: <pub-id pub-id-type="doi">10.1038/srep02758</pub-id>, PMID: <pub-id pub-id-type="pmid">24067890</pub-id></citation></ref>
<ref id="ref51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Trull</surname> <given-names>T. J.</given-names></name> <name><surname>Durrett</surname> <given-names>C. A.</given-names></name></person-group> (<year>2005</year>). <article-title>Categorical and dimensional models of personality disorder</article-title>. <source>Annu. Rev. Clin. Psychol.</source> <volume>1</volume>, <fpage>355</fpage>&#x2013;<lpage>380</lpage>. doi: <pub-id pub-id-type="doi">10.1146/annurev.clinpsy.1.102803.144009</pub-id></citation></ref>
<ref id="ref52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vallacher</surname> <given-names>R. R.</given-names></name> <name><surname>Wegner</surname> <given-names>D. M.</given-names></name></person-group> (<year>1989</year>). <article-title>Levels of personal agency: individual variation in action identification</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>57</volume>, <fpage>660</fpage>&#x2013;<lpage>671</lpage>. doi: <pub-id pub-id-type="doi">10.1037/0022-3514.57.4.660</pub-id></citation></ref>
<ref id="ref53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Zhu</surname> <given-names>C.</given-names></name> <name><surname>Zeng</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Want to reduce labeling cost? GPT-3 can help (arXiv:2108.13487)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2108.13487</pub-id></citation></ref>
<ref id="ref54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Webb</surname> <given-names>T.</given-names></name> <name><surname>Holyoak</surname> <given-names>K. J.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name></person-group> (<year>2023</year>). <article-title>Emergent analogical reasoning in large language models (arXiv:2212.09196)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2212.09196</pub-id></citation></ref>
<ref id="ref55"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wei</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Schuurmans</surname> <given-names>D.</given-names></name> <name><surname>Bosma</surname> <given-names>M.</given-names></name> <name><surname>Ichter</surname> <given-names>B.</given-names></name> <name><surname>Xia</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2023</year>). <source>Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 2022.</source></citation></ref>
<ref id="ref56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Jin</surname> <given-names>H.</given-names></name> <name><surname>Tang</surname> <given-names>R.</given-names></name> <name><surname>Han</surname> <given-names>X.</given-names></name> <name><surname>Feng</surname> <given-names>Q.</given-names></name> <name><surname>Jiang</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Harnessing the power of LLMs in practice: a survey on ChatGPT and beyond (arXiv:2304.13712)</article-title>. <source>arXiv</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2304.13712</pub-id></citation></ref>
<ref id="ref57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>K.</given-names></name></person-group> (<year>2022</year>). <article-title>How resource scarcity influences the preference for counterhedonic consumption</article-title>. <source>J. Consum. Res.</source> <volume>48</volume>, <fpage>904</fpage>&#x2013;<lpage>919</lpage>. doi: <pub-id pub-id-type="doi">10.1093/jcr/ucab024</pub-id></citation></ref>
</ref-list>
</back>
</article>