<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Sleep</journal-id>
<journal-title>Frontiers in Sleep</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Sleep</abbrev-journal-title>
<issn pub-type="epub">2813-2890</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frsle.2025.1625185</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Sleep</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Dreams are more &#x0201C;predictable&#x0201D; than you think</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Bertolini</surname> <given-names>Lorenzo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2927906/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Consoli</surname> <given-names>Sergio</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/3067416/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Weeds</surname> <given-names>Julie</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/557727/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>European Commission, Joint Research Centre (JRC)</institution>, <addr-line>Ispra</addr-line>, <country>Italy</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Informatics, University of Sussex</institution>, <addr-line>Brighton</addr-line>, <country>United Kingdom</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Claudia Picard-Deland, Montreal University, Canada</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Don Kuiken, University of Alberta, Canada</p>
<p>Kristoffer Appel, Institute of Sleep and Dream Technologies, Germany</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Lorenzo Bertolini <email>lorenzo.bertolini&#x00040;ec.europa.eu</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>07</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>4</volume>
<elocation-id>1625185</elocation-id>
<history>
<date date-type="received">
<day>08</day>
<month>05</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>06</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2025 Bertolini, Consoli and Weeds.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Bertolini, Consoli and Weeds</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>A growing body of work has used machine learning and AI tools to analyse dream reports, and compare them to other textual content. Since these tools are usually trained on text from the web, researchers have speculated they might not be suited to model dreams reports, often labeled as &#x0201C;unusual&#x0201D; and &#x0201C;bizarre&#x0201D; content.</p></sec>
<sec>
<title>Methods</title>
<p>We used a set of large language models (LLMs) to encode dream reports from DreamBank and Wikipedia. To estimate the ability of LLMs to model and predict textual reports we adopted perplexity, a measure based on entropy, formally, the exponentiated log-likelihood of a sequence. Intuitively, perplexity indicates how &#x0201C;surprising&#x0201D; a sequence of words is to a model.</p></sec>
<sec>
<title>Results</title>
<p>In most models, perplexity scores for dream reports were significantly lower than those for Wikipedia articles. Moreover, we found that perplexity scores were significantly different in reports produced by male vs female participants, and between blind and normally sighted individuals. In one case, we found this difference to be significant between clinical and healthy subjects.</p></sec>
<sec>
<title>Discussion</title>
<p>Dream reports were found to be generally easier to model and predict than Wikipedia articles. LLMs were also found to implicitly encode group differences previously observed in the literature based on gender, visual impairment, and clinical population.</p></sec></abstract>
<kwd-group>
<kwd>dream report analysis</kwd>
<kwd>dream reports modeling</kwd>
<kwd>gender difference</kwd>
<kwd>dreaming in blind participants</kwd>
<kwd>machine learning</kwd>
<kwd>large language models</kwd>
<kwd>natural language processing</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="2"/>
<equation-count count="1"/>
<ref-count count="61"/>
<page-count count="12"/>
<word-count count="9565"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Sleep, Behavior and Mental Health</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Dream reports describe the content of the conscious experiences we had while asleep. Through the years, researchers have used these transcripts to connect dreams with awakened states (Blagrove et al., <xref ref-type="bibr" rid="B6">2004</xref>; Skancke et al., <xref ref-type="bibr" rid="B50">2014</xref>; Andrews and Hanna, <xref ref-type="bibr" rid="B2">2020</xref>), and to study consciousness (Nir and Tononi, <xref ref-type="bibr" rid="B39">2010</xref>; Siclari et al., <xref ref-type="bibr" rid="B49">2017</xref>) and pathological conditions (Kobayashi et al., <xref ref-type="bibr" rid="B28">2008</xref>; Skancke et al., <xref ref-type="bibr" rid="B50">2014</xref>; Thompson et al., <xref ref-type="bibr" rid="B52">2015</xref>; Andrews and Hanna, <xref ref-type="bibr" rid="B2">2020</xref>). For these reasons, both researchers and practitioners have been consistently interested in dream reports, and have developed a variety of frameworks to study, analyse, and annotate their content in a systematic way (Hall and Van De Castle, <xref ref-type="bibr" rid="B21">1966</xref>; Hauri, <xref ref-type="bibr" rid="B22">1975</xref>; Schredl, <xref ref-type="bibr" rid="B48">2010</xref>).</p>
<p>The analysis and annotation processes of dream reports can be extremely time-consuming and rely upon human experts who usually undergo long training, which has limited the growth and reproducibility of research around dreams and dream reports (Elce et al., <xref ref-type="bibr" rid="B15">2021</xref>). As a result, researchers have shown a growing interest in adopting automatic analysis of dream reports&#x00027; content and structure, based on machine learning and natural language processing (NLP) (see Elce et al., <xref ref-type="bibr" rid="B15">2021</xref> for a review). Many of these approaches use models that have been fully, or partially, trained on large amounts of rather standardized text from the internet, such as Wikipedia (Nadeau et al., <xref ref-type="bibr" rid="B38">2006</xref>; Razavi et al., <xref ref-type="bibr" rid="B44">2013</xref>; Altszyler et al., <xref ref-type="bibr" rid="B1">2017</xref>; Sanz et al., <xref ref-type="bibr" rid="B46">2018</xref>; McNamara et al., <xref ref-type="bibr" rid="B32">2019</xref>; Bertolini et al., <xref ref-type="bibr" rid="B5">2024b</xref>,<xref ref-type="bibr" rid="B4">a</xref>; Cortal, <xref ref-type="bibr" rid="B10">2024</xref>).</p>
<p>Since a vast body of work identifies dream reports as being more bizarre than wakeful experience (Rosen, <xref ref-type="bibr" rid="B45">2018</xref>), one might assume that training a model on more structured and formal textual data might limit the ability of the said model to deal with reports from dreams&#x02014;a position informally held by multiple researchers in the community. While the extent to which dream reports quantitatively differ from other forms of textual transcripts remains a matter of significant debate (Kahan and LaBerge, <xref ref-type="bibr" rid="B26">2011</xref>; Domhoff, <xref ref-type="bibr" rid="B12">2017</xref>; Zheng and Schweickert, <xref ref-type="bibr" rid="B60">2023</xref>), multiple studies have indeed shown that their semantic content and word use can significantly diverge from other forms of textual items. Many of these studies are based on dictionary-based frequency analysis of content words (e.g., Bulkeley and Graves, <xref ref-type="bibr" rid="B8">2018</xref>; Mallett et al., <xref ref-type="bibr" rid="B30">2021</xref>; Zheng and Schweickert, <xref ref-type="bibr" rid="B59">2021</xref>; Yu, <xref ref-type="bibr" rid="B58">2022</xref>; Zheng and Schweickert, <xref ref-type="bibr" rid="B60">2023</xref>; Zheng et al., <xref ref-type="bibr" rid="B61">2024</xref>). While fully transparent and computationally efficient, dictionary-based approaches such as LIWC (Pennebaker et al., <xref ref-type="bibr" rid="B42">2015</xref>) do present some critical issues (Bulkeley and Graves, <xref ref-type="bibr" rid="B8">2018</xref>; Zheng and Schweickert, <xref ref-type="bibr" rid="B60">2023</xref>; Bertolini et al., <xref ref-type="bibr" rid="B4">2024a</xref>), such as typographical errors, or limited access to a broader context and syntactic structure. However, multiple works have shown how these methods could be used to discover differences between different types of dreams, such as nightmares, lucid dreams, and baseline dream reports (Bulkeley and Graves, <xref ref-type="bibr" rid="B8">2018</xref>; Zheng and Schweickert, <xref ref-type="bibr" rid="B60">2023</xref>). A partial solution was proposed by Zheng and Schweickert (<xref ref-type="bibr" rid="B60">2023</xref>), which expanded on the previous literature by studying the differences between dream reports and other types of textual transcripts, using both LIWC and support vector machines (SVM) (Cortes and Vapnik, <xref ref-type="bibr" rid="B11">1995</xref>). The LIWC approach found a large set of categories that significantly differ between dream and non-dream reports, and the proposed SVM approach could successfully discriminate between the two categories of reports. However, the adopted dataset was quite limited in magnitude&#x02014;around 800 instances, balanced between dream and non-dream reports. This constraints the generalisability of the findings, largely grounding the observed difference to the dataset of choice. Altszyler et al. (<xref ref-type="bibr" rid="B1">2017</xref>) introduced an approach more rooted in the overall semantic content of the textual report, by comparing two word-embedding approaches (namely Latent Semantic Analysis (LSA) (Landauer and Dumais, <xref ref-type="bibr" rid="B29">1997</xref>) and word2vec&#x00027;s skip-gram with negative samples (Mikolov et al., <xref ref-type="bibr" rid="B36">2013</xref>) to investigate how the relationship between a content word like <italic>run</italic> changes in large web corpora compared to a large collection of dream reports from DreamBank (Domhoff and Schneider, <xref ref-type="bibr" rid="B13">2008</xref>). In this work the authors discovered that LSA better encodes the difference in the type of contexts such words appearing in the two types of corpora.</p>
<p>Since sleep and dream research are witnessing an increasing amount of NLP-based approaches, investigating whether these qualitative differences might have a quantifiable impact on NLP models is of crucial importance, as it might limit the ability of such tools to model dream reports, particularly if these methodologies utilize unsupervised techniques. This work proposes to address this specific issue directly. Unlike previous work, which focused on qualitatively identifying <italic>what content</italic> makes a (limited set of) dream and waking reports different (Zheng and Schweickert, <xref ref-type="bibr" rid="B60">2023</xref>; Zheng et al., <xref ref-type="bibr" rid="B61">2024</xref>), we study in a quantitative manner <italic>how much</italic> a (large) set of dream reports appears to be &#x0201C;surprising&#x0201D; to a model that has seen a huge amount of non-dream-based text. To do so, we adopt a fully unsupervised solution based on pre-trained autoregressive large language models (LLMs), and on perplexity, a popular NLP metric, intuitively indicating how well an LLM can predict a sequence of words. The proposed approach has found similar application in Colla et al. (<xref ref-type="bibr" rid="B9">2022</xref>) work, where authors showed how perplexity scores from GPT2 (Radford et al., <xref ref-type="bibr" rid="B43">2019</xref>) and n-<italic>grams</italic> can be used to discriminate between healthy participants and patients with Alzheimer&#x00027;s disease.</p>
<p>This work makes four main contributions. First, it shows that, when considered as a continuous string of text, (a large proportion of) DreamBank is only marginally harder to predict than (a comparable section of) Wikipedia. Second, and most importantly, dream reports are on average significantly more predictable than Wikipedia articles when considered as single textual units. Third, it identifies a negative correlation between the number of words in a report/article and how &#x0201C;surprising&#x0201D; such a report/article appears to the model. Fourth, it provides preliminary evidence suggesting that gender and visual impairment can significantly impact how &#x0201C;surprising&#x0201D; a report appears to the model, providing the first evidence that modern NLP tools such as LLMs internally and implicitly replicate group differences previously observed in the literature.</p></sec>
<sec sec-type="materials and methods" id="s2">
<title>2 Materials and methods</title>
<sec>
<title>2.1 Metric and models</title>
<p>The primary interest of this work is to quantitatively assess whether dream reports are in fact harder to model and predict for a pre-trained large language model (LLM), the current tool of choice in most NLP research and applications. To measure this phenomenon, we adopt perplexity (PPL) (Huyen, <xref ref-type="bibr" rid="B25">2019</xref>). Intuitively, perplexity can be seen as a measures of how &#x0201C;unpredictable&#x0201D; or &#x0201C;surprising&#x0201D; a given string of text is for a model. In other words, given a target word <italic>i</italic>, and a sequence of words (<italic>c</italic>, for context) preceding <italic>i</italic>, perplexity measures the ability of an LLM to predict <italic>i</italic>, given its context <italic>c</italic>. Lower the perplexity scores, higher is the ability of a model to predicting how a sentence evolves. In other words, low perplexity indicates low surprisal. Formally speaking, perplexity is the exponentiated log-likelihood of a sequence <italic>X</italic> and is computed using <xref ref-type="disp-formula" rid="E1">Equation 1</xref>:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>e</mml:mi><mml:mi>p</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>X</italic> &#x0003D; (<italic>x</italic><sub>0</sub>, <italic>x</italic><sub>1</sub>, ..., <italic>x</italic><sub><italic>t</italic></sub>) is the sequence of words, <italic>logp</italic><sub>&#x003B8;</sub>(<italic>x</italic><sub><italic>i</italic></sub>|<italic>x</italic><sub> &#x0003C; <italic>i</italic></sub>) is the log-likelihood, of the <italic>i</italic><sup><italic>th</italic></sup> word conditioned by the preceding context (<italic>x</italic><sub> &#x0003C; <italic>i</italic></sub>). While many other solutions have been proposed to evaluate how well a language model can capture different linguistic phenomena, perplexity is still widely used and can inform us on how well a model reflects natural language by measuring how distant a string is to a more &#x0201C;natural&#x0201D; sequence (Meister and Cotterell, <xref ref-type="bibr" rid="B34">2021</xref>). Hence, our goal can be stated as understanding whether a machine trained on a very large amount of textual data &#x0201C;perceives&#x0201D; dream reports as &#x0201C;surprising&#x0201D; (i.e., as having a high perplexity).</p>
<p>For models without computational constraints, perplexity should be evaluated using a sliding-window approach. This method slides the context window across the text, ensuring the model has sufficient context for each prediction. The process sums negative log-likelihoods for all word-context pairs and averages across total words, as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Schematic representation of perplexity computation, using a sliding fixed window context of three words.</p></caption>
<alt-text>Text displays the phrase &#x0201C;This is a study on dreams and LLMs&#x0201D; repeated eight times. Words are highlighted in alternating blue and orange colors, shifting to emphasize different words in each repetition.</alt-text>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsle-04-1625185-g0001.tif"/>
</fig>
<p>This approach better approximates true sequence probability decomposition and typically produces more favorable scores. However, it requires a separate forward pass for each token, making it computationally expensive. A practical solution uses strided sliding windows, moving the context by larger steps rather than single tokens. This maintains a large context while significantly reducing computation time. Following Hugging Face implementation<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>, we use a stride of 512 tokens with each model&#x00027;s maximum sequence length as input size (context plus target word). These settings surpass the results reported in the original GPT-2 papers.</p>
<p>This strided approach efficiently computes perplexity for large datasets that cannot fit entirely in model memory. For shorter sequences that fit within the model limits, we can process them entirely at once, obtaining a single perplexity score per sequence. Our work primarily uses this single-sequence approach, focusing on individual dream reports and Wikipedia articles. As detailed below, these texts never exceed the maximum input length for any model investigated.</p>
<p>To model our textual data, we adopt models from two series of autoregressive pre-trained LLMs: GPT2 and OLM0 (Groeneveld et al., <xref ref-type="bibr" rid="B19">2024</xref>). The GPT2 family consists of GPT2 (137 million (M) parameters), GPT2-Medium (380 M), GPT2-Large (812 M), and GPT2-XL (1,610 M). On the other hand, the OLMo family presents two models: OLMo-1B (1,180 M), and OLMo-7B (6,890 M).</p>
<p>Indeed, the current landscape of autoregressive LLMs offers a suite of impressive alternatives, such as GPT-4 (OpenAI et al., <xref ref-type="bibr" rid="B40">2024</xref>), Gemini (Team, <xref ref-type="bibr" rid="B18">2024</xref>), or Llama 3 (Dubey et al., <xref ref-type="bibr" rid="B14">2024</xref>). However, our selection of models allows us to control for multiple interesting factors, namely the impact of model size, training data, and the evolving state of the art. While it might seem an era ago&#x02014;and certainly was in AI terms...&#x02014;GPT2 once was (at) the pinnacle of the LLM leader-board. Indeed, OLMo&#x00027;s performance is not <italic>extremely</italic> representative of the state of the art. However, at its release time, it was on par with the highest-end competitors, such as Llama 2. Aside from their performance, these two families share an important factor, which makes them more suitable for our experiments than more recent and powerful models: the extent to which we know their training data. Contrary to its more recent siblings, we have quite some information on the data used to train GPT2. Most importantly, on what was <italic>not</italic> used for its training, namely, Wikipedia (Radford et al., <xref ref-type="bibr" rid="B43">2019</xref>). On the other hand, and even more unusual for the current standard, OLMo&#x00027;s training data is <italic>fully</italic> open source. Not only do we know Wikipedia was used for its training, but we can search <italic>which</italic> articles were used in the model training. While documents from Wikipedia compose a little over .1% of the overall documents in training set (Soldaini et al., <xref ref-type="bibr" rid="B51">2024</xref>), this is extremely relevant to our experiments as it allows us to frame the results of the models with respect to their pre-training procedure. Lastly, both families, which have a convenient point of contact in the two one-billion-parameters models, present a set of models growing in size, which nicely reflects the capabilities and approach of the time frame they were built in, and can allow us to study how increasing the number of parameters in a model impacts its ability to model dream and other textual data, depending on its training data. In summary, if the hypothesis that dream reports are harder to model for LLMs, we should find that average PPL scores for dream reports should be, on average, significantly higher than PPL scores for Wikipedia articles, especially in LLMs exposed to Wikipedia&#x00027;s articles during training.</p></sec>
<sec>
<title>2.2 Dataset</title>
<sec>
<title>2.2.1 Dream dataset</title>
<p>Similarly to previous work (Fogli et al., <xref ref-type="bibr" rid="B17">2020</xref>; Gutman Music et al., <xref ref-type="bibr" rid="B20">2022</xref>; Bertolini et al., <xref ref-type="bibr" rid="B4">2024a</xref>,<xref ref-type="bibr" rid="B5">b</xref>; Cortal, <xref ref-type="bibr" rid="B10">2024</xref>), we adopt a set of dream reports extracted from DreamBank (Domhoff and Schneider, <xref ref-type="bibr" rid="B13">2008</xref>)<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref>, an online collection of dream reports from different people and scientific studies. The original dataset contains approximately 22k reports in the English language, annotated with respect to gender, year of collection, and series&#x02014;the specific subsets of DreamBank representing (groups of) individuals from which dreams are collected.</p></sec>
<sec>
<title>2.2.2 Text dataset</title>
<p>We use Wikipedia as the source of our baseline text. More specifically, we consider the WikiText2 dataset (Merity et al., <xref ref-type="bibr" rid="B35">2017</xref>)<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref>, an open-source dataset containing approximately 20k articles from Wikipedia. This specific baseline choice is motivated by two main reasons. First, and more specifically to our models of choice, (part of) Wikipedia is included of OLMo&#x00027;s training set. Moreover, WikiText2 was entirely excluded from GPT2 pre-training, and was instead used as one of the testing benchmarks in the original paper. Second, and on a general stance, adopting Wikipedia allows for a strict comparison with a standardized text, in terms of syntactic and semantic structure. This is due to the fact that large portions of Wikipedia are formally and heavily curated, and can hence work as a &#x0201C;stress&#x0201D; test for the hypothesis that dream reports are notably different.</p></sec>
<sec>
<title>2.2.3 Sampling</title>
<p>Given the discrepancy between the datasets&#x00027; magnitude and some of their specific content, we use a filtering and sampling procedure over the original datasets. We begin by filtering out from Wikipedia all those instances that do not contain an article&#x00027;s body&#x02014;that is, instances consisting of only titles or empty strings. To limit the possibility that a variable such as the number of words might impact our experiments, we further extract from both DreamBank and Wikipedia the set of items laying that contain between 30 and 250 words. The remaining datasets consist of approximately 13 k Wikipedia articles and 17 k dream reports. To generate a test set with a similar distribution in the number of words per instance, we interactively sample a subset of dream reports of the same magnitude as the remaining Wikipedia set (i.e., 13 k) for 250 iterations. We then run a random permutation test comparing the Wikipedia set against each sample dream set and select the least diverging one. The final distributions are described in <xref ref-type="fig" rid="F2">Figure 2</xref>, and are made freely available (see link in the &#x0201C;data availability statement&#x0201D; section).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Word-count distributions of the number of words (No. Words) per instance in the final Wikipedia (WikiText2) and dream (DreamBank) test sets used in the experiments.</p></caption>
<alt-text>Box plot comparing the number of words in two datasets DreamBank and WikiText2. DreamBank has a median around 100, with a range from approximately 50 to 200 words. WikiText2 has a similar range but the median is slightly higher. Both datasets show similar variability.</alt-text>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsle-04-1625185-g0002.tif"/>
</fig></sec></sec>
<sec>
<title>2.3 Statistical analyses</title>
<p>We compare one-dimensional distributions (e.g., how many words constitute each dream report) with a random permutation test. To assess whether two-dimensional distributions (e.g., the number of words <italic>and</italic> perplexity scores of each Wikipedia article) are significantly different from one another, we adopt the Peacock test, which is a two-dimensional non-parametric generalization of the Kolmogorov-Smirnov test (Peacock, <xref ref-type="bibr" rid="B41">1983</xref>; Fasano and Franceschini, <xref ref-type="bibr" rid="B16">1987</xref>). Correlation analyses are based on Spearman&#x00027;s coefficient. All <italic>p</italic> values in the work refer to scores obtained after applying Holm correction (Holm, <xref ref-type="bibr" rid="B23">1979</xref>), a method used to adjust p-values for multiple comparisons to minimize Type I errors, by sequentially adjusting the significance threshold as a function of the number of tests performed. Experiments were run with the support of an NVIDIA H100 80GB HBM3 GPU. The code and data to replicate the experiments are freely available at <ext-link ext-link-type="uri" xlink:href="https://github.com/jrcf7/report_perplexity">https://github.com/jrcf7/report_perplexity</ext-link>.</p></sec></sec>
<sec sec-type="results" id="s3">
<title>3 Results</title>
<sec>
<title>3.1 Comparing dream reports and Wikipedia articles</title>
<p><xref ref-type="table" rid="T1">Table 1</xref> gives an overview of the overall perplexities produced by the different versions of GPT2 and OLMo on the two test sets, namely DreamBank and WikiText2. The table further contains the respective lengths of the datasets, in terms of the total number of tokens, and the size of each model (in millions of learnable parameters). Based on the results in the table, we can make three main observations. First, the perplexity scores for WikiText2 from our experiments closely resemble those of the original paper that introduced the GPT2 models (Radford et al., <xref ref-type="bibr" rid="B43">2019</xref>). Second, compared to DreamBank, each model seems to produce lower perplexity scores for Wikipedia. Third, while the perplexity scores produced by GPT2 on WikiText2 and DreamBank appear close to each other (31.9 vs 27.4), the distance grows with model size. Although in a smaller magnitude, a similar trend is observed for the two variants of OLMo. Interestingly, in these models, the differences between datasets are not as marked as they are for GPT2 models. This behavior is unexpected since (part of) Wikipedia is included in OLMo&#x00027;s training set, and should hence have a significant advantage over out-of-distribution data like dream reports. This evidence could suggest that WikiText2, or part of it, might not be part of the Wikipedia subset used to train OLMo models. Overall, the differences remain relatively small across the board of the GPT2 models. Moreover, whilst the perplexity of DreamBank is overall higher than that of WikiText2, this discrepancy might be explained by the fact that DreamBank is built from collections of very different individuals, from (very) different time periods. In other words, while Wikipedia articles tend to follow a more unified language type and structure, DreamBank&#x00027;s reports can suddenly and significantly vary from one line to the other.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Whole corpora results.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Data</bold></th>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Perplexity</bold></th>
<th valign="top" align="center"><bold>Dataset length (M)</bold></th>
<th valign="top" align="center"><bold>Model size (M)</bold></th>
<th valign="top" align="center"><bold>Original PPL</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">DreamBank</td>
<td valign="top" align="left">GPT2</td>
<td valign="top" align="center">31.9</td>
<td valign="top" align="center">1.4</td>
<td valign="top" align="center">137</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">WikiText2</td>
<td valign="top" align="left">GPT2</td>
<td valign="top" align="center">27.4</td>
<td valign="top" align="center">1.7</td>
<td valign="top" align="center">137</td>
<td valign="top" align="center">29.41</td>
</tr> <tr>
<td valign="top" align="left">DreamBank</td>
<td valign="top" align="left">GPT2-Medium</td>
<td valign="top" align="center">26.7</td>
<td valign="top" align="center">1.4</td>
<td valign="top" align="center">380</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">WikiText2</td>
<td valign="top" align="left">GPT2-Medium</td>
<td valign="top" align="center">20.0</td>
<td valign="top" align="center">1.7</td>
<td valign="top" align="center">380</td>
<td valign="top" align="center">22.76</td>
</tr> <tr>
<td valign="top" align="left">DreamBank</td>
<td valign="top" align="left">GPT2-Large</td>
<td valign="top" align="center">24.2</td>
<td valign="top" align="center">1.4</td>
<td valign="top" align="center">812</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">WikiText2</td>
<td valign="top" align="left">GPT2-Large</td>
<td valign="top" align="center">17.2</td>
<td valign="top" align="center">1.7</td>
<td valign="top" align="center">812</td>
<td valign="top" align="center">19.93</td>
</tr> <tr>
<td valign="top" align="left">DreamBank</td>
<td valign="top" align="left">GPT2-XL</td>
<td valign="top" align="center">22.9</td>
<td valign="top" align="center">1.4</td>
<td valign="top" align="center">1610</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">WikiText2</td>
<td valign="top" align="left">GPT2-XL</td>
<td valign="top" align="center">15.5</td>
<td valign="top" align="center">1.7</td>
<td valign="top" align="center">1610</td>
<td valign="top" align="center">18.34</td>
</tr> <tr>
<td valign="top" align="left">DreamBank</td>
<td valign="top" align="left">OLMo-1B</td>
<td valign="top" align="center">19.9</td>
<td valign="top" align="center">1.4</td>
<td valign="top" align="center">1180</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">WikiText2</td>
<td valign="top" align="left">OLMo-1B</td>
<td valign="top" align="center">12.4</td>
<td valign="top" align="center">1.7</td>
<td valign="top" align="center">1180</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">DreamBank</td>
<td valign="top" align="left">OLMo-7B</td>
<td valign="top" align="center">16.5</td>
<td valign="top" align="center">1.4</td>
<td valign="top" align="center">6890</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">WikiText2</td>
<td valign="top" align="left">OLMo-7B</td>
<td valign="top" align="center">8.7</td>
<td valign="top" align="center">1.7</td>
<td valign="top" align="center">6890</td>
<td valign="top" align="center">-</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Analysis of the relation between length (number of tokens) and perplexity scores produced by GPT2 when considering DreamBank and Wikipedia data as a whole.</p>
</table-wrap-foot>
</table-wrap>
<p>These results suggest that, considered as a whole corpus (i.e., a subsequent and unique string of text), Wikipedia is slightly easier to predict for all selected models. However, the main focus of our work is to understand if <italic>single</italic> dream reports are harder to model&#x02014;i.e., are less <italic>predictable</italic>&#x02014;than <italic>single</italic> Wikipedia articles, as these would generally be the input to any given LLM. <xref ref-type="fig" rid="F3">Figure 3</xref> offers a rather intuitive and straightforward answer to this question, by plotting the average perplexity score produced by each model (Y axis), given an instance with a defined number of words (X axis). In each diagram, the continuous blue line represents dream reports from DreamBank, while the dashed orange line represents articles from Wikipedia. Our analysis reveals that for all GPT2 models, the two two-dimensional distributions are significantly different from one another (<italic>p</italic> &#x0003C; 0.0001), and a random permutation test further showed that the one-dimensional distribution of the perplexity scores alone is too (<italic>p</italic> &#x0003C; 0.01). As the figure intuitively suggests, our analysis also conforms to the fact that, while always significant, the differences tend to fade as the model size increases. Looking at the Peacock test (Peacock, <xref ref-type="bibr" rid="B41">1983</xref>; Fasano and Franceschini, <xref ref-type="bibr" rid="B16">1987</xref>), we see how the score of GPT2, <italic>D</italic>=.34, slowly reduces passing from GPT2-Medium, <italic>D</italic> = 0.23, GPT2-Large, <italic>D</italic> = 0.19, and reaching <italic>D</italic>=.15 for GPT2-XL.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Perplexities by number of words per single item. Visualization of the interaction (mean and standard error, described by the shades) between the number of words (x-axis) and the perplexity scores (y-axis) produced by GPT2 for single dream reports and WikiText2 articles.</p></caption>
<alt-text>Comparison of perplexity across six models GPT2, GPT2-Medium, GPT2-Large, GPT2-XL, OLMo-1B, and OLMo-7B, shown in line graphs. Each plot displays perplexity against the number of words, with two datasets, DreamBank and WikiText2, represented by solid blue and dashed orange lines, respectively. Perplexity generally decreases as the number of words increases.</alt-text>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsle-04-1625185-g0003.tif"/>
</fig>
<p>For the OLMo models, we observe a rather different trend. The two lines appear to overlap under OLMo-1B, and the GPT2 tendency seems inverted for OLMo-7B, with DreamBank&#x00027;s scores surpassing Wikipedia ones. This interpretation is confirmed by the statistical analysis. Under both models, we found a significant overall difference with the Peacock test (<italic>p</italic> &#x0003C; 0.0001). However, the random permutation analysis showed that the difference in perplexity scores is not significant for OLMo-1B. Moreover, the distance between the two distributions in OLMo-7B is notably small (<italic>D</italic> = 0.17), a rather surprising result considering that both OLMo have been exposed to (part of) Wikipedia during their training phase.</p>
<p>In summary, our results show that a large proportion of the LLMs under investigation found dream reports to be significantly more predictable than a more formal and structured text, such as Wikipedia articles. Most importantly, these models, namely the GPT2 ones, did not include any Wikipedia article in their training data, and hence have no advantages over dream reports. In the case of the two OLMo models, that did include some Wikipedia articles in their training data, and should hence have a clear advantage over dream reports, only the 7B version is overall better at modeling Wikipedia articles over dream reports.</p></sec>
<sec>
<title>3.2 Group analysis</title>
<p>The previous section provides consistent evidence that dream reports might be easier to model than more &#x0201C;standardised&#x0201D; strings of texts, such as Wikipedia articles. In this section, we study whether three directly measurable macro factors previously studied in the relevant literature also impact how well an LLM can model dream reports. The analysis takes into consideration five factors. The number of words per report (No. Words), year of collection, and three variables that were previously observed in the literature to produce qualitative changes in the content and structure of dream reports, namely gender, vision impairment, and clinical patients (Hall and Van De Castle, <xref ref-type="bibr" rid="B21">1966</xref>; Schrdel and Reinhard, <xref ref-type="bibr" rid="B47">2008</xref>; Wong et al., <xref ref-type="bibr" rid="B56">2016</xref>; Kirtley, <xref ref-type="bibr" rid="B27">1975</xref>; Hurovitz et al., <xref ref-type="bibr" rid="B24">1999</xref>; Meaidi et al., <xref ref-type="bibr" rid="B33">2014</xref>; Mota et al., <xref ref-type="bibr" rid="B37">2014</xref>; Zheng et al., <xref ref-type="bibr" rid="B61">2024</xref>). Lastly, this section focuses solely on GPT2 and OLMo-7B. This choice is motivated by the fact that they represent the models with the most marked preference for one of the two datasets.</p>
<p>As hinted by <xref ref-type="fig" rid="F3">Figure 3</xref>, our analysis has found a negative correlation between the number of words per report, and perplexity scores, both for GPT2 (&#x003C1; = - 0.33) and OLMo-7B (&#x003C1; = - 0.38), and both strongly significant (<italic>p</italic> &#x0003C; 0.0001). Observing lower perplexity scores for larger documents is not unexpected, since predicting a given word becomes easier as the context to guess said word becomes more abundant. While rather expected and explainable, this (co)relationship is likely more complicated than expected, as suggested by the relation between perplexity and word count in dream reports from participants of different genders. As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, the perplexity scores produced by participants who identify themselves as male are significantly (both <italic>p</italic> &#x0003C; 0.01) lower, and are hence easier to model and predict for both GPT2 and OLMo-7B. However, as clearly shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, these reports are also significantly (<italic>p</italic> &#x0003C; 0.01) shorter than the ones produced by participants who identify themself as female, as already observed in other work (e.g., Mathes and Schredl, <xref ref-type="bibr" rid="B31">2013</xref>). In other words, while from a general stance, shorter reports appear to entail higher perplexity, the trend seems to invert when taking into account the gender subgroup.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Per-report perplexities: gender. Distributions of perplexity scores obtained by GPT2 and OLMo-7B on DreamBank single dream reports grouped by the gender of the participants.</p></caption>
<alt-text>Two side-by-side bar charts compare the perplexity scores by gender for two language models. Chart A shows GPT2 with higher perplexity for females (around 43) than males (around 38). Chart B shows OLMo-7B with higher perplexity for females (around 23) than males (around 21). Error bars indicate variability.</alt-text>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsle-04-1625185-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Per-report number of words: gender. Distributions of the word count of DreamBank single dream reports divided by the gender of the participants.</p></caption>
<alt-text>Bar chart showing the average number of words spoken by gender. Males, in light pink, average around 115 words. Females, in darker pink, average about 125 words. Error bars are included.</alt-text>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsle-04-1625185-g0005.tif"/>
</fig>
<p>The effect of the year of data collection on the perplexity scores is also assessed with a correlation analysis. To this end, we converted the categorical framing of some instance (e.g., &#x0201C;1980s&#x02013;1990s&#x0201D;), by simply finding the average a time-span (e.g., <monospace>1985</monospace>). Instances with non-available dates were excluded from the analysis. The obtained dates, together with the original ones, are presented in <xref ref-type="table" rid="T2">Table 2</xref>. The result of the analysis suggests a negative correlation between the year of collection and the perplexity scores. In other words, as one might expect, reports produced in more recent years appear easier to model for GPT2, and hence tend to produce lower perplexity scores. However, while strongly significant (<italic>p</italic> &#x0003C; 0.0001), the effect was very weak for both GPT2 (&#x003C1; = -0.13) and OLMo-7B (&#x003C1; = -0.16).</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Conversion table for DreamBank&#x00027;s year of collection variable.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>DreamBank</bold></th>
<th valign="top" align="center"><bold>Integer conversion</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1897&#x02013;1918</td>
<td valign="top" align="center">1907</td>
</tr> <tr>
<td valign="top" align="left">1912&#x02013;1965</td>
<td valign="top" align="center">1938</td>
</tr> <tr>
<td valign="top" align="left">1939</td>
<td valign="top" align="center">1939</td>
</tr> <tr>
<td valign="top" align="left">1940&#x02013;1998</td>
<td valign="top" align="center">1969</td>
</tr> <tr>
<td valign="top" align="left">1940s&#x02013;1950s</td>
<td valign="top" align="center">1945</td>
</tr> <tr>
<td valign="top" align="left">1940s&#x02013;1950s &#x00026; 1990s</td>
<td valign="top" align="center">1960</td>
</tr> <tr>
<td valign="top" align="left">1946&#x02013;1950</td>
<td valign="top" align="center">1948</td>
</tr> <tr>
<td valign="top" align="left">1948&#x02013;1949</td>
<td valign="top" align="center">1948</td>
</tr> <tr>
<td valign="top" align="left">1949&#x02013;1964</td>
<td valign="top" align="center">1956</td>
</tr> <tr>
<td valign="top" align="left">1949&#x02013;1997</td>
<td valign="top" align="center">1973</td>
</tr> <tr>
<td valign="top" align="left">1957&#x02013;1959</td>
<td valign="top" align="center">1958</td>
</tr> <tr>
<td valign="top" align="left">1960&#x02013;1997</td>
<td valign="top" align="center">1978</td>
</tr> <tr>
<td valign="top" align="left">1960&#x02013;1999</td>
<td valign="top" align="center">1979</td>
</tr> <tr>
<td valign="top" align="left">1962</td>
<td valign="top" align="center">1962</td>
</tr> <tr>
<td valign="top" align="left">1963&#x02013;1965</td>
<td valign="top" align="center">1964</td>
</tr> <tr>
<td valign="top" align="left">1963&#x02013;1967</td>
<td valign="top" align="center">1965</td>
</tr> <tr>
<td valign="top" align="left">1964</td>
<td valign="top" align="center">1964</td>
</tr> <tr>
<td valign="top" align="left">1968</td>
<td valign="top" align="center">1968</td>
</tr> <tr>
<td valign="top" align="left">1970</td>
<td valign="top" align="center">1970</td>
</tr> <tr>
<td valign="top" align="left">1970&#x02013;2008</td>
<td valign="top" align="center">1989</td>
</tr> <tr>
<td valign="top" align="left">1971</td>
<td valign="top" align="center">1971</td>
</tr> <tr>
<td valign="top" align="left">1980&#x02013;2002</td>
<td valign="top" align="center">1991</td>
</tr> <tr>
<td valign="top" align="left">1985&#x02013;1997</td>
<td valign="top" align="center">1991</td>
</tr> <tr>
<td valign="top" align="left">1990&#x02013;1999</td>
<td valign="top" align="center">1994</td>
</tr> <tr>
<td valign="top" align="left">1990s</td>
<td valign="top" align="center">1990</td>
</tr> <tr>
<td valign="top" align="left">1991&#x02013;1993</td>
<td valign="top" align="center">1992</td>
</tr> <tr>
<td valign="top" align="left">1992&#x02013;1998</td>
<td valign="top" align="center">1995</td>
</tr> <tr>
<td valign="top" align="left">1992&#x02013;1999</td>
<td valign="top" align="center">1995</td>
</tr> <tr>
<td valign="top" align="left">1995</td>
<td valign="top" align="center">1995</td>
</tr> <tr>
<td valign="top" align="left">1996</td>
<td valign="top" align="center">1996</td>
</tr> <tr>
<td valign="top" align="left">1996&#x02013;1997</td>
<td valign="top" align="center">1996</td>
</tr> <tr>
<td valign="top" align="left">1996&#x02013;1998</td>
<td valign="top" align="center">1997</td>
</tr> <tr>
<td valign="top" align="left">1997</td>
<td valign="top" align="center">1997</td>
</tr> <tr>
<td valign="top" align="left">1997&#x02013;1999</td>
<td valign="top" align="center">1998</td>
</tr> <tr>
<td valign="top" align="left">1997&#x02013;2000</td>
<td valign="top" align="center">1998</td>
</tr> <tr>
<td valign="top" align="left">1997&#x02013;2001</td>
<td valign="top" align="center">1999</td>
</tr> <tr>
<td valign="top" align="left">1998</td>
<td valign="top" align="center">1998</td>
</tr> <tr>
<td valign="top" align="left">1998&#x02013;2000</td>
<td valign="top" align="center">1999</td>
</tr> <tr>
<td valign="top" align="left">1999</td>
<td valign="top" align="center">2010</td>
</tr> <tr>
<td valign="top" align="left">1999&#x02013;2000</td>
<td valign="top" align="center">1999</td>
</tr> <tr>
<td valign="top" align="left">1999&#x02013;2001</td>
<td valign="top" align="center">2000</td>
</tr> <tr>
<td valign="top" align="left">2000</td>
<td valign="top" align="center">2000</td>
</tr> <tr>
<td valign="top" align="left">2000&#x02013;2001</td>
<td valign="top" align="center">2000</td>
</tr> <tr>
<td valign="top" align="left">2001&#x02013;2003</td>
<td valign="top" align="center">2002</td>
</tr> <tr>
<td valign="top" align="left">2003&#x02013;2004</td>
<td valign="top" align="center">2003</td>
</tr> <tr>
<td valign="top" align="left">2003&#x02013;2005</td>
<td valign="top" align="center">2004</td>
</tr> <tr>
<td valign="top" align="left">2003&#x02013;2006</td>
<td valign="top" align="center">2004</td>
</tr> <tr>
<td valign="top" align="left">2004</td>
<td valign="top" align="center">2004</td>
</tr> <tr>
<td valign="top" align="left">2007&#x02013;2010</td>
<td valign="top" align="center">2008</td>
</tr> <tr>
<td valign="top" align="left">2009</td>
<td valign="top" align="center">2009</td>
</tr> <tr>
<td valign="top" align="left">2010&#x02013;2011</td>
<td valign="top" align="center">2010</td>
</tr> <tr>
<td valign="top" align="left">?</td>
<td valign="top" align="center">NaN</td>
</tr> <tr>
<td valign="top" align="left">Late 1990s</td>
<td valign="top" align="center">1998</td>
</tr> <tr>
<td valign="top" align="left">Mid-1980s</td>
<td valign="top" align="center">1985</td>
</tr> <tr>
<td valign="top" align="left">Mid-1990s</td>
<td valign="top" align="center">1995</td>
</tr></tbody>
</table>
</table-wrap>
<p>Among DreamBank&#x00027;s series, there are two that collect reports from several blind participants, both males and females, for a total of 285 dream reports. To compare this restricted set of reports with the one produced by normally-sighted individuals, we have sampled a set of reports from DreamBank that has the same range of perplexity scores observed for blind participants. Just like for the general and gender-based results, these two sets show to be significantly different (<italic>p</italic> &#x0003C; 0.0001) when considered as two-dimensional distributions (as in <xref ref-type="fig" rid="F3">Figure 3</xref>); however, when taken separately, only the perplexity score turned out to be significantly different (<italic>p</italic> &#x0003C; 0.01). <xref ref-type="fig" rid="F6">Figure 6</xref> summarizes the differences in the perplexity scores distributions obtained for reports produced by visually impaired and normally sighted participants. As shown, even when sampling from a limited range of items, perplexity scores for visually impaired participants are on average considerably lower and have a remarkably smaller variance, especially when encoded with GPT2.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Per-report perplexities: vision impairment. Distributions of perplexity scores obtained by GPT2 and OLM0-7B on DreamBank single dream reports grouped by vision impairment of the participants.</p></caption>
<alt-text>Box plots labeled A and B compare the perplexity of two models, GPT2 and OLMo-7B, for impaired and normal vision groups. Both plots show higher perplexity in impaired vision (yellow box) compared to normal vision (blue box). Outliers are present in both groups, with the impaired vision group displaying greater variance in perplexity.</alt-text>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsle-04-1625185-g0006.tif"/>
</fig>
<p>Lastly, we consider a very small set (circa 70 instances) of reports belonging to a subject diagnosed with post-traumatic stress disorder (PTSD), a veteran of the Vietnam War. We follow the same sampling procedure and overall analysis described in the previous paragraph for the visually impaired participants, summarized in <xref ref-type="fig" rid="F7">Figure 7</xref>. As suggested by the two diagrams, the difference in perplexity scores is significant only under the OLMo-7B model (<italic>p</italic> &#x0003C; .01). Indeed, these results are limited by the small size of sample, and are hence harder to frame and contextualize. We note that the results from GPT2 appear in line with the last experiment in Bertolini et al. (<xref ref-type="bibr" rid="B4">2024a</xref>), where the authors showed that a small LLM trained to classy dream for emotional content, using a report from healthy participants, performed well on this same set, despite being out of distribution. In contrast, the results from OLMo-7B appear in line with the work suggesting that clinical participants produce dream reports that significantly differ from healthy participants, as suggested by Mota et al. (<xref ref-type="bibr" rid="B37">2014</xref>). It is interesting to note that the work from Mota et al. (<xref ref-type="bibr" rid="B37">2014</xref>) largely relies on graphs-based analysis and patterns, and that transformers (Vaswani et al., <xref ref-type="bibr" rid="B53">2017</xref>), the neural network at the base of most LLMs, can be considered as a special case of graph neural networks (Veli&#x0010D;kovi&#x00107;, <xref ref-type="bibr" rid="B54">2023</xref>). It is possible that a large enough model could locally and implicitly represent the same type of graph that is useful to distinguish between clinical and healthy participants.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Per-report perplexities: clinical patient. Distributions of perplexity scores obtained by GPT2 and OLMo-7B on DreamBank single dream reports grouped by clinical condition.</p></caption>
<alt-text>Box plots comparing perplexity scores for two models, GPT2 (A) and OLMo-7B (B), based on PTSD clinical status. The &#x0201C;Yes&#x0201D; group is in orange, and the &#x0201C;No&#x0201D; group is in green. Data points show variability, with potential outliers present.</alt-text>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frsle-04-1625185-g0007.tif"/>
</fig>
</sec></sec>
<sec sec-type="discussion" id="s4">
<title>4 Discussion</title>
<p>A growing amount of work has adopted NLP tools to investigate and annotate dream reports (see Elce et al., <xref ref-type="bibr" rid="B15">2021</xref>; Bertolini et al., <xref ref-type="bibr" rid="B4">2024a</xref>; Cortal, <xref ref-type="bibr" rid="B10">2024</xref> for more details). Many of these approaches rely on neural models of various dimensions, trained on large text corpora scraped from the web (Radford et al., <xref ref-type="bibr" rid="B43">2019</xref>). Since a consistent body of evidence has shown that the structure and semantic content of dream reports can significantly differ from other types of textual transcripts (see Altszyler et al., <xref ref-type="bibr" rid="B1">2017</xref>; Bulkeley and Graves, <xref ref-type="bibr" rid="B8">2018</xref>; Zheng and Schweickert, <xref ref-type="bibr" rid="B60">2023</xref>, inter alia), it is important to understand, and possibly quantify, if and how much these differences impact the ability of NLP tools to model and interpret rather specific strings of text, such as dream reports. This is especially relevant and important when adopting off-the-shelf and unsupervised models and methods, as already hinted by Bertolini et al. (<xref ref-type="bibr" rid="B4">2024a</xref>).</p>
<p>In this work, we adopted a set of large language models (LLMs) from the GPT2 and OLMo family, to investigate these issues. More specifically, we studied how well they can model and predict dream reports, compared to a more &#x0201C;standard&#x0201D; text, like Wikipedia articles, using perplexity as a measure of uncertainty. Our results have shown how most LLMs produce significantly lower perplexity scores (hence better) for single dream reports than for single Wikipedia articles. The only exceptions to this trend were observed in the two OLMo models. However, these models contained part of Wikipedia in their training data, and hence had a notable advantage. Moreover, we found only a partial significance in the smaller model (OLMo-1B), and a marginal advantage for the larger model (OLMo-7B).</p>
<p>These findings paint a clear picture. A picture where LLMs, or at least the ones tested in this work, do not seem to struggle at all with processing dream reports, nor do they seem more &#x0201C;surprised&#x0201D; by dream reports than they are by Wikipedia articles. In the literature, a consistent research line tends to associate dreams and their reports with bizarreness (Rosen, <xref ref-type="bibr" rid="B45">2018</xref>), entailing a significant deviation from normal experience, whatever that might be. This view appears in clear contrast with our findings, as they indicate that for LLMs, dream reports are as &#x0201C;predictable&#x0201D; as the &#x0201C;the norm&#x0201D;, at least in the form of Wikipedia text. This might come as unexpected but it is likely due to what is <italic>our</italic> compass for dream reports bizarreness: reality. Aside very specific pathological cases and scenarios, we are generally capable of distinguishing a bizarre or absurd event from reality. LLMs, on the other hand, are machines designed and trained to encode or generate text, regardless of its truthiness or correctness, and can in fact frequently struggle even with identifying simple and established facts (Wang et al., <xref ref-type="bibr" rid="B55">2024</xref>). This does not mean that LLMs can not or should not be used in dream research. On the contrary, our work suggests they <italic>can</italic> handle these unique strings of text. However, these results show that when using LLMs in this line of research, we should be extremely careful in projecting <italic>our</italic> definition and understanding of the mind and world onto these tools. One of these definitions might in fact be bizarreness, for which humans and LLMs might have a very different &#x0201C;concept&#x0201D;. Future work will have to focus on providing more insight into the existing relation between perplexity, or other mathematically measurable metrics, to carefully operationalised human concepts such as bizarreness, surprisal, or predicability.</p>
<p>Indeed, the main findings of this work suggest that dream reports are not a unique and unpredictable class of textual strings per se. However, just like Wikipedia&#x00027;s articles, some <italic>can be</italic> harder to model and predict. The second part of the work has hence proposed a set of analyses to understand which features of a report&#x02014;or their author&#x02014;might have an impact on the model of choice. The focus was on two specific models, namley GPT2 and OLMo-7B, and five variables immediately measurable from the adopted dataset: word count (No.Words), gender, year of collection, visual impairments, and mental health. We focused our attention on these models since they produced the most marked preference for dream reports or Wikipedia, while the choice of variables was based on group differences previously found in the literature.</p>
<p>The correlation analysis found a (rather expected) negative interaction between the number of tokens contained in a report (No.Words), and the perplexity scores produced by each model, which was also found for Wikipedia articles. However, further analyses suggested that the observed effect might be largely influenced by a consistent set of outliers, with very low perplexity scores. In other words, there appears to be another mediating variable influencing how challenging is for an LLM to model a dream report. Overall, the analysis further weakened the hypothesis that dream reports are rather unique strings of texts. All DreamBank&#x00027;s results, from the negative correlation to the outliers&#x00027; effect and the shape of the distribution, found a strong match in the results produced by the models when tested on Wikipedia data.</p>
<p>The results based on gender and visual impairment further challenged the strength of the negative correlation between perplexity and word count. On the one hand, under both models, reports from blind participants did resulted in significantly lower perplexity, but no significant effect was found between the two groups in terms of reports&#x00027; length. Even more strikingly, in the case of gender, the group with significantly lower perplexity scores (i.e., male) turned out to produce also significantly shorter reports. Again, these results patterns were stable across the two models. These discrepancies suggest that what really has an impact on the ability of the model to process a given report might have less to do with the number of words and more with the <italic>type</italic> of words in a report. A similar conclusion was also proposed in Bertolini et al. (<xref ref-type="bibr" rid="B4">2024a</xref>). Using an out-of-distribution ablation experiment, it was shown that leaving a specific DreamBank series out of training made it difficult for the model to handle a specific emotion (e.g., &#x0201C;happiness&#x0201D; for the <monospace>Bea 1</monospace> series.). The authors noted that this could not be simply explained by the number of instances in the training data, and was likely related to the specific vocabulary used in that specific series to describe that particular emotion.</p>
<p>The work also adds more evidence to the existing body of scientific knowledge showing how the gender of a participant might impact the related dream report (Hall and Van De Castle, <xref ref-type="bibr" rid="B21">1966</xref>; Schrdel and Reinhard, <xref ref-type="bibr" rid="B47">2008</xref>; Wong et al., <xref ref-type="bibr" rid="B56">2016</xref>; Zheng et al., <xref ref-type="bibr" rid="B61">2024</xref>). While repeatedly observed, these differences were mainly constrained to a report&#x00027;s semantic content and/or grammatical structure, such as a reference to a specific emotion, use of violent language, or part-of-speech use. This work suggests that the observed distinction might have a very tangible effect since reports produced by male dreamers were found to be significantly easier, on average, to model by both GPT2 and OLMo. This likely suggests that the distinction is even deeper than previously noted, and might include a combination of content, vocabulary, and structure.</p>
<p>A possible explanation for the observed gender-based difference might come from the data used to train these models, which is largely scraped from the internet. Multiple reports and preliminary studies have identified a worldwide disproportion in internet usage that disadvantages female users (Breen et al., <xref ref-type="bibr" rid="B7">2025</xref>). This disproportion might not be limited to internet usage. For instance, in 2012, a Wikipedia blogpost estimated that up to 90% of its editors were men<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref>, a number later confirmed by a survey in 2018<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref>. More recently, researchers have used corpus-linguistics and word embedding to show that within (a large English-based corpus extracted from) the internet, the concepts of &#x0201C;people&#x0201D; and &#x0201C;person&#x0201D; do not appear to be gender neutral, but are more aligned with the concept of &#x0201C;men&#x0201D; (Bailey et al., <xref ref-type="bibr" rid="B3">2022</xref>). This misalignment was also observed in machine-human interaction. A preliminary work found that ChatGPT was more frequently perceived as male rather than female on a variety of tasks (Wong and Kim, <xref ref-type="bibr" rid="B57">2023</xref>). In other words, it is possible that LLMs might find male-generated dream reports easier to model and predict because they have been primarily trained on male-generated data.</p>
<p>Results suggesting that blind dreamers produced more predictable reports seem more difficult to frame in the current literature and knowledge. Multiple pieces of evidence across time have shown how blind participants express a significantly lower amount of visual features in their reports, predominantly presenting auditory, tactile and olfactory reference (Kirtley, <xref ref-type="bibr" rid="B27">1975</xref>; Hurovitz et al., <xref ref-type="bibr" rid="B24">1999</xref>; Meaidi et al., <xref ref-type="bibr" rid="B33">2014</xref>; Zheng et al., <xref ref-type="bibr" rid="B61">2024</xref>). However, Meaidi et al. (<xref ref-type="bibr" rid="B33">2014</xref>) showed that these differences can significantly vary between congenitally and late blind participants, and both series contain a mixture of congenitally and non-congenitally blind participants (although most have been for more than 20 years). A possibility might be that maintaining access to the visual modality while dreaming allows for a larger degree of abstraction and variance of dream content, leading sighted participants to generate more diverse reports, that can result in harder sequences to predict for the model. Regardless of this hypothesis, it is important to notice that, since the two series contain reports produced by several individuals&#x02014;approximately thirty&#x02014;with an age window spanning from 24 to 70, and remarkably different backgrounds, it is unlikely that a single participant drives the observed difference in perplexity scores.</p>
<p>Concerning the year of collection of each report, one might find the observed small effect as unexpected, considering that many reports were collected at a time when the internet existed only in the minds of visionary scientists and writers. However, this might be explained by the fact that the internet is a collection of extremely heterogeneous documents, that obviously include very old textual instances. It is hence possible that, while specific reports did not leak into the training data, their vocabulary and style might very well have. In other words, the model might have also been exposed to the form and vocabulary used in older reports.</p>
<p>Overall, we believe that this work adds an important piece of evidence to the literature investigating differences in dream experiences from different groups. We have long been aware that reports produced by participants with different gender or visual impairment tend to present significantly different content&#x02014;and hence different word distributions. The experiments proposed in this work, however, further suggest that these differences are not limited to <italic>which</italic> words these groups use, but also <italic>how</italic> these groups use words, and that these differences in word usage as a measurable impact on current NLP tools.</p>
<p>To conclude, it is important to notice that this work has three main limitations. First of all, while OLMo training set is fully open-source, WebText, the dataset used to train GPT2, is not, and it is hence harder to estimate possible data leakage from DreamBank. That is, whether a part of the test data used in this work was also included in the training data for the model. In their work, Radford et al. (<xref ref-type="bibr" rid="B43">2019</xref>) note that training text for GPT2 was scraped following outbound links from Reddit, with at least 3 karma, and one link connecting Reddit to DreamBank. However, the link reached the main page of DreamBank, which does not allow scraping dream reports. As shown by example codes (e.g., here<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref>), the main solution to acquire dream reports from DreamBank is to iteratively sample them via the <monospace>random sample</monospace> page, which requires actively entering specific settings&#x02014;such as series or number of words&#x02014;to print out a set of reports. In other words, it seems quite unlikely that a consistent part of the test data for this work was in fact also included in the training data for GPT2. Future work will have to focus on models like OLMO, where the full extent of the training data is available. This would ensure better comparison and understanding of other relevant phenomena, such as whether the difference in perlocutionary scores might be connected to a specific type of documents, like Wikipedia articles of web-scraped dialogue, and with what strength. Second, the language of tested items was limited to English. Third, the adopted dream report dataset, DreamBank, is not fully transparent about the extent to which the reports were manipulated. The extended amount of grammatical errors and informal structures/forms found upon a manual inspection of a (limited) set of reports suggested that the data went through a very limited manipulation, but this can not be widely confirmed. Future work will have to investigate how strongly these findings can be generalized to other languages and dream datasets, as well as to provide a more detailed explanation of what might make a report more complex to predict for a current LLMs, taking more into consideration semantic content and syntactic structures.</p></sec>
<sec sec-type="conclusions" id="s5">
<title>5 Conclusion</title>
<p>This study has provided compelling evidence that dream reports are not the unpredictable textual entities they were once thought to be. By employing a set of large language models to analyze and predict the textual content of dream reports and compare it with standardized texts from Wikipedia, the research has shown that dream reports are, on average, more predictable than Wikipedia articles. This finding challenges the assumption that dream content is too peculiar or bizarre for models trained on web-based corpora. Additionally, the study has uncovered intriguing differences in predictability related to the gender and visual impairment of dream report authors, suggesting that these factors significantly influence the language models&#x00027; performance. These results not only contribute to our understanding of dream report characteristics but also have implications for the use of natural language processing tools in dream research. The insights of the presented study into the predictability of dream reports and the factors that affect it open the path for future research into the complex ways in which different groups express their dream experiences.</p></sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="ethics-statement" id="s7">
<title>Ethics statement</title>
<p>Ethical approval was not required for the study involving humans in accordance with the local legislation and institutional requirements. Written informed consent to participate in this study was not required from the participants or the participants&#x00027; legal guardians/next of kin in accordance with the national legislation and the institutional requirements.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>LB: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Software, Validation, Visualization, Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. SC: Formal analysis, Methodology, Writing &#x02013; review &#x00026; editing. JW: Project administration, Supervision, Writing &#x02013; review &#x00026; editing.</p>
</sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research and/or publication of this article. This research was partially conducted while the LB was at the University of Sussex. This research was partially supported by the EU Horizon 2020 project HumanE-AI (grant no. 952026).</p>
</sec>
<ack><p>We would like to thank the colleagues of the Digital Health Unit (JRC.F7) at the Joint Research Centre of the European Commission for their helpful guidance and support.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s10">
<title>Generative AI statement</title>
<p>The author(s) declare that no Gen AI was used in the creation of this manuscript.</p></sec><sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="disclaimer" id="s12">
<title>Author disclaimer</title>
<p>The views expressed are purely those of the authors and may not in any circumstance be regarded as stating an official position of the European Commission.</p>
</sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://huggingface.co/docs/transformers/perplexity">https://huggingface.co/docs/transformers/perplexity</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://dreambank.net/">https://dreambank.net/</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup>We use the <monospace>wikitext2-v1-raw</monospace> subset from <ext-link ext-link-type="uri" xlink:href="https://huggingface.co/datasets/Salesforce/wikitext">https://huggingface.co/datasets/Salesforce/wikitext</ext-link>.</p></fn>
<fn id="fn0004"><p><sup>4</sup><ext-link ext-link-type="uri" xlink:href="https://diff.wikimedia.org/2012/04/27/nine-out-of-ten-wikipedians-continue-to-be-men/">https://diff.wikimedia.org/2012/04/27/nine-out-of-ten-wikipedians-continue-to-be-men/</ext-link></p></fn>
<fn id="fn0005"><p><sup>5</sup><ext-link ext-link-type="uri" xlink:href="https://meta.wikimedia.org/wiki/Community_Insights/2018_Report/Contributors">https://meta.wikimedia.org/wiki/Community_Insights/2018_Report/Contributors</ext-link></p></fn>
<fn id="fn0006"><p><sup>6</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/mattbierner/DreamScrape">https://github.com/mattbierner/DreamScrape</ext-link></p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altszyler</surname> <given-names>E.</given-names></name> <name><surname>Ribeiro</surname> <given-names>S.</given-names></name> <name><surname>Sigman</surname> <given-names>M.</given-names></name> <name><surname>Slezak</surname> <given-names>D. F.</given-names></name></person-group> (<year>2017</year>). <article-title>The interpretation of dream meaning: resolving ambiguity using latent semantic analysis in a small corpus of text</article-title>. <source>Consciousn. Cognit</source>. <volume>56</volume>, <fpage>178</fpage>&#x02013;<lpage>187</lpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2017.09.004</pub-id><pub-id pub-id-type="pmid">28943127</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andrews</surname> <given-names>S.</given-names></name> <name><surname>Hanna</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Investigating the psychological mechanisms underlying the relationship between nightmares, suicide and self-harm</article-title>. <source>Sleep Med. Rev</source>. <volume>54</volume>:<fpage>101352</fpage>. <pub-id pub-id-type="doi">10.1016/j.smrv.2020.101352</pub-id><pub-id pub-id-type="pmid">32739825</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bailey</surname> <given-names>A. H.</given-names></name> <name><surname>Williams</surname> <given-names>A.</given-names></name> <name><surname>Cimpian</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Based on billions of words on the internet, PEOPLE = MEN</article-title>. <source>Sci. Adv</source>. <volume>8</volume>:<fpage>eabm2463</fpage>. <pub-id pub-id-type="doi">10.1126/sciadv.abm2463</pub-id><pub-id pub-id-type="pmid">35363515</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bertolini</surname> <given-names>L.</given-names></name> <name><surname>Elce</surname> <given-names>V.</given-names></name> <name><surname>Michalak</surname> <given-names>A.</given-names></name> <name><surname>Widhoelzl</surname> <given-names>H.-S.</given-names></name> <name><surname>Bernardi</surname> <given-names>G.</given-names></name> <name><surname>Weeds</surname> <given-names>J.</given-names></name></person-group> (<year>2024a</year>). <article-title>&#x0201C;Automatic annotation of dream report&#x00027;s emotional content with large language models,&#x0201D;</article-title> in <source>Proceedings of the 9th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2024)</source>, eds. A. Yates, B. Desmet, E. Prud&#x00027;hommeaux, A. Zirikly, S. Bedrick, S. MacAvaney, et al. (St. Julians: Association for Computational Linguistics), <fpage>92</fpage>&#x02013;<lpage>107</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bertolini</surname> <given-names>L.</given-names></name> <name><surname>Michalak</surname> <given-names>A.</given-names></name> <name><surname>Weeds</surname> <given-names>J.</given-names></name></person-group> (<year>2024b</year>). <article-title>Dreamy: a library for the automatic analysis and annotation of dream reports with multilingual large language models</article-title>. <source>Sleep Med</source>. <volume>115</volume>, <fpage>406</fpage>&#x02013;<lpage>407</lpage>. <pub-id pub-id-type="doi">10.1016/j.sleep.2023.11.1092</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blagrove</surname> <given-names>M.</given-names></name> <name><surname>Farmer</surname> <given-names>L.</given-names></name> <name><surname>Williams</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>The relationship of nightmare frequency and nightmare distress to well-being</article-title>. <source>J. Sleep Res</source>. <volume>13</volume>, <fpage>129</fpage>&#x02013;<lpage>136</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2869.2004.00394.x</pub-id><pub-id pub-id-type="pmid">15175092</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breen</surname> <given-names>C.</given-names></name> <name><surname>Fatehkia</surname> <given-names>M.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name> <name><surname>Zhao</surname> <given-names>X.</given-names></name> <name><surname>Leasure</surname> <given-names>D. R.</given-names></name> <name><surname>Weber</surname> <given-names>I.</given-names></name> <name><surname>Kashyap</surname> <given-names>R.</given-names></name></person-group> (<year>2025</year>). <source>Mapping Subnational Gender Gaps in Internet and Mobile Adoption Using Social Media Data</source>. Center for Open Science. <pub-id pub-id-type="doi">10.31235/osf.io/qnzsw_v2</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bulkeley</surname> <given-names>K.</given-names></name> <name><surname>Graves</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Using the LIWC program to study dreams</article-title>. <source>Dreaming</source> <volume>28</volume>, <fpage>43</fpage>&#x02013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1037/drm0000071</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Colla</surname> <given-names>D.</given-names></name> <name><surname>Delsanto</surname> <given-names>M.</given-names></name> <name><surname>Agosto</surname> <given-names>M.</given-names></name> <name><surname>Vitiello</surname> <given-names>B.</given-names></name> <name><surname>Radicioni</surname> <given-names>D. P.</given-names></name></person-group> (<year>2022</year>). <article-title>Semantic coherence markers: The contribution of perplexity metrics</article-title>. <source>Artif. Intellig. Med</source>. <volume>134</volume>:<fpage>102393</fpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2022.102393</pub-id><pub-id pub-id-type="pmid">36462890</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cortal</surname> <given-names>G.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Sequence-to-sequence language models for character and emotion detection in dream narratives,&#x0201D;</article-title> in <source>Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)</source>, eds. N. Calzolari, M. Y. Kan, V. Hoste, A. Lenci, S. Sakti, and N. Xue (Torino: ELRA and ICCL), <fpage>14717</fpage>&#x02013;<lpage>14728</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cortes</surname> <given-names>C.</given-names></name> <name><surname>Vapnik</surname> <given-names>V.</given-names></name></person-group> (<year>1995</year>). <article-title>Support-vector networks</article-title>. <source>Mach. Learn</source>. <volume>20</volume>, <fpage>273</fpage>&#x02013;<lpage>297</lpage>. <pub-id pub-id-type="doi">10.1007/BF00994018</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Domhoff</surname> <given-names>G. W.</given-names></name></person-group> (<year>2017</year>). <source>The Emergence of Dreaming: Mind-Wandering, Embodied Simulation, and the Default Network</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Domhoff</surname> <given-names>G. W.</given-names></name> <name><surname>Schneider</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>Studying dream content using the archive and search engine on DreamBank.net</article-title>. <source>Consciousn. Cognit</source>. <volume>17</volume>, <fpage>1238</fpage>&#x02013;<lpage>1247</lpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2008.06.010</pub-id><pub-id pub-id-type="pmid">18682331</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dubey</surname> <given-names>A.</given-names></name> <name><surname>Jauhri</surname> <given-names>A.</given-names></name> <name><surname>Pandey</surname> <given-names>A.</given-names></name> <name><surname>Kadian</surname> <given-names>A.</given-names></name> <name><surname>Al-Dahle</surname> <given-names>A.</given-names></name> <name><surname>Letman</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>The LLAMA 3 Herd of Models</article-title>. <source>CoRR</source> abs/2407.21783. <pub-id pub-id-type="doi">10.48550/arXiv.2407.21783</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elce</surname> <given-names>V.</given-names></name> <name><surname>Handjaras</surname> <given-names>G.</given-names></name> <name><surname>Bernardi</surname> <given-names>G.</given-names></name></person-group> (<year>2021</year>). <article-title>The language of dreams: application of linguistics-based approaches for the automated analysis of dream experiences</article-title>. <source>Clocks &#x00026; Sleep</source> <volume>3</volume>, <fpage>495</fpage>&#x02013;<lpage>514</lpage>. <pub-id pub-id-type="doi">10.3390/clockssleep3030035</pub-id><pub-id pub-id-type="pmid">34563057</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fasano</surname> <given-names>G.</given-names></name> <name><surname>Franceschini</surname> <given-names>A.</given-names></name></person-group> (<year>1987</year>). <article-title>A multidimensional version of the kolmogorov test</article-title>. <source>Monthly Notices Royal Astronom. Soc</source>. <volume>225</volume>, <fpage>155</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/225.1.155</pub-id><pub-id pub-id-type="pmid">38859537</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fogli</surname> <given-names>A.</given-names></name> <name><surname>Aiello</surname> <given-names>L. M.</given-names></name> <name><surname>Quercia</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Our dreams, our selves: automatic analysis of dream reports</article-title>. <source>Royal Soc. Open Sci</source>.<volume>7</volume>:<fpage>192080</fpage>. <pub-id pub-id-type="doi">10.1098/rsos.192080</pub-id><pub-id pub-id-type="pmid">32968499</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gemini Team Google: Petko Georgiev</surname> <given-names>Lei, V. I.</given-names></name> <name><surname>Burnell</surname> <given-names>R.</given-names></name> <name><surname>Bai</surname> <given-names>L.</given-names></name> <name><surname>Gulati</surname> <given-names>A.</given-names></name> <name><surname>Tanzer</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Gemini 1.5: unlocking multimodal understanding across millions of tokens of context</article-title>. <source>CoRR</source> abs/2403.05530. <pub-id pub-id-type="doi">10.48550/arXiv.2403.05530</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Groeneveld</surname> <given-names>D.</given-names></name> <name><surname>Beltagy</surname> <given-names>I.</given-names></name> <name><surname>Walsh</surname> <given-names>P.</given-names></name> <name><surname>Bhagia</surname> <given-names>A.</given-names></name> <name><surname>Kinney</surname> <given-names>R.</given-names></name> <name><surname>Tafjord</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>OLMo: Accelerating the science of language models</article-title>. <source>arXiv [Preprint]</source>. <pub-id pub-id-type="doi">10.18653/v1/2024.acl-long.841</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gutman Music</surname> <given-names>M.</given-names></name> <name><surname>Holur</surname> <given-names>P.</given-names></name> <name><surname>Bulkeley</surname> <given-names>K.</given-names></name></person-group> (<year>2022</year>). <article-title>Mapping dreams in a computational space: A phrase-level model for analyzing fight/flight and other typical situations in dream reports</article-title>. <source>Consciousn. Cognit</source>. <volume>106</volume>:<fpage>103428</fpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2022.103428</pub-id><pub-id pub-id-type="pmid">36341867</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hall</surname> <given-names>C. S.</given-names></name> <name><surname>Van De Castle</surname> <given-names>R. L.</given-names></name></person-group> (<year>1966</year>). <source>The Content Analysis of Dreams</source>. <publisher-loc>Norwalk, CT</publisher-loc>: <publisher-name>Appleton-Century-Crofts</publisher-name>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hauri</surname> <given-names>P.</given-names></name></person-group> (<year>1975</year>). <article-title>&#x0201C;Categorization of sleep mental activity for psychophysiological studies,&#x0201D;</article-title> in <source>The Experimental Study of Sleep: Methodological Problems</source>, <fpage>271</fpage>&#x02013;<lpage>281</lpage>.<pub-id pub-id-type="pmid">40190605</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holm</surname> <given-names>S.</given-names></name></person-group> (<year>1979</year>). <article-title>A simple sequentially rejective multiple test procedure</article-title>. <source>Scand. J. Statist</source>. <volume>6</volume>, <fpage>65</fpage>&#x02013;<lpage>70</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hurovitz</surname> <given-names>C. S.</given-names></name> <name><surname>Dunn</surname> <given-names>S.</given-names></name> <name><surname>Domhoff</surname> <given-names>G. W.</given-names></name> <name><surname>Fiss</surname> <given-names>H.</given-names></name></person-group> (<year>1999</year>). <article-title>The dreams of blind men and women: A replication and extension of previous findings</article-title>. <source>Dreaming</source> <volume>9</volume>, <fpage>183</fpage>&#x02013;<lpage>193</lpage>. <pub-id pub-id-type="doi">10.1023/A:1021397817164</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huyen</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <source>Evaluation Metrics for Language Modeling</source>. <publisher-loc>Stanford, CA</publisher-loc>: <publisher-name>The Gradient</publisher-name>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kahan</surname> <given-names>T. L.</given-names></name> <name><surname>LaBerge</surname> <given-names>S. P.</given-names></name></person-group> (<year>2011</year>). <article-title>Dreaming and waking: similarities and differences revisited</article-title>. <source>Consciousn. Cognit</source>. <volume>20</volume>, <fpage>494</fpage>&#x02013;<lpage>514</lpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2010.09.002</pub-id><pub-id pub-id-type="pmid">20933437</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kirtley</surname> <given-names>D. D.</given-names></name></person-group> (<year>1975</year>). <source>The Psychology of Blindness</source>. <publisher-loc>Chicago, IL</publisher-loc>: <publisher-name>Nelson-Hall</publisher-name>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kobayashi</surname> <given-names>I.</given-names></name> <name><surname>Sledjeski</surname> <given-names>E. M.</given-names></name> <name><surname>Spoonster</surname> <given-names>E.</given-names></name> <name><surname>Fallon Jr</surname> <given-names>W. F.</given-names></name> <name><surname>Delahanty</surname> <given-names>D. L.</given-names></name></person-group> (<year>2008</year>). <article-title>Effects of early nightmares on the development of sleep disturbances in motor vehicle accident victims</article-title>. <source>J. Traumatic Stress</source> <volume>21</volume>, <fpage>548</fpage>&#x02013;<lpage>555</lpage>. <pub-id pub-id-type="doi">10.1002/jts.20368</pub-id><pub-id pub-id-type="pmid">19107721</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Landauer</surname> <given-names>T. K.</given-names></name> <name><surname>Dumais</surname> <given-names>S. T.</given-names></name></person-group> (<year>1997</year>). <article-title>A solution to plato&#x00027;s problem: the latent semantic analysis theory of acquisition, induction, and representation of knowledge</article-title>. <source>Psychol. Rev</source>. <volume>104</volume>, <fpage>211</fpage>&#x02013;<lpage>240</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.104.2.211</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mallett</surname> <given-names>R.</given-names></name> <name><surname>Picard-Deland</surname> <given-names>C.</given-names></name> <name><surname>Pigeon</surname> <given-names>W.</given-names></name> <name><surname>Wary</surname> <given-names>M.</given-names></name> <name><surname>Grewal</surname> <given-names>A.</given-names></name> <name><surname>Blagrove</surname> <given-names>M.</given-names></name> <name><surname>Carr</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>The relationship between dreams and subsequent morning mood using self-reports and text analysis</article-title>. <source>Affect. Sci</source>. <volume>3</volume>, <fpage>400</fpage>&#x02013;<lpage>405</lpage>. <pub-id pub-id-type="doi">10.1007/s42761-021-00080-8</pub-id><pub-id pub-id-type="pmid">36046002</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mathes</surname> <given-names>J.</given-names></name> <name><surname>Schredl</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Gender differences in dream content: Are they related to personality?</article-title> <source>Int. J. Dream Res</source>. <volume>6</volume>, <fpage>104</fpage>&#x02013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.11588/ijodr.2013.2.10954</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McNamara</surname> <given-names>P.</given-names></name> <name><surname>Duffy-Deno</surname> <given-names>K.</given-names></name> <name><surname>Marsh</surname> <given-names>T.</given-names></name> <name><surname>Marsh</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>Dream content analysis using artificial intelligence</article-title>. <source>Int. J. Dream Res</source>. <volume>12</volume>:<fpage>1</fpage>. <pub-id pub-id-type="doi">10.11588/ijodr.2019.1.48744</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meaidi</surname> <given-names>A.</given-names></name> <name><surname>Jennum</surname> <given-names>P.</given-names></name> <name><surname>Ptito</surname> <given-names>M.</given-names></name> <name><surname>Kupers</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>The sensory construction of dreams and nightmare frequency in congenitally blind and late blind individuals</article-title>. <source>Sleep Med</source>. <volume>15</volume>, <fpage>586</fpage>&#x02013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1016/j.sleep.2013.12.008</pub-id><pub-id pub-id-type="pmid">24709309</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><collab>Meister C. and Cotterell, R..</collab></person-group> (<year>2021</year>). Language model evaluation beyond perplexity. In <italic>Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)</italic> (Online: Association for Computational Linguistics), <fpage>5328</fpage>&#x02013;<lpage>5339</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2021.acl-long.414</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Merity</surname> <given-names>S.</given-names></name> <name><surname>Xiong</surname> <given-names>C.</given-names></name> <name><surname>Bradbury</surname> <given-names>J.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Pointer sentinel mixture models,&#x0201D;</article-title> in <source>5th International Conference on Learning Representations (ICLR)</source>, 1&#x02013;13. Available online at: <ext-link ext-link-type="uri" xlink:href="https://openreview.net/forum?id=Byj72udxe">https://openreview.net/forum?id=Byj72udxe</ext-link></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <source>Efficient Estimation of Word Representations in Vector Space</source>.<pub-id pub-id-type="pmid">31752376</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mota</surname> <given-names>N. B.</given-names></name> <name><surname>Furtado</surname> <given-names>R.</given-names></name> <name><surname>Maia</surname> <given-names>P. P. C.</given-names></name> <name><surname>Copelli</surname> <given-names>M.</given-names></name> <name><surname>Ribeiro</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>Graph analysis of dream reports is especially informative about psychosis</article-title>. <source>Scient. Reports</source> <volume>4</volume>:<fpage>1</fpage>. <pub-id pub-id-type="doi">10.1038/srep03691</pub-id><pub-id pub-id-type="pmid">24424108</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nadeau</surname> <given-names>D.</given-names></name> <name><surname>Sabourin</surname> <given-names>C.</given-names></name> <name><surname>Koninck</surname> <given-names>J. D.</given-names></name> <name><surname>Matwin</surname> <given-names>S.</given-names></name> <name><surname>Turney</surname> <given-names>P. D.</given-names></name></person-group> (<year>2006</year>). <article-title>&#x0201C;Automatic dream sentiment analysis,&#x0201D;</article-title> in <source>Proc. of the Workshop on Computational Aesthetics at the Twenty-First National Conf. on Artificial Intelligence</source> (<publisher-loc>Washington, DC</publisher-loc>: <publisher-name>AAAI</publisher-name>).</citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nir</surname> <given-names>Y.</given-names></name> <name><surname>Tononi</surname> <given-names>G.</given-names></name></person-group> (<year>2010</year>). <article-title>Dreaming and the brain: from phenomenology to neurophysiology</article-title>. <source>Trends Cognit. Sci</source>. <volume>14</volume>, <fpage>88</fpage>&#x02013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2009.12.001</pub-id><pub-id pub-id-type="pmid">20079677</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>OpenA</surname> <given-names>I.</given-names></name> <name><surname>Achiam</surname> <given-names>J.</given-names></name> <name><surname>Adler</surname> <given-names>S.</given-names></name> <name><surname>Agarwal</surname> <given-names>S.</given-names></name> <name><surname>Ahmad</surname> <given-names>L.</given-names></name> <name><surname>Akkaya</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2024</year>). <source>Gpt-4 Technical Report</source>.</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peacock</surname> <given-names>J. A.</given-names></name></person-group> (<year>1983</year>). <article-title>Two-dimensional goodness-of-fit testing in astronomy</article-title>. <source>Monthly Notices Royal Astronom. Soc</source>. <volume>202</volume>, <fpage>615</fpage>&#x02013;<lpage>627</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/202.3.615</pub-id><pub-id pub-id-type="pmid">38859537</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pennebaker</surname> <given-names>J. W.</given-names></name> <name><surname>Boyd</surname> <given-names>R. L.</given-names></name> <name><surname>Jordan</surname> <given-names>K.</given-names></name> <name><surname>Blackburn</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;The development and psychometric properties of LIWC2015,&#x0201D;</article-title> in <source>Technical Report</source>. <publisher-loc>Austin TX</publisher-loc>: <publisher-name>University of Texas at Austin</publisher-name>.<pub-id pub-id-type="pmid">35330723</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Child</surname> <given-names>R.</given-names></name> <name><surname>Luan</surname> <given-names>D.</given-names></name> <name><surname>Amodei</surname> <given-names>D.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Language models are unsupervised multitask learners,&#x0201D;</article-title> in <source>OpenAI Blog</source>.<pub-id pub-id-type="pmid">35637722</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Razavi</surname> <given-names>A. H.</given-names></name> <name><surname>Matwin</surname> <given-names>S.</given-names></name> <name><surname>Koninck</surname> <given-names>J. D.</given-names></name> <name><surname>Amini</surname> <given-names>R. R.</given-names></name></person-group> (<year>2013</year>). <article-title>Dream sentiment analysis using second order soft co-occurrences (SOSCO) and time course representations</article-title>. <source>J. Intellig. Inform. Syst</source>. <volume>42</volume>, <fpage>393</fpage>&#x02013;<lpage>413</lpage>. <pub-id pub-id-type="doi">10.1007/s10844-013-0273-4</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosen</surname> <given-names>M. G.</given-names></name></person-group> (<year>2018</year>). <article-title>How bizarre? A pluralist approach to dream content</article-title>. <source>Consciousn. Cognit</source>. <volume>62</volume>, <fpage>148</fpage>&#x02013;<lpage>162</lpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2018.03.009</pub-id><pub-id pub-id-type="pmid">29739723</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanz</surname> <given-names>C.</given-names></name> <name><surname>Zamberlan</surname> <given-names>F.</given-names></name> <name><surname>Erowid</surname> <given-names>E.</given-names></name> <name><surname>Erowid</surname> <given-names>F.</given-names></name> <name><surname>Tagliazucchi</surname> <given-names>E.</given-names></name></person-group> (<year>2018</year>). <article-title>The experience elicited by hallucinogens presents the highest similarity to dreaming within a large database of psychoactive substance reports</article-title>. <source>Front. Neurosci</source>. <volume>12</volume>:<fpage>7</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2018.00007</pub-id><pub-id pub-id-type="pmid">29403350</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schrdel</surname> <given-names>M.</given-names></name> <name><surname>Reinhard</surname> <given-names>I.</given-names></name></person-group> (<year>2008</year>). <article-title>Gender differences in dream recall: a meta-analysis</article-title>. <source>J. Sleep Res</source>. <volume>17</volume>, <fpage>125</fpage>&#x02013;<lpage>131</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2869.2008.00626.x</pub-id><pub-id pub-id-type="pmid">18355162</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schredl</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Dream content analysis: Basic principles</article-title>. <source>Int. J. Dream Res</source>. <volume>3</volume>:<fpage>1</fpage>. <pub-id pub-id-type="doi">10.11588/ijodr.2010.1.474</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siclari</surname> <given-names>F.</given-names></name> <name><surname>Baird</surname> <given-names>B.</given-names></name> <name><surname>Perogamvros</surname> <given-names>L.</given-names></name> <name><surname>Bernardi</surname> <given-names>G.</given-names></name> <name><surname>LaRocque</surname> <given-names>J. J.</given-names></name> <name><surname>Riedner</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>The neural correlates of dreaming</article-title>. <source>Nat. Neurosci</source>. <volume>20</volume>, <fpage>872</fpage>&#x02013;<lpage>878</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4545</pub-id><pub-id pub-id-type="pmid">28394322</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Skancke</surname> <given-names>J. F.</given-names></name> <name><surname>Holsen</surname> <given-names>I.</given-names></name> <name><surname>Schredl</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Continuity between waking life and dreams of psychiatric patients: a review and discussion of the implications for dream research</article-title>. <source>Int. J. Dream Res</source>. <volume>7</volume>, <fpage>39</fpage>&#x02013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.11588/ijodr.2014.1.12184</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soldaini</surname> <given-names>L.</given-names></name> <name><surname>Kinney</surname> <given-names>R.</given-names></name> <name><surname>Bhagia</surname> <given-names>A.</given-names></name> <name><surname>Schwenk</surname> <given-names>D.</given-names></name> <name><surname>Atkinson</surname> <given-names>D.</given-names></name> <name><surname>Authur</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2024</year>). Dolma: 559 an open corpus of three trillion tokens for language model pretraining research.</citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thompson</surname> <given-names>A.</given-names></name> <name><surname>Lereya</surname> <given-names>S. T.</given-names></name> <name><surname>Lewis</surname> <given-names>G.</given-names></name> <name><surname>Zammit</surname> <given-names>S.</given-names></name> <name><surname>Fisher</surname> <given-names>H. L.</given-names></name> <name><surname>Wolke</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>Childhood sleep disturbance and risk of psychotic experiences at 18: UK birth cohort</article-title>. <source>Br. J. Psychiat</source>. <volume>207</volume>, <fpage>23</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1192/bjp.bp.113.144089</pub-id><pub-id pub-id-type="pmid">25953892</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmar</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Attention is all you need,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, eds. I. Guyon, U. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Red Hook, NY: Curran Associates, Inc).</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Veli&#x0010D;kovi&#x00107;</surname> <given-names>P.</given-names></name></person-group> (<year>2023</year>). <article-title>Everything is connected: Graph neural networks</article-title>. <source>Curr. Opin. Struct. Biol</source>. <volume>79</volume>:<fpage>102538</fpage>. <pub-id pub-id-type="doi">10.1016/j.sbi.2023.102538</pub-id><pub-id pub-id-type="pmid">36764042</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Manzoor</surname> <given-names>M. A.</given-names></name> <name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Georgiev</surname> <given-names>G. N.</given-names></name> <name><surname>Das</surname> <given-names>R. J.</given-names></name> <name><surname>Nakov</surname> <given-names>P.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Factuality of large language models: A survey,&#x0201D;</article-title> in <source>Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing</source>, eds. Y. Al-Onaizan, M. Bansal, and Y. N. Chen (Miami, FL: Association for Computational Linguistics), <fpage>19519</fpage>&#x02013;<lpage>19529</lpage>.</citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wong</surname> <given-names>C.</given-names></name> <name><surname>Amini</surname> <given-names>R.</given-names></name> <name><surname>Koninck</surname> <given-names>J. D.</given-names></name></person-group> (<year>2016</year>). <article-title>Automatic gender detection of dream reports: A promising approach</article-title>. <source>Consciousn. Cognit</source>. <volume>44</volume>, <fpage>20</fpage>&#x02013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2016.06.004</pub-id><pub-id pub-id-type="pmid">27344136</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Wong</surname> <given-names>J.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>ChatGPT is more likely to be perceived as male than female</article-title>. <source>arXiv [Preprint]</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2305.12564">https://arxiv.org/abs/2305.12564</ext-link></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>C. K.-C.</given-names></name></person-group> (<year>2022</year>). <article-title>Automated analysis of dream sentiment royal road to dream dynamics?</article-title> <source>Dreaming</source> <volume>32</volume>, <fpage>33</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1037/drm0000189</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Schweickert</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>Comparing hall van de castle coding and linguistic inquiry and word count using canonical correlation analysis</article-title>. <source>Dreaming</source> <volume>31</volume>, <fpage>207</fpage>&#x02013;<lpage>224</lpage>. <pub-id pub-id-type="doi">10.1037/drm0000173</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Schweickert</surname> <given-names>R.</given-names></name></person-group> (<year>2023</year>). <article-title>Differentiating dreaming and waking reports with automatic text analysis and support vector machines</article-title>. <source>Conscious. Cognit</source>. <volume>107</volume>:<fpage>103439</fpage>. <pub-id pub-id-type="doi">10.1016/j.concog.2022.103439</pub-id><pub-id pub-id-type="pmid">36463797</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Schweickert</surname> <given-names>R.</given-names></name> <name><surname>Song</surname> <given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>Automatic dream content analysis finds effects of gender, age, and blindness on word use</article-title>. <source>Dreaming</source>. <volume>35</volume>, <fpage>68</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1037/drm0000287</pub-id></citation>
</ref>
</ref-list>
</back>
</article> 