<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="review-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2025.1654716</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Conceptual Analysis</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Artificial Creativity: from predictive AI to Generative System 3</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Ch&#x00E1;vez-Autor</surname>
<given-names>Juan Carlos</given-names>
</name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<xref ref-type="author-notes" rid="fn0003"><sup>&#x2020;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/3010806/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>College of Psychology, Keiser University</institution>, <addr-line>Fort Lauderdale, FL</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Facultad de Comunicaci&#x00F3;n, Universidad Panamericana</institution>, <addr-line>Mexico City</addr-line>, <country>Mexico</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Business, Economics and Law, University of Gothenburg</institution>, <addr-line>Gothenburg</addr-line>, <country>Sweden</country></aff>
<aff id="aff4"><sup>4</sup><institution>Bio-Intelligence and Creativity Institute</institution>, <addr-line>Playa del Carmen</addr-line>, <country>Mexico</country></aff>
<author-notes>
<fn fn-type="edited-by" id="fn0001">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2335590/overview">Eric Chalmers</ext-link>, Mount Royal University, Canada</p>
</fn>
<fn fn-type="edited-by" id="fn0002">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1111983/overview">Predrag K. Nikolic</ext-link>, Swinburne University of Technology Sarawak Campus, Malaysia</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2193683/overview">Iv&#x00E1;n Durango</ext-link>, University of Castilla La Mancha, Spain</p>
</fn>
<corresp id="c001">&#x002A;Correspondence: Juan Carlos Ch&#x00E1;vez-Autor, <email>jcchavez@up.edu.mx</email>; <email>jc@g-8d.com</email></corresp>
<fn fn-type="other" id="fn0003"><p><sup>&#x2020;</sup>ORCID: Juan Carlos Ch&#x00E1;vez-Autor, <ext-link ext-link-type="uri" xlink:href="https://orcid.org/0009-0009-3887-7880">orcid.org/0009-0009-3887-7880</ext-link></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>15</day>
<month>10</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>8</volume>
<elocation-id>1654716</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>07</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>09</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2025 Ch&#x00E1;vez-Autor.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Ch&#x00E1;vez-Autor</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Large language models generate fluent text yet often fail to sustain novelty, task relevance, and diversity across extended contexts. We argue this shortfall persists because current systems implement only fragments of a tri-process loop that supports human creativity: spontaneous ideation in the default-mode network (DMN; broadly <italic>System 1</italic>&#x2013;like), goal-directed evaluation in the central-executive network (CEN; broadly <italic>System 2</italic>&#x2013;like), and a metacognitive integrator&#x2014;<italic>System 3</italic>&#x2014;that, via neuromodulatory gain control, shifts between exploration and focused control. We introduce <italic>Generative System 3</italic> (GS-3), an architecture-agnostic design pattern with three roles: a high-entropy <italic>generator</italic>, a learned <italic>critic</italic>, and an adaptive <italic>gain controller</italic>. Beyond &#x201C;pure prediction&#x201D; and simple &#x201C;reflective prompting,&#x201D; GS-3 identifies the missing pieces for <italic>Artificial Creativity</italic>: an internal evaluator, endogenous control over sampling entropy, and adaptive priors maintained across extended contexts. This conceptual analysis (i) formalizes <italic>novelty</italic>, <italic>usefulness</italic>, and <italic>diversity</italic> with operational definitions; (ii) develops multiple gain-update policies (exponential, linear, logistic) with stability constraints and sensitivity expectations; (iii) derives falsifiable behavioral indices&#x2014;<italic>associative-distance density</italic>, <italic>analytic-verification ratio</italic>, and <italic>convergence latency</italic>&#x2014;with pass&#x2013;fail criteria; and (iv) provides a proof-of-concept blueprint and evaluation protocol (tasks, metrics, ablations, reproducibility kit). We position GS-3 relative to computational-creativity and co-creative frameworks, and delineate where brain&#x2013;model analogies are functional rather than literal. Ethical guidance addresses bias, cultural homogenization, and reward gaming of proxy objectives (often termed &#x201C;dopamine hacking&#x201D;) through plural critics, transparent logging, and outcome-tied entropy caps. The result is a testable roadmap for transitioning from regulated prediction to genuinely creative generative systems.</p>
</abstract>
<kwd-group>
<kwd>Artificial Creativity</kwd>
<kwd>Generative System 3 (GS-3)</kwd>
<kwd>large language models (LLMs)</kwd>
<kwd>adaptive gain control</kwd>
<kwd>computational creativity</kwd>
<kwd>System 3 (metacognitive control)</kwd>
<kwd>exploration&#x2013;exploitation trade-off</kwd>
<kwd>tri-process cognition</kwd>
</kwd-group>
<counts>
<fig-count count="0"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="31"/>
<page-count count="12"/>
<word-count count="8456"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Natural Language Processing</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="sec1">
<label>1</label>
<title>Introduction&#x2014;from fluent prediction to creative control</title>
<p>Large language models (LLMs) now produce remarkably fluent text, yet they often struggle to sustain novelty, task relevance, and diversity across extended contexts. We argue this shortfall persists because current systems implement only fragments of a tri-process loop that supports human creativity: spontaneous ideation in the default-mode network (DMN; broadly aligned with <italic>System 1</italic>), goal-directed evaluation in the central-executive network (CEN; broadly aligned with <italic>System 2</italic>), and a metacognitive integrator&#x2014;<italic>System 3</italic>&#x2014;that, supported by neuromodulatory gain, shifts the mind between exploration and focused control to achieve integrative creative outcomes. In this view, dual-process accounts are necessary but not sufficient; System 3 coordinates and regulates how ideas are generated and pruned, turning &#x201C;thoughts of thoughts&#x201D; into adaptive action. The convergence of these mechanisms in machine systems is what we call <italic>Artificial Creativity</italic>.</p>
<sec id="sec2">
<label>1.1</label>
<title>The missing link between predictive AI and creative cognition</title>
<p>From a neuroscience vantage, creative performance depends on flexible DMN&#x2013;CEN interaction, with dopaminergic signals modulating the exploration&#x2013;exploitation balance (<xref ref-type="bibr" rid="ref5">Chen et al., 2025</xref>; <xref ref-type="bibr" rid="ref28">Shine, 2019</xref>; <xref ref-type="bibr" rid="ref30">Westbrook et al., 2021</xref>). By contrast, most LLMs behave like DMN-only decoders: excellent at sequence extension, but lacking an internal evaluator and endogenous gain control to decide when to broaden or narrow the search. Bridging this gap requires importing System 3 principles into model design.</p>
</sec>
<sec id="sec3">
<label>1.2</label>
<title>Why a conceptual analysis now?</title>
<p>Evidence on both sides is converging. LLMs can match or exceed median human fluency on some divergent-thinking tasks, yet at scale, their outputs tend to homogenize, reducing collective diversity (<xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>). Human&#x2013;AI co-creation increases speed and fluency but, without structure, can dampen variety or drift from task goals (<xref ref-type="bibr" rid="ref4">Chen and Chan, 2024</xref>; <xref ref-type="bibr" rid="ref3">Chakrabarty et al., 2024</xref>). Meanwhile, covert neurofeedback that strengthens DMN&#x2013;CEN coupling elevates originality in human participants (<xref ref-type="bibr" rid="ref23">Luchini et al., 2025</xref>). Together, these findings motivate a synthesis that links cognitive theory, neural evidence, and generative-model engineering&#x2014;and states testable criteria for when an artificial system merits the label creative. Throughout, any brain&#x2013;model correspondences are treated as functional analogies, not biological isomorphisms.</p>
</sec>
<sec id="sec4">
<label>1.3</label>
<title>Contribution and scope</title>
<p>This conceptual analysis does not report new empirical data. Instead, it advances a falsifiable framework&#x2014;<italic>Generative System 3</italic> (GS-3)&#x2014;and a concrete evaluation program. GS-3 is an architecture-agnostic design pattern with three roles: a high-entropy <italic>generator</italic> (idea expansion), a learned <italic>critic</italic> (context-sensitive appraisal), and an adaptive <italic>gain controller</italic> (endogenous regulation of sampling entropy). We contribute four elements: (i) operational definitions of <italic>novelty</italic> (distributional distance to a baseline), <italic>usefulness</italic> (task-conditioned utility), and <italic>diversity</italic> (across-run dispersion); (ii) behavioral indices with pass&#x2013;fail criteria&#x2014;<italic>associative-distance density</italic>, <italic>analytic-verification ratio</italic>, and <italic>convergence latency</italic>&#x2014;so the theory can be falsified in practice; (iii) a mathematical treatment of gain policies (exponential, linear, logistic) with stability constraints and sensitivity expectations; and (iv) a proof-of-concept blueprint and evaluation protocol (tasks, metrics, ablations, reproducibility kit) that research groups can implement.</p>
<p>We situate these contributions along a predictive-to-generative continuum. <italic>Pure prediction</italic> extends sequences without internal evaluation. <italic>Regulated generation</italic> introduces external controls (e.g., temperature, top-<italic>k</italic>) but still lacks an inner judge. <italic>Reflective generation</italic> uses self-prompted critique yet remains scaffold-dependent. <italic>GS-3&#x2013;level creativity</italic> emerges only when a system (a) cycles autonomously between idea expansion and evaluative pruning, (b) adjusts its own sampling entropy in response to real-time reward-prediction error, and (c) maintains adaptive priors over long contexts.</p>
</sec>
<sec id="sec5">
<label>1.4</label>
<title>Roadmap</title>
<p>Section 2 positions GS-3 within computational-creativity traditions, LLM-based co-creation, and alternative theories. Section 3 summarizes the DMN/CEN/dopamine template and its limits (functional analogies, not isomorphisms). Section 4 presents the GS-3 architecture with formal definitions, falsification tests, and gain-policy mathematics. Section 5 offers a proof-of-concept blueprint and evaluation protocol (hypotheses, metrics, ablations). Section 6 locates today&#x2019;s systems on the continuum and identifies gaps GS-3 fills. Section 7 expands ethics and governance with concrete mitigation steps. Section 8 concludes with a Discussion that synthesizes contributions, states boundary conditions, and outlines future work. By unifying cognitive theory, network neuroscience, and AI engineering, we aim to establish Artificial Creativity as a testable construct and to provide a practical roadmap for building and auditing genuinely creative generative systems.</p>
</sec>
</sec>
<sec id="sec6">
<label>2</label>
<title>Computational creativity traditions</title>
<p>Early taxonomies emphasize what counts as creative behavior and how to evaluate it. Boden&#x2019;s distinctions between psychological and historical creativity (P- vs. H-creativity) foreground the mechanisms of combinational, exploratory, and transformational search (<xref ref-type="bibr" rid="ref1">Boden, 2004</xref>). Formal accounts characterize creative systems by their generative space and constraints (<xref ref-type="bibr" rid="ref31">Wiggins, 2006</xref>), while evaluation frameworks propose measurable criteria for attributing creativity to artifacts or systems (<xref ref-type="bibr" rid="ref26">Ritchie, 2007</xref>; <xref ref-type="bibr" rid="ref18">Jordanous, 2012</xref>). System exemplars, such as The Painting Fool, demonstrate end-to-end pipelines that produce artifacts and internal justifications (<xref ref-type="bibr" rid="ref8">Colton, 2012</xref>).</p>
<p>These traditions supply two pillars we retain: (a) creativity requires both a generator of candidates and a process that evaluates them in context, and (b) claims should be tied to operational criteria. GS-3 extends this foundation by adding an explicit, adaptive gain mechanism that regulates the breadth/depth of search online and by specifying falsifiable behavioral indices (novelty, usefulness, diversity) together with pass/fail thresholds. In short, GS-3 preserves the generator&#x2013;evaluator logic but makes the regulator a first-class component with its own dynamics.</p>
<sec id="sec7">
<label>2.1</label>
<title>LLM creativity and co-creation</title>
<p>Modern LLM pipelines span a pragmatic continuum. <italic>Pure prediction</italic> extends sequences using maximum-likelihood training (<xref ref-type="bibr" rid="ref29">Sutskever et al., 2014</xref>). <italic>Regulated generation</italic> introduces external controls&#x2014;temperature and top-<italic>k</italic>/top-<italic>p</italic> settings shape randomness, and beam search maintains several high-probability continuations&#x2014;which can prevent degeneracy but do not install an inner judge (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>). Steering methods alter token probabilities directly (e.g., plug-and-play controls for attributes without backbone retraining) (<xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>). <italic>Reflective prompting</italic> scaffolds brief internal critique (e.g., chain-of-thought) and retrieval-augmented generation injects external knowledge to improve coherence and factuality (<xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>; <xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>). <italic>Self-improvement/self-correction loops</italic> iteratively revise drafts with model feedback (<xref ref-type="bibr" rid="ref19">Kamoi et al., 2024</xref>; <xref ref-type="bibr" rid="ref9">Ding et al., 2024</xref>). <italic>Multi-agent set-ups</italic> coordinate multiple LLMs to critique and debate (e.g., generative agents) (<xref ref-type="bibr" rid="ref24">Park et al., 2023</xref>).</p>
<p>Empirically, co-creation studies show that LLM support often increases fluency and speed but can reduce variety without structured collaboration protocols (<xref ref-type="bibr" rid="ref4">Chen and Chan, 2024</xref>; <xref ref-type="bibr" rid="ref3">Chakrabarty et al., 2024</xref>). At the population scale, assistance can increase individual originality while decreasing collective diversity, consistent with homogenization risks (<xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>). Conceptually oriented analyses debate whether current LLMs meet creativity criteria and where limits remain (<xref ref-type="bibr" rid="ref13">Franceschelli and Musolesi, 2024</xref>; <xref ref-type="bibr" rid="ref12">Floridi and Chiriatti, 2020</xref>).</p>
<p>GS-3 aligns with these trajectories yet differs on one decisive point: the evaluator and the regulator are internal, learned, and adaptive. Rather than relying on hand-tuned temperatures, prompt engineering, or fixed debate scripts, GS-3 requires a critic that computes task-conditioned utility and a gain controller that adjusts sampling entropy based on reward-prediction error. This makes the exploration&#x2013;exploitation balance endogenous to the system and testable via ablations (remove critic or gain; swap update rules; vary the learning rate).</p>
</sec>
<sec id="sec8">
<label>2.2</label>
<title>Complementary theories and boundary conditions</title>
<p>Information-theoretic and intrinsic-motivation accounts explain why systems seek novelty or compressive structure (<xref ref-type="bibr" rid="ref27">Schmidhuber, 2010</xref>). Evolutionary approaches operationalize open-ended novelty (<xref ref-type="bibr" rid="ref21">Lehman and Stanley, 2011</xref>). Predictive-processing and free-energy views model perception and action as minimizing prediction error or free energy under learned priors (<xref ref-type="bibr" rid="ref7">Clark, 2013</xref>; <xref ref-type="bibr" rid="ref14">Friston, 2010</xref>). These perspectives illuminate why creative systems might alternate between broad exploration and tight verification.</p>
<p>GS-3 is compatible with these theories but adds a concrete control story: a generator produces candidates; a critic scores them relative to task and context; and a gain controller adjusts entropy and effort in real time, producing measurable signatures (e.g., shifts in associative-distance density and verification ratios). Where embodied and enactive views stress sensorimotor grounding, GS-3 can be seen as the cognitive-control core that any grounded agent still requires to manage the breadth and depth of search. Where information-theoretic approaches prize compression or novelty alone, GS-3 foregrounds usefulness by design through the critic&#x2019;s utility function. Finally, where predictive-processing emphasizes error minimization, GS-3 specifies when and how the system should temporarily widen its hypothesis space before re-engaging verification.</p>
<p>Together, these comparisons place GS-3 as a synthesis that retains the generator&#x2013;evaluator insight from computational creativity, adopts practical controls from contemporary LLM pipelines, and formalizes the missing adaptive regulator. Subsequent sections develop the architecture (Section 4), metrics, and gain policies with falsification tests (Section 4), and outline a proof-of-concept and evaluation protocol suitable for empirical validation (Section 5).</p>
</sec>
</sec>
<sec id="sec9">
<label>3</label>
<title>Neurobiological template</title>
<p>Creativity does not reside in a single cortical locus; it emerges from interactions among large-scale networks modulated by neuromodulatory systems. In broad terms, associative expansion is linked to the default-mode network (DMN), and evaluative control is associated with the central-executive network (CEN), while neuromodulators, such as dopamine, bias the system toward exploration or exploitation by altering integration and segregation dynamics. This section summarizes key findings and clarifies where brain&#x2013;model analogies are functional (useful for design) rather than literal (biological identity).</p>
<sec id="sec10">
<label>3.1</label>
<title>Large-scale network architecture: DMN&#x2013;CEN coupling</title>
<p>Resting-state and task-based studies converge on a picture in which creative performance is associated with flexible interaction between DMN hubs (e.g., medial prefrontal, posterior cingulate, temporoparietal regions) and CEN hubs (e.g., dorsolateral prefrontal, posterior parietal cortex). Using state-transition analyses of fMRI during divergent thinking, dynamic switching between DMN and executive-control states predicts higher originality and richer associative distance, consistent with the idea that creativity benefits from alternating expansion and evaluation rather than the dominance of either mode alone (<xref ref-type="bibr" rid="ref5">Chen et al., 2025</xref>). This dynamic view situates creativity as a property of network coupling over time, not a static activation pattern.</p>
</sec>
<sec id="sec11">
<label>3.2</label>
<title>Neuromodulatory gain and the exploration&#x2013;exploitation balance</title>
<p>Neuromodulatory accounts propose that large-scale brain dynamics shift between more integrated and more segregated network configurations as a function of arousal-linked chemical signals. Reviews of integration&#x2013;segregation emphasize that noradrenergic projections from locus coeruleus and cholinergic projections from basal forebrain are prominent levers for these state transitions: modest changes in their tone can reconfigure connectivity, biasing cognition toward either broad, globally integrated processing or more locally segregated, task-focused processing (<xref ref-type="bibr" rid="ref28">Shine, 2019</xref>). In parallel, work on dopamine and cognitive control links striatal D2 receptor availability to the subjective cost of exerting control and to cost&#x2013;benefit decisions about engaging effortful processing, consistent with a role for dopamine in setting how deeply and persistently goal-directed search is pursued (<xref ref-type="bibr" rid="ref30">Westbrook et al., 2021</xref>).</p>
<p>Together, these strands motivate an operational notion of gain: a control signal that widens or narrows the currently active hypothesis space. In integrated, high-gain states, the system samples more broadly (facilitating associative expansion); in segregated, low-gain states, it narrows and stabilizes processing (facilitating focused evaluation). We treat this as a functional template rather than a one-to-one biological mapping: multiple neuromodulators contribute to these shifts (<xref ref-type="bibr" rid="ref28">Shine, 2019</xref>), and dopamine&#x2019;s role is context dependent and tied to effort-related control policies (<xref ref-type="bibr" rid="ref30">Westbrook et al., 2021</xref>). In the Generative System 3 framework, the abstract gain controller corresponds to an endogenous mechanism that adjusts sampling entropy online (e.g., via temperature), thereby implementing the exploration&#x2013;exploitation trade-off that, in brains, is jointly shaped by neuromodulatory systems.</p>
</sec>
<sec id="sec12">
<label>3.3</label>
<title>Multiscale evidence: genetics, oscillations, and causal perturbation</title>
<p>Evidence for individual differences in creative cognition appears across levels of analysis. At the macroscale, large-cohort multimodal work shows that a neural pattern predicting divergent-thinking performance carries positive weights in default-mode and frontoparietal control networks and is linked to dopamine-related neurotransmitters and genes influencing neurotransmitter release, indicating a biological substrate for variability in network dynamics relevant to creativity (<xref ref-type="bibr" rid="ref22">Liu et al., 2024</xref>). At faster timescales, electrophysiological reviews report pre-solution modulations in low-frequency rhythms consistent with inwardly directed attention (alpha/theta changes) and brief gamma-band bursts localized to right anterior temporal cortex around the moment of insight&#x2014;together aligning with a generate-then-verify sequence (<xref ref-type="bibr" rid="ref20">Kounios and Beeman, 2014</xref>). Crucially, causal manipulations move beyond correlation: covert real-time fMRI neurofeedback that reinforces coactivation of default-mode and executive-control circuitry increases originality on divergent-thinking tasks relative to control conditions (<xref ref-type="bibr" rid="ref23">Luchini et al., 2025</xref>). Collectively, these findings support a cycle in which associative expansion and focused appraisal are coordinated by state-dependent control signals, providing a biologically grounded template for the alternation mechanisms formalized in GS-3.</p>
</sec>
<sec id="sec13">
<label>3.4</label>
<title>Translational lessons for artificial systems</title>
<p>Three design lessons follow for artificial systems seeking sustained novelty, usefulness, and diversity. First, at least two separable but re-entrant processing streams are required: a generator specialized for associative expansion and a critic specialized for task-conditioned evaluation. Second, a gain mechanism must adaptively regulate the breadth of search online; in practice, this means endogenously adjusting sampling entropy or effort as a function of a learned utility signal, rather than relying solely on fixed external controls. Third, the system should exhibit measurable signatures of alternation between expansion and verification over time. These lessons translate into testable predictions for Generative System 3: removing the critic should collapse usefulness at a fixed diversity level; removing the gain controller (freezing temperature) should eliminate alternation in associative-distance density and reduce across-run diversity, with within-run spread determined by the static decoding setting rather than by context; and reinforcing generator&#x2013;critic coactivation (e.g., by rewarding alternation) should increase originality without sacrificing task relevance.</p>
</sec>
<sec id="sec14">
<label>3.5</label>
<title>Boundary conditions and non-isomorphism</title>
<p>The DMN&#x2013;CEN&#x2013;dopamine template is a functional analogy, not a biological isomorphism. Biological networks operate with spiking dynamics, heterogeneous cell types, and complex neurochemical interactions; artificial networks are discrete symbol or vector processors trained under engineered objectives. Dopamine&#x2019;s roles are multifaceted and context dependent, extending beyond a simple exploration knob; likewise, temperature in a language model is only one of several ways to regulate uncertainty. The analogy is therefore limited to architectural roles and control functions: generator versus evaluator interactions and an adaptive gain that shifts the exploration&#x2013;exploitation balance. Our use of these mappings is pragmatic&#x2014;intended to generate falsifiable design claims&#x2014;rather than a claim of mechanistic identity.</p>
</sec>
</sec>
<sec id="sec15">
<label>4</label>
<title>GS-3 architecture: definitions, dynamics, and falsifiability</title>
<p>This section formalizes Generative System 3 (GS-3) while keeping implementation choices flexible. It specifies roles and interfaces, operational metrics, a bounded gain policy tied to a learning signal, behavioral indices with pass&#x2013;fail criteria, and ablations. Full task lists, hyperparameters, and pseudocode remain in <xref ref-type="sec" rid="sec122">Supplementary Data Sheet 1</xref>.</p>
<sec id="sec16">
<label>4.1</label>
<title>Roles and interfaces</title>
<p>The architecture comprises three roles with explicit interfaces.</p>
<sec id="sec17">
<label>4.1.1</label>
<title>Generator (G)</title>
<p>Proposes <italic>k</italic> candidates given a context, with a controllable sampling entropy (temperature <italic>T</italic>&#x208D;g&#x208E;). Output: a set of candidate continuations with scores from the base model (e.g., log probabilities).</p>
</sec>
<sec id="sec18">
<label>4.1.2</label>
<title>Critic (C)</title>
<p>Scores each candidate <italic>x</italic> with a task-conditioned utility <italic>U</italic>&#x208D;task&#x208E; (<italic>x</italic>|task, context) &#x2208; [0, 1], returning a real-valued score for each candidate and the index of the winner. The critic may be a rubric-based classifier or a preference model trained from human feedback; in the latter case, its design, data provenance, and validation should follow published guidance on feedback-driven NLG (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>).</p>
</sec>
<sec id="sec19">
<label>4.1.3</label>
<title>Gain controller (D)</title>
<p>Adjusts <italic>T</italic>&#x208D;g&#x208E; online as a function of recent reward-prediction error, using a bounded policy (Section 4.3) and a smoothed baseline of expected utility to calibrate expectations. Output: the next-step temperature <italic>T</italic>&#x208D;g, new&#x208E;.</p>
<p>Minimal message flow per cycle is G&#x202F;&#x2192;&#x202F;C&#x202F;&#x2192;&#x202F;D&#x202F;&#x2192;&#x202F;G, enabling alternating expansion and verification. This differs from externally tuned decoding (e.g., temperature sweeps or beam settings) and from pipelines that rely only on training-time preference alignment; here, usefulness is estimated by a learned critic active at inference time (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>). Regulated decoding still matters&#x2014;temperature/top-<italic>p</italic>/beam settings mitigate degeneracy (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>)&#x2014;but GS-3 requires an endogenous controller that adapts these levers during generation.</p>
</sec>
</sec>
<sec id="sec20">
<label>4.2</label>
<title>Operational definitions: novelty, usefulness, diversity</title>
<p>To permit falsification and fair comparison, we adopt simple, model-agnostic definitions.</p>
<sec id="sec21">
<label>4.2.1</label>
<title>Novelty</title>
<p>Represent an artifact <italic>x</italic> (e.g., a paragraph) with an embedding <italic>e</italic>(<italic>x</italic>) from a fixed, publicly documented encoder. Relative to a preregistered baseline corpus, novelty increases as the nearest neighbor to <italic>x</italic> becomes more distant in cosine space (the exact nearest-neighbor formula appears in <xref ref-type="sec" rid="sec122">Supplementary Data Sheet 1</xref>).</p>
</sec>
<sec id="sec22">
<label>4.2.2</label>
<title>Usefulness</title>
<p><italic>U</italic>&#x208D;task&#x208E; (<italic>x</italic>) is a task-conditioned score in [0, 1], produced either by a rubric-based human panel or by a separately validated reward model trained on task-specific preferences (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>). For experiments, preregister rubrics, rater training, and inter-rater reliability; when using reward models, report validation against held-out human judgments.</p>
</sec>
<sec id="sec23">
<label>4.2.3</label>
<title>Diversity</title>
<p>For a fixed prompt, report dispersion across independent runs (mean pairwise embedding distance). Exact formulas and encoder details are provided in <xref ref-type="sec" rid="sec122">Supplementary Data Sheet 1</xref>.</p>
<p>Together, novelty, usefulness, and diversity summarize originality, appropriateness, and dispersion and should be reported with confidence intervals.</p>
</sec>
</sec>
<sec id="sec24">
<label>4.3</label>
<title>Bounded gain policy and learning signal</title>
<p>Define the reward-prediction error as <italic>&#x03B4;</italic>&#x209C;&#x202F;= <italic>U</italic>&#x208D;best, t&#x208E;&#x202F;&#x2212; <italic>&#x016A;</italic>&#x209C;, where <italic>U</italic>&#x208D;best, t&#x208E; is the critic&#x2019;s top score at cycle <italic>t</italic> and <italic>&#x016A;</italic>&#x209C; is an exponentially weighted moving average of recent best scores (full expression in <xref ref-type="sec" rid="sec122">Supplementary Data Sheet 1</xref>). The controller updates sampling temperature with a bounded logistic rule: <italic>T</italic>&#x208D;g&#x208E;(<italic>t</italic> +&#x202F;1)&#x202F;= <italic>T</italic>&#x208D;min&#x208E;&#x202F;+&#x202F;(<italic>T</italic>&#x208D;max&#x208E;&#x202F;&#x2212; <italic>T</italic>&#x208D;min&#x208E;)&#x22EF;<italic>&#x03C3;</italic>(<italic>&#x03B1;</italic> + <italic>&#x03B7;</italic>&#x00B7;<italic>&#x03B4;</italic>&#x209C;), where &#x03C3; is the logistic function, &#x03B7; is a small learning-rate constant, and [<italic>T</italic>&#x208D;min&#x208E;, <italic>T</italic>&#x208D;max&#x208E;] are preregistered bounds. This yields smooth, monotone adjustments and prevents runaway entropy. Linear and exponential alternatives, stability notes, and sensitivity sweeps appear in <xref ref-type="sec" rid="sec122">Supplementary Data Sheet 1</xref>. The interpretation is consistent with accounts in which cost&#x2013;benefit control policies regulate effort allocation, while remaining an engineering&#x2014;not biological&#x2014;control law (<xref ref-type="bibr" rid="ref30">Westbrook et al., 2021</xref>).</p>
<p>Intuitively, each cycle compares the current best score to a smoothed baseline to obtain a &#x201C;surprise&#x201D; signal <italic>&#x03B4;</italic>. If performance is better than expected (<italic>&#x03B4;</italic> &#x003E;&#x202F;0), temperature increases smoothly; if worse (<italic>&#x03B4;</italic> &#x003C;&#x202F;0), it decreases. Logistic squashing keeps <italic>T</italic>&#x208D;g&#x208E; within preregistered bounds, making adjustments gradual and stable. The intercept &#x03B1; sets the default temperature when on trend, and the learning rate &#x03B7; controls how strongly surprises move it.</p>
</sec>
<sec id="sec25">
<label>4.4</label>
<title>Behavioral indices and pass&#x2013;fail criteria</title>
<sec id="sec26">
<label>4.4.1</label>
<title>Associative-distance density (ADD)</title>
<p>Distribution of cosine distances between successive idea units within a run (e.g., sentences or design sketches). GS-3 prediction: alternating wide&#x2013;narrow patterns reflecting expansion&#x2013;verification cycling; regulated baselines: unimodal, temperature-dependent spread.</p>
</sec>
<sec id="sec27">
<label>4.4.2</label>
<title>Analytic-verification ratio (AVR)</title>
<p>Proportion of cycles in which C vetoes G&#x2019;s top candidate and requests resampling at a lower <italic>T</italic>&#x208D;g&#x208E;. GS-3 prediction: AVR adapts to task difficulty; regulated baselines: AVR is fixed by external settings.</p>
</sec>
<sec id="sec28">
<label>4.4.3</label>
<title>Convergence latency (CL)</title>
<p>Cycles to meeting a preregistered success criterion (e.g., rubric score &#x2265; <italic>&#x03C4;</italic>). GS-3 prediction: CL decreases within a session as <italic>&#x016A;</italic> calibrates; reflective baselines show little within-session change.</p>
</sec>
<sec id="sec29">
<label>4.4.4</label>
<title>Pass&#x2013;fail criteria</title>
<p>Preregister that a GS-3 system must (a) exceed a temperature-matched baseline on usefulness at equal novelty (dominance on the novelty&#x2013;usefulness frontier), (b) achieve higher across-run diversity without external temperature sweeps, and (c) exhibit AVR and ADD signatures consistent with alternating exploration&#x2013;verification (e.g., significant periodicity by spectral analysis). Computation details appear in <xref ref-type="sec" rid="sec122">Supplementary Data Sheet 1</xref>.</p>
</sec>
</sec>
<sec id="sec30">
<label>4.5</label>
<title>Formal hypotheses and ablation tests</title>
<disp-quote>
<p>H1 (critic necessity). Removing C (scores replaced by random or constant) reduces usefulness at matched novelty, collapsing the novelty&#x2013;usefulness frontier.</p>
</disp-quote>
<disp-quote>
<p>H2 (gain necessity). Freezing <italic>T</italic>&#x208D;g&#x208E; (no D) eliminates ADD alternation and reduces across-run diversity; usefulness becomes more sensitive to the initial temperature setting.</p>
</disp-quote>
<disp-quote>
<p>H3 (policy sensitivity). Logistic-, linear-, and exponential-gain policies occupy distinct regions of the novelty&#x2013;usefulness&#x2013;diversity space; logistic yields the best stability at comparable usefulness.</p>
</disp-quote>
<disp-quote>
<p>H4 (memory horizon). Increasing <italic>&#x03B2;</italic> (longer <italic>&#x016A;</italic> memory) improves long-horizon coherence (e.g., cross-paragraph consistency) but slows adaptation after regime shifts.</p>
</disp-quote>
<p>Each hypothesis is falsifiable by implementing the corresponding ablation and reporting preregistered metrics with confidence intervals and effect sizes.</p>
</sec>
<sec id="sec31">
<label>4.6</label>
<title>Design space and non-isomorphism</title>
<p>G and C need not be separate models; they may be two modes of a single backbone, two cooperating agents, or a backbone plus a lightweight preference head (as in instruction-following systems informed by human feedback; <xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>). Likewise, D can be a small network conditioned on context features. Brain terms remain functional analogies: temperature is one of several levers (others include top-<italic>k</italic>/top-<italic>p</italic>, repetition penalties, and plug-and-play attribute controls) (<xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>). Retrieval can be added as an optional module to ground candidates; to avoid confounds, use the same retriever across GS-3 and RAG baselines (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>). The contribution here is to require that some endogenous gain exists, that it is coupled to a learned utility, and that its process-level signatures are measurable.</p>
</sec>
<sec id="sec32">
<label>4.7</label>
<title>Implementation notes and comparators</title>
<p>For completeness and parity, report the decoding settings (temperature/top-<italic>p</italic>/beam) and sampling budgets for all conditions, including pure prediction (<xref ref-type="bibr" rid="ref29">Sutskever et al., 2014</xref>) and regulated decoding baselines (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>). When including reflective prompting as a comparator, preregister the exact scaffolds (e.g., chain-of-thought prompts) to ensure fair budgets and to acknowledge that such reflectivity remains externally scaffolded (<xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>). If plug-and-play steering is used as a comparator, cite its peer-reviewed formulation and disclose active attribute controls (<xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>).</p>
</sec>
<sec id="sec33">
<label>4.8</label>
<title>Interim summary</title>
<p>GS-3 embeds a learned critic and a bounded, adaptive gain policy into the generation loop, evaluated with preregistered novelty, usefulness, and diversity metrics plus cycling signatures. These commitments turn a functional analogy into a testable engineering target while remaining agnostic to backbone choice and compatible with standard comparators such as RAG, plug-and-play steering, and reflective prompting (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>; <xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>; <xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>).</p>
</sec>
</sec>
<sec id="sec34">
<label>5</label>
<title>Proof-of-concept blueprint and evaluation protocol</title>
<p>This section describes how to implement and test a minimal instance of Generative System 3 (GS-3), specifying tasks, metrics, ablations, and analysis plans that allow other groups to falsify or support the framework.</p>
<sec id="sec35">
<label>5.1</label>
<title>Minimal implementation blueprint</title>
<sec id="sec36">
<label>5.1.1</label>
<title>Architecture</title>
<p>Use a single transformer backbone with two heads: a generator head for next-token prediction and a critic head that outputs a task-conditioned utility score <italic>U</italic>(<italic>x</italic>|task, context). A lightweight controller maps recent reward-prediction error <italic>&#x03B4;</italic> to an updated sampling temperature <italic>T</italic>&#x208D;g&#x208E; for the next generation step (see Section 4 for definitions and bounds). This single-backbone design enables shared representations while keeping roles separable for ablations.</p>
</sec>
<sec id="sec37">
<label>5.1.2</label>
<title>Training the critic</title>
<p>Collect paired or graded preferences for task outputs using a rubric aligned with usefulness (e.g., goal fit, coherence, constraint satisfaction). Train the critic with supervised regression to predict human utility or with a preference model trained on pairwise comparisons, as in established human-feedback pipelines and surveys of feedback integration (<xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>; <xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>). Keep evaluation sets disjoint from critic training data.</p>
</sec>
<sec id="sec38">
<label>5.1.3</label>
<title>Controller policy</title>
<p>Implement the bounded logistic gain policy described in Section 4 as the default; include linear and exponential variants in preregistered sensitivity analyses with clipped <italic>&#x03B4;</italic> and bounded <italic>T</italic>&#x208D;g&#x208E;. Preregister hyperparameter ranges and stopping rules.</p>
</sec>
<sec id="sec39">
<label>5.1.4</label>
<title>Baselines</title>
<p>Include three baselines: (a) pure prediction at multiple fixed temperatures; (b) regulated generation with beam search and temperature sweeps (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>); and (c) reflective prompting (e.g., chain-of-thought) without an internal learned critic, using a recent survey as the canonical reference (<xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>). Optional augmented baselines include retrieval-augmented generation using a published retrieval-and-generation pipeline (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>) and iterative self-refinement (<xref ref-type="bibr" rid="ref9">Ding et al., 2024</xref>; see also the self-correction survey, <xref ref-type="bibr" rid="ref19">Kamoi et al., 2024</xref>).</p>
</sec>
</sec>
<sec id="sec40">
<label>5.2</label>
<title>Tasks and datasets</title>
<sec id="sec41">
<label>5.2.1</label>
<title>Divergent thinking</title>
<p>Adapt the Alternate Uses Test (AUT) to text prompts (e.g., &#x201C;unusual uses for a paperclip&#x201D;), as used in neuroimaging work on creative switching, to allow comparison with network findings (<xref ref-type="bibr" rid="ref5">Chen et al., 2025</xref>). Score usefulness with a task rubric (plausibility under physical constraints), and compute novelty and diversity as in Section 4.</p>
</sec>
<sec id="sec42">
<label>5.2.2</label>
<title>Constrained creation</title>
<p>Short-form tasks, such as product names or headlines with explicit constraints (length, audience, and semantic cues), probe the critic&#x2019;s ability to trade novelty for goal fit. Retrieval-augmented variants test whether GS-3 maintains benefits when external knowledge is available (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>).</p>
</sec>
<sec id="sec43">
<label>5.2.3</label>
<title>Long-horizon composition</title>
<p>Multi-paragraph story or concept-expansion tasks assess maintenance of adaptive priors and coherence over extended contexts. Include checkpoints for mid-course critique and revision.</p>
</sec>
<sec id="sec44">
<label>5.2.4</label>
<title>Human&#x2013;AI co-creation</title>
<p>Writer-in-the-loop tasks mirror professional workflows and enable analysis of fluency&#x2013;variety trade-offs (<xref ref-type="bibr" rid="ref4">Chen and Chan, 2024</xref>; <xref ref-type="bibr" rid="ref3">Chakrabarty et al., 2024</xref>).</p>
</sec>
<sec id="sec45">
<label>5.2.5</label>
<title>Population-scale dispersion</title>
<p>To test homogenization risk, elicit many outputs per prompt and quantify across-run diversity and mode collapse, following concerns documented at scale (<xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>).</p>
</sec>
</sec>
<sec id="sec46">
<label>5.3</label>
<title>Metrics, reliability, and statistical analysis</title>
<sec id="sec47">
<label>5.3.1</label>
<title>Primary metrics</title>
<p>Use the operational definitions from Section 4: novelty <italic>N</italic> (embedding distance from a baseline corpus), usefulness <italic>U</italic>&#x208D;task&#x208E; (rubric or held-out reward model), and diversity <italic>D</italic> (mean pairwise distance across runs). Report within-run novelty and across-run diversity.</p>
</sec>
<sec id="sec48">
<label>5.3.2</label>
<title>Behavioral signatures</title>
<p>Compute associative-distance density (ADD) within runs, analytic-verification ratio (AVR; critic veto rate with resampling), and convergence latency (CL; cycles to reach a preregistered usefulness threshold). Assess periodicity in ADD to detect expansion&#x2013;verification alternation.</p>
</sec>
<sec id="sec49">
<label>5.3.3</label>
<title>Reliability</title>
<p>For human scoring, report inter-rater reliability (e.g., Krippendorff&#x2019;s alpha) and provide rater training materials. For model-based utility, validate the reward model against human judgments on a held-out set.</p>
</sec>
<sec id="sec50">
<label>5.3.4</label>
<title>Statistical plan</title>
<p>Preregister hypotheses, metrics, and analysis. Use hierarchical models or mixed-effects regressions to account for prompt and rater as random factors. Report effect sizes with confidence intervals and correct for multiple comparisons where applicable. Provide power analyses for planned contrasts (e.g., GS-3 vs. regulated baseline on usefulness at matched novelty).</p>
</sec>
</sec>
<sec id="sec51">
<label>5.4</label>
<title>Ablations and sensitivity</title>
<sec id="sec52">
<label>5.4.1</label>
<title>Critic removal (H1)</title>
<p>Replace <italic>U</italic> with random or constant scores and re-run; predict collapse of usefulness at matched novelty.</p>
</sec>
<sec id="sec53">
<label>5.4.2</label>
<title>Controller freeze (H2)</title>
<p>Hold <italic>T</italic>&#x208D;g&#x208E; constant; predict reduced across-run diversity and loss of ADD alternation.</p>
</sec>
<sec id="sec54">
<label>5.4.3</label>
<title>Policy comparison (H3)</title>
<p>Swap logistic (default), linear, and exponential policies while holding other components fixed; predict distinct novelty&#x2013;usefulness&#x2013;diversity trade-offs and fewer <italic>T</italic>&#x208D;g&#x208E; saturations for logistic.</p>
</sec>
<sec id="sec55">
<label>5.4.4</label>
<title>Memory horizon (H4)</title>
<p>Vary <italic>&#x03B2;</italic> in the baseline <italic>&#x016A;</italic>; predict improved long-horizon coherence at higher <italic>&#x03B2;</italic> but slower adaptation to shifts.</p>
</sec>
<sec id="sec56">
<label>5.4.5</label>
<title>Prompt perturbations</title>
<p>Vary prompt structure, length, and constraints to test robustness of gains. Include retrieval toggles to assess interaction with external knowledge (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>).</p>
</sec>
</sec>
<sec id="sec57">
<label>5.5</label>
<title>Reproducibility kit</title>
<p>Release code, model checkpoints (where licensing permits), exact prompts, rubrics, and analysis scripts. Fix random seeds; log <italic>T</italic>&#x208D;g&#x208E;, <italic>&#x03B4;</italic>, <italic>U</italic>&#x208D;best&#x208E;, and <italic>&#x016A;</italic> at each cycle for every run. Provide an audit sheet documenting compute budgets, training data used for the critic, and any human-in-the-loop procedures. For closed models, supply reproducible API settings and a synthetic variant using an open backbone.</p>
</sec>
<sec id="sec58">
<label>5.6</label>
<title>Risk controls and fairness checks</title>
<sec id="sec59">
<label>5.6.1</label>
<title>Homogenization audits</title>
<p>Track across-run diversity as a function of controller policy and dataset domain; include plural critics trained on diverse preference data to reduce mode collapse (<xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>).</p>
</sec>
<sec id="sec60">
<label>5.6.2</label>
<title>Bias and equity</title>
<p>Stratify usefulness and novelty by dialect, register, or cultural domain. If disparities emerge, retrain or reweight critic data and re-test.</p>
</sec>
<sec id="sec61">
<label>5.6.3</label>
<title>Overfitting to graders</title>
<p>When using reward or preference models, separate training, validation, and evaluation distributions; periodically cross-check with human ratings to prevent exploitation of grader idiosyncrasies (<xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>; <xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>).</p>
</sec>
<sec id="sec62">
<label>5.6.4</label>
<title>Safety valves</title>
<p>Bound <italic>T</italic>&#x208D;g&#x208E;, clip <italic>&#x03B4;</italic>, and cap cumulative entropy increases per session to prevent runaway exploration, consistent with concerns about degeneration under unbounded sampling (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>).</p>
</sec>
</sec>
<sec id="sec63">
<label>5.7</label>
<title>Decision rule</title>
<p>Declare GS-3 support only if, on preregistered tasks, the system (a) dominates regulated and reflective baselines on usefulness at matched novelty, (b) achieves higher across-run diversity without external temperature sweeps, and (c) exhibits cycling signatures in ADD and AVR consistent with alternating expansion and verification. Otherwise, the framework is falsified for that task setting, and ablations should identify which component failed to contribute.</p>
</sec>
</sec>
<sec id="sec64">
<label>6</label>
<title>Positioning current systems on the predictive&#x202F;&#x2192;&#x202F;generative continuum</title>
<p>This section locates prominent families of large language model (LLM) systems on a continuum from fluent prediction to partially reflective pipelines, and clarifies what each already achieves relative to Generative System 3 (GS-3). The emphasis is on the presence or absence of three ingredients that GS-3 treats as necessary for Artificial Creativity: (i) a generator capable of associative expansion, (ii) a learned, task-conditioned critic active during inference, and (iii) an endogenous gain controller that adaptively regulates sampling entropy during generation.</p>
<sec id="sec65">
<label>6.1</label>
<title>Pure prediction (decoder-only; no internal evaluation)</title>
<p>Autoregressive models trained for next-token prediction excel at fluent continuation but have no internal judge or regulator; behavior is largely governed by external decoding hyperparameters (e.g., temperature, top-<italic>p</italic>) (<xref ref-type="bibr" rid="ref29">Sutskever et al., 2014</xref>). Regulated decoding can mitigate repetition or dullness but remains an external knob rather than an internalized policy (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>). In GS-3 terms, this family has a generator but lacks an internal critic and lacks an endogenous gain controller.</p>
</sec>
<sec id="sec66">
<label>6.2</label>
<title>Prompt-scaffolded reflectivity</title>
<p>Prompting can scaffold brief internal critique&#x2014;e.g., chain-of-thought styles that elicit intermediate reasoning steps (<xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>). These strategies often improve reliability on structured tasks yet remain scaffold-dependent: the &#x201C;critic&#x201D; is effectively encoded in the prompt template, not learned as a task-conditioned utility model. Exploration&#x2013;exploitation is therefore not endogenously regulated, and adaptation across steps depends on the fixed script. In GS-3 terms, this family has a generator; the critic is externalized to prompts rather than learned and active during inference; there is no endogenous gain controller.</p>
</sec>
<sec id="sec67">
<label>6.3</label>
<title>Retrieval-augmented generation</title>
<p>Coupling generation to a retriever injects external knowledge and improves factual grounding on knowledge-intensive tasks (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>). Standard retrieval-augmented generation (RAG) pipelines still lack a learned internal critic that scores candidate continuations for task utility and a gain policy that adapts search breadth in real time. Breadth is set by retrieval depth and decoding parameters rather than updated by a live utility signal. In GS-3 terms, this family has a generator but lacks an internal critic and an endogenous gain controller.</p>
</sec>
<sec id="sec68">
<label>6.4</label>
<title>Plug-and-play steering at decoding time</title>
<p>Decoding-time &#x201C;plug-and-play&#x201D; controls can up- or down-weight attributes (e.g., sentiment, toxicity) on the fly without retraining the backbone (<xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>). Steering nudges the generator but does not maintain a persistent, task-conditioned evaluator nor an endogenous entropy controller tied to performance feedback. In GS-3 terms, this family has a generator but lacks an internal critic and an endogenous gain controller.</p>
</sec>
<sec id="sec69">
<label>6.5</label>
<title>Instruction-following with human feedback</title>
<p>Instruction-tuned models align behavior with human preferences via feedback pipelines. Surveys and analyses detail data collection, objectives, and limitations of feedback-integrated NLG (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>). These pipelines primarily externalize evaluation into the training data or reward modeling; at inference, most systems continue to rely on fixed decoding settings rather than a live gain policy tied to moment-to-moment utility. In GS-3 terms, this family has a generator; the critic is effectively baked in via training rather than active during inference; there is no endogenous gain controller.</p>
</sec>
<sec id="sec70">
<label>6.6</label>
<title>Self-correction and iterative refinement</title>
<p>Test-time self-correction mechanisms iteratively propose, critique, and revise drafts. A recent survey maps when such loops help or fail across tasks (<xref ref-type="bibr" rid="ref19">Kamoi et al., 2024</xref>), and domain-specific controllers demonstrate gains in code generation with explicit revise-and-retry cycles (<xref ref-type="bibr" rid="ref9">Ding et al., 2024</xref>). However, loop structure and revision depth are typically hand-designed; the exploration&#x2013;verification balance is not governed by an internal, learned gain signal that adapts step-to-step. In GS-3 terms, this family has a generator; the critic is scripted/self-referential rather than learned and general; there is no endogenous gain controller.</p>
</sec>
<sec id="sec71">
<label>6.7</label>
<title>Multi-agent orchestration</title>
<p>Agentic set-ups coordinate multiple LLMs (planner/critic/worker roles), sometimes with memory and tools, to simulate social feedback dynamics (<xref ref-type="bibr" rid="ref24">Park et al., 2023</xref>). While this can approximate a multi-perspective critique, policies are usually scripted; there is no single controller that adapts sampling entropy from reward-prediction error within a run. In GS-3 terms, this family has a generator; the critic role is scripted; there is no endogenous gain controller.</p>
</sec>
<sec id="sec72">
<label>6.8</label>
<title>Human&#x2013;AI co-creation and workflow integration</title>
<p>In professional settings, LLM support tends to increase throughput and fluency; without structured protocols, it can also reduce variety or drift from constraints (<xref ref-type="bibr" rid="ref4">Chen and Chan, 2024</xref>; <xref ref-type="bibr" rid="ref3">Chakrabarty et al., 2024</xref>). Effective workflows, therefore, need explicit mechanisms to preserve diversity while maintaining task fit&#x2014;precisely the trade-off that GS-3 formalizes via a learned critic and adaptive gain.</p>
</sec>
<sec id="sec73">
<label>6.9</label>
<title>Population-level effects and homogenization risk</title>
<p>At scale, generative assistance can raise individual originality while lowering collective diversity, indicating homogenization pressure when many users draw from similarly tuned models and prompts (<xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>). GS-3&#x2019;s evaluation emphasizes not only artifact-level usefulness and novelty but also across-run dispersion, making homogenization an explicit quantity to measure and manage via the gain policy and plural critics.</p>
</sec>
<sec id="sec74">
<label>6.10</label>
<title>Conceptual status of LLM creativity</title>
<p>Debates continue on whether contemporary LLMs meet criteria for creativity, where limits remain, and how to evaluate claims (<xref ref-type="bibr" rid="ref13">Franceschelli and Musolesi, 2024</xref>). GS-3 is positioned as a control-theoretic addition: it does not claim that steering, prompting, or retrieval alone are insufficient, but that creative competence requires an internalized evaluator and an adaptive regulator with falsifiable process-level signatures (e.g., alternating associative-distance density and adaptive verification rates).</p>
</sec>
<sec id="sec75">
<label>6.11</label>
<title>What is still missing (gap analysis)</title>
<p>Across these families, three ingredients remain only partially addressed:</p>
<list list-type="order">
<list-item>
<p>A learned, task-conditioned critic active during generation (not only at training time or via prompts).</p>
</list-item>
<list-item>
<p>An adaptive gain controller that smoothly adjusts sampling entropy from a simple learning signal within the session.</p>
</list-item>
<list-item>
<p>Process-level signatures (cycling in associative-distance density; adaptive verification rates) that make the mechanism auditable.</p>
</list-item>
<list-item>
<p>GS-3 contributes exactly these pieces while remaining architecture-agnostic and compatible with standard comparators (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>; <xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>; <xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>; <xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>).</p>
</list-item>
</list>
</sec>
<sec id="sec76">
<label>6.12</label>
<title>Summary</title>
<p>Current systems achieve parts of the creative loop&#x2014;fluent expansion, external steering, retrieval grounding, scripted reflection&#x2014;but lack an endogenous, learned mechanism that coordinates expansion with evaluation under adaptive gain. GS-3 specifies that mechanism and its signatures, providing clear ablations and pass&#x2013;fail criteria for empirical tests in Section 5.</p>
</sec>
</sec>
<sec id="sec77">
<label>7</label>
<title>Ethics, governance, and responsible deployment</title>
<p>GS-3 aims to operationalize creative generation while minimizing societal risk. This section outlines risks, design safeguards, reporting standards, and governance practices that make GS-3 auditable and alignable in real use.</p>
<sec id="sec78">
<label>7.1</label>
<title>Risk landscape</title>
<sec id="sec79">
<label>7.1.1</label>
<title>Bias and preference overfitting</title>
<p>Training or validating critics on narrow rater groups can encode majority preferences and crowd out minority aesthetics. Surveys and analyses of feedback-driven NLG document how data collection, rater instructions, objective choice, and optimization targets shape model behavior and can entrench unwanted preferences (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>).</p>
</sec>
<sec id="sec80">
<label>7.1.2</label>
<title>Homogenization</title>
<p>At the population level, assistance can raise individual originality while reducing collective diversity&#x2014;consistent with convergent styles and &#x201C;mode collapse&#x201D; at scale (<xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>). This risk is directly relevant to GS-3&#x2019;s diversity objective.</p>
</sec>
<sec id="sec81">
<label>7.1.3</label>
<title>Scaffold dependence</title>
<p>Prompted self-reflection (e.g., chain-of-thought styles) can improve reliability on some tasks yet remains externally scaffolded and can fail outside its design envelope (<xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>). GS-3 treats such reflectivity as a baseline, not a substitute for an internal critic and gain policy.</p>
</sec>
<sec id="sec82">
<label>7.1.4</label>
<title>Attribution and provenance</title>
<p>Use of external knowledge without source tracking can blur accountability. Retrieval-augmented generation highlights the need for explicit provenance trails (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>).</p>
</sec>
<sec id="sec83">
<label>7.1.5</label>
<title>Manipulation and reward gaming</title>
<p>Systems optimizing proxy rewards may learn to exploit engagement-like signals rather than usefulness; this motivates transparent utility functions, plural critics, and caps on entropy changes per cycle (<xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>).</p>
</sec>
</sec>
<sec id="sec84">
<label>7.2</label>
<title>Design safeguards</title>
<sec id="sec85">
<label>7.2.1</label>
<title>Plural critics and counterfactual scoring</title>
<p>Train multiple critics with diverse rater pools and aggregate via robust methods; monitor divergence to detect preference drift (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>).</p>
</sec>
<sec id="sec86">
<label>7.2.2</label>
<title>Telemetry for audit</title>
<p>Log candidate sets, critic scores, temperature trajectory, retrieval queries and sources, and rationale snippets. Release redacted logs for external auditing subject to privacy constraints.</p>
</sec>
<sec id="sec87">
<label>7.2.3</label>
<title>Entropy governance</title>
<p>Enforce bounded-logistic gain (Section 4) with rate-limiters on temperature change per cycle; preregister <italic>T</italic>&#x208D;min&#x208E;, <italic>T</italic>&#x208D;max&#x208E;, and learning-rate bounds.</p>
</sec>
<sec id="sec88">
<label>7.2.4</label>
<title>Attribute controls with disclosure</title>
<p>When using decoding-time steering, employ plug-and-play controls that nudge attributes without retraining, and disclose active controls in outputs (<xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>).</p>
</sec>
<sec id="sec89">
<label>7.2.5</label>
<title>Knowledge provenance</title>
<p>For any grounded claim, attach sources returned by the retriever and prefer evidence-linked output modes (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>).</p>
</sec>
<sec id="sec90">
<label>7.2.6</label>
<title>Co-creation protocols</title>
<p>In collaborative settings, use structured prompts and rubrics to preserve variety and constraint adherence (<xref ref-type="bibr" rid="ref4">Chen and Chan, 2024</xref>; <xref ref-type="bibr" rid="ref3">Chakrabarty et al., 2024</xref>).</p>
</sec>
</sec>
<sec id="sec91">
<label>7.3</label>
<title>Reporting standards</title>
<sec id="sec92">
<label>7.3.1</label>
<title>Preregistration</title>
<p>Publish prompts, success criteria, decoding budgets, and ablation plans.</p>
</sec>
<sec id="sec93">
<label>7.3.2</label>
<title>Human evaluation</title>
<p>Provide rater training materials and report inter-rater reliability; define task-conditioned rubrics. If using learned reward models, document data provenance, validation against held-out human judgments, and failure analyses (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>).</p>
</sec>
<sec id="sec94">
<label>7.3.3</label>
<title>Process-level signatures</title>
<p>Report associative-distance density, analytic-verification ratio, and convergence latency with confidence intervals, plus spectral/auto-correlation analyses evidencing cycling.</p>
</sec>
<sec id="sec95">
<label>7.3.4</label>
<title>Release materials</title>
<p>Share code for metrics, ablation toggles, seeds and decoding settings, and (where possible) a minimal GS-3 implementation to reproduce tables and figures.</p>
</sec>
</sec>
<sec id="sec96">
<label>7.4</label>
<title>Governance and oversight</title>
<sec id="sec97">
<label>7.4.1</label>
<title>Principle-guided constraints</title>
<p>Where high-stakes governance is required, adopt constitution-style rule sets derived from public input, and bind the critic&#x2019;s utility and admissible entropy range to these principles (<xref ref-type="bibr" rid="ref16">Huang et al., 2024</xref>). This layer is complementary to, not a replacement for, GS-3&#x2019;s endogenous regulation.</p>
</sec>
<sec id="sec98">
<label>7.4.2</label>
<title>Independent review</title>
<p>Establish review boards to audit data governance, preference diversity, impact on stakeholders, and telemetry practices; publish periodic system cards summarizing risks and mitigations.</p>
</sec>
<sec id="sec99">
<label>7.4.3</label>
<title>User agency and consent</title>
<p>Provide clear affordances to decline data use for feedback, select preference profiles, and request provenance for retrieved evidence.</p>
</sec>
</sec>
<sec id="sec100">
<label>7.5</label>
<title>Boundary conditions</title>
<p>GS-3 is a control-theoretic proposal for creative generation, not a normative theory of cultural value. It does not by itself resolve questions of authorship or intellectual property; rather, it supplies the mechanisms and measurements by which such policies can be evaluated.</p>
</sec>
</sec>
<sec id="sec101">
<label>8</label>
<title>Discussion and open problems</title>
<p>This section synthesizes the argument, states boundary conditions, and outlines priority experiments that could support or falsify Generative System 3 (GS-3). Emphasis is on what the framework adds beyond existing accounts, where it may fail, and how to test it with published, auditable methods.</p>
<sec id="sec102">
<label>8.1</label>
<title>What GS-3 adds</title>
<p>GS-3 contributes a concrete control story for moving beyond fluent prediction and scaffolded reflectivity: a generator for associative expansion, a learned critic for task-conditioned appraisal, and an endogenous gain controller that adjusts sampling entropy online from a reward-prediction error. In contrast to externally tuned decoding (e.g., temperature sweeps, beam width), the exploration&#x2013;exploitation balance becomes a learned, auditable policy with measurable signatures (Sections 4&#x2013;5). This reorients evaluation from static artifacts to process-level observables and ablation tests using preregistered metrics and baselines (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>; <xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>; <xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>).</p>
</sec>
<sec id="sec103">
<label>8.2</label>
<title>Boundary conditions and limitations</title>
<sec id="sec104">
<label>8.2.1</label>
<title>Non-isomorphism</title>
<p>The DMN&#x2013;CEN&#x2013;dopamine mapping is a functional analogy, not a claim of biological identity. Neuromodulators shape integration/segregation and effort allocation in flexible cognition, but their roles are contextual and multifaceted (<xref ref-type="bibr" rid="ref28">Shine, 2019</xref>; <xref ref-type="bibr" rid="ref30">Westbrook et al., 2021</xref>). Temperature and related decoding controls are only rough proxies for gain in artificial systems.</p>
</sec>
<sec id="sec105">
<label>8.2.2</label>
<title>Task domain and priors</title>
<p>Gains from an endogenous controller will depend on task structure. Problems with tight constraints may benefit more from strong critics and narrower entropy; open-ended ideation may require wider entropy and more permissive critics. Long-horizon composition introduces additional stability&#x2013;adaptation trade-offs (Section 4).</p>
</sec>
</sec>
<sec id="sec106">
<label>8.3</label>
<title>Proxy risks and evaluation pitfalls</title>
<sec id="sec107">
<label>8.3.1</label>
<title>Preference models</title>
<p>Utility models trained from narrow rater pools can encode unwanted biases or collapse diversity; the literature on feedback-integrated NLG documents these risks and recommended safeguards (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>). Accordingly, GS-3 advocates plural critics, provenance for feedback data, and validation of reward models against held-out human judgments.</p>
</sec>
<sec id="sec108">
<label>8.3.2</label>
<title>Measurement sensitivity</title>
<p>Novelty measured as embedding distance depends on the encoder and baseline corpus; conclusions should be cross-checked with human judgments and alternative encoders. Usefulness scores must report rater training and reliability; when model-based, they require external validation (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>).</p>
</sec>
</sec>
<sec id="sec109">
<label>8.4</label>
<title>Decoding pathologies and controller claims</title>
<p>High temperature can increase diversity at the expense of coherence; low temperature can induce repetition and dullness&#x2014;well-characterized failure modes under standard decoding (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>). GS-3 does not claim these trade-offs disappear, but rather that an online gain policy can steer them adaptively within a run; this remains an empirical question addressed by the pass&#x2013;fail criteria in Section 5.</p>
</sec>
<sec id="sec110">
<label>8.5</label>
<title>Priority experiments</title>
<p>Minimal single-backbone implementations should be compared against regulated and reflective baselines under matched compute, prompts, and retrieval settings (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>; <xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>; <xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>). Divergent-thinking tasks (e.g., alternative uses), constrained creation (e.g., headlines with requirements), and long-horizon composition provide complementary stress tests. Writer-in-the-loop tasks probe fluency&#x2013;variety trade-offs in professional workflows (<xref ref-type="bibr" rid="ref4">Chen and Chan, 2024</xref>; <xref ref-type="bibr" rid="ref3">Chakrabarty et al., 2024</xref>). At population scale, audits should test for homogenization (increases in individual usefulness/originality alongside decreases in collective diversity) and whether plural critics and gain policies mitigate it (<xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>). Preregistration should include hypotheses, ablations (remove critic; freeze gain; swap policies), stopping rules, and telemetry (candidate sets, critic scores, reward-prediction error, temperature trajectory) to support external audit.</p>
</sec>
<sec id="sec111">
<label>8.6</label>
<title>Open problems</title>
<sec id="sec112">
<label>8.6.1</label>
<title>Multimodal and embodied extensions</title>
<p>Extending the generator&#x2013;critic&#x2013;gain loop to vision, audio, and action raises questions about shared versus modality-specific critics and controllers, especially for agents that learn from interaction.</p>
</sec>
<sec id="sec113">
<label>8.6.2</label>
<title>Memory and priors</title>
<p>How should the running baseline of expected utility be maintained across chapters, sessions, or projects without inducing inertia or overfitting to early successes?</p>
</sec>
<sec id="sec114">
<label>8.6.3</label>
<title>Plural critics and value alignment</title>
<p>Aggregating diverse preference models may preserve diversity while maintaining task fit, but it complicates optimization and governance (<xref ref-type="bibr" rid="ref11">Fernandes et al., 2023</xref>; <xref ref-type="bibr" rid="ref2">Casper et al., 2024</xref>). What aggregation rules best handle disagreement without masking minority values? Can constitution-style, publicly derived principles provide guardrails without collapsing variety (<xref ref-type="bibr" rid="ref16">Huang et al., 2024</xref>)?</p>
</sec>
<sec id="sec115">
<label>8.6.4</label>
<title>Interaction with external tools</title>
<p>Retrieval and plug-and-play steering provide complementary control surfaces; their interaction with an internal gain policy requires systematic mapping to avoid redundant or destabilizing effects (<xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>; <xref ref-type="bibr" rid="ref25">Pascual et al., 2021</xref>).</p>
</sec>
</sec>
<sec id="sec116">
<label>8.7</label>
<title>Summary</title>
<p>GS-3 is a proposal to turn a functional analogy into a testable engineering target. Its value hinges on rigorous comparisons to strong baselines, preregistered metrics and ablations, and transparent reporting. If its predictions fail, that outcome is informative&#x2014;favoring alternative accounts such as scaffolded reflectivity or purely external regulation. If they succeed, they mark a step toward artificial systems that manage the tension between novelty, usefulness, and diversity by learning to regulate their own creative process (<xref ref-type="bibr" rid="ref15">Holtzman et al., 2020</xref>; <xref ref-type="bibr" rid="ref6">Chu et al., 2024</xref>; <xref ref-type="bibr" rid="ref17">Izacard and Grave, 2021</xref>; <xref ref-type="bibr" rid="ref4">Chen and Chan, 2024</xref>; <xref ref-type="bibr" rid="ref10">Doshi and Hauser, 2024</xref>).</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="author-contributions" id="sec117">
<title>Author contributions</title>
<p>JC-A: Conceptualization, Methodology, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec sec-type="funding-information" id="sec118">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research and/or publication of this article.</p>
</sec>
<sec sec-type="COI-statement" id="sec119">
<title>Conflict of interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="sec120">
<title>Generative AI statement</title>
<p>The author(s) declare that Gen AI was used in the creation of this manuscript. A generative AI system was used for language editing and formatting. Specifically, OpenAI o3 (June 2025 model; OpenAI, San Francisco, United States) assisted with phrasing, copyediting, and APA/Frontiers notation. The tool was not used to generate or analyze data, perform statistical procedures, or originate scientific claims. All content was reviewed and verified by the authors, who accept full responsibility for the manuscript. The AI system is not listed as an author.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<sec sec-type="disclaimer" id="sec121">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="sec122">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/frai.2025.1654716/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/frai.2025.1654716/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Boden</surname><given-names>M. A.</given-names></name></person-group> (<year>2004</year>). <source>The creative mind: myths and mechanisms</source>. <edition>2nd</edition> Edn. <publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>.</citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Casper</surname><given-names>S.</given-names></name> <name><surname>Hadfield</surname><given-names>G.</given-names></name> <name><surname>Leike</surname><given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>RLHF deciphered: a critical analysis of reinforcement learning from human feedback</article-title>. <source>Commun. ACM</source> <volume>67</volume>, <fpage>36</fpage>&#x2013;<lpage>47</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3743127</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Chakrabarty</surname><given-names>T.</given-names></name> <name><surname>Padmakumar</surname><given-names>V.</given-names></name> <name><surname>Brahman</surname><given-names>F.</given-names></name> <name><surname>Muresan</surname><given-names>S.</given-names></name></person-group> (<year>2024</year>). <article-title>Creativity support in the age of large language models: an empirical study involving professional writers</article-title>. <conf-name>Proceedings of the 16th Conference on Creativity &#x0026; Cognition</conf-name>. <fpage>132</fpage>&#x2013;<lpage>155</lpage>. <publisher-name>Association for Computing Machinery</publisher-name>: <publisher-loc>New York, NY</publisher-loc></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname><given-names>Z.</given-names></name> <name><surname>Chan</surname><given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>Large language model in creative work: the role of collaboration modality and user expertise</article-title>. <source>Manag. Sci.</source> <volume>70</volume>, <fpage>9101</fpage>&#x2013;<lpage>9117</lpage>. doi: <pub-id pub-id-type="doi">10.1287/mnsc.2023.03014</pub-id>, PMID: <pub-id pub-id-type="pmid">19642375</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname><given-names>Q.</given-names></name> <name><surname>Kenett</surname><given-names>Y. N.</given-names></name> <name><surname>Cui</surname><given-names>Z.</given-names></name> <name><surname>Takeuchi</surname><given-names>H.</given-names></name> <name><surname>Fink</surname><given-names>A.</given-names></name> <name><surname>Benedek</surname><given-names>M.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Dynamic switching between brain networks predicts creative ability</article-title>. <source>Commun. Biol.</source> <volume>8</volume>:<fpage>54</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s42003-025-07470-9</pub-id>, PMID: <pub-id pub-id-type="pmid">39809882</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Chu</surname><given-names>Z.</given-names></name> <name><surname>Chen</surname><given-names>J.</given-names></name> <name><surname>Chen</surname><given-names>Q.</given-names></name> <name><surname>Xu</surname><given-names>T.</given-names></name> <name><surname>Wu</surname><given-names>Z.</given-names></name> <name><surname>Li</surname><given-names>Y.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Navigate through enigmatic labyrinth&#x2014;a survey of chain-of-thought reasoning: advances, frontiers and future</article-title>. <conf-name>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024): Long Papers</conf-name>. <fpage>1173</fpage>&#x2013;<lpage>1203</lpage>. <publisher-name>Association for Computational Linguistics</publisher-name>: <publisher-loc>New York, NY</publisher-loc></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname><given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Whatever next? Predictive brains, situated agents, and the future of cognitive science</article-title>. <source>Behav. Brain Sci.</source> <volume>36</volume>, <fpage>181</fpage>&#x2013;<lpage>204</lpage>. doi: <pub-id pub-id-type="doi">10.1017/S0140525X12000477</pub-id>, PMID: <pub-id pub-id-type="pmid">23663408</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Colton</surname><given-names>S.</given-names></name></person-group> (<year>2012</year>). &#x201C;<article-title>The painting fool: stories from building an automated painter</article-title>&#x201D; in <source>Computers and creativity</source>. eds. <person-group person-group-type="editor"><name><surname>McCormack</surname><given-names>J.</given-names></name> <name><surname>d&#x2019;Inverno</surname><given-names>M.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>3</fpage>&#x2013;<lpage>38</lpage>.</citation></ref>
<ref id="ref9"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Ding</surname><given-names>Y.</given-names></name> <name><surname>Min</surname><given-names>M. J.</given-names></name> <name><surname>Kaiser</surname><given-names>G.</given-names></name> <name><surname>Ray</surname><given-names>B.</given-names></name></person-group> (<year>2024</year>). <article-title>CYCLE: learning to self-refine the code generation</article-title>. <conf-name>Proceedings of the ACM on Programming Languages</conf-name></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doshi</surname><given-names>A. R.</given-names></name> <name><surname>Hauser</surname><given-names>O. P.</given-names></name></person-group> (<year>2024</year>). <article-title>Generative AI enhances individual creativity but reduces the collective diversity of novel content</article-title>. <source>Sci. Adv.</source> <volume>10</volume>:<fpage>eadn5290</fpage>. doi: <pub-id pub-id-type="doi">10.1126/sciadv.adn5290</pub-id>, PMID: <pub-id pub-id-type="pmid">38996021</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fernandes</surname><given-names>P.</given-names></name> <name><surname>Ribeiro</surname><given-names>R.</given-names></name> <name><surname>Martins</surname><given-names>A. F. T.</given-names></name></person-group> (<year>2023</year>). <article-title>Bridging the gap: a survey on integrating (human) feedback for natural language generation</article-title>. <source>Trans. Assoc. Comput. Linguist.</source> <volume>11</volume>, <fpage>245</fpage>&#x2013;<lpage>270</lpage>. doi: <pub-id pub-id-type="doi">10.1162/tacl_a_00626</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Floridi</surname><given-names>L.</given-names></name> <name><surname>Chiriatti</surname><given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>GPT-3: its nature, scope, limits, and consequences</article-title>. <source>Minds Mach.</source> <volume>30</volume>, <fpage>681</fpage>&#x2013;<lpage>694</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11023-020-09548-1</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Franceschelli</surname><given-names>G.</given-names></name> <name><surname>Musolesi</surname><given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>On the creativity of large language models</article-title>. <source>AI Soc.</source> <volume>40</volume>, <fpage>3785</fpage>&#x2013;<lpage>3795</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s00146-024-02127-3</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname><given-names>K.</given-names></name></person-group> (<year>2010</year>). <article-title>The free-energy principle: a unified brain theory?</article-title> <source>Nat. Rev. Neurosci.</source> <volume>11</volume>, <fpage>127</fpage>&#x2013;<lpage>138</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nrn2787</pub-id>, PMID: <pub-id pub-id-type="pmid">20068583</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Holtzman</surname><given-names>A.</given-names></name> <name><surname>Buys</surname><given-names>J.</given-names></name> <name><surname>Du</surname><given-names>L.</given-names></name> <name><surname>Forbes</surname><given-names>M.</given-names></name> <name><surname>Choi</surname><given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>The curious case of neural text degeneration</article-title>. <conf-name>International Conference on Learning Representations (ICLR 2020)</conf-name></citation></ref>
<ref id="ref16"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Huang</surname><given-names>S.</given-names></name> <name><surname>Siddarth</surname><given-names>D.</given-names></name> <name><surname>Lovitt</surname><given-names>L.</given-names></name> <name><surname>Liao</surname><given-names>T. I.</given-names></name> <name><surname>Durmus</surname><given-names>E.</given-names></name> <name><surname>Tamkin</surname><given-names>A.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Collective constitutional AI: aligning a language model with public input</article-title>. <conf-name>Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT 2024)</conf-name>. <fpage>1395</fpage>&#x2013;<lpage>1417</lpage></citation></ref>
<ref id="ref17"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Izacard</surname><given-names>G.</given-names></name> <name><surname>Grave</surname><given-names>&#x00C9;.</given-names></name></person-group> (<year>2021</year>). <article-title>Leveraging passage retrieval with generative models for open-domain question answering</article-title>. <conf-name>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (EACL 2021)</conf-name>. <fpage>874</fpage>&#x2013;<lpage>880</lpage>. <publisher-name>Association for Computational Linguistics</publisher-name>: <publisher-loc>Stroudsburg, PA</publisher-loc></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jordanous</surname><given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>A standardised procedure for evaluating creative systems: computational creativity evaluation based on what it is to be creative</article-title>. <source>Cogn. Comput.</source> <volume>4</volume>, <fpage>246</fpage>&#x2013;<lpage>279</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s12559-012-9156-1</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kamoi</surname><given-names>R.</given-names></name> <name><surname>Wang</surname><given-names>Y.</given-names></name> <name><surname>Shi</surname><given-names>P.</given-names></name> <name><surname>Xu</surname><given-names>P.</given-names></name> <name><surname>Song</surname><given-names>H.</given-names></name> <name><surname>Zhang</surname><given-names>T.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>When can LLMs actually correct their own mistakes? A critical survey of self-correction of LLMs</article-title>. <source>Trans. Assoc. Comput. Linguist.</source> <volume>12</volume>, <fpage>1801</fpage>&#x2013;<lpage>1824</lpage>. doi: <pub-id pub-id-type="doi">10.1162/tacl_a_00713</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kounios</surname><given-names>J.</given-names></name> <name><surname>Beeman</surname><given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>The cognitive neuroscience of insight</article-title>. <source>Annu. Rev. Psychol.</source> <volume>65</volume>, <fpage>71</fpage>&#x2013;<lpage>93</lpage>. doi: <pub-id pub-id-type="doi">10.1146/annurev-psych-010213-115154</pub-id>, PMID: <pub-id pub-id-type="pmid">24405359</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lehman</surname><given-names>J.</given-names></name> <name><surname>Stanley</surname><given-names>K. O.</given-names></name></person-group> (<year>2011</year>). <article-title>Abandoning objectives: evolution through the search for novelty alone</article-title>. <source>Evol. Comput.</source> <volume>19</volume>, <fpage>189</fpage>&#x2013;<lpage>223</lpage>. doi: <pub-id pub-id-type="doi">10.1162/EVCO_a_00025</pub-id>, PMID: <pub-id pub-id-type="pmid">20868264</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname><given-names>C.</given-names></name> <name><surname>Zhuang</surname><given-names>K.</given-names></name> <name><surname>Zeitlen</surname><given-names>D. C.</given-names></name> <name><surname>Chen</surname><given-names>Q.</given-names></name> <name><surname>Wang</surname><given-names>X.</given-names></name> <name><surname>Feng</surname><given-names>Q.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Neural, genetic, and cognitive signatures of creativity</article-title>. <source>Commun. Biol.</source> <volume>7</volume>:<fpage>1324</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s42003-024-07007-6</pub-id>, PMID: <pub-id pub-id-type="pmid">39402209</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luchini</surname><given-names>S. A.</given-names></name> <name><surname>Zhang</surname><given-names>X.</given-names></name> <name><surname>White</surname><given-names>R. T.</given-names></name> <name><surname>L&#x00FC;hrs</surname><given-names>M.</given-names></name> <name><surname>Ramot</surname><given-names>M.</given-names></name> <name><surname>Beaty</surname><given-names>R. E.</given-names></name></person-group> (<year>2025</year>). <article-title>Enhancing creativity with covert neurofeedback: causal evidence for default&#x2013;executive network coupling in creative thinking</article-title>. <source>Cereb. Cortex</source> <volume>35</volume>:<fpage>bhaf065</fpage>. doi: <pub-id pub-id-type="doi">10.1093/cercor/bhaf065</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Park</surname><given-names>J. S.</given-names></name> <name><surname>O&#x2019;Brien</surname><given-names>J. C.</given-names></name> <name><surname>Cai</surname><given-names>C. J.</given-names></name> <name><surname>Morris</surname><given-names>M. R.</given-names></name> <name><surname>Liang</surname><given-names>P.</given-names></name> <name><surname>Bernstein</surname><given-names>M. S.</given-names></name></person-group> (<year>2023</year>). <article-title>Generative agents: interactive simulacra of human behavior</article-title>. <conf-name>Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST 2023)</conf-name>. <fpage>1</fpage>&#x2013;<lpage>22</lpage>. <publisher-name>Association for Computing Machinery</publisher-name>: <publisher-loc>New York, NY</publisher-loc></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pascual</surname><given-names>D.</given-names></name> <name><surname>Egressy</surname><given-names>B.</given-names></name> <name><surname>Meister</surname><given-names>C.</given-names></name> <name><surname>Cotterell</surname><given-names>R.</given-names></name> <name><surname>Wattenhofer</surname><given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>A plug-and-play method for controlled text generation</article-title>. <source>Find. Assoc. Comput. Linguist.</source> <volume>2021</volume>, <fpage>3973</fpage>&#x2013;<lpage>3997</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/2021.findings-emnlp.334</pub-id></citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ritchie</surname><given-names>G.</given-names></name></person-group> (<year>2007</year>). <article-title>Some empirical criteria for attributing creativity to a computer program</article-title>. <source>Minds Mach.</source> <volume>17</volume>, <fpage>67</fpage>&#x2013;<lpage>99</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11023-007-9066-2</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidhuber</surname><given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Formal theory of creativity, fun, and intrinsic motivation (1990&#x2013;2010)</article-title>. <source>IEEE Trans. Auton. Ment. Dev.</source> <volume>2</volume>, <fpage>230</fpage>&#x2013;<lpage>247</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAMD.2010.2056368</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shine</surname><given-names>J. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Neuromodulatory influences on integration and segregation in the brain</article-title>. <source>Trends Cogn. Sci.</source> <volume>23</volume>, <fpage>572</fpage>&#x2013;<lpage>583</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tics.2019.04.002</pub-id>, PMID: <pub-id pub-id-type="pmid">31076192</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Sutskever</surname><given-names>I.</given-names></name> <name><surname>Vinyals</surname><given-names>O.</given-names></name> <name><surname>Le</surname><given-names>Q. V.</given-names></name></person-group> (<year>2014</year>). <article-title>Sequence to sequence learning with neural networks</article-title>. <conf-name>Advances in Neural Information Processing Systems</conf-name>. <fpage>3104</fpage>&#x2013;<lpage>3112</lpage>. <publisher-name>Curran Associates, Inc.</publisher-name>: <publisher-loc>Red Hook, NY</publisher-loc></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Westbrook</surname><given-names>A.</given-names></name> <name><surname>Frank</surname><given-names>M. J.</given-names></name> <name><surname>Cools</surname><given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>A mosaic of cost&#x2013;benefit control over cortico-striatal circuitry</article-title>. <source>Trends Cogn. Sci.</source> <volume>25</volume>, <fpage>710</fpage>&#x2013;<lpage>721</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tics.2021.04.007</pub-id>, PMID: <pub-id pub-id-type="pmid">34120845</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wiggins</surname><given-names>G. A.</given-names></name></person-group> (<year>2006</year>). <article-title>A preliminary framework for description, analysis and comparison of creative systems</article-title>. <source>Knowl.-Based Syst.</source> <volume>19</volume>, <fpage>449</fpage>&#x2013;<lpage>458</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2006.04.009</pub-id></citation></ref>
</ref-list>
</back>
</article>