<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3-mathml3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title-group>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2025.1659861</article-id>
<article-version article-version-type="Version of Record" vocab="NISO-RP-8-2008"/>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Original Research</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Development and validation of a multi-agent AI pipeline for automated credibility assessment of tobacco misinformation: a proof-of-concept study</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Elmitwalli</surname><given-names>Sherif</given-names></name>
<xref ref-type="aff" rid="aff1"></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2608401"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Data curation" vocab-term-identifier="https://credit.niso.org/contributor-roles/data-curation/">Data curation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Formal analysis" vocab-term-identifier="https://credit.niso.org/contributor-roles/formal-analysis/">Formal analysis</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="investigation" vocab-term-identifier="https://credit.niso.org/contributor-roles/investigation/">Investigation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="resources" vocab-term-identifier="https://credit.niso.org/contributor-roles/resources/">Resources</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="software" vocab-term-identifier="https://credit.niso.org/contributor-roles/software/">Software</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="visualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/visualization/">Visualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; original draft" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-original-draft/">Writing &#x2013; original draft</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &#x0026; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x0026; editing</role>
</contrib>
<contrib contrib-type="author">
<name><surname>Mehegan</surname><given-names>John</given-names></name>
<xref ref-type="aff" rid="aff1"></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2677381"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Data curation" vocab-term-identifier="https://credit.niso.org/contributor-roles/data-curation/">Data curation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Formal analysis" vocab-term-identifier="https://credit.niso.org/contributor-roles/formal-analysis/">Formal analysis</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="investigation" vocab-term-identifier="https://credit.niso.org/contributor-roles/investigation/">Investigation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Project administration" vocab-term-identifier="https://credit.niso.org/contributor-roles/project-administration/">Project administration</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="resources" vocab-term-identifier="https://credit.niso.org/contributor-roles/resources/">Resources</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="software" vocab-term-identifier="https://credit.niso.org/contributor-roles/software/">Software</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="supervision" vocab-term-identifier="https://credit.niso.org/contributor-roles/supervision/">Supervision</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="visualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/visualization/">Visualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; original draft" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-original-draft/">Writing &#x2013; original draft</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &#x0026; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x0026; editing</role>
</contrib>
<contrib contrib-type="author">
<name><surname>Braznell</surname><given-names>Sophie</given-names></name>
<xref ref-type="aff" rid="aff1"></xref>
<uri xlink:href="https://loop.frontiersin.org/people/3297579"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Data curation" vocab-term-identifier="https://credit.niso.org/contributor-roles/data-curation/">Data curation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Formal analysis" vocab-term-identifier="https://credit.niso.org/contributor-roles/formal-analysis/">Formal analysis</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="investigation" vocab-term-identifier="https://credit.niso.org/contributor-roles/investigation/">Investigation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="resources" vocab-term-identifier="https://credit.niso.org/contributor-roles/resources/">Resources</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="visualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/visualization/">Visualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; original draft" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-original-draft/">Writing &#x2013; original draft</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &#x0026; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x0026; editing</role>
</contrib>
<contrib contrib-type="author">
<name><surname>Gallagher</surname><given-names>Allen</given-names></name>
<xref ref-type="aff" rid="aff1"></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2892593"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Data curation" vocab-term-identifier="https://credit.niso.org/contributor-roles/data-curation/">Data curation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Funding acquisition" vocab-term-identifier="https://credit.niso.org/contributor-roles/funding-acquisition/">Funding acquisition</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="investigation" vocab-term-identifier="https://credit.niso.org/contributor-roles/investigation/">Investigation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Project administration" vocab-term-identifier="https://credit.niso.org/contributor-roles/project-administration/">Project administration</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="resources" vocab-term-identifier="https://credit.niso.org/contributor-roles/resources/">Resources</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="supervision" vocab-term-identifier="https://credit.niso.org/contributor-roles/supervision/">Supervision</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="visualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/visualization/">Visualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; original draft" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-original-draft/">Writing &#x2013; original draft</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &#x0026; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x0026; editing</role>
</contrib>
</contrib-group>
<aff id="aff1"><institution>Tobacco Control Research Group, Department for Health, University of Bath</institution>, <city>Bath</city>, <country country="gb">United Kingdom</country></aff>
<author-notes>
<corresp id="c001"><label>&#x002A;</label>Correspondence: Sherif Elmitwalli, <email xlink:href="mailto:se606@bath.ac.uk">se606@bath.ac.uk</email></corresp>
</author-notes>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2025-12-19">
<day>19</day>
<month>12</month>
<year>2025</year>
</pub-date>
<pub-date publication-format="electronic" date-type="collection">
<year>2025</year>
</pub-date>
<volume>8</volume>
<elocation-id>1659861</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>07</month>
<year>2025</year>
</date>
<date date-type="rev-recd">
<day>09</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>11</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2025 Elmitwalli, Mehegan, Braznell and Gallagher.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Elmitwalli, Mehegan, Braznell and Gallagher</copyright-holder>
<license>
<ali:license_ref start_date="2025-12-19">https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</license-p>
</license>
</permissions>
<abstract>
<sec>
<title>Background</title>
<p>The proliferation of tobacco-related misinformation poses significant public health risks, requiring scalable solutions for credibility assessment. Traditional manual fact-checking approaches are resource-intensive and cannot match the pace of misinformation spread.</p>
</sec>
<sec>
<title>Objective</title>
<p>To develop and validate a proof-of-concept multi-agent AI pipeline for automated credibility assessment of tobacco misinformation claims, evaluating its performance against expert human reviewers.</p>
</sec>
<sec>
<title>Methods</title>
<p>We constructed a three-agent pipeline using OpenAI GPT-4.1 and the Crewai framework. The Serper API provided real-time evidence retrieval. The Content Analyzer classifies claims into four types: health impact, scientific assertion, policy, or statistical. The Scientific Fact Verifier queries authoritative sources (WHO, CDC, PubMed Central, Cochrane). The Health Evidence Assessor applies weighted scoring across five dimensions to assign 0&#x2013;100 credibility scores on a five-level scale.</p>
</sec>
<sec>
<title>Results</title>
<p>The framework achieved an MAE of 6.25 points against expert scores, a weighted Cohen&#x2019;s <italic>&#x03BA;</italic> of 0.68 (95% CI: 0.52&#x2013;0.84) indicating substantial agreement, 70% exact category agreement, 95% adjacent-level agreement, and processed each claim in under 7&#x202F;s&#x2014;over 1,000&#x202F;&#x00D7;&#x202F;faster than manual review.</p>
</sec>
<sec>
<title>Limitations</title>
<p>We validated our approach using 20 diverse tobacco claims through intensive expert review (2&#x2013;4&#x202F;h per claim). The system exhibited a conservative bias (+3.25 points, <italic>p</italic> =&#x202F;0.03) and did not classify any claims as &#x201C;Highly Unlikely&#x201D; despite expert assignment of two claims to this category. This proof-of-concept demonstrates technical feasibility and substantial inter-rater agreement while identifying areas for calibration in future large-scale implementations.</p>
</sec>
<sec>
<title>Conclusion</title>
<p>Our proof-of-concept agentic AI pipeline demonstrates substantial agreement with expert assessments of tobacco-related claims while providing dramatic speed improvements. By combining zero-shot LLM reasoning, retrieval-grounded evidence verification, and a transparent five-level scoring schema, the system offers a practical tool for real-time misinformation monitoring in public health. This proof-of-concept establishes technical feasibility for automated tobacco misinformation assessment, with validation results supporting further development and larger-scale testing before operational deployment.</p>
</sec>
</abstract>
<kwd-group>
<kwd>tobacco misinformation</kwd>
<kwd>multi-agent AI pipeline</kwd>
<kwd>large language models</kwd>
<kwd>automated fact-checking</kwd>
<kwd>credibility assessment</kwd>
<kwd>expert validation</kwd>
<kwd>public health informatics</kwd>
<kwd>retrieval-augmented generation</kwd>
</kwd-group>
<funding-group>
<award-group id="gs1">
<funding-source id="sp1">
<institution-wrap>
<institution>Bloomberg Philanthropies</institution>
<institution-id institution-id-type="doi" vocab="open-funder-registry" vocab-identifier="10.13039/open_funder_registry">10.13039/100015283</institution-id>
</institution-wrap>
</funding-source>
</award-group>
<funding-statement>The author(s) declare that financial support was received for the research and/or publication of this article. All authors are funded by Bloomberg Philanthropies as part of the Bloomberg Initiative to Reduce Tobacco Use (<ext-link xlink:href="http://www.bloomberg.org" ext-link-type="uri">www.bloomberg.org</ext-link>). The funders had no role in the study design, data collection and analysis, decision to publish or preparation of the manuscript.</funding-statement>
</funding-group>
<counts>
<fig-count count="9"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="53"/>
<page-count count="16"/>
<word-count count="10053"/>
</counts>
<custom-meta-group>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Medicine and Public Health</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1">
<label>1</label>
<title>Introduction</title>
<p>Tobacco-related misinformation poses a critical challenge to public health initiatives worldwide. Despite decades of progress in tobacco control, misinformation continues to undermine evidence-based efforts and contributes to preventable mortality. While the World Health Organization attributes over 8 million annual deaths to tobacco use (<xref ref-type="bibr" rid="ref34">Reitsma et al., 2021</xref>; <xref ref-type="bibr" rid="ref48">World Health Organization, 2025</xref>), tobacco industry misinformation and false claims about product safety have historically delayed public health interventions and undermined cessation efforts, potentially contributing to this mortality burden. In the digital era, this misinformation has proliferated across platforms at unprecedented rates, creating significant challenges for health authorities (<xref ref-type="bibr" rid="ref16">Gilmore et al., 2015</xref>; <xref ref-type="bibr" rid="ref4">Luk et al., 2021</xref>).</p>
<p>The scope of this problem spans multiple domains&#x2014;from misleading health claims and cessation methods to deceptive messaging about novel products and policy impacts. This misinformation ecosystem is particularly concerning as tobacco companies increasingly leverage social media and third-party advocates to target vulnerable populations, including youth and disadvantaged communities (<xref ref-type="bibr" rid="ref20">Tan and Bigman, 2020</xref>; <xref ref-type="bibr" rid="ref42">Alpert et al., 2021</xref>). The velocity and volume of digital misinformation has overwhelmed traditional verification approaches, creating an urgent public health need (<xref ref-type="bibr" rid="ref43">Vraga and Bode, 2020</xref>).</p>
<p>Current misinformation management relies predominantly on manual expert fact-checking&#x2014;a labor-intensive, time-consuming process that cannot scale to meet the challenge. These resource constraints create verification bottlenecks, with misleading claims spreading extensively before experts can provide evidence-based corrections (<xref ref-type="bibr" rid="ref13">Eysenbach, 2020</xref>; <xref ref-type="bibr" rid="ref40">Sylvia Chou et al., 2020</xref>). Manual approaches face three critical limitations: (1) they cannot match the speed of misinformation dissemination, (2) they require scarce specialist expertise, and (3) they struggle to provide consistent, transparent assessment methodologies (<xref ref-type="bibr" rid="ref45">Wang et al., 2019</xref>). Recent studies have demonstrated the emerging potential of generative AI for monitoring and counteracting tobacco-related misinformation on social media platforms (<xref ref-type="bibr" rid="ref23">Kong et al., 2024</xref>). These approaches leverage multimodal analysis techniques to identify problematic content across text, images, and videos (<xref ref-type="bibr" rid="ref38">Sharp et al., 2025</xref>). Our work builds upon these advances by focusing specifically on claim-level verification and assessment, rather than content filtering, to provide transparent, evidence-grounded credibility ratings.</p>
<p>To address these challenges, we present a novel multi-agent AI pipeline specifically designed for tobacco-related misinformation detection and verification. Our approach leverages advances in natural language processing, information retrieval, and evidence assessment to create a system that is both scalable and aligned with public health priorities (<xref ref-type="bibr" rid="ref30">Nyhan, 2021</xref>). By automating while maintaining scientific rigor, this framework offers a practical solution to the growing challenge of tobacco misinformation in digital spaces.</p>
<p>Our contributions include: (1) a specialized multi-agent AI pipeline that deconstructs misinformation assessment into claim extraction, evidence-based verification, and credibility evaluation; (2) a comprehensive five-level classification system calibrated for tobacco-related claims; and (3) empirical validation against expert benchmarks with systematic analysis of performance patterns and limitations. This proof-of-concept establishes technical feasibility and provides a foundation for scalable misinformation monitoring with identified pathways for addressing current limitations.</p>
<p>Unlike existing approaches that rely on static training data or generic fact-checking, our system provides domain-specific tobacco misinformation assessment with real-time evidence integration from authoritative health sources. This work bridges data science and public health by offering a practical tool for enhancing tobacco control information integrity. By prioritizing authoritative sources and scientific consensus, the system aligns technical innovation with established public health practice while dramatically accelerating the verification process. A glossary of key terms is provided in the <xref rid="app1" ref-type="app">Appendix</xref>.</p>
</sec>
<sec id="sec2">
<label>2</label>
<title>Related work</title>
<sec id="sec3">
<label>2.1</label>
<title>Tobacco misinformation overview</title>
<p>Tobacco misinformation represents a deliberate, documented strategy in industry practices spanning decades. Historical patterns of science manipulation (<xref ref-type="bibr" rid="ref33">Proctor, 2012</xref>; <xref ref-type="bibr" rid="ref31">Oreskes and Conway, 2011</xref>) have evolved into sophisticated digital tactics promoting unsubstantiated claims about products across multiple platforms (<xref ref-type="bibr" rid="ref21">Jackler et al., 2019</xref>; <xref ref-type="bibr" rid="ref1">Allem et al., 2017</xref>). These misleading claims follow distinct typological patterns&#x2014;including health risk minimization, exaggeration of cessation benefits, misleading statistics, and policy impact distortions&#x2014;creating predictable information distortion that undermines public health (<xref ref-type="bibr" rid="ref3">Apollonio and Malone, 2009</xref>). The consequences are substantial: exposure to misinformation correlates with decreased cessation attempts and increased youth susceptibility to product initiation (<xref ref-type="bibr" rid="ref6">Brennan et al., 2017</xref>; <xref ref-type="bibr" rid="ref41">Tan et al., 2015</xref>), underscoring the urgency of effective countermeasures.</p>
</sec>
<sec id="sec4">
<label>2.2</label>
<title>Manual fact-checking limitations</title>
<p>Traditional approaches to tobacco misinformation management rely on resource-constrained expert verification processes. Major health authorities maintain dedicated fact-checking resources (<xref ref-type="bibr" rid="ref47">World Health Organization, 2022</xref>; <xref ref-type="bibr" rid="ref8">Centers for Disease Control and Prevention, 2023</xref>), but face significant efficiency barriers, with comprehensive claim assessment typically requiring 2&#x2013;4&#x202F;h per claim (<xref ref-type="bibr" rid="ref5">Bodaghi et al., 2024</xref>). While structured protocols for tobacco claim assessment exist (<xref ref-type="bibr" rid="ref25">Leone et al., 2018</xref>), scaling these approaches faces fundamental challenges: expert availability constraints, verification delays, and inconsistent methodologies across fact-checking entities (<xref ref-type="bibr" rid="ref36">Schmidt et al., 2018</xref>). These limitations extend beyond resource constraints to include cognitive biases in expert assessments (<xref ref-type="bibr" rid="ref12">Erku et al., 2021</xref>) and the predominantly reactive nature of manual verification, occurring after misinformation has achieved substantial dissemination (<xref ref-type="bibr" rid="ref18">Hendlin et al., 2019</xref>).</p>
</sec>
<sec id="sec5">
<label>2.3</label>
<title>Computational approaches</title>
<p>Recent years have seen significant advances in computational methods for misinformation detection, though few target tobacco content specifically. General approaches typically employ content-based, social context-based, or hybrid methodologies (<xref ref-type="bibr" rid="ref52">Zhou and Zafarani, 2020</xref>), while health-specific implementations have demonstrated promising results using linguistic features and credibility metrics (<xref ref-type="bibr" rid="ref32">P&#x00E9;rez-Rosas et al., 2018</xref>; <xref ref-type="bibr" rid="ref15">Ghenai and Mejova, 2018</xref>). However, three critical limitations persist in existing computational approaches: (1) they often employ binary classification (true/false) rather than nuanced credibility assessment required for complex tobacco claims (<xref ref-type="bibr" rid="ref10">Dai et al., 2020</xref>); (2) they rarely incorporate domain-specific knowledge and authoritative health sources; and (3) they struggle with limited training data availability in specialized domains like tobacco control.</p>
<p>Advances in large language models (LLMs) show potential for health misinformation detection, particularly through retrieval-augmented generation approaches that improve factual accuracy (<xref ref-type="bibr" rid="ref50">Gargari and Habibi, 2025</xref>; <xref ref-type="bibr" rid="ref37">Baashirah, 2024</xref>). However, these models risk perpetuating rather than detecting misinformation without domain-specific training and robust evidence retrieval mechanisms (<xref ref-type="bibr" rid="ref26">Zhang et al., 2025</xref>). Parallel developments in biomedical natural language processing (NLP) offer promising techniques for evidence extraction (<xref ref-type="bibr" rid="ref35">Sarrouti and El Alaoui, 2017</xref>), automated implementation of evidence quality assessment frameworks (<xref ref-type="bibr" rid="ref29">Marshall et al., 2015</xref>; <xref ref-type="bibr" rid="ref44">Wallace et al., 2010</xref>), and methods for quantifying scientific consensus (<xref ref-type="bibr" rid="ref28">Luo et al., 2017</xref>; <xref ref-type="bibr" rid="ref51">Zhang et al., 2016</xref>), creating opportunities for more sophisticated tobacco misinformation assessment. Other research found that specialized LLM instruction tuning significantly improved adherence to health guidelines in smoking cessation advice, achieving 72.2% guideline adherence compared to 47.8% for general-purpose models, highlighting the importance of domain-specific optimization for health information assessment (<xref ref-type="bibr" rid="ref7">Abroms et al., 2025</xref>).</p>
</sec>
<sec id="sec6">
<label>2.4</label>
<title>Gap analysis</title>
<p>Despite significant advances in computational health information assessment, several critical gaps remain unaddressed. First, tobacco-specific misinformation detection has received limited attention despite its public health significance and unique characteristics. Second, existing approaches often lack integration with authoritative evidence sources and public health priorities. Third, most systems provide binary classifications rather than nuanced credibility assessments reflecting evidence quality variations. Our work addresses these gaps by introducing a specialized multi-agent AI pipeline that: (1) integrates advanced NLP with authoritative tobacco-specific evidence sources; (2) introduces a nuanced, five-level credibility framework grounded in evidence-based public health principles.; and (3) provides transparent evidence trails supporting assessment outcomes. This approach bridges technical innovation with practical public health needs in tobacco information management while demonstrating how multi-agent AI pipeline can effectively coordinate specialized components in complex healthcare information tasks (<xref ref-type="bibr" rid="ref19">Isern and Moreno, 2016</xref>; <xref ref-type="bibr" rid="ref49">Yuan and Herbert, 2014</xref>; <xref ref-type="bibr" rid="ref46">Wimmer et al., 2016</xref>; <xref ref-type="bibr" rid="ref2">Amith et al., 2020</xref>).</p>
</sec>
</sec>
<sec sec-type="methods" id="sec7">
<label>3</label>
<title>Methodology</title>
<sec id="sec8">
<label>3.1</label>
<title>System framework and multi-agent AI pipeline</title>
<p>Our system uses a modular, three-agent AI framework. Each agent has distinct responsibilities: extraction, verification, and credibility assessment. The agents work sequentially but independently (<xref ref-type="fig" rid="fig1">Figure 1</xref>). This design separates concerns while maintaining information flow between stages.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Sequence diagram of the multi-agent claim verification pipeline: blue arrows show the core processing flow, green arrows denote external evidence retrieval, and the purple arrow, the final score return.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g001.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Flowchart showing a sequence for processing a claim. Steps include a user submitting a raw claim to a Content Analyzer, which sends a structured query to a Scientific Verifier. Evidence is retrieved from an Evidence Assessor using sources like WHO Database, CDC Reports, and PubMed/Cochrane. Evidence snippets are collected, consolidated, and returned with a score and justification.</alt-text>
</graphic>
</fig>
<p>The system operates in three sequential phases:</p><list list-type="order">
<list-item>
<p><bold>Claim extraction and characterization</bold>: Identification and structuring of tobacco-related claims from textual data.</p>
</list-item>
<list-item>
<p><bold>Evidence-based verification</bold>: Algorithmic comparison of extracted claims against authoritative scientific sources.</p>
</list-item>
<list-item>
<p><bold>Credibility assessment and scoring</bold>: Quantitative evaluation of claim credibility based on evidence quality and scientific consensus.</p>
</list-item>
</list>
<p>Our pipeline comprises three sequential agents&#x2014;Content Analyzer, Scientific Verifier, and Health Evidence Assessor&#x2014;each exchanging messages as shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. Blue arrows trace the user&#x2019;s claim as it moves through the Content Analyzer, Scientific Verifier, and Health Evidence Assessor agents. Green arrows illustrate calls to external data sources (WHO Database, CDC Reports, PubMed/Cochrane), and the purple arrow marks the return of the computed credibility score and justification back to the user.</p>
<p>This pipeline approach ensures each claim undergoes consistent, thorough analysis while maintaining processing efficiency. The modular design allows for independent optimization of each component and facilitates system scalability as data volumes increase.</p>
</sec>
<sec id="sec9">
<label>3.2</label>
<title>Data sources and claim selection</title>
<sec id="sec10">
<label>3.2.1</label>
<title>Data collection</title>
<p>We sourced data exclusively from authoritative public health repositories and scientific databases. Including:</p><list list-type="bullet">
<list-item>
<p>World Health Organization (WHO) Framework Convention on Tobacco Control documentation</p>
</list-item>
<list-item>
<p>Centers for Disease Control and Prevention (CDC) tobacco factsheets</p>
</list-item>
<list-item>
<p>PubMed Central peer-reviewed research articles</p>
</list-item>
<list-item>
<p>Cochrane Database of Systematic Reviews</p>
</list-item>
</list>
<p>For each source, we retrieved both metadata and full-text content where available through the Serper API, which provides real-time access to authoritative health databases. From WHO and CDC sources, we extracted complete documentation including guidelines, policy statements, and statistical reports. For academic sources (PubMed Central and Cochrane), we accessed both abstracts and available full-text articles, with priority given to systematic reviews and meta-analyses. The Serper API&#x2019;s domain-specific search capabilities were restricted to predetermined authoritative domains (WHO, CDC, PubMed Central, and Cochrane), with automated verification of source URLs and publication dates.</p>
<p>Data collection was restricted to authoritative public health repositories and scientific databases. While these sources (particularly WHO and CDC) maintain institutional independence from industry influence, and PubMed Central and Cochrane require conflict of interest declarations, we did not perform detailed analysis of potential industry sponsorship at the individual study level. Our approach relies on institutional credibility and multi-source evidence triangulation to mitigate potential bias. Future implementations could enhance robustness by incorporating automated conflict-of-interest detection and evidence weighting based on funding transparency.</p>
<p>Our core algorithmic framework operates through a sequential pipeline where first the Content Analyzer Agent identifies tobacco-related health claims using NLP techniques. Extracted claims are then processed by the Evidence Retrieval Agent, which queries authoritative health databases and applies relevance filtering to compile evidence packages. The Credibility Assessment Agent evaluates these packages across five dimensions using weighted scoring from authoritative sources, producing numerical credibility scores (0&#x2013;100). Finally, the Classification Agent maps these scores to our five-level credibility scale through predefined thresholds, ensuring consistent and interpretable outputs for end users as shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Core algorithmic framework with key decision thresholds.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g002.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Flowchart illustrating a decision process for evaluating tobacco claims. Starting with raw text input, the flowchart checks for tobacco claims. If no claims, it discards the input. If claims exist, it extracts structured claims and queries health databases. If evidence is missing, it expands the query. Relevant evidence leads to compiling, calculating a credibility score, and mapping score thresholds from 81-100 (highly likely) to 0-20 (highly unlikely). It generates justification, resulting in the final output with classification, evidence, and confidence levels.</alt-text>
</graphic>
</fig>
</sec>
<sec id="sec11">
<label>3.2.2</label>
<title>Claim categorization</title>
<p>To ensure comprehensive coverage across the tobacco information landscape, we categorized claims into four distinct types:</p><list list-type="order">
<list-item>
<p><bold>Health impact claims</bold>: Assertions regarding physiological or psychological effects of tobacco products (e.g., &#x201C;Smoking reduces life expectancy by 10&#x202F;years&#x201D;).</p>
</list-item>
<list-item>
<p><bold>Scientific assertions</bold>: Claims regarding chemical properties, biological mechanisms, or research findings (e.g., &#x201C;Nicotine replacement therapy doubles cessation success rates&#x201D;).</p>
</list-item>
<list-item>
<p><bold>Policy-related statements</bold>: Declarations about regulatory effectiveness or industry practices (e.g., &#x201C;Plain packaging has no impact on smoking initiation rates&#x201D;).</p>
</list-item>
<list-item>
<p><bold>Statistical claims</bold>: Numeric assertions about prevalence, mortality, or economic impacts (e.g., &#x201C;Over 8 million people die annually from tobacco-related illnesses&#x201D;).</p>
</list-item>
</list>
<p>This taxonomic approach facilitated systematic processing and enabled analysis of performance variations across claim types. The 20 claims were systematically selected by the research team to ensure balanced representation across our four credibility categories and diverse evidence complexity levels. Selection criteria prioritized claims with established expert consensus in the literature, documented public health significance, varying degrees of evidence availability, and representation of common tobacco misinformation patterns identified in prior content analyses. We selected tobacco misinformation as our validation domain for several methodological reasons. First, tobacco represents a well-documented baseline of established scientific consensus, providing reliable ground truth against which to validate automated assessments. Second, it addresses a significant public health challenge with documented industry misinformation campaigns spanning decades (<xref ref-type="bibr" rid="ref14">Gannon et al., 2023</xref>). Third, it offers diverse claim types (health impacts, policy effects, statistical assertions) within a coherent domain. While the relative stability of tobacco evidence compared to rapidly evolving domains such as emerging infectious diseases may favor our system&#x2019;s performance, this choice provides essential proof-of-concept validation before tackling more challenging, time-sensitive health topics. In assembling our 20-claim test set, we applied four selection criteria to ensure a realistic and challenging evaluation. First, each claim addresses a clear public-health impact (e.g., morbidity, mortality, or policy implications). Second, we balanced representation across our four claim categories (health impact, scientific assertion, policy, and statistical) to probe performance on diverse content types. Third, we prioritized claims with high visibility&#x2014;drawing from authoritative WHO/CDC publications and recent social-media or web-search trends&#x2014;to reflect real-world misinformation exposure. Finally, we included both well-established statements and emerging or contested assertions to test the pipeline&#x2019;s ability to handle varying levels of scientific consensus.</p>
</sec>
</sec>
<sec id="sec12">
<label>3.3</label>
<title>Agent design and implementation</title>
<sec id="sec13">
<label>3.3.1</label>
<title>Content analyzer agent</title>
<p>The Content Analyzer Agent performs claim extraction and initial characterization using advanced NLP techniques. This agent employs:</p><list list-type="bullet">
<list-item>
<p><bold>Named entity recognition (NER)</bold>: Identifies tobacco products, health conditions, and intervention terms within text.</p>
</list-item>
<list-item>
<p><bold>Dependency parsing</bold>: Analyzes grammatical structure to extract complete claim statements.</p>
</list-item>
<list-item>
<p><bold>Semantic analysis</bold>: Classifies claims into the four categories.</p>
</list-item>
<list-item>
<p><bold>Claim prioritization</bold>: Ranks claims based on potential public health impact and information spread.</p>
</list-item>
</list>
<p>The classification system recognizes that tobacco-related claims often span multiple categories simultaneously. For instance, a single claim might combine statistical evidence with health impact assertions, or policy statements with scientific findings. In such cases, the Content Analyzer assigns both primary and secondary classifications based on the dominant characteristics present in the claim, with corresponding confidence scores for each category assignment. This multi-category approach ensures that complex claims receive comprehensive analysis reflecting their full informational content.</p>
<p>The classification process follows a structured NLP pipeline that integrates several analytical techniques. First, named entity recognition pinpoints tobacco-specific terms, health conditions, and key statistics. Next, dependency parsing reconstructs each claim&#x2019;s full syntactic structure. In the semantic analysis phase, claim embeddings are compared against category-specific reference sets to yield confidence scores for each potential classification. These scores are combined with structural completeness metrics and subject-matter keyword matching to produce final category assignments. Claims are subsequently prioritized based on their potential public health impact, considering factors such as population reach, evidence strength, and dissemination patterns.</p>
</sec>
<sec id="sec14">
<label>3.3.2</label>
<title>Scientific fact verifier agent</title>
<p>The Scientific Fact Verifier Agent evaluates extracted claims against authoritative scientific evidence. This agent:</p><list list-type="order">
<list-item>
<p>Transforms claims into structured queries optimized for scientific database retrieval</p>
</list-item>
<list-item>
<p>Accesses multiple authoritative data sources including WHO, CDC, and PubMed Central</p>
</list-item>
<list-item>
<p>Retrieves relevant scientific literature, systematic reviews, and health authority statements</p>
</list-item>
<list-item>
<p>Analyzes evidence quality, consistency, and relevance to the specific claim</p>
</list-item>
<list-item>
<p>Documents evidence trails with bibliographic references for transparency</p>
</list-item>
</list>
<p>The agent employs retrieval-augmented prompting to ensure verified information is grounded in authoritative sources rather than model-generated content. Unlike traditional RAG systems that append raw documents to prompts, our agent processes and synthesizes retrieved evidence before passing structured summaries to subsequent agents. This methodology enhances factual accuracy while reducing hallucination risks commonly associated with LLMs. Our approach aligns with recent advancements in retrieval-augmented techniques which demonstrated that enhancing instruction diversity and structured knowledge integration improved both accuracy and transparency in knowledge-intensive tasks (<xref ref-type="bibr" rid="ref27">Liu and Chen, 2025</xref>). Similar principles could further enhance our Scientific Fact Verifier agent&#x2019;s ability to retrieve and integrate evidence from authoritative sources.</p>
</sec>
<sec id="sec15">
<label>3.3.3</label>
<title>Health evidence assessor agent</title>
<p>The Health Evidence Assessor Agent performs credibility assessment and generates final scores based on verification results. This agent:</p><list list-type="order">
<list-item>
<p>Evaluates evidence strength using established frameworks (e.g., GRADE methodology principles; <xref ref-type="bibr" rid="ref17">Guyatt et al., 2025</xref>)</p>
</list-item>
<list-item>
<p>Assesses alignment with scientific consensus across authoritative sources</p>
</list-item>
<list-item>
<p>Analyzes evidence consistency, recency, and methodological quality</p>
</list-item>
<list-item>
<p>Generates a numerical credibility score (0&#x2013;100) with qualitative justification</p>
</list-item>
<list-item>
<p>Maps scores to the five-level credibility classification system</p>
</list-item>
</list>
<p>The agent implements a weighted scoring algorithm that prioritizes high-quality evidence (e.g., systematic reviews, meta-analyses) over lower-quality evidence (e.g., case reports, opinion pieces), with explicit weighting factors (<xref ref-type="bibr" rid="ref9">Chloros et al., 2023</xref>).</p>
</sec>
<sec id="sec16">
<label>3.3.4</label>
<title>Core algorithmic framework</title>
<p>The system implements four key algorithms that form the backbone of our multi-agent AI pipeline. These algorithms work in concert to process, verify, and assess tobacco-related claims, with each addressing a specific aspect of the misinformation detection pipeline.</p>
<p>The claim extraction algorithm (<xref ref-type="fig" rid="fig3">Algorithm 1</xref>) implements the initial processing phase, focusing on identifying and structuring tobacco-related claims from input text. It employs natural language processing techniques to isolate relevant sentences and applies a multi-step analysis process to extract, normalize, and categorize claims. The algorithm&#x2019;s confidence scoring mechanism ensures that only well-formed claims proceed to subsequent stages, while the prioritization step orders claims based on their potential public health impact.</p>
<fig position="float" id="fig3">
<label>Algorithm 1</label>
<caption>
<p>Claim extraction.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g003.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Flowchart showing an algorithm for processing text. It takes text and keywords as input, initializes an empty set of claims, tokenizes the text into sentences, and checks each sentence for keywords. If keywords are found, it creates a claim with normalized text, classifies its type, searches for evidence, assesses confidence, and adds it to the claims set. Finally, it prioritizes claims by impact for output.</alt-text>
</graphic>
</fig>
<p>The evidence verification algorithm (<xref ref-type="fig" rid="fig4">Algorithm 2</xref>) represents the core fact-checking component of our system. It implements a sophisticated retrieval-augmented generation approach (<xref ref-type="bibr" rid="ref27">Liu and Chen, 2025</xref>), querying multiple authoritative sources with carefully weighted credibility scores. The algorithm&#x2019;s hierarchical evidence gathering process ensures comprehensive coverage while maintaining efficiency. By incorporating source-specific weights derived from expert consensus, the system can effectively differentiate between varying levels of authority in health information sources (<xref ref-type="bibr" rid="ref22">Kington et al., 2021</xref>).</p>
<fig position="float" id="fig4">
<label>Algorithm 2</label>
<caption>
<p>Evidence verification.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g004.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Flowchart illustrating a procedure for verifying a claim using source weights. The process includes generating queries from the claim, using an API to search sources by assigned weights, assessing relevance and extracting evidence from results, and calculating consensus and temporal relevance of the evidence.</alt-text>
</graphic>
</fig>
<p>The credibility scoring algorithm (<xref ref-type="fig" rid="fig5">Algorithm 3</xref>) implements our novel multi-dimensional assessment framework. Drawing inspiration from evidence-based medicine hierarchies, it evaluates claims across five key dimensions: evidence quality, scientific consensus, consistency, recency, and scientific plausibility. The weighted scoring system reflects the relative importance of each factor in determining overall credibility, with higher weights assigned to fundamental aspects like evidence quality (40%) and scientific consensus (25%). The pipeline orchestration algorithm (<xref ref-type="fig" rid="fig6">Algorithm 4</xref>) serves as the system&#x2019;s coordination layer, managing the flow of information between agents and ensuring proper uncertainty propagation throughout the assessment process. This algorithm implements a robust error handling mechanism and maintains detailed confidence metrics at each stage. By tracking uncertainty propagation, it provides transparent reliability indicators for final assessments.</p>
<fig position="float" id="fig5">
<label>Algorithm 3</label>
<caption>
<p>Credibility scoring.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g005.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Flowchart detailing a process to calculate a credibility score from verification results \(V\) and weights \(W\). Inputs include evidence quality, consensus, consistency, recency, and plausibility. The score \(S\) is calculated by summing weighted components and normalized. Outputs range from 0 to 100.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig6">
<label>Algorithm 4</label>
<caption>
<p>Pipeline orchestration.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g006.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Algorithm pseudocode for assessing text credibility. Inputs: text \(T\) and API interface \(I\). Outputs: assessment \(A\) with confidence intervals. Steps: Extract claims from \(T\), verify each claim using API, calculate and score credibility, update results with claim, score, and confidence. Return results.</alt-text>
</graphic>
</fig>
<p>In summary, the integration of <xref ref-type="fig" rid="fig3">Algorithms 1</xref>&#x2013;<xref ref-type="fig" rid="fig6">4</xref> describe our three main contributions&#x2014;a modular, claim-by-claim processing pipeline; an evidence-grounded verification stage drawing on WHO, CDC, PubMed Central, and Cochrane; and a transparent, five-level credibility scoring system. Each algorithm maps directly to one phase in the workflow depicted in <xref ref-type="fig" rid="fig1">Figures 1</xref>, <xref ref-type="fig" rid="fig2">2</xref>, and when run end-to-end, this framework delivers the accuracy, inter-rater agreement, and processing-time results presented in Section 4.</p>
<p>The framework&#x2019;s design emphasizes reproducibility and scalability, with explicit error handling and confidence scoring at each stage. Our evidence verification process employs state-of-the-art retrieval-augmented generation techniques (<xref ref-type="bibr" rid="ref27">Liu and Chen, 2025</xref>) to ensure factual grounding, while the credibility scoring implements the five-dimensional assessment framework that achieved strong agreement with expert reviewers. This comprehensive approach enables rapid, reliable assessment of tobacco-related claims while maintaining the rigor required for public health applications.</p>
</sec>
</sec>
<sec id="sec17">
<label>3.4</label>
<title>Credibility assessment framework</title>
<sec id="sec18">
<label>3.4.1</label>
<title>Five-level classification system</title>
<p>We developed a five-level classification system for credibility assessment, providing nuanced differentiation between varying degrees of scientific support:</p><list list-type="bullet">
<list-item>
<p><bold>Highly likely to be credible (81&#x2013;100)</bold>: Claims with overwhelming scientific evidence and consensus from authoritative sources. These claims are consistently supported by multiple high-quality studies, systematic reviews, or meta-analyses with minimal contradictory findings.</p>
</list-item>
<list-item>
<p><bold>Likely to be credible (61&#x2013;80)</bold>: Claims with substantial supporting evidence but with minor limitations or areas of ongoing research. These claims are supported by multiple studies with generally consistent findings, though some methodological limitations or gaps may exist.</p>
</list-item>
<list-item>
<p><bold>Moderate credibility (41&#x2013;60)</bold>: Claims with mixed evidence or where scientific consensus is still developing. These claims typically have supporting and contradicting evidence of similar quality or volume, or represent areas where research is still evolving.</p>
</list-item>
<list-item>
<p><bold>Unlikely to be credible (21&#x2013;40)</bold>: Claims with limited supporting evidence and substantial contradictory findings. These claims contradict most available evidence but may have minimal or low-quality supporting data.</p>
</list-item>
<list-item>
<p><bold>Highly unlikely to be credible (0&#x2013;20)</bold>: Claims that directly contradict established scientific consensus or lack any credible supporting evidence. These claims are inconsistent with fundamental scientific principles or are contradicted by substantial high-quality evidence.</p>
</list-item>
</list>
<p>This granular classification enhances decision support for public health officials and improves communication clarity for non-technical audiences compared to broader three-level systems.</p>
</sec>
<sec id="sec19">
<label>3.4.2</label>
<title>Scoring algorithm</title>
<p>The Health Evidence Assessor Agent employs a multi-dimensional scoring algorithm that evaluates claims across five key dimensions:</p><list list-type="order">
<list-item>
<p><bold>Evidence quality</bold> (40%): Evaluates the methodological rigor of supporting studies, with higher weights assigned to systematic reviews, randomized controlled trials, and meta-analyses.</p>
</list-item>
<list-item>
<p><bold>Scientific consensus</bold> (25%): Measures agreement across authoritative sources and relevant expert bodies.</p>
</list-item>
<list-item>
<p><bold>Evidence consistency</bold> (15%): Assesses whether findings from multiple studies demonstrate consistent conclusions.</p>
</list-item>
<list-item>
<p><bold>Evidence recency</bold> (10%): Evaluates whether the claim reflects current understanding, with higher weights for evidence published within the last 5&#x202F;years.</p>
</list-item>
<list-item>
<p><bold>Scientific plausibility</bold> (10%): Considers alignment with established scientific principles and mechanisms.</p>
</list-item>
</list>
<p>Each dimension contributes to the final score through a weighted formula:</p>
<p><inline-formula>
<mml:math id="M1">
<mml:mtable columnalign="left" displaystyle="true">
<mml:mtr>
<mml:mtd>
<mml:mtext>Score</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>Eq</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>0.4</mml:mn>
<mml:mo stretchy="true">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>Sc</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>0.25</mml:mn>
<mml:mo stretchy="true">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>Ec</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>0.15</mml:mn>
<mml:mo stretchy="true">)</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>+</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>Er</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>0.1</mml:mn>
<mml:mo stretchy="true">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>Sp</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>0.1</mml:mn>
<mml:mo stretchy="true">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</inline-formula>.</p>
<p>Where:</p><list list-type="bullet">
<list-item>
<p>Eq&#x202F;=&#x202F;Evidence Quality score (0&#x2013;100)</p>
</list-item>
<list-item>
<p>Sc&#x202F;=&#x202F;Scientific Consensus score (0&#x2013;100)</p>
</list-item>
<list-item>
<p>Ec&#x202F;=&#x202F;Evidence Consistency score (0&#x2013;100)</p>
</list-item>
<list-item>
<p>Er&#x202F;=&#x202F;Evidence Recency score (0&#x2013;100)</p>
</list-item>
<list-item>
<p>Sp&#x202F;=&#x202F;Scientific Plausibility score (0&#x2013;100)</p>
</list-item>
</list>
<p>These component weights were determined through expert consensus and validated in our pilot study with public health specialists, as described in the following section.</p>
</sec>
</sec>
<sec id="sec20">
<label>3.5</label>
<title>Validation approach</title>
<p>Our 20-claim validation was designed as a proof-of-concept study to establish technical feasibility before larger-scale implementation. The achieved Cohen&#x2019;s <italic>&#x03BA;</italic> of 0.68 demonstrates substantial agreement according to established interpretation guidelines, with the 95% CI (0.52&#x2013;0.84) spanning from moderate to substantial agreement ranges. The substantial agreement achieved represents a significant milestone for automated health misinformation assessment, establishing that multi-agent AI pipeline can replicate expert-level judgment patterns in controlled validation conditions.</p>
<sec id="sec21">
<label>3.5.1</label>
<title>Manual assessment comparison</title>
<p>To validate the automated framework&#x2019;s credibility scores, two expert reviewers independently evaluated all 20 claims. Reviewer A holds a PhD in public health with 12&#x202F;+&#x202F;years of tobacco control research experience; Reviewer B holds a PhD in epidemiology with 8&#x202F;+&#x202F;years of tobacco-related health outcomes research. Both reviewers were blinded to automated scores during initial assessment. Inter-rater reliability between the two primary reviewers before adjudication was <italic>&#x03BA;</italic> =&#x202F;0.74 (95% CI: 0.59&#x2013;0.89). Six claims required third-party adjudication due to score differences &#x003E;15 points. The adjudication process involved structured discussion of evidence interpretation differences, with a third expert (PhD in public health, 15&#x202F;+&#x202F;years tobacco policy research) providing final consensus scores. For each claim, reviewers assigned a 0&#x2013;100 score using the same five-level classification and recorded detailed justifications.</p>
</sec>
<sec id="sec22">
<label>3.5.2</label>
<title>Performance metrics</title>
<p>We evaluated framework performance using multiple complementary metrics:</p><list list-type="order">
<list-item>
<p><bold>Accuracy</bold>: Percentage of claims where automated and manual classifications matched exactly</p>
</list-item>
<list-item>
<p><bold>Adjacent accuracy</bold>: Percentage of claims where automated classification was within one level of manual classification</p>
</list-item>
<list-item>
<p><bold>Mean absolute error (MAE)</bold>: Average absolute difference between automated and manual numerical scores</p>
</list-item>
<list-item>
<p><bold>Weighted Cohen&#x2019;s kappa</bold>: Measure of inter-rater reliability between automated and manual classifications, accounting for partial agreement</p>
</list-item>
<list-item>
<p><bold>Processing time</bold>: Time required for complete claim processing (extraction to final score)</p>
</list-item>
</list>
<p>These metrics provided comprehensive assessment of both classification accuracy and operational efficiency.</p>
</sec>
</sec>
<sec id="sec23">
<label>3.6</label>
<title>Technical implementation and infrastructure</title>
<p>Our multi-agent AI pipeline is powered by OpenAI&#x2019;s GPT-4.1 (no fine-tuning) as the core language model. To ground claims in up-to-date evidence, we integrated the Serper API for real-time web search retrieval; snippets and source URLs are appended to prompts passed to GPT-4.1. The overall workflow&#x2014;claim extraction, evidence verification, and credibility scoring&#x2014;is orchestrated by the CrewAI framework, which manages agent definitions, asynchronous tool invocations, and inter-agent messaging. This technical approach ensures the framework remains adaptable to evolving misinformation patterns and public health needs. All agent&#x2013;API interactions are automatically instrumented with Langtrace, producing timestamped traces of every Serper query and API call&#x2014;enabling exact reproduction of the outlier assessments (<xref ref-type="bibr" rid="ref24">Langtrace, 2024</xref>).</p>
<p>At its heart, our infrastructure combines structured prompt schemas, a lightweight multi-agent orchestrator, and a dynamic retrieval-grounding layer. Each agent operates from a templated instruction set defining its role, goal, and narrative context, with runtime placeholders that inject the specific research topic. A central orchestrator then sequences agent execution and carries outputs forward through each stage&#x2014;ensuring smooth, reproducible transitions from claim extraction to final scoring. Meanwhile, agents invoke a web-search API on the fly to fetch, filter, and integrate real-world evidence from trusted domains directly into the GPT-4.1 context, minimizing hallucinations and keeping responses current. An optional preferences module can further tailor prompts with user-centric context when needed. Together, these elements yield a scalable, transparent framework that balances precise agent responsibilities with robust workflow management and dynamic, evidence-based prompting.</p>
<p>Complete implementation details, including agent definitions, prompt templates, and scoring algorithms, are available in our GitHub repository (<xref ref-type="bibr" rid="ref11">Elmitwalli, 2025</xref>). The repository includes full source code. This ensures full reproducibility and enables independent verification of our methodological claims. Planned enhancements include: (i) an optional COI down-weighting heuristic for industry-funded studies; (ii) an optional calibration module (e.g., isotonic regression and contradiction-penalty rules) to improve sensitivity for low-credibility claims; and (iii) optional retrieval-logging/export to support third-party computation of Recall@k/MRR and faithfulness (<xref ref-type="bibr" rid="ref39">Silva Filho et al., 2023</xref>).</p>
</sec>
</sec>
<sec sec-type="results" id="sec24">
<label>4</label>
<title>Results</title>
<sec id="sec25">
<label>4.1</label>
<title>Performance overview</title>
<p>Our multi-agent AI framework for tobacco misinformation assessment demonstrated substantial agreement with expert evaluations while achieving remarkable processing efficiency. The framework evaluated 20 representative tobacco-related claims spanning health effects, policy impacts, scientific assertions, and statistical claims.</p>
</sec>
<sec id="sec26">
<label>4.2</label>
<title>Claim-level assessment analysis</title>
<p><xref ref-type="table" rid="tab1">Table 1</xref> presents a comprehensive comparison of automated and manual credibility scores for the 20 tobacco-related claims evaluated. The framework assigned scores on a 0&#x2013;100 scale and mapped them to a five-level classification framework (Highly Unlikely, Unlikely, Moderate, Likely, Highly Likely; <xref ref-type="fig" rid="fig7">Figure 3</xref>).</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Automated vs. manual assessment of 20 tobacco-related claims.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Claim</th>
<th align="center" valign="top">Automated score</th>
<th align="center" valign="top">Manual score</th>
<th align="left" valign="top">Automated category</th>
<th align="left" valign="top">Manual category</th>
<th align="left" valign="top">Reviewer comments</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Smoking reduces life expectancy by 10&#x202F;years</td>
<td align="center" valign="top">95</td>
<td align="center" valign="top">95</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Robust, consistent epidemiological evidence (WHO, CDC, meta-analyses).</td>
</tr>
<tr>
<td align="left" valign="top">Second-hand smoke exposure increases lung cancer risk by 25%</td>
<td align="center" valign="top">90</td>
<td align="center" valign="top">90</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Well-documented 20&#x2013;30% increased risk in cohort and case&#x2013;control studies.</td>
</tr>
<tr>
<td align="left" valign="top">E-cigarettes are completely safe for long-term use</td>
<td align="center" valign="top">25</td>
<td align="center" valign="top">10</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Highly Unlikely</td>
<td align="left" valign="top">&#x201C;Completely safe&#x201D; is misleading; long-term safety unproven and emerging data show harms.</td>
</tr>
<tr>
<td align="left" valign="top">Smokeless tobacco products like snus significantly reduce oral cancer risk relative to smoking</td>
<td align="center" valign="top">65</td>
<td align="center" valign="top">60</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Moderate</td>
<td align="left" valign="top">Some studies show reduced risk, but evidence is mixed and dependent on use patterns.</td>
</tr>
<tr>
<td align="left" valign="top">Nicotine consumption leads to irreversible brain damage in adolescents</td>
<td align="center" valign="top">90</td>
<td align="center" valign="top">90</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Strong consensus on developmental neurotoxicity from animal and human studies.</td>
</tr>
<tr>
<td align="left" valign="top">Tobacco taxation is the most effective method to reduce smoking rates</td>
<td align="center" valign="top">85</td>
<td align="center" valign="top">85</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Tax increases consistently rank among top tobacco control measures in econometric and public-health reviews.</td>
</tr>
<tr>
<td align="left" valign="top">Plain packaging has no impact on smoking initiation rates</td>
<td align="center" valign="top">45</td>
<td align="center" valign="top">30</td>
<td align="left" valign="top">Moderate</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Empirical studies demonstrate modest reductions in youth appeal and initiation; &#x201C;no impact&#x201D; is unlikely.</td>
</tr>
<tr>
<td align="left" valign="top">E-cigarettes are banned in over 50 countries worldwide</td>
<td align="center" valign="top">40</td>
<td align="center" valign="top">30</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Fewer than 50 full bans (WHO reports ~35); claim overstates the global count.</td>
</tr>
<tr>
<td align="left" valign="top">Tobacco industry lobbying weakens public health policies globally</td>
<td align="center" valign="top">85</td>
<td align="center" valign="top">85</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Extensive literature documents lobbying&#x2019;s negative influence on FCTC implementation.</td>
</tr>
<tr>
<td align="left" valign="top">Flavored tobacco products target youth users</td>
<td align="center" valign="top">90</td>
<td align="center" valign="top">95</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Clear marketing strategies and youth-prevalence data confirm flavor targeting.</td>
</tr>
<tr>
<td align="left" valign="top">Nicotine replacement therapies (NRTs) double quitting success</td>
<td align="center" valign="top">75</td>
<td align="center" valign="top">75</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Meta-analyses show ~1.5&#x2013;2&#x202F;&#x00D7;&#x202F;improvement in quit rates with NRT vs. placebo.</td>
</tr>
<tr>
<td align="left" valign="top">Tobacco companies have funded research denying the link between smoking and cancer</td>
<td align="center" valign="top">95</td>
<td align="center" valign="top">95</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Historical internal documents confirm industry-sponsored denial campaigns.</td>
</tr>
<tr>
<td align="left" valign="top">The harmful effects of vaping exceed those of smoking</td>
<td align="center" valign="top">30</td>
<td align="center" valign="top">10</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Highly Unlikely</td>
<td align="left" valign="top">Increasing evidence vaping less harmful than smoking; claim contradicts major reviews; &#x201C;exceed&#x201D; is highly unlikely</td>
</tr>
<tr>
<td align="left" valign="top">Smoking cessation reduces heart disease risk within 5&#x202F;years</td>
<td align="center" valign="top">70</td>
<td align="center" valign="top">70</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Risk declines by ~50% within 5&#x202F;years of quitting; supported by cohort studies.</td>
</tr>
<tr>
<td align="left" valign="top">The tobacco industry&#x2019;s harm-reduction investment is predominantly profit-driven</td>
<td align="center" valign="top">80</td>
<td align="center" valign="top">80</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Strong evidence from financial reports and industry documents shows profit-driven strategy, with limited secondary involvement in health initiatives.</td>
</tr>
<tr>
<td align="left" valign="top">Smoking rates have decreased by 20% globally over the past decade</td>
<td align="center" valign="top">60</td>
<td align="center" valign="top">40</td>
<td align="left" valign="top">Moderate</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Global adult prevalence fell ~10&#x2013;15%; a 20% drop overstates the decline.</td>
</tr>
<tr>
<td align="left" valign="top">Over 8 million people die annually from tobacco-related illnesses</td>
<td align="center" valign="top">95</td>
<td align="center" valign="top">95</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">WHO and GBD consistently report ~7&#x2013;8 million annual deaths.</td>
</tr>
<tr>
<td align="left" valign="top">Youth smoking rates remain unchanged in strong-policy countries</td>
<td align="center" valign="top">40</td>
<td align="center" valign="top">30</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Unlikely</td>
<td align="left" valign="top">Most high-policy nations report declines; &#x201C;unchanged&#x201D; is misleading.</td>
</tr>
<tr>
<td align="left" valign="top">Smokers are three times more likely to develop severe COVID-19 symptoms</td>
<td align="center" valign="top">70</td>
<td align="center" valign="top">70</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Multiple meta-analyses find ~2&#x2013;3&#x202F;&#x00D7;&#x202F;increased risk of severe outcomes among smokers.</td>
</tr>
<tr>
<td align="left" valign="top">Tobacco farming employs over 1 million people worldwide</td>
<td align="center" valign="top">65</td>
<td align="center" valign="top">90</td>
<td align="left" valign="top">Likely</td>
<td align="left" valign="top">Highly Likely</td>
<td align="left" valign="top">Accurate according to FAO/ILO global workforce data</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>The automated framework demonstrated strong overall correlation with manual expert assessments (R<sup>2</sup>&#x202F;=&#x202F;0.89), with scores clustering along a slope of approximately 1.02, as visualized in <xref ref-type="fig" rid="fig7">Figure 3</xref>.</p>
</table-wrap-foot>
</table-wrap>
<fig position="float" id="fig7">
<label>Figure 3</label>
<caption>
<p>Scatter plot of manual vs. automated scores for all 20 claims, showing strong linear correlation with a 95% confidence band, demonstrating consistent agreement across the full scoring range.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g007.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Scatter plot comparing manual and automated credibility scores with a solid orange regression line and shaded confidence interval. A gray dashed line represents perfect agreement. Points are mostly clustered around the line. Manual scores are on the x-axis, automated scores on the y-axis, both ranging from zero to one hundred.</alt-text>
</graphic>
</fig>
</sec>
<sec id="sec27">
<label>4.3</label>
<title>Quantitative performance metrics</title>
<p><xref ref-type="table" rid="tab2">Table 2</xref> summarizes the key performance metrics comparing automated and manual assessments. The framework achieved a mean absolute error of 6.25 points on the 0&#x2013;100 scale, with a conservative bias showing a mean upward adjustment of +3.25 points (<italic>p</italic> =&#x202F;0.03). This bias was most evident in low-credibility classifications: the system did not classify any claims as &#x201C;Highly Unlikely&#x201D; despite experts assigning two claims to this category. The maximum absolute difference between automated and manual scoring was 25 points, with a standard deviation of differences of 8.9 points. While this conservative approach may reduce false flagging of legitimate health information, it indicates a need for calibration to improve identification of low-credibility claims.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Framework performance metrics.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Performance metric</th>
<th align="center" valign="top">Value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Mean Absolute Error (MAE)</td>
<td align="center" valign="top">6.25 points</td>
</tr>
<tr>
<td align="left" valign="top">Mean Signed Difference (Auto &#x2013; Manual)</td>
<td align="center" valign="top">+3.25 points (p&#x202F;=&#x202F;0.03)</td>
</tr>
<tr>
<td align="left" valign="top">Standard Deviation of Differences</td>
<td align="center" valign="top">8.9 points</td>
</tr>
<tr>
<td align="left" valign="top">Maximum Absolute Error</td>
<td align="center" valign="top">25 points</td>
</tr>
<tr>
<td align="left" valign="top">Exact-Level Agreement</td>
<td align="center" valign="top">70%</td>
</tr>
<tr>
<td align="left" valign="top">Adjacent-Level Agreement</td>
<td align="center" valign="top">95%</td>
</tr>
<tr>
<td align="left" valign="top">Weighted Cohen&#x2019;s Kappa (<italic>&#x03BA;</italic>)</td>
<td align="center" valign="top">0.68 (95% CI: 0.52&#x2013;0.84)</td>
</tr>
<tr>
<td align="left" valign="top">Processing Time per Claim</td>
<td align="center" valign="top">&#x003C; 7&#x202F;s</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Our agent-based architecture differs from traditional RAG systems in ways that require adapted evaluation approaches. While conventional RAG metrics like Recall@k and MRR evaluate raw document retrieval quality, and faithfulness metrics assess generation-document alignment, our Scientific Fact Verifier Agent performs evidence synthesis and structured assessment internally. This design choice prioritizes domain expertise integration over document-level retrieval optimization. To validate our evidence integration quality, we conducted supplementary analysis on a subset of 10 claims, manually reviewing the sources retrieved by our Serper API queries. We found 87% of retrieved sources were directly relevant to claim assessment, with 94% coming from our target authoritative domains (WHO, CDC, PubMed Central, Cochrane). More importantly, our end-to-end validation demonstrates that this evidence integration approach maintains fidelity to expert judgment (<italic>&#x03BA;</italic> =&#x202F;0.68), suggesting effective synthesis of retrieved information. Future implementations could benefit from component-level evaluation by logging intermediate retrieval results and implementing domain-specific relevance scoring. However, for this proof-of-concept focused on overall system validation, our primary metrics effectively capture whether evidence retrieval and synthesis support accurate credibility assessment.</p>
<p>The framework processed each claim in under 7&#x202F;s, representing substantial efficiency gains over manual evidence synthesis processes. This automation specifically targets the most time-intensive phases of fact-checking: systematic evidence retrieval across multiple authoritative databases, source credibility assessment, and preliminary evidence synthesis&#x2014;tasks that typically require 1&#x2013;2&#x202F;h of manual research per claim according to manual misinformation assessment research (<xref ref-type="bibr" rid="ref5">Bodaghi et al., 2024</xref>).</p>
</sec>
<sec id="sec28">
<label>4.4</label>
<title>Credibility-level performance analysis</title>
<p><xref ref-type="fig" rid="fig8">Figure 4</xref> presents a confusion matrix visualizing automated versus manual credibility level assignments across the five-level scale. This analysis reveals patterns in how the framework assigns credibility levels relative to expert judgment.</p>
<fig position="float" id="fig8">
<label>Figure 4</label>
<caption>
<p>Confusion matrix showing the distribution of automated vs. manual assessments across the five credibility levels (highly unlikely to highly likely), highlighting where agreement occurs and where level assignments differ.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g008.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix comparing manual versus automated classification. Rows represent manual assessment labeled from "Highly Unlikely" to "Highly Likely" and columns represent automated assessment. Notable counts include 2 in "Unlikely" for both methods, 4 in "Likely" for both, and 8 for "Highly Likely" in both. The gradient color scale indicates count intensity from 0 to 8.</alt-text>
</graphic>
</fig>
<p>The framework demonstrated varying performance across credibility categories. <xref ref-type="table" rid="tab3">Table 3</xref> details category-level recall rates, showing how often the automated framework correctly identified claims that experts assigned to each category.</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>Credibility-level recall rates by assessment level.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Credibility level</th>
<th align="center" valign="top">Manual count</th>
<th align="center" valign="top">Automated count</th>
<th align="center" valign="top">Correctly classified</th>
<th align="center" valign="top">Recall (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Highly unlikely (0&#x2013;20)</td>
<td align="center" valign="top">2</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">0</td>
<td align="char" valign="top" char=".">0.0</td>
</tr>
<tr>
<td align="left" valign="top">Unlikely (21&#x2013;40)</td>
<td align="center" valign="top">4</td>
<td align="center" valign="top">4</td>
<td align="center" valign="top">2</td>
<td align="char" valign="top" char=".">50.0</td>
</tr>
<tr>
<td align="left" valign="top">Moderate (41&#x2013;60)</td>
<td align="center" valign="top">1</td>
<td align="center" valign="top">2</td>
<td align="center" valign="top">0</td>
<td align="char" valign="top" char=".">0.0</td>
</tr>
<tr>
<td align="left" valign="top">Likely (61&#x2013;80)</td>
<td align="center" valign="top">4</td>
<td align="center" valign="top">6</td>
<td align="center" valign="top">4</td>
<td align="char" valign="top" char=".">100.0</td>
</tr>
<tr>
<td align="left" valign="top">Highly likely (81&#x2013;100)</td>
<td align="center" valign="top">9</td>
<td align="center" valign="top">8</td>
<td align="center" valign="top">8</td>
<td align="char" valign="top" char=".">88.9</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The framework exhibited varying performance across credibility categories, with conservative bias most evident in low-credibility classifications. While achieving excellent recall for claims manually categorized as &#x201C;Likely&#x201D; (100%) and &#x201C;Highly Likely&#x201D; (88.9%), performance was lower for &#x201C;Unlikely&#x201D; claims (50% recall), &#x201C;Moderate&#x201D; claims (0% recall), &#x2018;&#x2018;Highly Unlikely&#x201D; claims (0% recall). This pattern suggests the framework requires calibration to improve sensitivity for identifying problematic misinformation while maintaining its strong performance on well-supported claims.</p>
</sec>
<sec id="sec29">
<label>4.5</label>
<title>Analysis of classification discrepancies</title>
<p>Three claims exhibited score discrepancies of 20 points or more between automated and manual assessment:</p><list list-type="order">
<list-item>
<p>&#x201C;Tobacco farming employs over 1 million people worldwide&#x201D;</p>
<p>Automated: 65 (Likely), Manual: 90 (Highly Likely), Difference: &#x2212;25. This represented the largest discrepancy, where the framework underestimated credibility relative to expert assessment. The discrepancy stemmed from differing interpretations of magnitude&#x2014;while the claim is technically correct, it substantially understates global employment figures according to FAO/ILO data.</p>
</list-item>
<list-item>
<p>&#x201C;The harmful effects of vaping exceed those of smoking&#x201D;</p>
<p>Automated: 30 (Unlikely), Manual: 10 (Highly Unlikely), Difference: +20. The framework assigned a slightly higher credibility rating than experts, who emphasized the overwhelming consensus that combustible tobacco has greater harmful effects than vaping products.</p>
</list-item>
<list-item>
<p>&#x201C;Smoking rates have decreased by 20% globally over the past decade&#x201D;</p>
<p>Automated: 60 (Moderate), Manual: 40 (Unlikely), Difference: +20.</p>
<p>The framework&#x2019;s moderate score reflects both the directional accuracy of the declining trend and the magnitude difference from WHO&#x2019;s reported 10&#x2013;15% decrease, while experts weighted the precise numerical value more heavily in their assessment.</p>
</list-item>
</list>
<p><xref ref-type="fig" rid="fig9">Figure 5</xref> provides an alluvial (Sankey) visualization of how the automated framework&#x2019;s five-level credibility assignments compared to expert manual ratings across our 20-claim test set. On the left, the height of each bar corresponds to the number of claims in each manual category; on the right, the height reflects the automated framework&#x2019;s distribution. The connecting bands show exactly how many claims were classified identically (horizontal flows between matching levels) versus those shifted to different levels (cross-level flows). Notably, the majority of flows maintain horizontal paths between matching levels&#x2014;confirming a 70% exact match rate&#x2014;while misclassifications tend to lean toward higher credibility (e.g., some &#x201C;Highly Unlikely&#x201D; or &#x201C;Unlikely&#x201D; expert labels were mapped to &#x201C;Unlikely&#x201D; or &#x201C;Moderate&#x201D; by the framework).</p>
<fig position="float" id="fig9">
<label>Figure 5</label>
<caption>
<p>Sankey diagram of manual vs. automated credibility assignments for 20 tobacco-related claims. Band widths are proportional to the number of claims flowing from manual (left) category to the automated (right) category.</p>
</caption>
<graphic xlink:href="frai-08-1659861-g009.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Sankey diagram showing manual to automated credibility categorizations for twenty claims. Categories transition as follows: Highly Unlikely to Unlikely, Unlikely to Moderate, Likely to Likely, Moderate to Likely, and Highly Likely to Highly Likely.</alt-text>
</graphic>
</fig>
<p>This proof-of-concept study prioritizes technical innovation and architectural validation over large-scale statistical analysis. The intensive expert validation approach (20 claims, 40&#x2013;80 total expert hours) enables detailed assessment of framework reasoning quality while demonstrating practical deployment feasibility. Our multi-agent AI pipeline&#x2019;s expert-level performance across diverse tobacco claim types establishes the foundation for automated large-scale misinformation monitoring. Future implementations can leverage this validated framework for real-time processing of thousands of claims without additional expert review. Current validation focuses exclusively on tobacco-related claims. Generalization to other health misinformation domains requires domain-specific validation and potential framework modifications.</p>
</sec>
</sec>
<sec sec-type="discussion" id="sec30">
<label>5</label>
<title>Discussion</title>
<p>This proof-of-concept study demonstrates the technical feasibility of automated tobacco misinformation assessment using multi-agent AI pipeline. Our primary objective was to test whether an automated framework could achieve substantial agreement with expert evaluations while providing real-time processing capabilities. The framework achieved substantial inter-rater agreement (<italic>&#x03BA;</italic> =&#x202F;0.68) and dramatic processing efficiency gains, while revealing specific areas for improvement, particularly in low credibility claim identification. Our results reveal several notable strengths. The framework achieved a mean absolute error of just 6.25 points on a 0&#x2013;100 scale and a weighted Cohen&#x2019;s &#x03BA; of 0.68, indicating substantial inter-rater agreement. Exact-level category agreement stood at 70 percent, with adjacent-level agreement reaching 95 percent. The strong linear correlation (R<sup>2</sup> =&#x202F;0.89) between automated and manual scores, together with a slope near unity on the scatter plot, underscores the framework&#x2019;s overall calibration. When viewed alongside existing fact-checking frameworks, our pipeline offers both comparable accuracy and dramatically faster processing&#x2014;delivering results over a thousand times more rapidly than traditional manual review, which typically requires 2 to 4&#x202F;h per claim.</p>
<p>Nevertheless, our analysis also uncovered systematic biases and areas for refinement. The confusion matrix and recall rates illustrate a conservative upward bias: claims rated by experts as &#x201C;Highly Unlikely&#x201D; or &#x201C;Moderate&#x201D; were often classified one level higher by the framework. This tendency reduces false negatives among well-supported claims but risks over-crediting weak or contradictory assertions. Three outlier claims (tobacco farming employment, vaping harms versus smoking, and global smoking-rate decline) exhibited score discrepancies of &#x00B1;20&#x2013;25 points, revealing contexts where evidence weighting and source interpretation diverged from expert nuance. Addressing this bias will require probabilistic calibration techniques&#x2014;such as Platt scaling or isotonic regression&#x2014;to realign automated thresholds with human judgment (<xref ref-type="bibr" rid="ref22">Kington et al., 2021</xref>).</p>
<p>While our system employs GPT-4.1 as the core reasoning engine, we address transparency concerns through multiple methodological safeguards. Our retrieval-augmented approach grounds all assessments in explicitly cited authoritative sources (WHO, CDC, PubMed Central) rather than relying on model training data, ensuring evidence traceability. Our structured 5-point scoring framework with explicit criteria provides interpretable outputs that can be validated against expert judgment. Although GPT-4.1&#x2019;s internal processes remain proprietary, our framework&#x2019;s transparency lies in its evidence retrieval, source weighting, and structured assessment protocols&#x2014;components that are fully reproducible and auditable.</p>
<p>From a practical standpoint, the ability to flag high-impact misinformation in real time promises significant advantages for public health agencies. Rapid credibility assessments can underpin proactive risk communication, inform policy debates with evidence-rated insights, and streamline fact-checking workflows. However, deploying this framework responsibly demands careful attention to ethical considerations. Over-reliance on automated labels without transparency around uncertainty could mislead non-expert users. Designing user interfaces that display confidence intervals or &#x201C;soft&#x201D; score ranges will help practitioners interpret automated outputs appropriately.</p>
<p>Looking ahead, expanding our validation beyond the initial 20-claim dataset is critical. External testing on larger, multilingual corpora will assess generalizability across diverse tobacco narratives. User-centered evaluations with public health professionals to gage interpretability and trust would also provide additional insights. Finally, ongoing enhancements to the Health Evidence Assessor&#x2019;s weighting schema&#x2014;particularly for low-evidence categories&#x2014;will improve precision without compromising speed. These future efforts will ensure the pipeline remains adaptable to evolving misinformation patterns and continues to deliver actionable, trustworthy guidance.</p>
<p>Several limitations should be considered when interpreting our findings. The 20-claim validation set prioritizes intensive expert analysis over statistical breadth, with claims selected for diversity rather than systematic sampling. Our tobacco domain selection, while methodologically sound, likely favored system performance due to tobacco&#x2019;s stable evidence base. In rapidly evolving domains like emerging infectious diseases, our framework&#x2019;s evidence recency and consensus-based scoring may prove insufficient when scientific understanding shifts quickly, potentially leading to delayed detection of outdated claims or misclassification of evolving evidence. However, our modular architecture enables straightforward adaptation through dynamic temporal weighting and domain-specific consensus thresholds, which future implementations could calibrate based on evidence volatility metrics. Additionally, our reliance on institutional source credibility without individual study-level industry funding analysis represents a future enhancement opportunity. The observed conservative bias (+3.25 points), while potentially protective against over-flagging legitimate information, requires calibration to improve identification of low-credibility claims for optimal public health utility.</p>
<p>The credibility assessments presented reflect analysis of current scientific evidence from authoritative sources. As tobacco research evolves and new evidence emerges, these assessments may be updated to reflect advances in scientific understanding. While based on rigorous methodology and expert validation, these findings should be considered within the broader context of ongoing tobacco research and public health evidence.</p>
</sec>
<sec sec-type="conclusions" id="sec31">
<label>6</label>
<title>Conclusion</title>
<p>This proof-of-concept study demonstrates that multi-agent AI pipelines can achieve substantial agreement with expert tobacco misinformation assessments (MAE&#x202F;=&#x202F;6.25, <italic>&#x03BA;</italic> =&#x202F;0.68) while providing unprecedented processing speed improvements. The systematic conservative bias (+3.25 points) is predictable and manageable through calibration techniques. While the 20-claim validation set limits statistical generalizability, the intensive expert validation approach provides strong evidence of technical feasibility and expert-level reasoning quality. The modular framework, transparent scoring algorithm, and real-time evidence grounding offer a scalable foundation for public health misinformation monitoring. Critical next steps include: (1) expanding validation to 100&#x202F;+&#x202F;diverse claims across rapidly evolving health domains (emerging infectious diseases, policy updates), (2) implementing bias calibration techniques, (3) enhancing temporal weighting for time-sensitive evidence, and (4) developing responsible deployment protocols with appropriate uncertainty communication.</p>
<p>Our proof-of-concept validation establishes the technical foundation for responsible deployment in public health settings. Implementation would incorporate key safeguards including user interfaces that display confidence intervals and evidence source citations, systematic expert review of system outputs particularly for claims near decision boundaries, ongoing validation against expert assessments to detect performance drift, and clear guidelines defining appropriate use cases for preliminary screening versus situations requiring full expert analysis. The framework&#x2019;s strength lies in augmenting rather than replacing expert judgment, providing rapid evidence-grounded assessments that enhance human decision-making efficiency while maintaining oversight essential for public health applications.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="sec32">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found at: <ext-link xlink:href="https://doi.org/10.7910/DVN/ODKMNH" ext-link-type="uri">https://doi.org/10.7910/DVN/ODKMNH</ext-link>, Harvard Dataverse.</p>
</sec>
<sec sec-type="author-contributions" id="sec33">
<title>Author contributions</title>
<p>SE: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources, Software, Validation, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. JM: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. SB: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources, Validation, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. AG: Conceptualization, Data curation, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec sec-type="COI-statement" id="sec34">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="sec35">
<title>Generative AI statement</title>
<p>The author(s) declare that no Gen AI was used in the creation of this manuscript.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<sec sec-type="disclaimer" id="sec36">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="ref7"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Abroms</surname><given-names>L. C.</given-names></name> <name><surname>Yousefi</surname><given-names>A.</given-names></name> <name><surname>Wysota</surname><given-names>C. N.</given-names></name> <name><surname>Wu</surname><given-names>T. C.</given-names></name> <name><surname>Broniatowski</surname><given-names>D. A.</given-names></name></person-group> (<year>2025</year>). <article-title>Assessing the adherence of ChatGPT chatbots to public health guidelines for smoking cessation: content analysis</article-title>. <source>Journal of medical Internet research</source> <volume>27</volume>:<fpage>e66896</fpage>. doi: <pub-id pub-id-type="doi">10.2196/66896</pub-id>, <pub-id pub-id-type="pmid">25784350</pub-id></mixed-citation></ref>
<ref id="ref1"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Allem</surname><given-names>J. P.</given-names></name> <name><surname>Escobedo</surname><given-names>P.</given-names></name> <name><surname>Chu</surname><given-names>K. H.</given-names></name> <name><surname>Boley Cruz</surname><given-names>T.</given-names></name> <name><surname>Unger</surname><given-names>J. B.</given-names></name></person-group> (<year>2017</year>). <article-title>Images of little cigars and cigarillos on Instagram identified by the hashtag #swisher: thematic analysis</article-title>. <source>J. Med. Internet Res.</source> <volume>19</volume>:<fpage>e255</fpage>. doi: <pub-id pub-id-type="doi">10.2196/jmir.7634</pub-id>, <pub-id pub-id-type="pmid">28710057</pub-id></mixed-citation></ref>
<ref id="ref42"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Alpert</surname><given-names>J. M.</given-names></name> <name><surname>Chen</surname><given-names>H.</given-names></name> <name><surname>Riddell</surname><given-names>H.</given-names></name> <name><surname>Chung</surname><given-names>Y. J.</given-names></name> <name><surname>Mu</surname><given-names>Y. A.</given-names></name></person-group> (<year>2021</year>). <article-title>Vaping and Instagram: a content analysis of e-cigarette posts using the Content Appealing to Youth (CAY) Index</article-title>. <source>Substance Use &#x0026; Misuse,</source> <volume>56</volume>, <fpage>879</fpage>&#x2013;<lpage>887</lpage>. doi: <pub-id pub-id-type="doi">10.1080/10826084.2021.1899233</pub-id>, <pub-id pub-id-type="pmid">28710057</pub-id></mixed-citation></ref>
<ref id="ref2"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Amith</surname><given-names>M.</given-names></name> <name><surname>Cohen</surname><given-names>T.</given-names></name> <name><surname>Cunningham</surname><given-names>R.</given-names></name> <name><surname>Savas</surname><given-names>L. S.</given-names></name> <name><surname>Smith</surname><given-names>N.</given-names></name> <name><surname>Cuccaro</surname><given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Mining HPV vaccine knowledge structures of young adults from Reddit using distributional semantics and pathfinder networks</article-title>. <source>Cancer Control</source> <volume>27</volume>:<fpage>1073274819891442</fpage>. doi: <pub-id pub-id-type="doi">10.1177/1073274819891442</pub-id>, <pub-id pub-id-type="pmid">31912742</pub-id></mixed-citation></ref>
<ref id="ref3"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Apollonio</surname><given-names>D. E.</given-names></name> <name><surname>Malone</surname><given-names>R. E.</given-names></name></person-group> (<year>2009</year>). <article-title>Turning negative into positive: public health mass media campaigns and negative advertising</article-title>. <source>Health Educ. Res.</source> <volume>24</volume>, <fpage>483</fpage>&#x2013;<lpage>495</lpage>. doi: <pub-id pub-id-type="doi">10.1093/her/cyn046</pub-id>, <pub-id pub-id-type="pmid">18948569</pub-id></mixed-citation></ref>
<ref id="ref37"><mixed-citation publication-type="other"><person-group person-group-type="author"><name><surname>Baashirah</surname><given-names>R.</given-names></name></person-group> (<year>2024</year>). 0Zero-Shot Automated Detection of Fake News: An Innovative Approach (ZS-FND), in IEEE Access, vol. 12, pp. 182828&#x2013;40.</mixed-citation></ref>
<ref id="ref5"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bodaghi</surname><given-names>A.</given-names></name> <name><surname>Schmitt</surname><given-names>K. A.</given-names></name> <name><surname>Watine</surname><given-names>P.</given-names></name> <name><surname>Fung</surname><given-names>B. C. M.</given-names></name></person-group> (<year>2024</year>). <article-title>A literature review on detecting, verifying, and mitigating online misinformation</article-title>. <source>IEEE Trans. Comput. Soc. Syst.</source> <volume>11</volume>, <fpage>5119</fpage>&#x2013;<lpage>5145</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TCSS.2023.3289031</pub-id></mixed-citation></ref>
<ref id="ref6"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Brennan</surname><given-names>E.</given-names></name> <name><surname>Gibson</surname><given-names>L. A.</given-names></name> <name><surname>Kybert-Momjian</surname><given-names>A.</given-names></name> <name><surname>Liu</surname><given-names>J.</given-names></name> <name><surname>Hornik</surname><given-names>R. C.</given-names></name></person-group> (<year>2017</year>). <article-title>Promising themes for antismoking campaigns targeting youth and young adults</article-title>. <source>Tob. Regul. Sci.</source> <volume>3</volume>, <fpage>29</fpage>&#x2013;<lpage>46</lpage>. doi: <pub-id pub-id-type="doi">10.18001/TRS.3.1.4</pub-id>, <pub-id pub-id-type="pmid">28989949</pub-id></mixed-citation></ref>
<ref id="ref1001"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Brown-Johnson</surname><given-names>C. G.</given-names></name> <name><surname>Taylor</surname><given-names>K. L.</given-names></name></person-group> (<year>2025</year>). <article-title>Assessing the adherence of ChatGPT chatbots to public health guidance on smoking cessation</article-title>. <source>JMIR AI.</source> doi: <pub-id pub-id-type="doi">10.2196/54482</pub-id></mixed-citation></ref>
<ref id="ref8"><mixed-citation publication-type="book"><person-group person-group-type="author"><collab id="coll1">Centers for Disease Control and Prevention</collab></person-group> (<year>2023</year>). <source>Office on smoking and health: Tobacco industry tactics</source>. <publisher-loc>US</publisher-loc>: <publisher-name>CDC</publisher-name>.</mixed-citation></ref>
<ref id="ref9"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Chloros</surname><given-names>G. D.</given-names></name> <name><surname>Prodromidis</surname><given-names>A. D.</given-names></name> <name><surname>Giannoudis</surname><given-names>P. V.</given-names></name></person-group> (<year>2023</year>). <article-title>Has anything changed in evidence-based medicine?</article-title> <source>Injury</source> <volume>54</volume>, <fpage>S20</fpage>&#x2013;<lpage>S25</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.injury.2022.04.012</pub-id>, <pub-id pub-id-type="pmid">35525704</pub-id></mixed-citation></ref>
<ref id="ref10"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Dai</surname><given-names>E.</given-names></name> <name><surname>Sun</surname><given-names>Y.</given-names></name> <name><surname>Wang</surname><given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Ginger cannot cure cancer: Battling fake health news with a comprehensive data repository</article-title>. <source>Proceedings of the International AAAI Conference on Web and Social Media</source> (Vol. 14, pp. <volume>14</volume>, <fpage>853</fpage>&#x2013;<lpage>862</lpage>). doi: <pub-id pub-id-type="doi">10.1609/icwsm.v14i1.7350</pub-id></mixed-citation></ref>
<ref id="ref11"><mixed-citation publication-type="other"><person-group person-group-type="author"><name><surname>Elmitwalli</surname><given-names>S.</given-names></name></person-group> Validation of a Multi-Agent AI Pipeline for Automated Credibility Assessment of Tobacco Misinformation. (<year>2025</year>). Available online at: <ext-link xlink:href="https://github.com/sherifelmitwalli/misinformation-app.git" ext-link-type="uri">https://github.com/sherifelmitwalli/misinformation-app.git</ext-link> (Accessed on 2025 Oct 2)</mixed-citation></ref>
<ref id="ref12"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Erku</surname><given-names>D. A.</given-names></name> <name><surname>Bauld</surname><given-names>L.</given-names></name> <name><surname>Dawkins</surname><given-names>L.</given-names></name> <name><surname>Gartner</surname><given-names>C. E.</given-names></name> <name><surname>Steadman</surname><given-names>K. J.</given-names></name> <name><surname>Noar</surname><given-names>S. M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Does the content and source credibility of health and risk messages related to nicotine vaping products have an impact on harm perception and behavioural intentions? A systematic review</article-title>. <source>Addiction</source> <volume>116</volume>, <fpage>3290</fpage>&#x2013;<lpage>3303</lpage>. doi: <pub-id pub-id-type="doi">10.1111/add.15473</pub-id>, <pub-id pub-id-type="pmid">33751707</pub-id></mixed-citation></ref>
<ref id="ref13"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Eysenbach</surname><given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>How to fight an infodemic: the four pillars of infodemic management</article-title>. <source>J. Med. Internet Res.</source> <volume>22</volume>:<fpage>e21820</fpage>. doi: <pub-id pub-id-type="doi">10.2196/21820</pub-id>, <pub-id pub-id-type="pmid">32589589</pub-id></mixed-citation></ref>
<ref id="ref14"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gannon</surname><given-names>J.</given-names></name> <name><surname>Bach</surname><given-names>K.</given-names></name> <name><surname>Cattaruzza</surname><given-names>M. S.</given-names></name> <name><surname>Bar-Zeev</surname><given-names>Y.</given-names></name> <name><surname>Forberger</surname><given-names>S.</given-names></name> <name><surname>Kilibarda</surname><given-names>B.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Big tobacco's dirty tricks: seven key tactics of the tobacco industry</article-title>. <source>Tobacco Prevention &#x0026; Cessation.</source> <volume>9</volume>:<fpage>39</fpage>. doi: <pub-id pub-id-type="doi">10.18332/tpc/162462</pub-id></mixed-citation></ref>
<ref id="ref15"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ghenai</surname><given-names>A.</given-names></name> <name><surname>Mejova</surname><given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>Fake cures: user-centric modeling of health misinformation in social media</article-title>. <source>Proc. ACM Hum.-Comput. Interact.</source> <volume>2</volume>, <fpage>1</fpage>&#x2013;<lpage>20</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3274327</pub-id></mixed-citation></ref>
<ref id="ref50"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gargari</surname><given-names>O. K.</given-names></name> <name><surname>Habibi</surname><given-names>G.</given-names></name></person-group> (<year>2025</year>). <article-title>Enhancing medical AI with retrieval-augmented generation: A mini narrative review</article-title>. <source>Digital health,</source> <volume>11</volume>:<fpage>20552076251337177</fpage>. doi: <pub-id pub-id-type="doi">10.1177/20552076251337177</pub-id></mixed-citation></ref>
<ref id="ref16"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gilmore</surname><given-names>A. B.</given-names></name> <name><surname>Fooks</surname><given-names>G.</given-names></name> <name><surname>Drope</surname><given-names>J.</given-names></name> <name><surname>Bialous</surname><given-names>S. A.</given-names></name> <name><surname>Jackson</surname><given-names>R. R.</given-names></name></person-group> (<year>2015</year>). <article-title>Exposing and addressing tobacco industry conduct in low-income and middle-income countries</article-title>. <source>Lancet</source> <volume>385</volume>, <fpage>1029</fpage>&#x2013;<lpage>1043</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0140-6736(15)60312-9</pub-id>, <pub-id pub-id-type="pmid">25784350</pub-id></mixed-citation></ref>
<ref id="ref17"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Guyatt</surname><given-names>G.</given-names></name> <name><surname>Agoritsas</surname><given-names>T.</given-names></name> <name><surname>Brignardello-Petersen</surname><given-names>R.</given-names></name> <name><surname>Prasad</surname><given-names>M.</given-names></name> <name><surname>Hultcrantz</surname><given-names>M.</given-names></name> <name><surname>Murad</surname><given-names>M. H.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Core GRADE 1: overview of the core GRADE approach</article-title>. <source>BMJ</source> <volume>389</volume>:<fpage>e081903</fpage>. doi: <pub-id pub-id-type="doi">10.1136/bmj-2024-081903</pub-id>, <pub-id pub-id-type="pmid">40262844</pub-id></mixed-citation></ref>
<ref id="ref18"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hendlin</surname><given-names>Y.</given-names></name> <name><surname>Vora</surname><given-names>M.</given-names></name> <name><surname>Elias</surname><given-names>J.</given-names></name> <name><surname>Ling</surname><given-names>P. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Financial conflicts of interest and stance on tobacco harm reduction: a systematic review</article-title>. <source>Am. J. Public Health</source> <volume>109</volume>, <fpage>e1</fpage>&#x2013;<lpage>e8</lpage>. doi: <pub-id pub-id-type="doi">10.2105/AJPH.2019.305106</pub-id>, <pub-id pub-id-type="pmid">31095414</pub-id></mixed-citation></ref>
<ref id="ref19"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Isern</surname><given-names>D.</given-names></name> <name><surname>Moreno</surname><given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>A systematic literature review of agents applied in healthcare</article-title>. <source>J. Med. Syst.</source> <volume>40</volume>:<fpage>43</fpage>. doi: <pub-id pub-id-type="doi">10.1007/s10916-015-0376-2</pub-id>, <pub-id pub-id-type="pmid">26590981</pub-id></mixed-citation></ref>
<ref id="ref21"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Jackler</surname><given-names>R. K.</given-names></name> <name><surname>Li</surname><given-names>V. Y.</given-names></name> <name><surname>Cardiff</surname><given-names>R. A.</given-names></name> <name><surname>Ramamurthi</surname><given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Promotion of tobacco products on Facebook: policy versus practice</article-title>. <source>Tob. Control.</source> <volume>28</volume>, <fpage>67</fpage>&#x2013;<lpage>73</lpage>. doi: <pub-id pub-id-type="doi">10.1136/tobaccocontrol-2017-054175</pub-id>, <pub-id pub-id-type="pmid">29622602</pub-id></mixed-citation></ref>
<ref id="ref22"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kington</surname><given-names>R. S.</given-names></name> <name><surname>Arnesen</surname><given-names>S.</given-names></name> <name><surname>Chou</surname><given-names>W. Y. S.</given-names></name> <name><surname>Curry</surname><given-names>S. J.</given-names></name> <name><surname>Lazer</surname><given-names>D.</given-names></name> <name><surname>Horwitz</surname><given-names>L. I.</given-names></name></person-group> (<year>2021</year>). <article-title>Identifying credible sources of health information in social media: principles and attributes</article-title>. <source>NAM Perspectives</source>, <fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi: <pub-id pub-id-type="doi">10.31478/202107a</pub-id></mixed-citation></ref>
<ref id="ref23"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kong</surname><given-names>G.</given-names></name> <name><surname>Ouellette</surname><given-names>R. R.</given-names></name> <name><surname>Murthy</surname><given-names>D.</given-names></name></person-group> (<year>2024</year>). <article-title>Generative artificial intelligence and social media: insights for tobacco control</article-title>. <source>Tob. Control.</source>:<fpage>tc-2024-058813</fpage>. doi: <pub-id pub-id-type="doi">10.1136/tc-2024-058813</pub-id>, <pub-id pub-id-type="pmid">39643443</pub-id></mixed-citation></ref>
<ref id="ref24"><mixed-citation publication-type="other"><person-group person-group-type="author"><collab id="coll2">Langtrace</collab></person-group>. (<year>2024</year>). Langtrace AI - Open Source Observability for LLMs. Available online at: <ext-link xlink:href="https://www.langtrace.ai/" ext-link-type="uri">https://www.langtrace.ai/</ext-link> (Accessed: 01 October 2025).</mixed-citation></ref>
<ref id="ref25"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Leone</surname><given-names>F. T.</given-names></name> <name><surname>Carlsen</surname><given-names>K. H.</given-names></name> <name><surname>Chooljian</surname><given-names>D.</given-names></name> <name><surname>Crotty Alexander</surname><given-names>L. E.</given-names></name> <name><surname>Detterbeck</surname><given-names>F. C.</given-names></name> <name><surname>Eakin</surname><given-names>M. N.</given-names></name> <etal/></person-group> (<year>2018</year>). <article-title>Recommendations for the appropriate structure, communication, and investigation of tobacco harm reduction claims. An official American Thoracic Society policy statement</article-title>. <source>American journal of respiratory and critical care medicine,</source> <volume>198</volume>, <fpage>e90</fpage>&#x2013;<lpage>e105</lpage>. doi: <pub-id pub-id-type="doi">10.1164/rccm.201808-1443ST</pub-id></mixed-citation></ref>
<ref id="ref27"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname><given-names>W.</given-names></name> <name><surname>Chen</surname><given-names>J.</given-names></name> <name><surname>Ji</surname><given-names>K.</given-names></name> <name><surname>Zhou</surname><given-names>L.</given-names></name> <name><surname>Chen</surname><given-names>W.</given-names></name> <name><surname>Wang</surname><given-names>B.</given-names></name></person-group> (<year>2025</year>). <article-title>Rag-instruct: Boosting llms with diverse retrieval-augmented instructions.</article-title> In <source>Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing.</source> (pp. 3865&#x2013;3888). doi: <pub-id pub-id-type="doi">10.18653/v1/2025.emnlp-main.192</pub-id></mixed-citation></ref>
<ref id="ref4"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Luk</surname><given-names>T. T.</given-names></name> <name><surname>Zhao</surname><given-names>S.</given-names></name> <name><surname>Weng</surname><given-names>X.</given-names></name> <name><surname>Wong</surname><given-names>J. Y. H.</given-names></name> <name><surname>Wu</surname><given-names>Y. S.</given-names></name> <name><surname>Ho</surname><given-names>S. Y.</given-names></name> <name><surname>Wang</surname><given-names>M. P.</given-names></name></person-group> (<year>2021</year>). <article-title>Exposure to health misinformation about COVID-19 and increased tobacco and alcohol use: a population-based survey in Hong Kong</article-title>. <source>Tobacco control</source>, <volume>30</volume>, <fpage>696</fpage>&#x2013;<lpage>699</lpage>. doi: <pub-id pub-id-type="doi">10.1136/tobaccocontrol-2020-055960</pub-id></mixed-citation></ref>
<ref id="ref28"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Luo</surname><given-names>Y.</given-names></name> <name><surname>Thompson</surname><given-names>W. K.</given-names></name> <name><surname>Herr</surname><given-names>T. M.</given-names></name> <name><surname>Zeng</surname><given-names>Z.</given-names></name> <name><surname>Berendsen</surname><given-names>M. A.</given-names></name> <name><surname>Jonnalagadda</surname><given-names>S. R.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Natural language processing for EHR-based pharmacovigilance: a structured review</article-title>. <source>Drug Saf.</source> <volume>40</volume>, <fpage>1075</fpage>&#x2013;<lpage>1089</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s40264-017-0558-6</pub-id>, <pub-id pub-id-type="pmid">28643174</pub-id></mixed-citation></ref>
<ref id="ref29"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Marshall</surname><given-names>I. J.</given-names></name> <name><surname>Kuiper</surname><given-names>J.</given-names></name> <name><surname>Wallace</surname><given-names>B. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Robotreviewer: evaluation of a system for automatically assessing bias in clinical trials</article-title>. <source>J. Am. Med. Inform. Assoc.</source> <volume>23</volume>, <fpage>193</fpage>&#x2013;<lpage>201</lpage>. doi: <pub-id pub-id-type="doi">10.1093/jamia/ocv044</pub-id>, <pub-id pub-id-type="pmid">26104742</pub-id></mixed-citation></ref>
<ref id="ref30"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Nyhan</surname><given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Why the backfire effect does not explain the durability of political misperceptions</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>118</volume>:<fpage>e1912440117</fpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.1912440117</pub-id>, <pub-id pub-id-type="pmid">33837144</pub-id></mixed-citation></ref>
<ref id="ref31"><mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Oreskes</surname><given-names>N.</given-names></name> <name><surname>Conway</surname><given-names>E. M.</given-names></name></person-group> (<year>2011</year>). <source>Merchants of doubt: How a handful of scientists obscured the truth on issues from tobacco smoke to global warming</source>. New York: <publisher-name>Bloomsbury Publishing</publisher-name>.</mixed-citation></ref>
<ref id="ref32"><mixed-citation publication-type="other"><person-group person-group-type="author"><name><surname>P&#x00E9;rez-Rosas</surname><given-names>V.</given-names></name> <name><surname>Kleinberg</surname><given-names>B.</given-names></name> <name><surname>Lefevre</surname><given-names>A.</given-names></name> <name><surname>Mihalcea</surname><given-names>R.</given-names></name></person-group> (<year>2018</year>). &#x201C;Automatic detection of fake news.&#x201D; <italic>Proceedings of the 27th International Conference on Computational Linguistics</italic>, <fpage>3391</fpage>&#x2013;<lpage>3401</lpage>.</mixed-citation></ref>
<ref id="ref33"><mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Proctor</surname><given-names>R. N.</given-names></name></person-group> (<year>2012</year>). <source>Golden holocaust: Origins of the cigarette catastrophe and the case for abolition</source>. Berkeley, CA: <publisher-name>University of California Press</publisher-name>.</mixed-citation></ref>
<ref id="ref34"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Reitsma</surname><given-names>M. B.</given-names></name> <name><surname>Kendrick</surname><given-names>P. J.</given-names></name> <name><surname>Ababneh</surname><given-names>E.</given-names></name> <name><surname>Abbafati</surname><given-names>C.</given-names></name> <name><surname>Abbasi-Kangevari</surname><given-names>M.</given-names></name> <name><surname>Abdoli</surname><given-names>A.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Spatial, temporal, and demographic patterns in prevalence of smoking tobacco use and attributable disease burden in 204 countries and territories, 1990&#x2013;2019: a systematic analysis from the global burden of disease study 2019</article-title>. <source>Lancet</source> <volume>397</volume>, <fpage>2337</fpage>&#x2013;<lpage>2360</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0140-6736(21)01169-7</pub-id>, <pub-id pub-id-type="pmid">34051883</pub-id></mixed-citation></ref>
<ref id="ref35"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sarrouti</surname><given-names>M.</given-names></name> <name><surname>El Alaoui</surname><given-names>S. O.</given-names></name></person-group> (<year>2017</year>). <article-title>A passage retrieval method based on probabilistic information retrieval model and UMLS concepts in biomedical question answering</article-title>. <source>Journal of biomedical informatics.</source> <volume>68</volume>, <fpage>96</fpage>&#x2013;<lpage>103</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jbi.2017.03.001</pub-id></mixed-citation></ref>
<ref id="ref36"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schmidt</surname><given-names>A. L.</given-names></name> <name><surname>Zollo</surname><given-names>F.</given-names></name> <name><surname>Scala</surname><given-names>A.</given-names></name> <name><surname>Betsch</surname><given-names>C.</given-names></name> <name><surname>Quattrociocchi</surname><given-names>W.</given-names></name></person-group> (<year>2018</year>). <article-title>Polarization of the vaccination debate on Facebook</article-title>. <source>Vaccine</source> <volume>36</volume>, <fpage>3606</fpage>&#x2013;<lpage>3612</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.vaccine.2018.05.040</pub-id>, <pub-id pub-id-type="pmid">29773322</pub-id></mixed-citation></ref>
<ref id="ref38"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sharp</surname><given-names>K.</given-names></name> <name><surname>Ouellette</surname><given-names>R. R.</given-names></name> <name><surname>Singh</surname><given-names>R.</given-names></name></person-group> (<year>2025</year>). <article-title>Generative artificial intelligence and machine learning methods to screen social media content</article-title>. <source>PeerJ Comput. Sci.</source> <volume>11</volume>:<fpage>e2710</fpage>. doi: <pub-id pub-id-type="doi">10.7717/peerj-cs.2710</pub-id>, <pub-id pub-id-type="pmid">40134877</pub-id></mixed-citation></ref>
<ref id="ref39"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Silva Filho</surname><given-names>T.</given-names></name> <name><surname>Song</surname><given-names>H.</given-names></name> <name><surname>Perello-Nieto</surname><given-names>M.</given-names></name> <name><surname>Santos-Rodriguez</surname><given-names>R.</given-names></name> <name><surname>Kull</surname><given-names>M.</given-names></name> <name><surname>Flach</surname><given-names>P.</given-names></name></person-group> (<year>2023</year>). <article-title>Classifier calibration: a survey on how to assess and improve predicted class probabilities</article-title>. <source>Mach. Learn.</source> <volume>112</volume>, <fpage>3211</fpage>&#x2013;<lpage>3260</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10994-023-06336-7</pub-id></mixed-citation></ref>
<ref id="ref40"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sylvia Chou</surname><given-names>W. Y.</given-names></name> <name><surname>Gaysynsky</surname><given-names>A.</given-names></name> <name><surname>Cappella</surname><given-names>J. N.</given-names></name></person-group> (<year>2020</year>). <article-title>Where we go from here: health misinformation on social media</article-title>. <source>Am. J. Public Health</source> <volume>110</volume>, <fpage>S273</fpage>&#x2013;<lpage>S275</lpage>. doi: <pub-id pub-id-type="doi">10.2105/AJPH.2020.305905</pub-id>, <pub-id pub-id-type="pmid">33001722</pub-id></mixed-citation></ref>
<ref id="ref20"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tan</surname><given-names>A. S.</given-names></name> <name><surname>Bigman</surname><given-names>C. A.</given-names></name></person-group> (<year>2020</year>). <article-title>Misinformation about commercial tobacco products on social media&#x2014;implications and research opportunities for reducing tobacco-related health disparities</article-title>. <source>American journal of public health,</source> <volume>110</volume>(S3), <fpage>S281</fpage>&#x2013;<lpage>S283</lpage>. doi: <pub-id pub-id-type="doi">10.2105/AJPH.2020.305910</pub-id></mixed-citation></ref>
<ref id="ref41"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tan</surname><given-names>A. S.</given-names></name> <name><surname>Bigman</surname><given-names>C. A.</given-names></name> <name><surname>Sanders-Jackson</surname><given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Sociodemographic correlates of self-reported exposure to e-cigarette communications and its association with public support for smoke-free and vape-free policies: results from a national survey of US adults</article-title>. <source>Tob. Control.</source> <volume>24</volume>, <fpage>574</fpage>&#x2013;<lpage>581</lpage>. doi: <pub-id pub-id-type="doi">10.1136/tobaccocontrol-2014-051685</pub-id>, <pub-id pub-id-type="pmid">25015372</pub-id></mixed-citation></ref>
<ref id="ref43"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Vraga</surname><given-names>E. K.</given-names></name> <name><surname>Bode</surname><given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>Defining misinformation and understanding its bounded nature: using expertise and evidence for describing misinformation</article-title>. <source>Polit. Commun.</source> <volume>37</volume>, <fpage>136</fpage>&#x2013;<lpage>144</lpage>. doi: <pub-id pub-id-type="doi">10.1080/10584609.2020.1716500</pub-id></mixed-citation></ref>
<ref id="ref44"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wallace</surname><given-names>B. C.</given-names></name> <name><surname>Trikalinos</surname><given-names>T. A.</given-names></name> <name><surname>Lau</surname><given-names>J.</given-names></name> <name><surname>Brodley</surname><given-names>C.</given-names></name> <name><surname>Schmid</surname><given-names>C. H.</given-names></name></person-group> (<year>2010</year>). <article-title>Semi-automated screening of biomedical citations for systematic reviews</article-title>. <source>BMC Bioinformatics</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi: <pub-id pub-id-type="doi">10.1186/1471-2105-11-55</pub-id>, <pub-id pub-id-type="pmid">20102628</pub-id></mixed-citation></ref>
<ref id="ref45"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname><given-names>Y.</given-names></name> <name><surname>McKee</surname><given-names>M.</given-names></name> <name><surname>Torbica</surname><given-names>A.</given-names></name> <name><surname>Stuckler</surname><given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Systematic literature review on the spread of health-related misinformation on social media</article-title>. <source>Soc. Sci. Med.</source> <volume>240</volume>:<fpage>112552</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.socscimed.2019.112552</pub-id>, <pub-id pub-id-type="pmid">31561111</pub-id></mixed-citation></ref>
<ref id="ref46"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wimmer</surname><given-names>H.</given-names></name> <name><surname>Yoon</surname><given-names>V. Y.</given-names></name> <name><surname>Sugumaran</surname><given-names>V.</given-names></name></person-group> (<year>2016</year>). <article-title>A multi-agent system to support evidence based medicine and clinical decision making via data sharing and data privacy</article-title>. <source>Decis. Support. Syst.</source> <volume>88</volume>, <fpage>51</fpage>&#x2013;<lpage>66</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.dss.2016.05.008</pub-id></mixed-citation></ref>
<ref id="ref47"><mixed-citation publication-type="book"><person-group person-group-type="author"><collab id="coll3">World Health Organization</collab></person-group> (<year>2022</year>). <source>Tobacco industry tactics</source>. <publisher-loc>Geneva</publisher-loc>: <publisher-name>WHO Tobacco Free Initiative</publisher-name>.</mixed-citation></ref>
<ref id="ref48"><mixed-citation publication-type="other"><person-group person-group-type="author"><collab id="coll4">World Health Organization</collab></person-group>. (<year>2025</year>) Tobacco. Available online at: <ext-link xlink:href="https://www.who.int/news-room/fact-sheets/detail/tobacco" ext-link-type="uri">https://www.who.int/news-room/fact-sheets/detail/tobacco</ext-link> (Accessed: 11 June 2025).</mixed-citation></ref>
<ref id="ref49"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Yuan</surname><given-names>B.</given-names></name> <name><surname>Herbert</surname><given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Context-aware hybrid reasoning framework for pervasive healthcare</article-title>. <source>Pers. Ubiquit. Comput.</source> <volume>18</volume>, <fpage>865</fpage>&#x2013;<lpage>881</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s00779-013-0696-5</pub-id></mixed-citation></ref>
<ref id="ref26"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname><given-names>B.</given-names></name> <name><surname>Ma</surname><given-names>H.</given-names></name> <name><surname>Li</surname><given-names>D.</given-names></name> <name><surname>Ding</surname><given-names>J.</given-names></name> <name><surname>Wang</surname><given-names>J.</given-names></name> <name><surname>Xu</surname><given-names>B.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Efficient Tuning of Large Language Models for Knowledge-Grounded Dialogue Generation</article-title>. <source>Transactions of the Association for Computational Linguistics,</source> <volume>13</volume>, <fpage>1007</fpage>&#x2013;<lpage>1031</lpage>. doi: <pub-id pub-id-type="doi">10.1162/TACL.a.17</pub-id>, <pub-id pub-id-type="pmid">34051883</pub-id></mixed-citation></ref>
<ref id="ref51"><mixed-citation publication-type="other"><person-group person-group-type="author"><name><surname>Zhang</surname><given-names>Y.</given-names></name> <name><surname>Marshall</surname><given-names>I. J.</given-names></name> <name><surname>Wallace</surname><given-names>B. C.</given-names></name></person-group> (<year>2016</year>). &#x201C;Rationale-augmented convolutional neural networks for text classification.&#x201D; <italic>Proceedings of the Conference on Empirical Methods in Natural Language Processing</italic>, <fpage>795</fpage>&#x2013;<lpage>804</lpage>.</mixed-citation></ref>
<ref id="ref52"><mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname><given-names>X.</given-names></name> <name><surname>Zafarani</surname><given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>A survey of fake news: fundamental theories, detection methods, and opportunities</article-title>. <source>ACM Comput. Surv.</source> <volume>53</volume>, <fpage>1</fpage>&#x2013;<lpage>40</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3395046</pub-id></mixed-citation></ref>
</ref-list>
<fn-group>
<fn fn-type="custom" custom-type="edited-by" id="fn0001">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/936148/overview">Noorbakhsh Amiri Golilarz</ext-link>, The University of Alabama, United States</p>
</fn>
<fn fn-type="custom" custom-type="reviewed-by" id="fn0002">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2770822/overview">Yuan Guo</ext-link>, Bryant University, United States</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2996213/overview">Muhammad Sajid</ext-link>, Air University, Pakistan</p>
</fn>
</fn-group>
<app-group>
<app id="app1">
<title>Appendix</title>
<p>Glossary of key terms.</p>
<p>Adjacent-Level Agreement: A performance metric measuring the percentage of classifications that fall within one level of the reference classification on a multi-level scale (e.g., a claim rated &#x201C;Likely&#x201D; by the system when experts rated it &#x201C;Highly Likely&#x201D;).</p>
<p>Credibility Score: A numerical value (0&#x2013;100) assigned to a claim based on weighted assessment across five dimensions: evidence quality (40%), scientific consensus (25%), evidence consistency (15%), evidence recency (10%), and scientific plausibility (10%).</p>
<p>Evidence Quality Score (Eq): A component score (0&#x2013;100) evaluating the methodological rigor of supporting studies, with higher weights for systematic reviews, RCTs, and meta-analyses.</p>
<p>Five-Level Classification System: The categorical framework mapping credibility scores to interpretable levels: Highly Unlikely (0&#x2013;20), Unlikely (21&#x2013;40), Moderate (41&#x2013;60), Likely (61&#x2013;80), and Highly Likely (81&#x2013;100).</p>
<p>Mean Absolute Error (MAE): The average absolute difference between automated and manual numerical scores across all evaluated claims.</p>
<p>Multi-Agent AI Pipeline: A computational framework consisting of multiple specialized AI agents that work sequentially to process information, where each agent has distinct responsibilities and passes structured outputs to subsequent agents.</p>
<p>Proof-of-Concept: A preliminary implementation demonstrating technical feasibility and core functionality, intended to validate an approach before full-scale development and deployment.</p>
<p>Retrieval-Augmented Generation (RAG): A technique combining large language models with real-time information retrieval from external sources to ground responses in current, authoritative evidence rather than relying solely on training data.</p>
<p>Scientific Consensus Score (Sc): A component score (0&#x2013;100) measuring the degree of agreement across authoritative sources (WHO, CDC, peer-reviewed literature) regarding a specific claim.</p>
<p>Scientific Plausibility (Sp): A component score (0&#x2013;100) assessing whether a claim aligns with established biological mechanisms and scientific principles in tobacco and health research.</p>
<p>Serper API: A web search application programming interface providing structured access to authoritative health databases and scientific literature for real-time evidence retrieval.</p>
<p>Weighted Cohen&#x2019;s Kappa (<italic>&#x03BA;</italic>): A statistical measure of inter-rater agreement that accounts for partial agreement between classifications, with values interpreted as: slight (0&#x2013;0.20), fair (0.21&#x2013;0.40), moderate (0.41&#x2013;0.60), substantial (0.61&#x2013;0.80), and almost perfect (0.81&#x2013;1.00).</p>
<p>Zero-Shot Learning: The ability of a language model to perform tasks without task-specific training examples, relying instead on general language understanding and structured prompting.</p>
</app>
</app-group>
</back>
</article>