<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3-mathml3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="review-article" dtd-version="1.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Drug Saf. Regul.</journal-id>
<journal-title-group>
<journal-title>Frontiers in Drug Safety and Regulation</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Drug Saf. Regul.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2674-0869</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1599217</article-id>
<article-id pub-id-type="doi">10.3389/fdsfr.2025.1599217</article-id>
<article-version article-version-type="Version of Record" vocab="NISO-RP-8-2008"/>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Mini Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Tokenization techniques for privacy-preserving healthcare data: tokenization nuts and bolts</article-title>
<alt-title alt-title-type="left-running-head">Cook</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fdsfr.2025.1599217">10.3389/fdsfr.2025.1599217</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Cook</surname>
<given-names>Camille V.</given-names>
</name>
<xref ref-type="aff" rid="aff1"/>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2860319"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; original draft" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-original-draft/">Writing &#x2013; original draft</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="resources" vocab-term-identifier="https://credit.niso.org/contributor-roles/resources/">Resources</role>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>LexisNexis Risk Solutions</institution>, <city>Alpharetta</city>, <state>GA</state>, <country country="US">United States</country>
</aff>
<author-notes>
<corresp id="c001">
<label>&#x2a;</label>Correspondence: Camille V. Cook, <email xlink:href="cambam85@gmail.com">cambam85@gmail.com</email>
</corresp>
</author-notes>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2025-12-18">
<day>18</day>
<month>12</month>
<year>2025</year>
</pub-date>
<pub-date publication-format="electronic" date-type="collection">
<year>2025</year>
</pub-date>
<volume>5</volume>
<elocation-id>1599217</elocation-id>
<history>
<date date-type="received">
<day>24</day>
<month>03</month>
<year>2025</year>
</date>
<date date-type="rev-recd">
<day>04</day>
<month>08</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>08</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Cook.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Cook</copyright-holder>
<license>
<ali:license_ref start_date="2025-12-18">https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</license-p>
</license>
</permissions>
<abstract>
<p>Tokenization is a crucial technology for ensuring the security and privacy of patient data in clinical research, pharmacovigilance, and drug safety monitoring. As healthcare increasingly integrates diverse data sources-ranging from clinical records to non-clinical data such as social determinants of health (SDOH)-it is essential to protect sensitive patient information while improving data quality and analysis (National Institutes of Health, 2006). This article emphasizes tokenization&#x2019;s critical role in safeguarding privacy, particularly in pharmacovigilance activities including safety monitoring, risk assessment, and post-market surveillance. Beyond security, tokenization enriches research datasets by enabling integration of external information, thereby enhancing the rigor and reliability of pharmacovigilance outcomes. With effective tokenization, researchers can better protect patients while gaining deeper insights into clinical and pharmacological research (Cruz et al., 2024). Recent global applications validate tokenization as a foundational privacy-preserving technology in pharmacovigilance. An applied example from a psoriasis clinical trial demonstrated referential tokenization&#x2019;s capacity to securely link electronic health records (EHRs) and claims data across systems with greater than 99% linkage precision while maintaining privacy standards (D&#x27;Andrea et al., 2024). These capabilities align with emerging international frameworks, including the European Health Data Space (2025), reinforcing tokenization&#x2019;s value in generating regulatory-grade evidence for pharmacovigilance across national and multinational research environments. In many jurisdictions, tokenization that meets de-identification or pseudonymization standards may not require individual patient consent, though this varies based on data sensitivity, jurisdictional law, and the study&#x2019;s intent (Office for Civil Rights, 2023; EDPB, 2021).</p>
</abstract>
<kwd-group>
<kwd>tokenization</kwd>
<kwd>real-world data (RWD)</kwd>
<kwd>linking</kwd>
<kwd>privacy</kwd>
<kwd>pharmocovigilance</kwd>
<kwd>health economics</kwd>
</kwd-group>
<funding-group>
<funding-statement>The author(s) declare that no financial support was received for the research and/or publication of this article.</funding-statement>
</funding-group>
<counts>
<fig-count count="0"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="18"/>
<page-count count="6"/>
</counts>
<custom-meta-group>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Advanced Methods in Pharmacovigilance and Pharmacoepidemiology</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>The modern healthcare landscape increasingly relies on integrating heterogeneous datasets to advance clinical research and pharmacovigilance activities. Patient-level data&#x2014;encompassing clinical observations, medication histories, laboratory results, and socio-environmental contexts&#x2014;are critical for real-world evidence (RWE), personalized medicine, regulatory decision-making, and comprehensive pharmacovigilance. However, the sensitivity of personally identifiable information (PII) necessitates rigorous privacy protection. Tokenization addresses this imperative by substituting sensitive identifiers with pseudonymous tokens. While referential and salted tokenization methods are generally irreversible, deterministic or non-salted tokenization may permit re-identification under specific conditions (e.g., access to keys or mapping tables). Thus, tokenization&#x2019;s irreversibility depends on implementation context and security controls (<xref ref-type="bibr" rid="B2">Bernstam et al., 2023</xref>; <xref ref-type="bibr" rid="B18">Yue, 2022</xref>).</p>
<p>Decentralized clinical trials, multi-institutional observational studies, and regulatory-grade safety surveillance increasingly rely on robust tokenization techniques to enable privacy-compliant linkage across diverse data sources&#x2014;including electronic health records (EHRs), insurance claims, mortality registries, and emerging domains such as social determinants of health (SDOH) and genomics&#x2014;supporting high-quality evidence generation for regulatory decision-making. This integration fosters comprehensive longitudinal patient profiles essential for pharmacovigilance, enabling detection of adverse drug reactions, safety signal validation, and the study of treatment pathways and health disparities (<xref ref-type="bibr" rid="B3">Cruz et al., 2024</xref>). Subsequent sections examine tokenization methodologies, workflows, and regulatory considerations fundamental to privacy-preserving data linkage for pharmacovigilance.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Methods of tokenization</title>
<sec id="s2-1">
<label>2.1</label>
<title>Tokenization methodologies and their characteristics</title>
<p>Tokenization replaces direct identifiers&#x2014;such as names, dates of birth, and social security numbers&#x2014;with tokens that allow cross-dataset linkage without exposing raw PII. Three primary tokenization paradigms dominate:<list list-type="order">
<list-item>
<p>Deterministic tokenization applies a consistent algorithm producing the same token for identical inputs, facilitating reproducible linkage across data sources. This method is highly effective for structured datasets with stable identifiers, enabling longitudinal patient tracking and outcome analyses vital for pharmacovigilance. However, deterministic tokens may be vulnerable to re-identification if underlying algorithms are reverse-engineered or keys compromised (<xref ref-type="bibr" rid="B2">Bernstam et al., 2023</xref>).</p>
</list-item>
<list-item>
<p>Randomized tokenization generates unique tokens upon each application, maximizing privacy through obfuscation. While reducing re-identification risk, randomized tokens preclude direct record linkage, limiting utility for longitudinal or multi-source data integration critical in pharmacovigilance. Thus, this method suits anonymized datasets intended for aggregate safety analyses or exploratory research where linkage is unnecessary (<xref ref-type="bibr" rid="B2">Bernstam et al., 2023</xref>).</p>
</list-item>
<list-item>
<p>Referential tokenization leverages a secure reference repository and cryptographic hashing of demographic elements and identifiers to generate tokens. By abstracting identifiers into irreversible pseudonyms and managing mappings within secure environments, referential tokenization enables precise, privacy-compliant linkage of heterogeneous datasets. This approach is widely adopted for regulatory-grade pharmacoepidemiology and pharmacovigilance studies, supporting integration of EHR, claims, mortality, and SDOH data at scale (<xref ref-type="bibr" rid="B5">D&#x2019;Andrea et al., 2024</xref>; <xref ref-type="bibr" rid="B6">Eckrote et al., 2024</xref>).</p>
</list-item>
</list>
</p>
<p>Referential tokenization&#x2019;s scalability and accuracy have been demonstrated in real-world applications, such as the linkage of large payer and non-payer datasets, with linkage precision exceeding 99% in high-fidelity settings (<xref ref-type="bibr" rid="B5">D&#x2019;Andrea et al., 2024</xref>). This level of performance facilitates comprehensive safety signal detection, benefit-risk assessment, and population health research essential to pharmacovigilance.</p>
<p>Note: The overlap with non-payer (open) claims data was approximately 90%, based on linkage to clearinghouse data representing &#x3e;70 million lives. By contrast, overlap with payer-specific datasets was &#x223c;40%, though with higher precision in capturing longitudinal medical claims.</p>
<p>
<xref ref-type="table" rid="T1">Table 1</xref> summarizes the three primary tokenization methods alongside their key advantages for pharmacovigilance and pharmacoepidemiology.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Tokenization methods and their applications.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Tokenization method</th>
<th align="left">Description</th>
<th align="left">Advantages in pharmacovigilance and pharmacoepidemiology</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Deterministic</td>
<td align="left">Generates a fixed token for each data element, ensuring that the same input data always produces the same token. Ideal for structured data integration, ensuring consistent and reliable data linkage.</td>
<td align="left">Enables longitudinal tracking of patient outcomes; efficient analysis across datasets.</td>
</tr>
<tr>
<td align="left">Randomized</td>
<td align="left">Generates a different token each time for the same data element, increasing privacy by obscuring relationships. This makes it more difficult to infer useful information from tokenized data but complicates data queries and may reduce match rates across datasets.</td>
<td align="left">Enhances patient privacy; prevents re-identification, especially in real-world data studies.</td>
</tr>
<tr>
<td align="left">Referential</td>
<td align="left">Combines deterministic and randomized tokenization. Uses tokens referring to an external &#x201c;reference&#x201d; or lookup table, balancing flexibility and privacy. Ensures high match rates between datasets, making it ideal for integrating external data sources.</td>
<td align="left">Allows secure, dynamic access to original data for analysis; flexible for cross-dataset linkage.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2-2">
<label>2.2</label>
<title>Tokenization workflow and global regulatory framing</title>
<p>A robust tokenization workflow typically involves several key stages that ensure privacy preservation in pharmacovigilance as seen in <xref ref-type="table" rid="T2">Table 2</xref>:<list list-type="order">
<list-item>
<p>Identification of Sensitive Data Elements: Initially, data custodians identify direct identifiers (e.g., patient names, medical record numbers, exact dates of birth) and quasi-identifiers (e.g., ZIP codes, gender) subject to privacy protection (<xref ref-type="bibr" rid="B2">Bernstam et al., 2023</xref>).</p>
</list-item>
<list-item>
<p>Token Generation: Cryptographic algorithms, often based on Secure Hash Algorithm (SHA) variants, produce fixed-length tokens that are deterministic and non-reversible without access to secure keys or lookup tables (<xref ref-type="bibr" rid="B18">Yue, 2022</xref>; <xref ref-type="bibr" rid="B6">Eckrote et al., 2024</xref>). The selection of algorithmic parameters is critical to balancing linkage fidelity and privacy.</p>
</list-item>
<list-item>
<p>Secure Mapping and Storage: The token-to-identifier mappings are stored separately in encrypted, access-controlled repositories, minimizing risks of data leakage. Best practices mandate AES-256 encryption and stringent role-based access controls to enforce separation of duties (<xref ref-type="bibr" rid="B17">Walters et al., 2025</xref>).</p>
</list-item>
<list-item>
<p>Linkage and Translation Across Datasets: Tokenized identifiers enable &#x201c;cross-walking&#x201d; between datasets originating from disparate systems, often with heterogeneous data standards and terminologies. Referential tokenization supports the use of a &#x201c;golden record&#x201d; as a canonical reference to resolve discrepancies and enhance matching accuracy (<xref ref-type="bibr" rid="B4">Dunne, 2023</xref>).</p>
</list-item>
<list-item>
<p>Data Enrichment: Tokenized datasets can be securely enriched with auxiliary data streams such as SDOH, genomic profiles, and mortality records, adding contextual layers that enhance RWE analyses without compromising privacy (<xref ref-type="bibr" rid="B12">Li et al., 2024</xref>; <xref ref-type="bibr" rid="B7">Elhussein et al., 2024</xref>).</p>
</list-item>
</list>
</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Tokenization workflow.</p>
</caption>
<table>
<tbody valign="top">
<tr>
<td align="left">
<inline-graphic xlink:href="fdsfr-05-1599217-fx1.tif"/>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The diagram above shows the process of identifying sensitive data, generating and securely storing tokens, enabling their use while maintaining privacy, managing the token lifecycle for compliance, and enriching the dataset with secondary data elements.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s2-3">
<label>2.3</label>
<title>Regulatory context and global privacy frameworks</title>
<p>The adoption of tokenization is tightly linked to regulatory frameworks governing patient privacy and pharmacovigilance. The Health Insurance Portability and Accountability Act (HIPAA) in the United States, the General Data Protection Regulation (GDPR) in the European Union, and the European Health Data Space Regulation (EU 2025/327) collectively emphasize pseudonymization and de-identification as cornerstones for lawful data reuse in safety surveillance (European Parliament, 2025; <xref ref-type="bibr" rid="B9">European Data Protection Board, 2021</xref>; <xref ref-type="bibr" rid="B15">Office for Civil Rights, 2023</xref>).</p>
<p>Referential tokenization aligns with these frameworks by enabling deterministic linkage via pseudonymous tokens while preventing the exchange or exposure of PII. The EHDS explicitly endorses token-based data linkage for secondary use of health data across EU member states, contingent upon compliance with GDPR principles (<xref ref-type="bibr" rid="B8">European Commission, 2025</xref>). Such alignment facilitates multinational clinical trials and pharmacovigilance studies that require longitudinal datasets while respecting data sovereignty and patient rights.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Technical challenges</title>
<sec id="s3-1">
<label>3.1</label>
<title>Data variability and patient-centricity</title>
<p>Clinical trial and pharmacovigilance datasets are inherently dynamic, reflecting ongoing changes in patient health status, treatment responses, and personal circumstances (<xref ref-type="bibr" rid="B12">Li et al., 2024</xref>). Tokenization frameworks must account for this variability to preserve critical longitudinal information needed for personalized interventions and accurate safety assessments. Failure to accommodate these fluctuations risks oversimplifying patient trajectories, diminishing the clinical and scientific value of pharmacovigilance data. Overreliance on aggregated population-level data risks masking heterogeneity in treatment effects and adverse event profiles, which are essential considerations in pharmacovigilance and precision medicine (<xref ref-type="bibr" rid="B1">Abul-Husn and Kenny, 2019</xref>). Consequently, tokenization methodologies should maintain sufficient granularity to support nuanced, patient-centric pharmacovigilance analyses throughout the clinical research lifecycle.</p>
</sec>
<sec id="s3-2">
<label>3.2</label>
<title>Linkage precision: balancing over-linking and under-linking</title>
<p>Accurate linkage of tokenized data is fundamental in pharmacovigilance and health economic studies. Over-linking, when unrelated patient records are mistakenly connected, can introduce confounding bias by falsely aggregating clinical events, exposures, or outcomes across distinct individuals. This misattribution may distort incidence rates, mask true safety signals, or generate spurious associations, ultimately compromising the validity of pharmacovigilance and real-world evidence analyses (<xref ref-type="bibr" rid="B16">Velummailum et al., 2023</xref>). Conversely, under-linking&#x2014;failure to match related records&#x2014;results in incomplete patient profiles and potential missed adverse drug reactions or treatment outcomes. Referential tokenization employs advanced matching algorithms supported by &#x201c;golden records&#x201d; or trusted data sources to mitigate these risks, enhancing the precision of pharmacovigilance cross-dataset linkage and minimizing both false positives and false negatives (<xref ref-type="bibr" rid="B4">Dunne, 2023</xref>).</p>
</sec>
<sec id="s3-3">
<label>3.3</label>
<title>Data harmonization and integrity</title>
<p>Tokenization&#x2019;s effectiveness depends heavily on the harmonization of data across diverse sources with varying standards, formats, and completeness (<xref ref-type="bibr" rid="B16">Velummailum et al., 2023</xref>). Clinical trial and real-world pharmacovigilance data often originate from heterogeneous environments, complicating direct linkage. Without rigorous data preprocessing and standardization, tokenization can generate mismatches or fragmentations, threatening data integrity and skewing pharmacovigilance outcomes (<xref ref-type="bibr" rid="B17">Walters et al., 2025</xref>). Incorporating privacy-preserving record linkage (PPRL) mechanisms alongside tokenization further ensures secure integration of datasets without exposing sensitive patient identifiers. However, weak encryption or lax access controls risk breaching patient confidentiality, undermining trust in the entire pharmacovigilance data ecosystem.</p>
</sec>
<sec id="s3-4">
<label>3.4</label>
<title>Linking Economical and clinical data</title>
<p>In health economics and safety studies, linking clinical trial and pharmacovigilance data to economic datasets is pivotal for accurate cost-effectiveness and safety analyses. Inaccuracies in tokenization or linkage processes can yield incomplete or biased datasets, obscuring critical cost drivers or patient-level treatment responses. This threatens the reliability of economic evaluations that inform reimbursement and regulatory decisions, potentially affecting trial feasibility and therapeutic innovation (<xref ref-type="bibr" rid="B12">Li et al., 2024</xref>).</p>
</sec>
<sec id="s3-5">
<label>3.5</label>
<title>Integrating primary and secondary data sources</title>
<p>Merging primary clinical trial and pharmacovigilance datasets with secondary sources&#x2014;such as claims, social determinants of health (SDOH), genomic, and mortality data&#x2014;introduces additional complexities. The distinction between primary and secondary data depends on collection context. Primary data refers to information collected for the specific study at hand (e.g., clinical trial data), whereas secondary data (e.g., claims or SDOH) is repurposed for analysis. Depending on how pharmacovigilance data is generated, it may fall into either category (<xref ref-type="bibr" rid="B11">Gliklich et al., 2014</xref>). These datasets often differ in structure and standards, necessitating robust preprocessing and referential matching algorithms to ensure consistent token mapping across all sources (<xref ref-type="bibr" rid="B6">Eckrote et al., 2024</xref>). Failure to properly align these data can degrade the accuracy of pharmacovigilance real-world evidence (RWE) and impede valid interpretation. Furthermore, enriching tokenized datasets with external data must be carefully managed within regulatory frameworks to prevent inadvertent exposure of protected health information (<xref ref-type="bibr" rid="B15">Office for Civil Rights, 2023</xref>).</p>
</sec>
<sec id="s3-6">
<label>3.6</label>
<title>Controlled re-identification for safety signal follow-up</title>
<p>A critical challenge involves enabling controlled, secure re-identification for pharmacovigilance activities such as safety signal validation and adverse event follow-up. Tokenization frameworks in pharmacoepidemiology often include tightly regulated re-identification protocols managed by neutral data custodians. These custodians maintain encrypted mappings between tokens and original identifiers within secure environments, allowing re-identification only under ethically approved, regulatorily compliant conditions (<xref ref-type="bibr" rid="B10">ENCePP, 2020</xref>). This approach balances stringent privacy protection with the operational necessity of patient-level follow-up in pharmacovigilance, supported by access controls, audit trails, and regulatory oversight.</p>
</sec>
<sec id="s3-7">
<label>3.7</label>
<title>Applied scenarios and criteria for tokenization in pharmacoepidemiology and safety surveillance</title>
<p>Although tokenization is gaining adoption in regulatory-grade data linkage, explicit examples in published pharmacoepidemiology studies are limited due to commercial and privacy constraints. The following three applied scenarios illustrate how tokenization enables privacy-preserving linkage for post-market safety research and real-world evidence (RWE) generation:</p>
<p>Scenario 1: Linking Claims and EHRs to Detect Adverse Events Post-Approval.</p>
<p>After FDA approval of a new anticoagulant, a manufacturer initiates post-marketing surveillance. Using referential tokenization, patient-level claims data (hospitalizations, diagnoses, procedures) are linked to EHRs (lab values, medication reconciliation, bleeding risk scores). This enables early detection of adverse drug reactions (e.g., GI bleeding) that may not be fully captured in claims alone fulfilling a regulatory risk minimization requirement (<xref ref-type="bibr" rid="B6">Eckrote et al., 2024</xref>).<list list-type="bullet">
<list-item>
<p>Tokenization Use: Demographic fields (e.g., full DOB, sex, ZIP code) are cryptographically hashed and tokenized by a third-party tokenization network.</p>
</list-item>
<list-item>
<p>Outcome: Combined datasets support a retrospective cohort study on bleeding events stratified by renal function, fulfilling a regulatory risk minimization requirement.</p>
</list-item>
</list>
</p>
<p>Scenario 2: Linking Registry and Claims Data for Long-Term Drug Safety Monitoring.</p>
<p>A national rheumatoid arthritis (RA) registry collects longitudinal clinical data on disease activity, biologic use, and side effects. Researchers seek to evaluate long-term malignancy risk associated with a specific TNF inhibitor informing black-box warning updates (<xref ref-type="bibr" rid="B5">D&#x2019;Andrea et al., 2024</xref>).<list list-type="bullet">
<list-item>
<p>Tokenization Use: Patients in the registry are tokenized using referential methods. Claims data from Medicare and commercial payers are then linked to capture incident cancer diagnoses, procedures, and prescription fills.</p>
</list-item>
<list-item>
<p>Outcome: The linked dataset allows evaluation of rare, long-latency adverse events over multi-year periods, providing evidence for regulatory signal refinement and black-box warning updates.</p>
</list-item>
</list>
</p>
<p>Scenario 3: Linking Clinical Trial Participants to Claims for Safety Follow-up.</p>
<p>A pivotal oncology trial identifies a rare cardiovascular signal in interim analysis. To assess whether this signal emerges in broader populations, the sponsor uses tokenization to link trial participants to U.S. claims data post-trial enabling regulatory submission of updated safety data (DIA RWE, 2023).<list list-type="bullet">
<list-item>
<p>Tokenization Use: At consent, participant demographics are tokenized. Tokens are matched with commercial claims data over a 2-year period post-trial.</p>
</list-item>
</list>
</p>
<p>Outcome: Analysis identifies real-world incidence rates of cardiac events, enabling regulatory submission of updated safety data and refined risk communication in product labeling.</p>
<p>
<xref ref-type="table" rid="T3">Table 3</xref> summarizes pharmacoepidemiology-specific scenarios underscore tokenization&#x2019;s pivotal role in enabling safety signal detection, longitudinal outcomes research, and regulatory-grade data integration&#x2014;without compromising patient privacy. As tokenized infrastructures mature, their adoption will expand across global post-marketing surveillance networks and distributed data networks.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Minimum criteria for tokenization in pharmacoepidemiology safety studies.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Category</th>
<th align="left">Minimum requirement</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Identifier Inputs</td>
<td align="left">Full DOB, ZIP code, sex/gender, first &#x26; last name preferred (if available).</td>
</tr>
<tr>
<td align="left">Tokenization Method</td>
<td align="left">Referential tokenization with SHA-2 or SHA-3 algorithms, salted, non-reversible hashes.<break/>
<italic>Salted: Salted tokens involve adding a random value (a &#x201c;salt&#x201d;) to the identifier before hashing, improving security by preventing reverse-engineering of the token.</italic>
</td>
</tr>
<tr>
<td align="left">Linkage Environment</td>
<td align="left">Clean room or enclave with enforced access controls, audit logs, and role-based permissions.<break/>
<italic>Clean room or enclave: A highly secure computing environment that enforces role-based access, audit logs, and strict governance for analyzing sensitive data without exposing identifiable elements.</italic> See <xref ref-type="bibr" rid="B17">Walters et al., 2025</xref>.</td>
</tr>
<tr>
<td align="left">Linkage Accuracy</td>
<td align="left">Match precision &#x2265;95% validated via gold-standard test sets or sample clerical review.</td>
</tr>
<tr>
<td align="left">Re-identification Protocols</td>
<td align="left">Restricted to neutral data custodian under IRB- or regulator-approved protocol.</td>
</tr>
<tr>
<td align="left">Regulatory Alignment</td>
<td align="left">Must comply with HIPAA Expert Determination or GDPR pseudonymization guidelines.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<label>4</label>
<title>Discussion</title>
<sec id="s4-1">
<label>4.1</label>
<title>Conclusion: optimizing tokenization for clinical trial applications</title>
<p>In conclusion, tokenization serves as an essential enabler for privacy-preserving data integration in clinical research and pharmacovigilance. By securely replacing sensitive patient identifiers with irreversible tokens, tokenization enables comprehensive analysis of complex, multi-source datasets&#x2014;ranging from electronic health records (EHRs) and claims to genomic and patient-reported outcomes&#x2014;while maintaining patient confidentiality throughout the clinical trial and post-market pharmacovigilance lifecycle.</p>
<p>Referential tokenization, in particular, supports scalable, high-precision linkage of disparate datasets, ensuring compliance with evolving global privacy frameworks such as GDPR and the European Health Data Space (EHDS). This approach facilitates the construction of richer, longitudinal patient profiles essential for safety monitoring, benefit-risk assessment, and health equity research, all while minimizing the risk of data exposure through adherence to de-identification and pseudonymization standards.</p>
<p>To fully realize tokenization&#x2019;s transformative potential in pharmacovigilance, key technical challenges must be addressed. These include ensuring data integrity through rigorous standardization and harmonization, managing token lifecycles effectively, maintaining linkage accuracy to avoid both over-linking and under-linking, and implementing controlled re-identification mechanisms for authorized pharmacovigilance follow-up.</p>
<p>Additionally, enriching tokenized datasets with secondary sources such as claims, social determinants of health (SDOH), and mortality data offers valuable clinical insights but demands strict governance to uphold data alignment and regulatory compliance.</p>
<p>Empirical evidence underscores tokenization&#x2019;s impact, as demonstrated in applied clinical trial settings with secure linkage of real-world data achieving high precision, enabling robust longitudinal analyses and international collaboration under stringent privacy constraints (<xref ref-type="bibr" rid="B5">D&#x2019;Andrea et al., 2024</xref>). As clinical research and pharmacovigilance increasingly rely on interoperable real-world data, tokenization will remain central to unlocking the full analytic value of healthcare datasets while safeguarding patient privacy.</p>
<p>Future directions should focus on refining tokenization algorithms, establishing international standards for tokenized data reporting, and advancing trusted data custodian models that balance patient confidentiality with investigational imperatives. These efforts will underpin the continued evolution of clinical research toward resilient, patient-centric paradigms that harmonize privacy preservation with data-driven innovation and regulatory rigor.</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="author-contributions" id="s5">
<title>Author contributions</title>
<p>CC: Methodology, Writing &#x2013; original draft, Resources.</p>
</sec>
<sec sec-type="COI-statement" id="s7">
<title>Conflict of interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s8">
<title>Generative AI statement</title>
<p>The author(s) declare that Generative AI was used in the creation of this manuscript. AI tools were used to assist in refining grammar, punctuation, and clarity of language throughout the manuscript. AI was also leveraged to support chronological ordering of references, streamline scenario articulation, and ensure consistency in technical terminology. All content was reviewed, edited, and approved by the author to ensure accuracy, originality, and integrity.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s10">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fdsfr.2025.1599217/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fdsfr.2025.1599217/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Image1.png" id="SM1" mimetype="application/png" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<fn-group>
<fn fn-type="custom" custom-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1735689/overview">Simone Pinheiro</ext-link>, AbbVie, United States</p>
</fn>
<fn fn-type="custom" custom-type="reviewed-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1538459/overview">Efe Eworuke</ext-link>, United States Food and Drug Administration, United States</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abul-Husn</surname>
<given-names>N. S.</given-names>
</name>
<name>
<surname>Kenny</surname>
<given-names>E. E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Personalized medicine and the power of electronic health records</article-title>. <source>Cell</source> <volume>177</volume> (<issue>1</issue>), <fpage>58</fpage>&#x2013;<lpage>69</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2019.02.039</pub-id>
<pub-id pub-id-type="pmid">30901549</pub-id>
</mixed-citation>
</ref>
<ref id="B2">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Bernstam</surname>
<given-names>E. V.</given-names>
</name>
<name>
<surname>Applegate</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chaudhari</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Coda</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2023</year>). <source>Real-world matching performance of de-identified record-linking tokens</source>. <publisher-name>ResearchGate</publisher-name>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.researchgate.net/publication/362313925_Real-world_matching_performance_of_de-identified_record_linking_tokens">https://www.researchgate.net/publication/362313925_Real-world_matching_performance_of_de-identified_record_linking_tokens</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B3">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cruz</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Guimar&#xe3;es</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Santos</surname>
<given-names>M. F.</given-names>
</name>
<name>
<surname>Machado</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Decentralize healthcare marketplace</article-title>. <source>Procedia Comput. Sci.</source> <volume>231</volume>, <fpage>439</fpage>&#x2013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2023.12.231</pub-id>
</mixed-citation>
</ref>
<ref id="B4">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dunne</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>How to transition RWE studies to non-clinical regulatory grade</article-title>. <source>Appl. Clin. Trials</source>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.appliedclinicaltrialsonline.com/view/how-transition-rwe-studies-non-clinical-regulatory-grade">https://www.appliedclinicaltrialsonline.com/view/how-transition-rwe-studies-non-clinical-regulatory-grade</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B5">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>D&#x2019;Andrea</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>Y. C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2024</year>). &#x201c;<article-title>Unlocking insights: lessons learned from the tokenization of a psoriasis trial</article-title>,&#x201d; in <source>International conference on pharmacoepidemiology (ICPE)</source>. <publisher-loc>Berlin, Germany</publisher-loc>. <comment>Abstract &#x26; Poster</comment>.</mixed-citation>
</ref>
<ref id="B6">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eckrote</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Nielson</surname>
<given-names>C. M.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Alexander</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>Low</surname>
<given-names>K. W.</given-names>
</name>
<etal/>
</person-group> (<year>2024</year>). <article-title>Linking clinical trial participants to their U.S. real-world data through tokenization: a practical guide</article-title>. <source>Contemp. Clin. Trials Commun.</source> <volume>41</volume>, <fpage>101354</fpage>. <pub-id pub-id-type="doi">10.1016/j.conctc.2024.101354</pub-id>
<pub-id pub-id-type="pmid">39280783</pub-id>
</mixed-citation>
</ref>
<ref id="B7">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Elhussein</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Baymuradov</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Elhadad</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Natarajan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>G&#xfc;rsoy</surname>
<given-names>G.</given-names>
</name>
</person-group>
<collab>NYGC ALS Consortium</collab> (<year>2024</year>). <article-title>A framework for sharing of clinical and genetic data for precision medicine applications</article-title>. <source>Nat. Med.</source> <volume>30</volume>, <fpage>3578</fpage>&#x2013;<lpage>3589</lpage>. <pub-id pub-id-type="doi">10.1038/s41591-024-03239-5</pub-id>
<pub-id pub-id-type="pmid">39227443</pub-id>
</mixed-citation>
</ref>
<ref id="B8">
<mixed-citation publication-type="journal">
<collab>European Commission</collab> (<year>2025</year>). <article-title>Regulation (EU) 2025/327 of the European parliament and of the council on the European health data space</article-title>. <source>Official J. Eur. Union</source>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=OJ:L:2025:327">https://eur-lex.europa.eu/legal-content/EN/TXT/?uri&#x3d;OJ:L:2025:327</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B9">
<mixed-citation publication-type="web">
<collab>European Data Protection Board</collab> (<year>2021</year>). <article-title>Guidelines on pseudonymisation under GDPR</article-title>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://edpb.europa.eu/sites/default/files/files/file1/edpb_guidelines_202010_pseudonymisation_en.pdf">https://edpb.europa.eu/sites/default/files/files/file1/edpb_guidelines_202010_pseudonymisation_en.pdf</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B10">
<mixed-citation publication-type="book">
<collab>European Network of Centres for Pharmacoepidemiology and Pharmacovigilance (ENCePP)</collab> (<year>2020</year>). <source>ENCePP methodological guide (Version 5)</source>. <publisher-name>European Medicines Agency</publisher-name>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://encepp.europa.eu/document/download/7f94d0ab-aa8f-40f2-b22b-189b10489583_en?filename=methodologicalGuideReferences.pdf">https://encepp.europa.eu/document/download/7f94d0ab-aa8f-40f2-b22b-189b10489583_en?filename&#x3d;methodologicalGuideReferences.pdf</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B11">
<mixed-citation publication-type="book">
<person-group person-group-type="editor">
<name>
<surname>Gliklich</surname>
<given-names>R. E.</given-names>
</name>
<name>
<surname>Dreyer</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Leavy</surname>
<given-names>M. B.</given-names>
</name>
</person-group> (<year>2014</year>). <source>Registries for evaluating patient outcomes: a user&#x2019;s guide</source>. <edition>3rd ed.</edition> (<publisher-name>Agency for Healthcare Research and Quality</publisher-name>). <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/books/NBK208616/">https://www.ncbi.nlm.nih.gov/books/NBK208616/</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B12">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Mowery</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Vurgun</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Hwang</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2024</year>). <article-title>Realizing the potential of social determinants data in EHR systems: a scoping review of approaches for screening, linkage, extraction, analysis, and interventions</article-title>. <source>J. Clin. Transl. Sci.</source> <volume>8</volume> (<issue>1</issue>), <fpage>e147</fpage>. <pub-id pub-id-type="doi">10.1017/cts.2024.571</pub-id>
<pub-id pub-id-type="pmid">39478779</pub-id>
</mixed-citation>
</ref>
<ref id="B13">
<mixed-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Liede</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>&#x201c;Tokenization&#x201d;: privacy-Preserving data integration to enhance clinical trials and real-world evidence studies</article-title>,&#x201d; in <source>DIA real-world evidence conference</source>. <publisher-loc>Philadelphia</publisher-loc>.</mixed-citation>
</ref>
<ref id="B14">
<mixed-citation publication-type="book">
<collab>National Institutes of Health</collab> (<year>2006</year>). <source>MedlinePlus medical encyclopedia</source>. <publisher-name>U.S. National Library of Medicine</publisher-name>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/books/NBK9579/">https://www.ncbi.nlm.nih.gov/books/NBK9579/</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B15">
<mixed-citation publication-type="book">
<collab>Office for Civil Rights</collab> (<year>2023</year>). <source>HIPAA compliance and tokenization: enhancing patient privacy</source>. <publisher-name>U.S. Department of Health and Human Services</publisher-name>. <comment>Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/examples/how-ocr-enforces-the-hipaa-privacy-and-security-rules/index.html">https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/examples/how-ocr-enforces-the-hipaa-privacy-and-security-rules/index.html</ext-link>.</comment>
</mixed-citation>
</ref>
<ref id="B16">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Velummailum</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Flannelly</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Desai</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Matsumoto</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kulkarni</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Privacy-preserving record linkage: a scoping review</article-title>. <source>Int. J. Med. Inf.</source> <volume>169</volume>, <fpage>105191</fpage>. <pub-id pub-id-type="doi">10.1016/j.ijmedinf.2023.105191</pub-id>
</mixed-citation>
</ref>
<ref id="B17">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Walters</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ridley</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Papavero</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kurland</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Improving data security in healthcare: practical approaches to tokenization and encryption</article-title>. <source>J. Healthc. Inf. Manag.</source> <volume>39</volume> (<issue>2</issue>), <fpage>34</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation>
</ref>
<ref id="B18">
<mixed-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yue</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Privacy-preserving methods for medical data sharing: tokenization and beyond</article-title>. <source>Health Data Sci. J.</source> <volume>6</volume> (<issue>1</issue>), <fpage>12</fpage>&#x2013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1016/j.hdsj.2022.01.005</pub-id>
</mixed-citation>
</ref>
</ref-list>
</back>
</article>