<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Digit. Health</journal-id>
<journal-title>Frontiers in Digital Health</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Digit. Health</abbrev-journal-title>
<issn pub-type="epub">2673-253X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdgth.2025.1604001</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Digital Health</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Challenges of identification and anonymity in time-continuous data from medical environments</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes"><name><surname>Hammer</surname><given-names>Freimut</given-names></name>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref><uri xlink:href="https://loop.frontiersin.org/people/2910378/overview"/><role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/><role content-type="https://credit.niso.org/contributor-roles/methodology/"/><role content-type="https://credit.niso.org/contributor-roles/visualization/"/><role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/><role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/></contrib>
<contrib contrib-type="author"><name><surname>Strufe</surname><given-names>Thorsten</given-names></name><uri xlink:href="https://loop.frontiersin.org/people/2909324/overview" /><role content-type="https://credit.niso.org/contributor-roles/methodology/"/><role content-type="https://credit.niso.org/contributor-roles/supervision/"/><role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/></contrib>
</contrib-group>
<aff><institution>KASTEL Security Research Labs, Karlsruhe Institute of Technology (KIT)</institution>, <addr-line>Karlsruhe</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p><bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2784832/overview">Thijs Veugen</ext-link>, Netherlands Organisation for Applied Scientific Research, Netherlands</p></fn>
<fn fn-type="edited-by"><p><bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2962798/overview">Mahesh Kumar Goyal</ext-link>, Google, United States</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/3085915/overview">Samsad Jahan</ext-link>, Victoria University, Australia</p></fn>
<corresp id="cor1"><label>&#x002A;</label><bold>Correspondence:</bold> Freimut Hammer <email>hammerfreimut@gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub"><day>20</day><month>08</month><year>2025</year></pub-date>
<pub-date pub-type="collection"><year>2025</year></pub-date>
<volume>7</volume><elocation-id>1604001</elocation-id>
<history>
<date date-type="received"><day>01</day><month>04</month><year>2025</year></date>
<date date-type="accepted"><day>10</day><month>07</month><year>2025</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 Hammer and Strufe.</copyright-statement>
<copyright-year>2025</copyright-year><copyright-holder>Hammer and Strufe</copyright-holder><license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>In medical environments, time-continuous data, such as electrocardiographic records, necessitates a distinct approach to anonymization due to the paramount importance of preserving its spatio-temporal integrity for optimal utility. A wide array of data types, characterized by their high sensitivity to the patient&#x2019;s well-being and their substantial interest to researchers, are generated. A significant proportion of this data may be of interest to researchers beyond the original purposes for which it was collected. This necessity underscores the pressing need for effective anonymization methods, a challenge that existing approaches often fail to adequately address. Robust privacy mechanisms are essential to uphold patient rights and ensure informed consent, particularly within the framework of the European Health Data Space. This paper explores the challenges and opportunities inherent in developing a novel approach to anonymize such data and devise suitable metrics to assess the efficacy of anonymization. One promising approach is the adoption of differential privacy to account for temporal context and correlations, making it suitable for time-continuous data.</p>
</abstract>
<kwd-group>
<kwd>privacy</kwd>
<kwd>medical data</kwd>
<kwd>differential privacy (DP)</kwd>
<kwd>anonymity</kwd>
<kwd>data sharing and reuse</kwd>
</kwd-group><counts>
<fig-count count="9"/>
<table-count count="1"/><equation-count count="275"/><ref-count count="61"/><page-count count="19"/><word-count count="0"/></counts><custom-meta-wrap><custom-meta><meta-name>section-at-acceptance</meta-name><meta-value>Health Informatics</meta-value></custom-meta></custom-meta-wrap>
</article-meta>
</front>
<body><sec id="s1" sec-type="intro"><label>1</label><title>Introduction</title>
<p>Patient data is defined as physiological data and metadata collected in a medical environment. This paper concentrates on this subject, with implications that extend to time-continuous data in general. Such data can be recorded in a variety of settings, including hospitals, research facilities, and clinical practices. In a broader sense, physiological data collected by wearable devices such as smartwatches can also be considered patient data. This data encompasses information related to an individual&#x2019;s physiology, psychology, and overall health status, facilitating a unique identification. Due to its sensitivity, such data necessitates a regulatory framework that ensures it is handled with particular care.</p>
<p>The release or sharing of such data is essential for the generation of knowledge with mutual control and replicability and is therefore fundamental to the ethics and progress of science.</p>
<p>The utilization of patient data can facilitate the detection of rare diseases or serve as realistic training material for medical professionals. Additionally, it facilitates the training and verification of machine-learned models for medical data and enables data-driven research in this field in general.</p>
<p>These objectives are often in stark contrast with the sensitivity of the data, especially in medical environments. This underscores the necessity for a mechanism to safeguard the individual behind the data. A mechanism that effectively anonymizes data would be an optimal solution, enabling straightforward sharing of data among medical professionals or researchers, as well as its publication, without compromising the rights of the individual to whom the data pertains.</p>
<p>A distinguishing feature of numerous categories of medical data is their time-continuous nature. The spatio-temporal relationship within such data is of paramount importance to their utility. An example of such data is an ECG, which can be seen in <xref ref-type="fig" rid="F4">Figures&#x00A0;4</xref>&#x2013;<xref ref-type="fig" rid="F6">6</xref>. There are plenty of different types of such data, like SPO<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM1"><mml:msub><mml:mi></mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> saturation, electroencephalogram, and blood pressure. In <xref ref-type="fig" rid="F1">Figure&#x00A0;1</xref>, two physiological signals&#x2014;photoplethysmogram (PPG) and arterial blood pressure (BP)&#x2014;are plotted over time to enable direct visual comparison of their temporal dynamics. The data is derived from Wikimedia Commons (<xref ref-type="bibr" rid="B1">1</xref>). To facilitate this comparison, both signals have been normalized to a common scale between 0 and 1 along the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM2"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis. This normalization preserves the relative shape and timing of signal fluctuations while removing the influence of differing absolute amplitudes and units. The blue line represents the PPG signal, which reflects blood volume changes in the microvascular bed of tissue, typically measured using optical sensors. It is characterized by sharp upstrokes corresponding to systolic blood ejection, followed by a slower decay during diastole. The red line represents the arterial blood pressure waveform, which reflects pressure changes within the arteries during the cardiac cycle. It features broader peaks and lower frequency content compared to the PPG signal. Plotting both signals on the same normalized axis reveals their phase relationship, morphological similarity, and temporal alignment, features that are especially useful in multimodal cardiovascular analysis, where understanding the interaction between pressure and volume dynamics is critical. Although normalization removes units (e.g., mmHg or arbitrary light absorption units), it retains the essential temporal characteristics needed for waveform analysis, including peak timing, rise and fall slopes, and periodicity.</p>
<fig id="F1" position="float"><label>Figure 1</label>
<caption><p>Time-continuous medical data example.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g001.tif"><alt-text content-type="machine-generated">Graph comparing photoplethysmogram and blood pressure signals over five seconds. Both signals are normalized and display periodic peaks. The photoplethysmogram is shown in blue and blood pressure in red, with similar wave patterns.</alt-text>
</graphic>
</fig>
<p>Both ECG and the signals shown in <xref ref-type="fig" rid="F1">Figures&#x00A0;1</xref>, <xref ref-type="fig" rid="F4">4</xref>&#x2013;<xref ref-type="fig" rid="F6">6</xref> are time-continuous, as they reflect physiological processes&#x2014;electrical activity, blood volume, and pressure&#x2014;that evolve smoothly over time without discrete jumps. These signals are typically sampled at high frequencies to capture their continuous nature and preserve critical temporal features such as waveforms and phase relationships.</p>
<p>Formalizing and analyzing mechanisms to ensure privacy and utility is crucial for sensitive applications. Existing methods such as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM3"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity (<xref ref-type="bibr" rid="B2">2</xref>), <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM4"><mml:mi>t</mml:mi></mml:math></inline-formula>-closeness (<xref ref-type="bibr" rid="B3">3</xref>), and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM5"><mml:mi>l</mml:mi></mml:math></inline-formula>-diversity (<xref ref-type="bibr" rid="B4">4</xref>) do not provide any information-theoretic guarantees or statements regarding the level of privacy achieved.</p>
<p>Implementations such as CASTLE (<xref ref-type="bibr" rid="B5">5</xref>, <xref ref-type="bibr" rid="B6">6</xref>), which employs <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM6"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity to continuous streams of non-continuous data, and SABRE (<xref ref-type="bibr" rid="B7">7</xref>), which utilizes a <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM7"><mml:mi>t</mml:mi></mml:math></inline-formula>-closeness-centered bucketization approach to achieve <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM8"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity, are limited in their ability to preserve the spatio-temporal relation by compressing or stretching the data. Other approaches that preserve spatio-temporal relationships, like the group- and link-based approach by Nergiz et al. (<xref ref-type="bibr" rid="B8">8</xref>), fail to maintain the continuity, as the generation of representative data might not preserve the order or the step size. A notable shortcoming of the aforementioned approaches is the absence of formal statements about the level of privacy achieved.</p>
<p>Differential privacy (DP) stands out as a notable exception, offering a formal framework that provides robust guarantees and quantifies the level of privacy attained (<xref ref-type="bibr" rid="B9">9</xref>). The application of DP to health data has been examined by Dankar and El Emam (<xref ref-type="bibr" rid="B10">10</xref>). However, their analysis does not address the specific challenges posed by time-continuous data or the necessity of preserving spatio-temporal dependencies. Instead, their focus lies in the realm of categorical or numerical point-based measurements. In contrast, Olawoyin et al. (<xref ref-type="bibr" rid="B11">11</xref>) introduced a bottom-up generalization approach for temporal data. After the generalization of the time dimension, a DP mechanism is employed to ensure privacy along the value axis. This approach, while preserving the spatio-temporal relation, destroys the continuity of the data along the time axis.</p>
<p>This means that current approaches are not suitable for this task, as they fail to preserve the spatio-temporal relationships and the continuity of the data. Additionally, most fail to provide strong formal guarantees.</p>
<p>The ARX anonymization framework, developed at the Charit&#x00E9; Berlin, is an open-source tool designed to protect sensitive data. It implements a wide range of de-identification algorithms and risk analysis methods (<xref ref-type="bibr" rid="B12">12</xref>). The ARX model has emerged as a promising candidate for seamlessly integrating a developed anonymization mechanism for time-continuous medical data. Formalizing a utility that is suitable for most, if not all, kinds of diagnoses and analyses that can be performed on the data can be challenging, as different diagnostic or analytical tasks depend on vastly different characteristics of the data. For instance, an ECG can be used by a medical professional to not only inspect various infractions by analyzing multiple parameters such as the lengths, amplitudes, and distances between recorded waveform complexes but also to assess ventricular depolarization by analyzing the electrical axis of the heart. This can be approximated by calculating the area under the curve of the QRS complex (<xref ref-type="bibr" rid="B13">13</xref>).</p>
<p>In this paper, we explore the intricate field of anonymizing time-continuous medical data. We emphasize the importance of addressing this issue and highlight the limitations of current methods. We elaborate the specific challenges a suitable mechanism must solve, the opportunities that arise from such a solution, and the requirements for potential quality metrics. To the best of our knowledge, no existing approach effectively anonymizes time-continuous medical data while preserving its utility for diagnostic and analytical purposes, such as investigating correlations between diagnoses and patient characteristics. This paper introduces and formalizes candidate metrics for utility and privacy. It also emphasizes the challenges a potential solution must address and the opportunities that arise from overcoming them.</p>
</sec>
<sec id="s2" sec-type="background"><label>2</label><title>Background</title>
<p>In this section, the foundation is established by formalizing the fundamental properties required for the anonymization of time-continuous medical data. This facilitates formal analysis of such properties and supports the evaluation of the usefulness of different classes of anonymization approaches, which are addressed in <xref ref-type="sec" rid="s3">Section 3</xref>.</p>
<p>The notation is used to define the requirements, challenges, and opportunities for an anonymization approach targeting the class of data defined in <xref ref-type="statement" rid="st6">Definition 2.6</xref>, which preserves the time continuity.</p>
<sec id="s2a"><label>2.1</label><title>Notation</title>
<p>The date is considered part of a domain, where a point within this domain represents a data value or attribute, as formalized in <xref ref-type="statement" rid="st1">Definition 2.1</xref>.</p><statement id="st1"><label>DEFINITION 2.1</label><title>(Data value)</title>
<p>A value <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM9"><mml:mi>v</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM10"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> is the domain of this value, is called a data value. It is sometimes called an attribute.</p></statement>
<p>An example of a data value is the height of the patient, e.g., 183 cm, or a continuous measurement of blood pressure.</p>
<p>Multiple data values corresponding to the same individual or case make up a data record, as formalized in <xref ref-type="statement" rid="st2">Definition 2.2</xref>.</p><statement id="st2"><label>DEFINITION 2.2</label><title>(Data record)</title>
<p>A tuple <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM11"><mml:mi>d</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>n</mml:mi></mml:msub></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM12"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the domain of the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM13"><mml:mi>i</mml:mi></mml:math></inline-formula>th element and the tuple refers to exactly one individual, is called a data record. A data record is a cross-product of multiple data values. That is, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM14"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM15"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is a data value.</p></statement>
<p>The patientID, e.g., the hexadecimal ID <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM16"><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mn>1</mml:mn><mml:mi>d</mml:mi><mml:mi>c</mml:mi><mml:mn>2</mml:mn><mml:mi>e</mml:mi><mml:mn>6</mml:mn><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mn>6</mml:mn></mml:math></inline-formula>, together with two data values from the previous example, i.e., height and continuous blood pressure, forms a simple data record.</p>
<p>Usually, multiple data records are processed together, forming a database, as formalized in <xref ref-type="statement" rid="st3">Definition 2.3</xref>.</p><statement id="st3"><label>DEFINITION 2.3</label><title>(Database)</title>
<p>A set of data records <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM17"><mml:mi>S</mml:mi><mml:mo>&#x2286;</mml:mo><mml:mrow><mml:mo fence="false" stretchy="false">{</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> is called a database if each <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM18"><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> is a data record.</p></statement>
<p>An example of such a database could be the following set of patients, where each patient is a data record:
<list list-type="simple">
<list-item><label>&#x2022;</label>
<p>patientID: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM19"><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mn>1</mml:mn><mml:mi>d</mml:mi><mml:mi>c</mml:mi><mml:mn>2</mml:mn><mml:mi>e</mml:mi><mml:mn>6</mml:mn><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mn>6</mml:mn></mml:math></inline-formula>
<list list-type="simple">
<list-item><label>&#x2022;</label>
<p>height: 183 cm</p></list-item>
<list-item><label>&#x2022;</label>
<p>continuous blood pressure: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM20"><mml:mo>&#x003C;</mml:mo><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula></p></list-item>
</list></p></list-item>
<list-item><label>&#x2022;</label>
<p>patientID: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM21"><mml:mi>b</mml:mi><mml:mn>1</mml:mn><mml:mi>f</mml:mi><mml:mn>63</mml:mn><mml:mi>e</mml:mi><mml:mn>6</mml:mn><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mn>60</mml:mn><mml:mi>d</mml:mi></mml:math></inline-formula>
<list list-type="simple">
<list-item><label>&#x2022;</label>
<p>height: 159 cm</p></list-item>
<list-item><label>&#x2022;</label>
<p>weight: 93 kg</p></list-item>
<list-item><label>&#x2022;</label>
<p>continuous blood pressure: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM22"><mml:mo>&#x003C;</mml:mo><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula></p></list-item>
<list-item><label>&#x2022;</label>
<p>email: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM23"><mml:mi>j</mml:mi><mml:mi>o</mml:mi><mml:mi>h</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mo>.</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula></p></list-item>
</list></p></list-item>
<list-item><label>&#x2022;</label>
<p>patientID: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM24"><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi><mml:mi>f</mml:mi><mml:mn>2267</mml:mn><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:math></inline-formula>
<list list-type="simple">
<list-item><label>&#x2022;</label>
<p>continuous blood pressure: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM25"><mml:mo>&#x003C;</mml:mo><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula></p></list-item>
</list></p></list-item>
</list>Not all data records are required to contain the same types of data values, even though it is practical for them to do so in practice.</p>
</sec>
<sec id="s2b"><label>2.2</label><title>Continuity</title>
<p>Continuity can be seen as a property of the generating process. For example, the electrical activity of the heart, as measured by an ECG, is continuous according to <xref ref-type="statement" rid="st5">Definition 2.5</xref>. However, it can also be a property of the recording process. When both the generating and recording processes are continuous, the resulting data value can also be considered truly continuous. As all data recorded by a system is somehow sampled and not truly continuous, another distinction is needed. This is provided in <xref ref-type="statement" rid="st4">Definition 2.4</xref>.</p><statement id="st4"><label>DEFINITION 2.4</label><title>(Variable and equidistant step time)</title>
<p>Considering <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM26"><mml:mi>m</mml:mi></mml:math></inline-formula> data records of possibly different recordings <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM27"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM28"><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the time and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM29"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the value, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM30"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> for <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM31"><mml:mi>n</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">N</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>If <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM32"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:math></inline-formula>, then the time component of this data record has equidistant time steps; otherwise, it is considered a time attribute with variable step size. For multiple devices or recordings, the time attribute is said to have an equidistant step size if the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM33"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> is equal for all devices or recordings.</p></statement>
<p><xref ref-type="fig" rid="F2">Figure&#x00A0;2</xref> illustrates <xref ref-type="statement" rid="st4">Definition 2.4</xref>. It shows two P-waves of an ECG: the blue curve has the same time intervals between measurements and thus exhibits equidistant time steps, whereas the red curve shows slight variations in the sampling rate and thus exhibits a variable time step.</p>
<fig id="F2" position="float"><label>Figure 2</label>
<caption><p>Variable and equidistant step time.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g002.tif"><alt-text content-type="machine-generated">Graph depicting changes in millivolts (mV) over time in seconds. It features two lines: a blue line with circles for equidistant step time and a red line with triangles for variable step time, both peaking at 0.2 mV.</alt-text>
</graphic>
</fig>
<p>We define continuity for data records and their databases with with a single time axis based on the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM34"><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> criterion for continuous functions by Weierstra&#x00DF; and Jordan, as follows:</p><statement id="st5"><label>DEFINITION 2.5</label><title>(Continuity)</title>
<p>Given a data record <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM35"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM36"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> is the time domain and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM37"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> is some value domain. <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM38"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> must have subtraction and absolute operation defined on its elements. Additionally, the order operator has to define a total order on all non-negative elements of the domain.<xref ref-type="fn" rid="FN0001"><sup>1</sup></xref> For every time entry of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM39"><mml:mi>d</mml:mi></mml:math></inline-formula>, it holds that <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM40"><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>We write that <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM41"><mml:mi>d</mml:mi></mml:math></inline-formula> contains <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM42"><mml:mi>t</mml:mi></mml:math></inline-formula>, or <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM43"><mml:mi>t</mml:mi><mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:msub><mml:mi>d</mml:mi></mml:math></inline-formula>, if <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM44"><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> for some 
<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM45"><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo></mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>Let <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM46"><mml:mi>f</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo>;</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mtext>if</mml:mtext><mml:mspace width="thinmathspace"/><mml:mi>t</mml:mi><mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:mspace width="2em" /></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>i</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mtext>else</mml:mtext><mml:mspace width="2em" /></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM47"><mml:mi>i</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> 
is some sensible, domain- and application-specific interpolate. The data record <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM48"><mml:mi>d</mml:mi></mml:math></inline-formula> is continuous if <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM49"><mml:mi>f</mml:mi></mml:math></inline-formula> fulfills the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM50"><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> criterion. A database is called continuous if every record it contains is continuous.</p></statement>
<p><xref ref-type="fig" rid="F3">Figure&#x00A0;3</xref> shows an example with two QRS complexes. The blue one contains a spot, where the value jumps from 1 mV to another at the same timepoint and includes a gap in the data, where no sensible interpolation is possible. The other, in red, is continuous.</p>
<fig id="F3" position="float"><label>Figure 3</label>
<caption><p>Continuity.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g003.tif"><alt-text content-type="machine-generated">Graph comparing non-continuous (blue) and continuous (red) signals over time in seconds, measured in millivolts. The non-continuous graph shows a jump and a gap, while the continuous graph is smooth.</alt-text>
</graphic>
</fig>
<p>For most practical applications, especially the class of data considered in this work, introduced formally in <xref ref-type="statement" rid="st6">Definition 2.6</xref>, <xref ref-type="statement" rid="st5">Definition 2.5</xref> leads to <xref ref-type="statement" rid="co1">Corollary 2.1</xref>.</p><statement id="co1"><label>COROLLARY 2.1</label><title>(Practical continuity)</title>
<p>To be considered continuous according to <xref ref-type="statement" rid="st5">Definition 2.5</xref>, a data record&#x2014;and by extension, a database&#x2014;has to fulfill the following conditions:
<list list-type="simple">
<list-item><label>1.</label>
<p>The sampling rate needs to be high enough; i.e., the step size needs to be sufficiently small.</p></list-item>
<list-item><label>2.</label>
<p>A heuristic to interpolate values between every possible pair of samples can be defined.</p></list-item>
<list-item><label>3.</label>
<p>This interpolate introduces no discontinuities between the samples.</p></list-item>
</list></p></statement>
<p>Conditions 2 and 3 depend on the recorded process, whereas Condition 1 depends on the recording process. This means that someone checking the requirement can look at the recording process and decide whether Condition 1 is fulfilled. However, for Conditions 2 and 3, it is necessary to find a suitable interpolate or provide proof of its absence. If <xref ref-type="statement" rid="co1">Corollary 2.1</xref> cannot be fulfilled, the data is not considered continuous.</p>
<p>Two exemplary interpolates are introduced in <xref ref-type="disp-formula" rid="disp-formula1">Equations 1</xref> and <xref ref-type="disp-formula" rid="disp-formula2">2</xref>. The interpolate <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM51"><mml:msub><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="disp-formula1">Equation 1</xref> holds the last valid value until a new one is reached, whereas the interpolate <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM52"><mml:msub><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="disp-formula2">Equation 2</xref> applies linear interpolation between two samples. <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM53"><mml:msub><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> is not a suitable candidate for any dataset that could potentially fulfill <xref ref-type="statement" rid="co1">Corollary 2.1</xref>, as it produces jumps into the record, thus violating Condition 3 and the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM54"><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> criterion of Weierstra&#x00DF; and Jordan. In contrast, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM55"><mml:msub><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> does not produce such jumps and is thus a valid candidate.</p>
<p><xref ref-type="disp-formula" rid="disp-formula1">Equation 1</xref> is well-defined for all time entries after the first sample and is undefined for all prior time points. In contrast, <xref ref-type="disp-formula" rid="disp-formula2">Equation 2</xref> is well-defined for all time entries between the first and last sample, provided that at least two samples exist; otherwise, it is undefined.<disp-formula id="disp-formula1"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM1"><mml:mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo>;</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mtext>if</mml:mtext><mml:mspace width="thinmathspace"/><mml:mi>t</mml:mi><mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:mspace width="2em" /></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mtext>&#xA0;else&#xA0;</mml:mtext><mml:mspace width="2em" /></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula><disp-formula id="disp-formula2"><label>(2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM2"><mml:mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo>;</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mtext>if</mml:mtext><mml:mspace width="thinmathspace"/><mml:mi>t</mml:mi><mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:mspace width="2em" /></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:mrow><mml:mtext>&#xA0;else&#xA0;</mml:mtext><mml:mspace width="2em" /></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s2c"><label>2.3</label><title>Class of data</title>
<p>There is a huge variety in medical time-continuous data types, many of which share some similar properties. As the general properties and requirements are identical, we focus on a small subset of such data without the loss of generality. Thus, the data used for the proposed approach in this work is that induced by a patient monitor and formalized in <xref ref-type="statement" rid="st6">Definition 2.6</xref>. Patient monitor data is chosen due to its prevalence in both clinical setting and research databases like MIMIC IV (<xref ref-type="bibr" rid="B14">14</xref>) and eICU (<xref ref-type="bibr" rid="B15">15</xref>).</p><statement id="st6"><label>DEFINITION 2.6</label><title>(Considered class of data)</title>
<p>The class of data considered in this work is a database consisting of multiple domains, each linked by an identifier. Every domain consists of a time dimension and a value dimension, i.e., the domain <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM56"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula>. Both dimensions have to fulfill continuity criteria according to <xref ref-type="statement" rid="st5">Definition 2.5</xref>. The value dimension typically consists of a series of real values, which are generated by events resampling a waveform, e.g., an ECG, and sometimes represents a numeric value, e.g., SPO<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM57"><mml:msub><mml:mi></mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>.</p></statement>
<p>Many databases of this type have already been published. For example, MIMIC IV v.3.1 contains data from 364,627 individuals and 546,028 hospitalizations from emergency and intensive care units (<xref ref-type="bibr" rid="B14">14</xref>). Another example is the PTB-XL database, which contains 21,799 clinical 12-lead ECGs from 18,869 individuals, each lasting 10&#x2009;s. Every ECG in the PTB-XL database is annotated by one or two cardiologists (<xref ref-type="bibr" rid="B16">16</xref>, <xref ref-type="bibr" rid="B17">17</xref>). One example of a database with a broader scope is eICU. It collects data from Philips eICU monitors in multiple intensive care units across the United States. It contains more than 200,000 admissions (<xref ref-type="bibr" rid="B15">15</xref>), with information from drug admission records, standardized care-taker notes, ventilation data, and aperiodic vital parameters, such as pulmonary artery occlusion pressure (<xref ref-type="bibr" rid="B18">18</xref>).</p>
<p>The data captured by a patient monitor is summarized in <xref ref-type="table" rid="T1">Table&#x00A0;1</xref>, and we describe its details in the following. A patient monitor usually displays an ECG, which records the electrical activity in the heart as electrical potential on the orders of a few millivolts. The ECG is measured using electrodes placed on the skin, whose number and position may vary. Depending on the electrode&#x2019;s configuration, it is possible to have different leads, which can sometimes be monitored simultaneously. This allows detailed monitoring of the entire cardiac cycle. Additionally, heart frequency (HF), measured in bpm, reflects the fluctuations in the electrical activity of the heart using an ECG. HF should not be confused with heart rate (HR), which measures the actual contractions of the heart. HF is mostly identical to the pulse, measured in bpm, and can be detected at any peripheral body point using a pulse oximeter. This sensor also measures the peripheral oxygen saturation (SPO<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM58"><mml:msub><mml:mi></mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>) as a percentage. Breathing frequency (BF), expressed in <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM59"><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mfrac></mml:mrow></mml:math></inline-formula>, represents the respiratory rate and can also be estimated by an ECG. Furthermore, both non-invasive blood pressure (NIBP) and invasive blood pressure (IBP) can be displayed. In medical settings, both pressures are expressed in mmHg as the unit rather than the SI unit, Pa. Both include systolic pressure, i.e., the peak pressure during a heartbeat, and diastolic pressure, i.e., the minimal pressure between two heartbeats. IBP can be measured continuously, whereas NIBP is recorded only a few times per hour.</p>
<table-wrap id="T1" position="float"><label>Table 1</label>
<caption><p>Patient monitor data types.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Abbreviation</th>
<th valign="top" align="left">Unit</th>
<th valign="top" align="left">Name/description</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ECG</td>
<td valign="top" align="left">mV</td>
<td valign="top" align="left">Electrocardiogram</td>
</tr>
<tr>
<td valign="top" align="left">HF</td>
<td valign="top" align="left">bpm</td>
<td valign="top" align="left">Heart frequency</td>
</tr>
<tr>
<td valign="top" align="left">Pulse</td>
<td valign="top" align="left">bpm</td>
<td valign="top" align="left">Peripheral pulse</td>
</tr>
<tr>
<td valign="top" align="left">SPO<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM60"><mml:msub><mml:mi></mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left">&#x0025;</td>
<td valign="top" align="left">Peripheral oxygen saturation</td>
</tr>
<tr>
<td valign="top" align="left">BF</td>
<td valign="top" align="left"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM61"><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mfrac></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">Breathing frequency</td>
</tr>
<tr>
<td valign="top" align="left">NIBP</td>
<td valign="top" align="left">Pa or mmHg</td>
<td valign="top" align="left">Non-invasive blood pressure</td>
</tr>
<tr>
<td valign="top" align="left">IBP</td>
<td valign="top" align="left">Pa or mmHg</td>
<td valign="top" align="left">Invasive blood pressure</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As many of the examples in this paper are based on ECG data, the fundamental components of a physiological ECG are described in the following paragraph, as introduced in Kardiovaskul&#x00E4;re Medizin Online (<xref ref-type="bibr" rid="B19">19</xref>) and Cardiovascular Medicine (<xref ref-type="bibr" rid="B20">20</xref>). An ECG waveform can be split into multiple segments, which are usually labeled P, Q, R, S, T, and sometimes U. These segments are shown in <xref ref-type="fig" rid="F5">Figure&#x00A0;5</xref>. First, a small &#x201C;hill&#x201D; appears, called the P-wave, lasting roughly 0.1 s, followed by a horizontal line referred to as the PQ segment. The PQ segment serves as the reference or baseline for the ECG. The interval from the beginning of the P-wave to the end of the PQ segment is called the PQ interval. Afterward, a small, sharp dip labeled Q is observed, followed by a large, sharp spike, called the R-wave. This is followed by the S-wave, a downward deflection that is usually slightly deeper than Q. The two downward deflections Q and S below the baseline together with the large spike R dominate the appearance of an ECG and are often referenced together as the QRS complex. After S, the signal rises relatively smoothly to the baseline and is followed by a horizontal line called the ST segment. The beginning of this line, i.e., the point at which the ECG returns to the baseline, is called the J (junction) point. The ST segment is then followed by a T-wave. After returning to the baseline, a small U-wave might appear. The U-wave does not appear for all individuals, and its cause is not well known. The QT interval is measured from the start of the QRS complex to the end of the T-wave. It represents the time from ventricular depolarization to full repolarization. As it depends on the current heart frequency, it is usually corrected for heart frequency, which is called the QTc interval. The segment from the end of the T-wave to the start of the next Q-wave is called the TP interval. These segments and waves can be mapped onto the myocardial action potential and thus provides insights into the various stages of the cardiac cycle.</p>
</sec>
<sec id="s2d"><label>2.4</label><title>Anonymity</title>
<p>Anonymity can be seen as the absence of information that would otherwise allow the linkage of a record to an individual. One of the main tasks of anonymization is to ensure privacy by breaking such linkability while preserving the relationship between the time and value domains. Additionally, correlations between multiple domains, e.g., between blood pressure and certain cardiovascular diseases, should also be preserved. Data collected from a patient monitor exhibits strong inter-domain connections and possibly correlations. For instance, an obstruction in a tube during mechanical ventilation may be noticeable in both peripheral oxygen saturation (SPO<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM62"><mml:msub><mml:mi></mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>) and inspiratory pressure (P<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM63"><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>).</p>
<p>Even after the removal of direct identifiers, linking may of course be possible using information about an individual from the remaining data in a database record. Oprea et al. (<xref ref-type="bibr" rid="B21">21</xref>) demonstrated that classifications on time-continuous medical data are possible with a reasonable degree of accuracy, with a particular focus on explainable AI (XAI). In their study, the type of breath (spontaneous, mechanical, triggered) was classified using the airway flow and pressure, achieving accuracies of 79.78&#x0025; for spontaneous, 81.69&#x0025; for mechanical, and 77.05&#x0025; for triggered breath. This automatic detection and classification can be used to extract information about an individual that is not directly provided, e.g., whether a patient was dependent on mechanical ventilation. Such a classification could be used to extract discrete information for an attack on the privacy of the individual.</p>
<p>To address such risks, anonymity must first be formalized, beginning with a clear definition of the goals. A key goal is to limit the risk of disclosure, as defined by Samarati and Sweeney (<xref ref-type="bibr" rid="B2">2</xref>) and Sweeney (<xref ref-type="bibr" rid="B22">22</xref>), i.e., the unintended release of explicit or inferable information about a person.</p>
<p>Disclosure risk can be categorized into three levels: identity disclosure, attribute disclosure, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM64"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> breaches. Identity disclosure occurs when explicit identifiers, e.g., a patient&#x2019;s name or address, are present in the database or can be inferred. This inferability of identifiers might happen through external data sources, which can be linked to the record.</p>
<p>Another form of disclosure is attribute disclosure, which occurs when the precise value of an attribute, e.g., a patient&#x2019;s diagnosis, is disclosed to an attacker. The attacker does not need to identify the victim&#x2019;s record within the database to achieve this. For example, if all individuals of an equivalence class of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM65"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity share the same value for one attribute, an attacker only needs to know that the victim must be within that equivalence class to achieve attribute disclosure. Similarly, the value can be reconstructed for deterministically created overlapping regions of equivalence classes, as mentioned in Cao et al. (<xref ref-type="bibr" rid="B5">5</xref>).</p>
<p>Partial attribute disclosure can occur when the distribution of attributes in an equivalence class does not follow the distribution in the population closely. This means that an attacker learns a certain value is far more (or less) likely than others for the attribute of the victim. This was a key consideration behind the development of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM66"><mml:mi>l</mml:mi></mml:math></inline-formula>-diversity by Machanavajjhala et al. (<xref ref-type="bibr" rid="B4">4</xref>) and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM67"><mml:mi>t</mml:mi></mml:math></inline-formula>-closeness by Li et al. (<xref ref-type="bibr" rid="B3">3</xref>). In a <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM68"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> breach, neither the identity nor an attribute is disclosed to an adversary, but the belief about the victim changes fundamentally after inspecting the data, as shown by Evfimievski et al. (<xref ref-type="bibr" rid="B23">23</xref>).</p>
<p>The ambition of absolute disclosure prevention, i.e., achieving anonymity through semantic security as foreshadowed by Dalenius (<xref ref-type="bibr" rid="B24">24</xref>) and formalized by Goldwasser and Micali (<xref ref-type="bibr" rid="B25">25</xref>), is not very useful, as this anonymity cannot be achieved when additional data sources can be used by the attacker, as shown by Dwork (<xref ref-type="bibr" rid="B9">9</xref>). Semantic security in the context of encryption, as defined by Goldwasser and Micali (<xref ref-type="bibr" rid="B25">25</xref>), means that the encryption scheme is secure if no information about the plaintext can be learned by looking at the ciphertext. Similarly, Dwork (<xref ref-type="bibr" rid="B9">9</xref>) proposed the concept of semantic security in a privacy setting: nothing about an individual should be learnable by looking at the results of queries on an anonymized database. Thus, the corresponding privacy notion of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM69"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula>-DP, as defined by Dwork (<xref ref-type="bibr" rid="B9">9</xref>), is used in this paper. The intuition is that at most <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM70"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> is learnable about an individual when looking at the protected responses to queries on the database containing sensitive information.</p>
</sec>
<sec id="s2e"><label>2.5</label><title>Measuring utility</title>
<p>Utility quantifies the quality with which the target operation can be performed on the protected data after anonymization. More formally, it measures how good cross-attribute correlations work on an anonymized database [cf. (<xref ref-type="bibr" rid="B26">26</xref>)]. The necessity for an empirical metric for the utility was highlighted by Iyengar (<xref ref-type="bibr" rid="B27">27</xref>), LeFevre et al. (<xref ref-type="bibr" rid="B28">28</xref>), and Wang et al. (<xref ref-type="bibr" rid="B29">29</xref>). Achieving &#x201C;good&#x201D; utility is hard, in general, as the work to be performed on the data is often not known beforehand. If the work is known in advance, the institution responsible for anonymization can simply execute this work on the unanonymized data and release the results instead. Thus, the goal is to develop a workload-independent metric, which works for a wide range of applications. This requirement for the utility was already highlighted by Brickell and Shmatikov (<xref ref-type="bibr" rid="B26">26</xref>).</p>
<p>Utility can be measured by syntactic and semantic metrics. An example of a syntactic metric is the information loss, measured by the amount of generalization or suppression applied, as presented by Hammer et al. (<xref ref-type="bibr" rid="B30">30</xref>). Other syntactic metrics are presented by Machanavajjhala et al. (<xref ref-type="bibr" rid="B4">4</xref>), which include the mean size of quasi-identifier equivalence classes, the sum of squares of the class sizes, and the number of generalization steps performed.</p>
<p>Semantic metrics can be categorized as either workload-independent or workload-dependent. A workload-independent metric quantifies the &#x201C;harm&#x201D; done to the data by the anonymization process itself, but it does not measure the remaining utility of the data. As already mentioned, the precise workload is not known to the anonymizing party. Thus, it cannot optimize against the application-specific utility. Nevertheless, LeFevre et al. (<xref ref-type="bibr" rid="B28">28</xref>) proposed a workload-specific anonymization approach that aims to optimize utility for one or multiple classes of workloads.</p>
<p>This leads to a general definition of utility, as formalized in <xref ref-type="statement" rid="st7">Definition 2.7</xref>. The methods to quantify the probabilities depend on the specific workload class; in many cases, these probabilities cannot be determined beforehand or need manual evaluation by a professional. Thus, approximations of these definitions will be needed in many instances.</p><statement id="st7"><label>DEFINITION 2.7</label><title>(Utility)</title>
<p>Given the data record <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM71"><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi></mml:math></inline-formula> of database <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM72"><mml:mi>S</mml:mi></mml:math></inline-formula> and the anonymization function <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM73"><mml:mi>&#x03B1;</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:msub></mml:math></inline-formula>, which maps a data record to the anonymized data record. Due to the lack of an objectively correct decision, a decision is treated as correct if it is made based on the raw data. Respectively, any deviation between a decision made on the anonymized data and the decision made on the raw data is treated as an error. Making a decision <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM74"><mml:mi>d</mml:mi></mml:math></inline-formula> based on <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM75"><mml:mi>e</mml:mi></mml:math></inline-formula> is denoted <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM76"><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Let <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM77"><mml:mi>f</mml:mi></mml:math></inline-formula> be the decision-correctness function, which returns <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM78"><mml:mn>1</mml:mn></mml:math></inline-formula> if a decision based on the raw data is the same as the decision based on the anonymized data. Thus,<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM1"><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mspace width="2em" /></mml:mtd><mml:mtd><mml:mtext>&#xA0;if&#xA0;</mml:mtext><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>e</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mspace width="2em" /></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mspace width="2em" /></mml:mtd><mml:mtd><mml:mtext>&#xA0;else&#xA0;</mml:mtext><mml:mspace width="2em" /></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>This leads to the definition of the utility <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM79"><mml:mi>U</mml:mi></mml:math></inline-formula> for a set of decisions <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM80"><mml:mi>D</mml:mi></mml:math></inline-formula><disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM2"><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>D</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:munder><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mi>D</mml:mi><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula></p></statement>
<p>Syntactic methods for privacy, such as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM81"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity (<xref ref-type="bibr" rid="B2">2</xref>), <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM82"><mml:mi>l</mml:mi></mml:math></inline-formula>-diversity [cf. (<xref ref-type="bibr" rid="B4">4</xref>)], or <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM83"><mml:mi>t</mml:mi></mml:math></inline-formula>-closeness (<xref ref-type="bibr" rid="B3">3</xref>), focus on obscuring the individual&#x2019;s identity in the dataset by making it indistinguishable from several others. In contrast, semantic methods aim to protect the sensitive attributes themselves. For example, in the case of differential privacy (<xref ref-type="bibr" rid="B9">9</xref>), this is done by adding random noise to obscure any information that is specific to an individual in the dataset.</p>
<p>Similar to privacy, utility can also be defined syntactically or semantically. For example, the information loss is syntactical, whereas a metric that measures the accuracy of a query is semantic. In time-continuous medical applications, the syntax-based utility metric from Hammer et al. (<xref ref-type="bibr" rid="B30">30</xref>) measures information loss as the average width of anonymized data. This is the range of each equivalence class; it is zero for raw data and grows with anonymization unless the value fits perfectly. The raw data, e.g., an ECG, is recorded and treated as the ground truth. The information loss then quantifies the new thickness of the ECG in relation to the overall domain width. If the data is anonymized using <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM84"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity for time-continuous data, as presented in Hammer et al. (<xref ref-type="bibr" rid="B30">30</xref>), the information loss is the average width of the equivalence class divided by the domain width.</p>
<p>There are usually two types of privacy models for DP: interactive and non-interactive. An interactive privacy model allows the user to submit queries to the system, which are then answered in a privacy-preserving manner. As the query can be performed on both the raw and sanitized data, it enables calculating the accuracy of each answer. The user is then given a measure of utility along with the answer to the query based on the sanitized data. In the non-interactive privacy model, the system sanitizes and releases the data to the public, thus allowing the uncontrolled execution of queries by the user post-protection.</p>
<p>As the releasing party does not anticipate the queries to be performed and users do not have access to the raw data, the accuracy cannot be calculated directly. Thus, only an exemplary measure of the accuracy of some previously known queries could be given to users who need access to the data. Alternatively, a utility metric can measure the preservation of patterns and trends in the data, which are needed to perform meaningful analyses and derive insight. This is highly application-specific, as the relevance of patterns depends not only on the analyzed data but also on the questions a user wants to address through the analysis.</p>
<p>One could think about alternatives to the proposed <xref ref-type="statement" rid="st7">Definition 2.7</xref> and come up with syntactic utility metrics. However, these approaches do not consider the underlying structure and meaning of the data. While syntactic metrics are relatively easy to implement and use, they offer limited insight into the usefulness of the anonymized data, which is the goal of a utility metric. Thus, such metrics are not well-suitable for the intended purpose. <xref ref-type="statement" rid="st7">Definition 2.7</xref> can only be applied to known and formalized queries. Furthermore, the same set of queries needs to be performed on the raw data; thus, such a definition is only feasible for interactive models or a selected set of benchmark queries performed before data release.</p>
<p>Non-interactive models have somewhat different requirements. As the workload is not known beforehand and <xref ref-type="statement" rid="st7">Definition 2.7</xref> can only be applied a posteriori, an alternative approach is needed for some applications. A precise calculation of the usefulness of the data to all possible classes of workload is not possible. As a result, only an estimate of the usefulness of the data is possible, which is introduced in <xref ref-type="statement" rid="st8">Definition 2.8</xref>. Such a metric can also be relevant for interactive approaches, as it provides an estimation of the utility before executing queries, thus reducing the possibility of spending time and resources on the work with an unsuitable database.</p><statement id="st8"><label>DEFINITION 2.8</label><title>(Heuristic utility)</title>
<p>This definition aims to approximate utility according to <xref ref-type="statement" rid="st7">Definition 2.7</xref>. Given a database <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM85"><mml:mi>S</mml:mi><mml:mo>&#x2282;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> with functions <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM86"><mml:msub><mml:mi>U</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>S</mml:mi><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM87"><mml:msub><mml:mi>U</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> is the utility according to <xref ref-type="statement" rid="st7">Definition 2.7</xref> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM88"><mml:msub><mml:mi>U</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula> is the utility using a heuristic; then, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM89"><mml:msub><mml:mi>U</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x2248;</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>x</mml:mi><mml:mo>&#x2282;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>.</p>
<p>The precise relation is often application-specific, as is the heuristic utility itself. <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM90"><mml:msub><mml:mi>U</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula> is determined by the underlying heuristic <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM91"><mml:mi>h</mml:mi></mml:math></inline-formula>; however, in general, the result must be normalized similar to <xref ref-type="statement" rid="st7">Definition 2.7</xref>; thus, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM92"><mml:msub><mml:mi>U</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo></mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:math></inline-formula>.</p></statement>
<p><xref ref-type="statement" rid="st8">Definition 2.8</xref> can use semantic or syntactic heuristics. Promising approaches for the approximation of utility include metrics such as accuracy, precision, and recall of a query, mean relative error (MRE), and the preservation of patterns.</p>
<p>The MRE is commonly defined for recorded values <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM93"><mml:mi>y</mml:mi></mml:math></inline-formula> and the predicted or, in our case, sanitized values <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM94"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>:<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM3"><mml:mtext>MRE</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>For non-continuous data, accuracy and other relevance metrics, like precision and recall, are relatively straightforward to calculate. In the following, all necessary metrics required to adapt the MRE for use with continuous data are introduced. This adaptation could then be used to define accuracy, precision, recall, or other candidate metrics for utility.</p>
<p>One challenge is the quantification of distances between continuous data values, e.g., between curves. A possible candidate for such a metric is the Hausdorff distance <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM95"><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Intuitively, it measures how much two subsets of a metric space <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM96"><mml:mi>X</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM97"><mml:mi>Y</mml:mi></mml:math></inline-formula> must be thickened to fully contain the other (<xref ref-type="bibr" rid="B31">31</xref>). As an example, consider the domain of ECGs, defined over time <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM98"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> voltage. Let there be two ECGs, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM99"><mml:msub><mml:mi>E</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM100"><mml:msub><mml:mi>E</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, then <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM101"><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the minimal <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM102"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> such that making <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM103"><mml:msub><mml:mi>E</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> wider by <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM104"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> contains <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM105"><mml:msub><mml:mi>E</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, and vice versa. This can lead to very large values in the case of misaligned or otherwise transformed subsets, even when the general shapes are similar.</p>
<p>The Gromov&#x2013;Hausdorff distance expands upon the concept of the Hausdorff distance by iterating over all possible isometric embeddings of two subsets into a common third space, which allows for transformations of the subsets (<xref ref-type="bibr" rid="B31">31</xref>). In the context of the ECG example from the Hausdorff distance, both <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM106"><mml:msub><mml:mi>E</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM107"><mml:msub><mml:mi>E</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> would be transformed into a common isometric embedding. <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM108"><mml:msub><mml:mi>E</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM109"><mml:msub><mml:mi>E</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> could be rotated or shifted. Afterward, the Hausdorff distance is calculated on these embedded <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM110"><mml:msub><mml:mi>E</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM111"><mml:msub><mml:mi>E</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>. The Gromov&#x2013;Hausdorff distance is defined as the infimum of these calculated Hausdorff distances over all possible embeddings.</p>
<p>So far, we only measure the distance between two subsets, without considering the general shape and other properties. This can be done when using the Fr&#x00E9;chet distance, which calculates the similarity of two curves (<xref ref-type="bibr" rid="B32">32</xref>). Intuitively, one can imagine a dog and its owner, connected via a leash. The dog walks along curve <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM112"><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and its owner along <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM113"><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>; both can independently walk forward or stop, but neither can go back. The shortest possible leash length that allows both to reach the end of their curve is the Fr&#x00E9;chet distance (<xref ref-type="bibr" rid="B33">33</xref>). This quantifies the similarity of curves.</p>
<p>Both the Hausdorff and Fr&#x00E9;chet distances measure similarity, or better the lack thereof, and thus their measurements carry semantic meaning. However, they approach this from different perspectives: while the Hausdorff distance measures the proximity between sets in a metric space, the Fr&#x00E9;chet distance measures the similarity of shape and trend between curves or trajectories. Thus, the latter is far more suitable for the intended application. However, calculating the Fr&#x00E9;chet distance is computationally expensive. Alt and Godau (<xref ref-type="bibr" rid="B32">32</xref>) mentioned the runtime for an exact calculation with two polygons with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM114"><mml:mi>p</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM115"><mml:mi>q</mml:mi></mml:math></inline-formula> segments as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM116"><mml:mrow><mml:mi mathvariant="script">O</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mi>q</mml:mi><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Even the approximate version presented by Eiter and Mannila (<xref ref-type="bibr" rid="B33">33</xref>) has a runtime of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM117"><mml:mrow><mml:mn mathvariant="script">0</mml:mn></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>.</p>
<p>The Fr&#x00E9;chet distance is still prone to producing high measurements due to misalignments or other transformation-induced errors. One could investigate the possibility of extending it similarly to the Gromov&#x2013;Hausdorff distance, i.e., by taking the minimum overall isometric embeddings for all Fr&#x00E9;chet distances. Still, a small number of spikes or valleys, which do not match from one curve to the other, can drastically inflate the result. This can be addressed by detecting such erroneous regions and omitting them up to a certain threshold. Then, the Fr&#x00E9;chet distance can be calculated independently on each of the resulting segments, with the results combined to form a final similarity score. The detection of such regions and the calculation of the corresponding threshold for omission could be performed using an approximation for the geometric edit distance, similar to the one proposed by Andoni and Onak (<xref ref-type="bibr" rid="B34">34</xref>), in a sliding window manner or through shape matching based on the skeleton of the shape using the edit distance, as introduced by Klein et al. (<xref ref-type="bibr" rid="B35">35</xref>). While the approximation by Andoni and Onak (<xref ref-type="bibr" rid="B34">34</xref>) operates in near-linear time, it is only intended for strings. An approximation of the geometric edit distance is still somewhat expensive with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM118"><mml:mrow><mml:mi mathvariant="script">O</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> (<xref ref-type="bibr" rid="B36">36</xref>). The shape matching variant by Klein et al. (<xref ref-type="bibr" rid="B35">35</xref>) is even more expensive. Thus, both the Fr&#x00E9;chet distance and the Edit distance are computationally expensive.</p>
<p>For comparisons among multiple pairs of curves, both the basic Fr&#x00E9;chet distance and its proposed extensions require normalization. An intuitive way to normalize for a given set of curve pairs is to calculate the sum of all distances and divide each distance by this sum.</p>
<p>Suitable candidate metrics to quantify the distance between trajectories along time are identified by Su et al. (<xref ref-type="bibr" rid="B37">37</xref>). This paper evaluates and categorizes multiple metrics for trajectories. It identifies three of them as spatio-temporal and continuous metrics. These metrics are suitable candidates for the challenges presented in this paper. Besides the Fr&#x00E9;chet distance, the spatio-temporal Euclidean distance measure (STED), as proposed by Nanni and Pedreschi (<xref ref-type="bibr" rid="B38">38</xref>), was evaluated by Su et al. (<xref ref-type="bibr" rid="B37">37</xref>). STED quantifies the Euclidean distance between the curves along time, normalized by their length. However, this metric does not consider the shape of the curves or potential alignment or transformation issues.</p>
<p>In the following, we formalize one approach rigorously as an example. For this, it is assumed that the answer to a query for interactive anonymization, e.g., DP, is made based on the recorded data, and the sanitization is applied afterward. This means that both a query without anonymization and one with it applied contain the same set of curves; however, in the case of DP, the later set has added some carefully tuned noise to each curve.</p>
<p>In situations where this assumption is not given, the following calculation would only hold on the intersection of the sanitized and unsanitized queries. One could then provide accuracy as a tuple of this result and the number of curves that are missing in one of the sets.</p>
<p>To define the MRE for interactive approaches on time-continuous data, first, the difference between the curves is formalized according to <xref ref-type="statement" rid="st9">Definition 2.9</xref>.</p><statement id="st9"><label>DEFINITION 2.9</label><title>(Curve difference)</title>
<p>Given a query <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM119"><mml:mi>q</mml:mi></mml:math></inline-formula>, the database <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM120"><mml:mi>S</mml:mi></mml:math></inline-formula>, the anonymization function <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM121"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, and the Fr&#x00E9;chet distance <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM122"><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Additionally, let <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM123"><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> be the transformation analog to the one proposed for the Gromov&#x2013;Hausdorff distance, which transforms <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM124"><mml:mi>x</mml:mi></mml:math></inline-formula> as close to <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM125"><mml:mi>y</mml:mi></mml:math></inline-formula> as possible. Let <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM126"><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> be the mapping, which was informally introduced in the previous paragraph. It corrects<xref ref-type="fn" rid="FN0002"><sup>2</sup></xref> smaller regions of a mismatch from both curves <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM127"><mml:mi>x</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM128"><mml:mi>y</mml:mi></mml:math></inline-formula>, with a maximum size of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM129"><mml:mi>&#x03C9;</mml:mi></mml:math></inline-formula>, up to a given threshold <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM130"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>, using the edit distance. This leads to the definition of the difference between two curves:<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM4"><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></disp-formula></p></statement>
<p>A norm <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM131"><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo>&#x2218;</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo></mml:math></inline-formula> on curves has to be formalized to be able to adapt the MRE to time-continuous medical data. This is done in <xref ref-type="statement" rid="st10">Definition 2.10</xref>. The fast Fourier transformation (FFT) maps curves to frequency space, decomposing complex waves into a structured vector of frequencies. Applying a <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM132"><mml:mi>p</mml:mi></mml:math></inline-formula>-norm to this vector provides a stable, mathematically consistent curve norm. Other norms could also be used; the given definition is just one candidate.</p><statement id="st10"><label>DEFINITION 2.10</label><title>(Curve norm)</title>
<p>First, the FFT is used to map the curve to a vector of its frequencies. Afterward, the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM133"><mml:mi>p</mml:mi></mml:math></inline-formula>-norm for some <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM134"><mml:mn>2</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula> is applied. If <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM135"><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the function that maps a given curve to the vector of its frequencies, then the curve norm is<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM5"><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mi>c</mml:mi><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mi>p</mml:mi></mml:msub></mml:math></disp-formula></p></statement>
<p>This allows defining the MRE for time-continuous data according to <xref ref-type="statement" rid="st11">Definition 2.11</xref>.</p><statement id="st11"><label>DEFINITION 2.11</label><title>(Time-continuous MRE)</title>
<p>Instead of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM136"><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM137"><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> is used as a notation. The results of a query are noted as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM138"><mml:msub><mml:mi>R</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM139"><mml:msub><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. A normalization term for the summation is needed:<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM6"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:munder><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></disp-formula>This is used to define the normalized time-continuous MRE:<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM7"><mml:mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>M</mml:mi><mml:mi>R</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msub></mml:mfrac></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mspace width="1em" /><mml:mo>&#x22C5;</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></statement>
<p>The components used to define the MRE on curves can also be used for accuracy, precision, and recall.</p>
<p>A potential approach to reduce computational complexity is to map the data into a lower-dimensional space; this is similar to SABRE-AK how uses the Hilbert space-filling curve to approximate nearest neighbors in a multidimensional feature space (<xref ref-type="bibr" rid="B7">7</xref>).</p>
<p>All these candidate metrics do not consider the semantic importance of missing or artificially introduced features. For example, for some metrics, many small changes might lead to a similar utility score, as if only a single large change, which impacts the deduced knowledge, would occur. Thus, a semantic metric should be considered, preferably one that considers application-specific requirements.</p>
</sec>
<sec id="s2f"><label>2.6</label><title>Measuring privacy</title>
<p>Besides utility, privacy is another important property of an anonymization mechanism. It quantifies the level of security an individual in the database achieves. <xref ref-type="statement" rid="st12">Definition 2.12</xref> considers not only reidentification but also the general risk an individual faces by being part of a database. According to Dwork (<xref ref-type="bibr" rid="B9">9</xref>), privacy is quantifies by what can be learned about an individual through data analysis, or more precisely the lack thereof. This privacy loss, e.g., with DP, is usually represented as a function of the privacy budget <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM140"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> and the number of allowed queries to the data <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM141"><mml:mi>a</mml:mi></mml:math></inline-formula>.</p><statement id="st12"><label>DEFINITION 2.12</label><title>(Privacy (<xref ref-type="bibr" rid="B9">9</xref>))</title>
<p>For some <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM142"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM143"><mml:mi>a</mml:mi></mml:math></inline-formula>, the probability of making a decision based on any attribute of the database <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM144"><mml:mi>S</mml:mi></mml:math></inline-formula> and an individual <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM145"><mml:mi>i</mml:mi></mml:math></inline-formula> is quantified by <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM146"><mml:mi>P</mml:mi><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. Let the database <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM147"><mml:mi>S</mml:mi><mml:mo>&#x2216;</mml:mo><mml:mi>i</mml:mi></mml:math></inline-formula> be the same as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM148"><mml:mi>S</mml:mi></mml:math></inline-formula>, but <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM149"><mml:mi>i</mml:mi></mml:math></inline-formula> is not included. The privacy level is<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM8"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mo>&#x2216;</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula></p></statement>
<p>When looking at <xref ref-type="statement" rid="st12">Definition 2.12</xref>, one can note that removing the individual of interest cannot add information about them to the dataset. Therefore, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM150"><mml:mi>P</mml:mi><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2265;</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo>&#x2216;</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>; thus, the privacy <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM151"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. The closer the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM152"><mml:mi>P</mml:mi></mml:math></inline-formula> is to <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM153"><mml:mn>1</mml:mn></mml:math></inline-formula>, the less information can be gained about an individual by analyzing the data. Quantifying such probabilities remains an application-specific open task.</p>
<p>Now, all basic concepts and necessary notations for time-continuous anonymization have been introduced. Thus, these can be used in the following section to deduce relevant properties and analyze them formally for multiple classes of mechanisms.</p>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Applying syntactic mechanisms to medical data</title>
<p>With the definitions from <xref ref-type="sec" rid="s2">Section 2</xref> and the intended application proposed in <xref ref-type="sec" rid="s1">Section 1</xref>, we can now derive requirements for any anonymization mechanism intended for time-continuous medical data. <xref ref-type="sec" rid="s3a">Section 3.1</xref> presents some general threats to privacy and a more detailed look into potential attacks specific to time-continuous data. The requirements are formalized to properties of the mechanisms and investigated in <xref ref-type="sec" rid="s3b">Section 3.2</xref>. These are then used to define some properties of anonymization mechanisms, which are investigated and proven for multiple classes of mechanisms in <xref ref-type="sec" rid="s3c">Section 3.3</xref>. This provides a strong formal set of properties such a mechanism must fulfill.</p>
<sec id="s3a"><label>3.1</label><title>Threats to anonymity</title>
<p>This section discusses some general types of attacks on privacy and some specific to time-continuous data. There are multiple types of privacy threats. Reidentification is a general attack vector for anonymized data, achievable through linking or background knowledge attacks. This applies not only to time-continuous medical scenarios.</p>
<p><xref ref-type="fig" rid="F4">Figure&#x00A0;4</xref> shows multiple cycles of an ECG from a patient with a pacemaker. The data is taken from Wikimedia Commons (<xref ref-type="bibr" rid="B39">39</xref>) and adapted to include a spike before the Q-wave.</p>
<fig id="F4" position="float"><label>Figure 4</label>
<caption><p>Example&#x2014;ECG reidentification.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g004.tif"><alt-text content-type="machine-generated">Electrocardiogram (ECG) trace showing voltage in millivolts (mV) on the vertical axis and time in seconds on the horizontal axis. Peaks and troughs represent heartbeats over a span of 2.5 seconds.</alt-text>
</graphic>
</fig>
<p>The ECG in <xref ref-type="fig" rid="F4">Figure&#x00A0;4</xref> looks rather similar to the optimal one, i.e., a typical physiological ECG in <xref ref-type="fig" rid="F5">Figure&#x00A0;5</xref>, with most of the differences attributable to physiological and interpersonal variations. However, the additional spikes before the Q-waves are identifiable as a pacemaker with ventricular stimulation. Even without this spike, an attacker with background knowledge about their victim could identify the ECG or, at least, narrow down the set of potential targets.</p>
<fig id="F5" position="float"><label>Figure 5</label>
<caption><p>Example&#x2014;<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM154"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> linking.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g005.tif"><alt-text content-type="machine-generated">Electrocardiogram (ECG) graph displaying characteristic waves: P, Q, R, S, T, and U against time in seconds and voltage in millivolts. The R peaks are prominent, with labeled intervals showing &#x0394;t&#x02081; between the two R waves.</alt-text>
</graphic>
</fig>
<p>Moss (<xref ref-type="bibr" rid="B40">40</xref>) explained how the biological sex of a patient can be determined using the ECG, which reflects the morphology of the heart. Thus, the R-wave can be used to classify ECGs by sex, as done by Nolin-Lapalme et al. (<xref ref-type="bibr" rid="B41">41</xref>). This can be done without any background knowledge about the target. If background knowledge is present, even more attack vectors become viable. For example, an attacker knows that their victim has a pacemaker implanted and suffers from a decreased cardiac output and feels constantly exhausted; this leads to an S&#x2013;T ratio close to one or larger.</p>
<p>Additionally, an inference attack can be performed to statistically deduct behavioral information about an individual; again, this works for most types of data. Such an attack can be used to infer the value of an attribute or the membership of the victim to the database (<xref ref-type="bibr" rid="B42">42</xref>).</p>
<p>In the context of time-continuous medical data, reidentification could be performed by correlating different detected events or through longitudinal analysis linking multiple treatments to datasets. This enables temporal or treatment event linkage. As an example, database <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM155"><mml:mi>S</mml:mi></mml:math></inline-formula> with multiple records <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM156"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is given. Some records might belong to the same person using the aforementioned approaches, and an attacker might be able to link those records and potentially reidentify the individual.</p>
<p>In the following, a record <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM157"><mml:mi>i</mml:mi></mml:math></inline-formula> belonging to person <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM158"><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> will now be labeled as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM159"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>. The first time-continuous data-specific approach allows correlating different detected events.</p>
<p>The attacker has to detect events in the record. An event is not necessarily a complex medical diagnosis but potentially something like a recurring spike in the data or the QRS part of an ECG.</p>
<p>An assumption is that the time spans between similar recurring events can be used to identify an individual across multiple data records. In the simplest case, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM160"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x2248;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> for two detected events in <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM161"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:msub><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM162"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:msub><mml:mn>2</mml:mn><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>. Of course, more complex patterns of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM163"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> values can occur, which might not fully match. Even with some uncertainty, a linkage might still be possible. Thus, this temporal linkage via correlation of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM164"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mspace width="thinmathspace"/><mml:mi>t</mml:mi></mml:math></inline-formula> for events presents a viable attack vector.</p>
<p><xref ref-type="fig" rid="F5">Figure&#x00A0;5</xref> shows an exemplary ECG curve. It is based on the data from Commons (<xref ref-type="bibr" rid="B43">43</xref>), but a U-wave was added under the assumption that if a U-wave occurs it is usually 25&#x0025; of the T-wave in amplitude (<xref ref-type="bibr" rid="B19">19</xref>).</p>
<p>In this example, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM165"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the difference between the two R-waves of the ECG. If the attacker can find the same <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM166"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> in another record, they can link those records together. Additionally, such a linkage can be verified or narrowed down, especially if multiple candidates for the linkage exist, using specific features. The existence of the U-wave, which does not occur for every person, would be good verification for the linkage.</p>
<p>In other words, for a database of ECGs <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM167"><mml:mi>S</mml:mi></mml:math></inline-formula> containing ECG <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM168"><mml:mi>e</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi></mml:math></inline-formula> and given a reference ECG <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM169"><mml:msub><mml:mi>e</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> known to belong to the victim, an attacker narrows down the set of potential ECGs linked to the victim. This is done by creating the set <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM170"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mspace width="thinmathspace" /><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM171"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> is the previously defined difference between events and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM172"><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:msub></mml:math></inline-formula> is a tolerance, which allows some error to account for sampling variances, rounding, or potential perturbation. From these candidate ECGs, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM173"><mml:mi>E</mml:mi></mml:math></inline-formula> is the subset containing the verification feature; here, the U-wave is assumed to belong to the victim. This subset is <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM174"><mml:msub><mml:mi>E</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mtext>&#xA0;has U-wave</mml:mtext><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. Depending on the chosen events and verification features, this could pose a potent attack vector.</p>
<p>The second time-continuous data-specific approach works similarly to the first one but detects more complex events to construct a characteristic treatment or illness pattern for a patient. This will most likely work best for more severe illnesses, with more individualized treatments, but could be combined with the first approach to work on more standard treatments.</p>
<p>Here, the attacker detects and classifies the events <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM175"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:msub><mml:mn>1</mml:mn><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:msub><mml:mn>1</mml:mn><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> for record 1 and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM176"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:msub><mml:mn>2</mml:mn><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:msub><mml:mn>2</mml:mn><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> for record 2. In the best case, these sequences would match; more realistically, a metric like the edit distance must be used to determine their similarity. This allows an attacker to link multiple records together via event-based linkage and potentially track or even reidentify an individual.</p>
<p><xref ref-type="fig" rid="F6">Figure&#x00A0;6</xref> shows four cycles of an ECG. The data is based on Commons (<xref ref-type="bibr" rid="B43">43</xref>). It shows four markers named <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM177"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>&#x2013;<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM178"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>4</mml:mn></mml:msub></mml:math></inline-formula>, each highlighting a single event.</p>
<fig id="F6" position="float"><label>Figure 6</label>
<caption><p>Example&#x2014;event sequence linking.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g006.tif"><alt-text content-type="machine-generated">Electrocardiogram (ECG) trace displaying voltage in millivolts over time in seconds. The graph shows repeated spikes with annotated markers &#x03C4;1, &#x03C4;2, &#x03C4;3, and &#x03C4;4 indicating specific points in the cardiac cycle.</alt-text>
</graphic>
</fig>
<p><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM179"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> marks a particularly low Q-wave, which could be an indicator for a myocardial infarction, also named Q-wave infarction (<xref ref-type="bibr" rid="B44">44</xref>). The second marker <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM180"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> highlights a particularly shallow S-wave, as the ST complex is still facing upward; this can be seen as a non-pathological finding but is a notable event nevertheless. The highlighted event of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM181"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:math></inline-formula> is the lack of the U-wave; all cycles before and after show a significant U-wave, which is notable in itself, but its lack in just one cycle provides a strong event for identification. During the last cycle, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM182"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>4</mml:mn></mml:msub></mml:math></inline-formula> points at a shallow R-wave, thus hinting at a previous myocardial infarction (<xref ref-type="bibr" rid="B13">13</xref>).</p>
<p>The sequence of these events could be used to form a signature for the patient, which is named <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM183"><mml:mi>&#x03C3;</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2026;</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:math></inline-formula>. Linkage of multiple ECGs to a single known patient could then be performed by calculating the edit distance between the signatures of the ECGs and assuming them to belong to the same patient up to a certain threshold.</p>
<p>The feasibility of this attack vector still has to be shown and a reasonable threshold has to be found. To do this, a well-known set of events must be recognizable, which ideally should be automated.</p>
<p>Additionally, time patterns can be recognizable for a specific individual if only local noise is applied on the time axis. This means that the patterns are only slightly perturbed; thus, the rough date and distances between events stay intact. This could serve both as an identifier and a sensitive attribute, depending on the event and pattern.</p>
<p>Another possible way to infer knowledge specific to an individual from the data would be to detect certain events across multiple types of data, e.g., ECG and HF. These events are then aligned, as the detected cause for the event will happen roughly simultaneously for all types. This can lead to a restriction of the plausible parameters along the time axis and thus result in a breach of privacy for the time axis or at least increase the chances of other linking attacks.</p>
<p>Furthermore, biometric integration using the medical data, i.e., the use of recorded curves as a biometric identifier, could be used. The mapping of an ECG to the cardiac cycle of a person is a suitable example, as it can be used to infer characteristics about one&#x2019;s cardiac parameters.</p>
<p>The possibility of this mapping has been demonstrated and proven successful for the authentication of users in a lab environment (<xref ref-type="bibr" rid="B45">45</xref>, <xref ref-type="bibr" rid="B46">46</xref>). Morphological features of an identifiable in the ECG of an individual, as proposed for sex-specific differences by Moss (<xref ref-type="bibr" rid="B40">40</xref>), have been successfully implemented in a classifier by Attia et al. (<xref ref-type="bibr" rid="B47">47</xref>).</p>
<p>Similarly, detection mechanisms can be used to reduce the time-continuous data to discrete data. The detection of breathing cycles and breath classification provide a good example. This allows for the inference of breathing frequency and a diagnosis, which could be used in a linking attack. The feasibility of automatic detection and classification of breath types has been demonstrated by Oprea et al. (<xref ref-type="bibr" rid="B21">21</xref>).</p>
<p>Depending on the chosen noise and anonymized data, some noise might be removable, or at least the noise level could be reduced; this is particularly true if biological boundaries are not considered when applying the noise.</p>
<p>A wide range of potential attack vectors exist for time-continuous medical data. Some have already proven feasible, while others are introduced in this paper, requiring further investigation into their implementation and feasibility.</p>
</sec>
<sec id="s3b"><label>3.2</label><title>Requirements of time-continuous medical data</title>
<p>Time-continuous medical data comes with specific requirements w.r.t. anonymization. In the following, the requirements, which can be derived from the sanitized data alone, i.e., without inspecting the inner workings of the mechanism, are presented. To remain usable for diagnostic and research applications, the data should remain continuous and differentiable across all axes. High jumps within the data must not be introduced by any anonymization, as this will be implausible for any attacker with application-specific knowledge and thus potentially render the anonymization less powerful. The precise height of an implausible jump in the data is highly application- and data-specific and open for discussion. The sampling characteristics should be preserved. That is, any trajectory with a variable sampling rate should not be sampled with equidistant step size after anonymization. The opposite direction will most likely not be achieved, but the mean sampling rate should be preserved. For any curve, the order of the data points and the continuity of the data must be preserved. All mechanisms have to provide strong, formal privacy guarantees with a quantifiable level of protection and utility.</p>
<p>There are several classes of anonymization mechanisms, which are referenced later in this section and shortly introduced in the following. Additionally, the most important requirements for time-continuous medical data are evaluated. They have been compared to one another in a non-medical environment by Murthy et al. (<xref ref-type="bibr" rid="B48">48</xref>). The evaluated classes include generalization, suppression, masking, swapping, distortion, and perturbation. All these classes are syntactic protection methods. For some more promising classes of mechanisms, the properties forming the requirements are evaluated formally in <xref ref-type="sec" rid="s3c">Section 3.3</xref>.</p>
<p><bold>Generalization</bold> replaces the existing values with semantically consistent values. It works rather intuitively for most numeric domains, as generalization to a range domain is intuitive for most applications. For categorical domains, a domain generalization hierarchy is needed, which can become unfeasible for large domains and potentially impossible if the domain has cardinality <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM184"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula> and no inherent hierarchic structure. Such mechanisms are not well suited for the tasks of anonymizing time-continuous medical data, as they fail to preserve continuity and the sampling characteristics. Additionally, the order of the data values is not preserved for such mechanisms. This is proven in <xref ref-type="statement" rid="th3">Proofs 3.1.a</xref> and <xref ref-type="statement" rid="th9">3.2.a</xref>. Furthermore, generalization-based mechanisms like <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM185"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity fail to provide strong formal guarantees. To our knowledge, there does not exist a generalization-based mechanism with formal, quantifiable guarantees comparable to, e.g., DP.</p>
<p><bold>Suppression</bold>, i.e., the deletion of certain data, can be used to completely deny the usage of sensitive data. This can conceal sensitive attributes or tuples from both an attacker and a valid user. This may be a good solution if one or few records pose a very high risk of identifying an individual, as the omission of this individual can reduce the noise needed to achieve a sufficient privacy level and thus improve the utility of the anonymized data. For example, if a very rare disease is not of interest to the research question, it might be a good idea to suppress the records with this disease, as protecting them would most likely require large amounts of noise and would increase the information loss unnecessarily. This is also true for attributes of a database. That is, the removal of less relevant attributes from all records can improve utility by making the anonymization easier on the remaining database. Suppression of records alone does not provide privacy for those records that are released, but it does not affect the continuity and order of the data values, as the data itself is not changed. If attributes are suppressed, this can improve the privacy of all individuals, without affecting other attributes. Thus, such mechanisms are not suitable as a standalone solution.</p>
<p><bold>Masking</bold>, according to Murthy et al. (<xref ref-type="bibr" rid="B48">48</xref>), provides similar protection to generalization and shares the same problems w.r.t. the introduced requirements. It is similar to suppression, but only parts of the data are made unusable. That is, a fraction of each attribute is replaced by a placeholder. For example, zip codes 52070 and 52062 are replaced by 520XX. This also highlights the similarity to generalization in some cases, as the two zip codes could also be generalized to 520*, where * is a wildcard character. It is easier to apply; for instance, no domain generalization hierarchy is needed to increase data protection. Additionally, each record can be masked individually and independently. Not every part of the data carries the same amount of information; for example, some regions are of high interest to identify the biological sex, as highlighted by Moss (<xref ref-type="bibr" rid="B40">40</xref>). Depending on the data and record, these regions vary significantly and are also likely to be interesting to the researcher. Thus, removing them is either not possible, as they would need to be identified first, or not purposeful, as these regions are the reason for releasing the data in the first place. Similar to generalization, masking does not preserve continuity and order, and to our knowledge, no approach provides strong formal privacy guarantees, with a quantifiable level of protection.</p>
<p>Another approach is <bold>swapping</bold>, which rearranges the values in each column randomly, as described by Murthy et al. (<xref ref-type="bibr" rid="B48">48</xref>). This works entirely independent of the type of data and without much overhead. This is easy to implement and preserves continuity and order, as the data itself is not changed. It fails to provide strong formal guarantees. The successful usage of a swapping-based approach is unlikely because the correlations across multiple domains, e.g., between ECG and diagnosis, must be preserved to maintain the usefulness of the data.</p>
<p>A remaining class includes <bold>distortion</bold>-based mechanisms. These mechanisms change the value of an attribute to something else, either by adding some noise <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM186"><mml:msub><mml:mi>V</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula> to the value <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM187"><mml:msub><mml:mi>V</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula>, also called perturbation, or, for an explicit identifier, through a hash function (<xref ref-type="bibr" rid="B48">48</xref>). The latter allows for a unique reidentification within the anonymized database while not giving away the explicit identifier.</p>
<p>It is the most promising class for the intended purpose of this paper. Such mechanisms preserve continuity and can preserve the order if certain bounds are guaranteed. <bold>Perturbation</bold> is a special form of distortion and works by adding noise to the data values to ensure privacy. If such a mechanism-based approach is used, the noise must be bounded, with the bounds given by the type of data that is anonymized. For example, the noise applied to an ECG can never be in the range of full volts and seconds, as this would make it unusable. For the detection of myocardial infarctions, differences in elevation <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM188"><mml:mo>&#x2265;</mml:mo><mml:mn>0.1</mml:mn></mml:math></inline-formula>&#x2009;mV and 80&#x2009;ms width are used for the diagnosis (<xref ref-type="bibr" rid="B49">49</xref>, <xref ref-type="bibr" rid="B50">50</xref>). Thus, the applied noise should be substantially lower. Additionally, the added noise is not independent for each point, as highlighted by Cao et al. (<xref ref-type="bibr" rid="B51">51</xref>). Therefore, the context of the curve must be considered when adding noise. Perturbation does preserve the continuity of the data. The preservation of the order is not guaranteed unless the noise is sufficiently bounded. This is proven in <xref ref-type="statement" rid="th5">Proofs 3.1.b</xref> and <xref ref-type="statement" rid="th12">3.2.b</xref>.</p>
<p>The two perturbed ECG plots in <xref ref-type="fig" rid="F7">Figures&#x00A0;7</xref> and <xref ref-type="fig" rid="F8">8</xref> illustrate the privacy&#x2013;utility tradeoff in time-continuous data and problems arising when applying perturbation-based approaches to such data. The data segment is loosely based on ExCard Research (<xref ref-type="bibr" rid="B52">52</xref>). <xref ref-type="fig" rid="F7">Figure&#x00A0;7</xref> adds noise only to the value axis, while <xref ref-type="fig" rid="F8">Figure&#x00A0;8</xref> adds noise to both time and value axes. In both figures, the red curve is the original, the blue curve has a small amount of added noise, and the brown curve has a larger amount of noise added. Laplacian noise with an absolute maximum of 0.5&#x2009;s, 1.9&#x2009;mV perturbation and a mean perturbation of 0.289&#x2009;s, 0.75&#x2009;mV is added to the low-noise version. For the high-noise version, up to 1.5&#x2009;s, 9.5&#x2009;mV is added, and on average, each sample is perturbed by 1.05&#x2009;S, 2.86&#x2009;mV.</p>
<fig id="F7" position="float"><label>Figure 7</label>
<caption><p>ECG perturbation&#x2014;value axis.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g007.tif"><alt-text content-type="machine-generated">Line graph showing three curves labeled: original curve (red), low noise curve (blue), and high noise curve (brown). The y-axis represents millivolts (mV) and the x-axis represents seconds. Original and low noise curves closely follow each other, peaking sharply at 0.2 seconds, while the high noise curve shows less variation.</alt-text>
</graphic>
</fig>
<fig id="F8" position="float"><label>Figure 8</label>
<caption><p>ECG perturbation&#x2014;time and value axes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g008.tif"><alt-text content-type="machine-generated">A line graph depicting four curves: the original curve in red, low noise curve in blue, high noise curve in beige, and order perturbed points in green. The x-axis represents time in seconds, and the y-axis represents millivolts (mV). The curves illustrate variations and noise effects on voltage measurements over time.</alt-text>
</graphic>
</fig>
<p>Even when perturbing both axes, the slight distorted shape of the signal in the blue curve retains its fundamental diagnostic features and temporal order. In contrast to this, the highly perturbed version&#x2014;displayed in blue&#x2014;is usable in <xref ref-type="fig" rid="F8">Figure&#x00A0;8</xref>. The order is retained, but the characteristically raised T-wave is drastically damped in this example. Still, the elevated T-wave visible hints at signs of subacute myocardial infarction, even though it is an unusual shape, potentially obscuring. For the version in <xref ref-type="fig" rid="F8">Figure&#x00A0;8</xref>, the ECG becomes completely unusable, as the order is destroyed and no meaningful insight can be drawn from the remaining data. The highly perturbed version in <xref ref-type="fig" rid="F8">Figure&#x00A0;8</xref> lead to a change in the order of the samples. This example highlights that an independent, pointwise application of perturbation is not useful for such data.</p>
<p>Multiple perturbation-based approaches exist that provide strong formal guarantees, e.g., DP. It is based on carefully adding crafted noise to the data to distort it as much as needed to fulfill the privacy guarantees. That is, the approach must stay within the budget of leaked information, thus reducing the risk to each individual to a quantifiable amount. This can be done both interactively and non-interactively. To our knowledge, there is no DP-based approach that focuses on time-continuous data, especially in a medical environment, even though the risk of privacy leakage through time-continuous data was identified and quantified by Cao et al. (<xref ref-type="bibr" rid="B51">51</xref>) as temporal privacy leakage. The properties of DP are formally evaluated in <xref ref-type="statement" rid="th7">Proofs 3.1.c</xref> and <xref ref-type="statement" rid="th14">3.2.c</xref>.</p>
<p>Alternative approaches like PrivECG (<xref ref-type="bibr" rid="B41">41</xref>), which employs a GAN to improve the privacy of an ECG w.r.t. the classifiable of the biological sex, are promising but fall short of formally quantifying the level of privacy. The only privacy metrics evaluated against this approach are chosen accuracy metrics of a classifier with and without PrivECG in place. While this provides good initial intuition, it makes no statement about the information-theoretic level of privacy, as DP would, and only considers one specific classifier.</p>
<p>In summary, a mechanism addressing the anonymization of time-continuous medical data must preserve continuity and order. Where possible, the step-size characteristics should be preserved. Spatio-temporal correlation is of high importance to most analyses and must therefore be preserved. Such a mechanism must provide a strong formal foundation and a quantifiable level of privacy. None of the syntactic approaches provide such guarantees. Thus, a DP-based approach seems to be the most promising candidate, even though it introduces potentially fake data.</p>
</sec>
<sec id="s3c"><label>3.3</label><title>Properties of anonymization mechanisms for time-continuous medical data</title>
<p>This subsection defines some properties of anonymization mechanisms that are relevant for time-continuous medical data and proves whether these properties hold for certain classes of mechanisms.</p>
<p>If the requirements for continuity are fulfilled, a dataset is continuous according to <xref ref-type="statement" rid="st5">Definition 2.5</xref>; this property should also hold after applying anonymization. This might be true for some classes of anonymization approaches, while others might fail to provide such a guarantee, as proposed in <xref ref-type="statement" rid="th1">Theorem 3.1</xref>.</p><statement id="th1"><label>THEOREM 3.1</label><title>(Continuity preservation of time-continuous anonymization)</title>
<p>Some classes of anonymization algorithms are inherently safe w.r.t. the continuity, as defined in <xref ref-type="statement" rid="st5">Definition 2.5</xref> and <xref ref-type="statement" rid="co1">Corollary 2.1</xref>, while others might lose the continuity through anonymization. The proof of this theorem follows from <xref ref-type="statement" rid="th2">Theorems 3.1.a</xref> and <xref ref-type="statement" rid="th4">3.1.b</xref> and could be extended to other classes.</p></statement>
<p>The intuition behind <xref ref-type="statement" rid="th2">Theorem 3.1.a</xref> and the corresponding proof is that the generalization of a time-continuous domain is not necessarily an endomorphism. That is, the domain after the anonymization can differ from that of the raw data. This can cause a loss of continuity, as shown in the following.</p><statement id="th2"><label>THEOREM 3.1.a</label><title>(Continuity preservation of time-continuous generalization)</title>
<p>Generalization-based approaches are not inherently safe w.r.t. preserving the continuity of the anonymized data.</p></statement><statement id="th3"><label>PROOF OF THEOREM 3.1.a.</label>
<p>Given a domain <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM189"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, its generalized domain <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM190"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula>, and the anonymization function <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM191"><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula>. If a dataset in <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM192"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> is continuous according to <xref ref-type="statement" rid="st5">Definition 2.5</xref>, there must exist the subtraction <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM193"><mml:mo>&#x2299;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula>, absolute <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM194"><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mo>&#x2299;</mml:mo><mml:mo fence="false" stretchy="false">|</mml:mo></mml:math></inline-formula>, and order <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM195"><mml:mo>&#x2299;</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula> operator on its values, and the last has to define a total order on all non-negative elements of the domain. Additionally, there exists the interpolate <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM196"><mml:mi>f</mml:mi></mml:math></inline-formula>, as defined in <xref ref-type="statement" rid="st5">Definition 2.5</xref>, such that, for every point in the time domain, there is a mapping to some value in <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM197"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>.</p>
<p>For <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM198"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula>, these operators must also exist; otherwise, <xref ref-type="statement" rid="st5">Definition 2.5</xref> is not fulfillable. This is not necessarily the case for every generalized domain.</p>
<p>There indeed exists such a domain that <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM199"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> is continuous, and after applying <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM200"><mml:mi>g</mml:mi></mml:math></inline-formula>, the continuity is lost. This means that <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM201"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> is not continuous according to <xref ref-type="statement" rid="st5">Definition 2.5</xref>.</p>
<p>Let <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM202"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> be the domain of an ECG. That is, the domain of each data point is <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM203"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">Q</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, where the unit of the second element of the tuple is mV. Both seconds and rational numbers have a straightforward subtraction, absolute, and order operation. Additionally, both can be interpolated, e.g., using <xref ref-type="disp-formula" rid="disp-formula2">Equation 2</xref>, such that the domains are indeed continuous.</p>
<p>Let <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM204"><mml:mi>g</mml:mi></mml:math></inline-formula> map each data entry, that is, every curve, to the corresponding diagnosis. Then, the value part of the generalized domain, called <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM205"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> for <bold>I</bold>llnesses <bold>g</bold>eneralized, is a nominal. While it could be the ICD-Coding, the following simplification is used for the sake of readability:<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="UDM9"><mml:mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>M</mml:mi><mml:mi>y</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>M</mml:mi><mml:mi>I</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mspace width="1em" /><mml:mi>C</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>h</mml:mi><mml:mi>y</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mspace width="1em" /><mml:mi>O</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>I</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>O</mml:mi><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>H</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>Then, the generalized domain is <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM206"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM207"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">T</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> is the generalization of the time domain, e.g., up to a day. In such a case, there exists no interpolate between values nor is there a meaningful subtraction or absolute operator in <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM208"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula>, as these are nominal values. Thus, the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM209"><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> criterion, according to <xref ref-type="statement" rid="st5">Definition 2.5</xref>, is not fulfillable, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM210"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> is not continuous, even though the non-anonymized domain <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM211"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> has this property. Therefore, generalization-based approaches are not inherently safe w.r.t. the preservation of the continuity of the anonymized data.</p></statement>
<p>Intuitively, <xref ref-type="statement" rid="th4">Theorem 3.1.b</xref> means that perturbation does not change the domain of the data. This is rather straightforward, as adding some noise to data of this domain cannot change the domain but only the data value. Thus, continuity will always be preserved.</p><statement id="th4"><label>THEOREM 3.1.b</label><title>(Continuity preservation of time-continuous perturbation)</title>
<p>In contrast to <xref ref-type="statement" rid="th2">Theorem 3.1.a</xref>, the anonymization mapping of a perturbation-based approach is the endomorphism <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM212"><mml:mi>p</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> with domain <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM213"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, as anonymization only applies some form of noise to the data.</p></statement><statement id="th5"><label>PROOF OF THEOREM 3.1.b.</label>
<p>As the domain is continuous before anonymization and does not change through the mechanism, the dataset remains continuous after applying <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM214"><mml:mi>p</mml:mi></mml:math></inline-formula>. If dataset <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM215"><mml:mi>d</mml:mi></mml:math></inline-formula> from <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM216"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> is continuous, the perturbed dataset <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM217"><mml:msub><mml:mi>d</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>, which is produced by applying <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM218"><mml:mi>p</mml:mi></mml:math></inline-formula> to every element of the dataset, is also continuous. This is because continuity is an attribute of the domain, which has not changed. Thus, perturbation is inherently continuity preserving w.r.t. <xref ref-type="statement" rid="st5">Definition 2.5</xref>.</p></statement>
<p>As an example of semantic approaches, differential privacy is investigated. For the properties of preservation of continuity and order, it does not matter if the approach is interactive or non-interactive, as the mechanisms that provide privacy are the same. In interactive differential privacy, a user of the mechanism submits a query, and a custom-tailored response is given to that query, which represents the database w.r.t. the statistical properties, adheres to the privacy budget, and tries to answer the query as well as possible. Thus, some carefully crafted noise is added to some subset of the database; the amount of noise might vary depending on the attribute. That is, for a query interested in heart-related issues, adding larger amounts of noise to the age might be acceptable, while the ECG values are less perturbed. For the non-interactive variant, the amount of noise is decided beforehand and the database is released as a whole. This can lead to different distributions and levels of noise, but the general mechanism is the same and thus both variants can be investigated together. It should be noted that DP, as described in <xref ref-type="statement" rid="th6">Theorem 3.1.c</xref>, means the application of DP to time-continuous data that each time entry is treated independently. This comes with certain privacy implications, as demonstrated by Cao et al. (<xref ref-type="bibr" rid="B51">51</xref>).</p>
<p>Similar to perturbation-based approaches, <xref ref-type="statement" rid="th6">Theorem 3.1.c</xref> can be explained intuitively by the fact that the domain itself is not changed by anonymization. This kind of noise is not important for the preservation of continuity.</p><statement id="th6"><label>THEOREM 3.1.c</label><title>(Continuity preservation of differential privacy)</title>
<p>Similar to <xref ref-type="statement" rid="th4">Theorem 3.1.b</xref>, anonymization mapping of a differential privacy-based approach is endomorphism <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM219"><mml:mi>p</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> with domain <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM220"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">D</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, as anonymization applies carefully crafted noise to the data.</p></statement><statement id="th7"><label>PROOF OF THEOREM 3.1.c.</label>
<p><xref ref-type="statement" rid="th5">Proof 3.1.b</xref> also applies to <xref ref-type="statement" rid="th6">Theorem 3.1.c</xref>, as both classes of privacy mechanisms have an endomorphism as their anonymization, and this is the only prerequisite for this proof. Thus, DP is inherently continuity preserving w.r.t. <xref ref-type="statement" rid="st5">Definition 2.5</xref>.</p></statement>
<p>For many applications, especially in a medical environment, time serves as more than just a means of ordering data points. It functions as a continuous axis, which contributes to the joint meaning of the data. For example, the spacing between measurements carries additional information regarding the level of trust one can have within a given segment in the curve; thus, <xref ref-type="statement" rid="st4">Definition 2.4</xref> is needed.</p>
<p>A generalization along the time axis would lead not only to information loss due to reduced precision regarding the time value but also possibly to loss of information about the local sampling rate. Such an approach still preserves the ordering of the data. A perturbation-based approach will threaten not only the precision and local sampling rate but also the ordering. Thus, further restrictions are needed for the intended usage to ensure that the order is preserved. For example, a bound could be introduced to the level of noise, which is applied to the time axis. The ordering property is investigated in <xref ref-type="statement" rid="th8">Theorem 3.2</xref>.</p><statement id="th8"><label>THEOREM 3.2</label><title>(Order of time-continuous anonymization)</title>
<p>Some classes of anonymizing algorithms pose a threat to the order of time-continuous data, whereas others are inherently safe in this regard. The proof follows from <xref ref-type="statement" rid="th9">Theorems 3.2.a</xref> and <xref ref-type="statement" rid="th11">3.2.b</xref> and could be extended to other classes.</p></statement><statement id="th9"><label>THEOREM 3.2.a</label><title>(Order of time-continuous generalization)</title>
<p>Generalization-based approaches are inherently safe w.r.t. the order of time-continuous data.</p></statement><statement id="th10"><label>PROOF OF THEOREM 3.2.a.</label>
<p>Given the records <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM221"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM222"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM223"><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> is the total order w.r.t. the time entry. Then, a generalization-based approach might group together records <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM224"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM225"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2026;</mml:mo><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> into the non-overlapping generalizations <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM226"><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM227"><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM228"><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x22EF;</mml:mo><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>d</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> to <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM229"><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Any generalization based on the closeness of the time value would lead to <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM230"><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> if and only if <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM231"><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM232"><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> otherwise, thus preserving the ordering.</p></statement>
<p>Intuitively, if too much noise is added to the time axis, some samples might be swapped. Thus, the order is not preserved; this is visualized by an example in <xref ref-type="fig" rid="F9">Figure&#x00A0;9</xref> and is formalized in <xref ref-type="statement" rid="th11">Theorem 3.2.b</xref> and <xref ref-type="statement" rid="th12">Proof 3.2.b</xref>.</p>
<fig id="F9" position="float"><label>Figure 9</label>
<caption><p>Order perturbation example.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fdgth-07-1604001-g009.tif"><alt-text content-type="machine-generated">Line graph showing voltage in millivolts over time in seconds. The blue line represents original data, the red line shows perturbed data with crosses, and green diamonds indicate swapped samples. The curves depict variations around 0.05 seconds.</alt-text>
</graphic>
</fig><statement id="th11"><label>THEOREM 3.2.b</label><title>(Order of time-continuous perturbation)</title>
<p>Perturbation-based approaches are not inherently safe w.r.t. the order of time-continuous data.</p></statement><statement id="th12"><label>PROOF OF THEOREM 3.2.b.</label>
<p>Given the records <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM233"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM234"><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> and the corresponding perturbation values <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM235"><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM236"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM237"><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> produced by the perturbation mechanism. Then, the perturbed data points are <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM238"><mml:msubsup><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM239"><mml:msubsup><mml:mi>d</mml:mi><mml:mi>k</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msubsup></mml:math></inline-formula> with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM240"><mml:msubsup><mml:mi>d</mml:mi><mml:mi>j</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. If <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM241"><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM242"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>, then <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM243"><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>. Thus, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM244"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> holds, but the data gets perturbed such that <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM245"><mml:msubsup><mml:mi>d</mml:mi><mml:mi>k</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msubsup><mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>t</mml:mi></mml:msub><mml:msubsup><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msubsup></mml:math></inline-formula> holds. That is, the ordering is not preserved.</p></statement>
<p>As <xref ref-type="statement" rid="th11">Theorem 3.2.b</xref> highlights the threat to the order of the data, any perturbation-based approach intended for time-continuous data has to employ a process of adding noise to the data designed with this safety in mind to preserve the order of time-continuous data. That is, the added noise <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM246"><mml:mi>p</mml:mi></mml:math></inline-formula> has to be sufficiently bounded to preserve the order.</p><statement id="th13"><label>THEOREM 3.2.c</label><title>(Order of differential privacy)</title>
<p>Differential privacy-based approaches are not inherently safe w.r.t. the order of time-continuous data.</p></statement><statement id="th14"><label>PROOF OF THEOREM 3.2.c.</label>
<p>Similar to continuity, the order property of DP can also be proven by referencing the perturbation-based approaches. This is the case, as the proof only depends on the mechanism adding noise along both the value and time axes, with no hard guarantee for the amount of noise being added on a single point. Thus, by extension of <xref ref-type="statement" rid="th12">Proof 3.2.b</xref>, the ordering is not preserved.</p></statement>
<p>Similar to perturbation-based approaches in general, <xref ref-type="statement" rid="th13">Theorem 3.2.c</xref> shows the threat to the data order. Adding noise along both axes can lead to a loss of order; thus, a carefully crafted mechanism for time-continuous data has to be designed with this safety in mind to preserve the order of time-continuous data for DP.</p>
<p>The intrinsic dependency on time-continuous data implies that any approach aiming to anonymize such data should not only preserve the ordering, as shown using <xref ref-type="statement" rid="th8">Theorem 3.2</xref>, but also the continuityas defined in <xref ref-type="statement" rid="st5">Definition 2.5</xref>, and possibly the characteristics of the step time, as defined in <xref ref-type="statement" rid="st4">Definition 2.4</xref>.</p>
</sec>
</sec>
<sec id="s4"><label>4</label><title>State of the art</title>
<p>Samarati and Sweeney (<xref ref-type="bibr" rid="B2">2</xref>) proposed the distinction between explicit identifiers, quasi-identifiers, and sensitive attributes and the mitigation of certain privacy attacks using <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM247"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity. It aims to reduce this risk by grouping together <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM248"><mml:mi>k</mml:mi></mml:math></inline-formula> data records and generalizing the quasi-identifiers (<xref ref-type="bibr" rid="B2">2</xref>, <xref ref-type="bibr" rid="B22">22</xref>). <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM249"><mml:mi>T</mml:mi></mml:math></inline-formula>-closeness reduces this risk of distribution attacks by introducing a constraint that forces the distribution of sensitive attributes within a grouping to be at most <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM250"><mml:mi>t</mml:mi></mml:math></inline-formula> apart from the population distribution (<xref ref-type="bibr" rid="B3">3</xref>). This makes the usage of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM251"><mml:mi>l</mml:mi></mml:math></inline-formula>-diversity, as proposed by Machanavajjhala et al. (<xref ref-type="bibr" rid="B4">4</xref>), obsolete, as pointed out by Li et al. (<xref ref-type="bibr" rid="B3">3</xref>).</p>
<p>CASTLE introduces an anonymization scheme based on <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM252"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity for continuous streams of discrete data, see Cao et al. (<xref ref-type="bibr" rid="B5">5</xref>, <xref ref-type="bibr" rid="B6">6</xref>). This was adapted to a <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM253"><mml:mi>t</mml:mi></mml:math></inline-formula>-closeness first approach with SABRE by Cao et al. (<xref ref-type="bibr" rid="B7">7</xref>). Both approaches fail to preserve the continuity of the data. All approaches based on <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM254"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity share the same basic problem: There is no way to mathematically determine which attribute is a quasi-identifier and which is a non-identifying sensitive attribute (<xref ref-type="bibr" rid="B26">26</xref>). This leads to a lack of provable privacy, making them unsuitable for time-continuous medical data. Information-theoretic approaches like DP (<xref ref-type="bibr" rid="B9">9</xref>) achieve this formal level of anonymization and do not build upon a vague definition of quasi-identifiers.</p>
<p>Cao et al. (<xref ref-type="bibr" rid="B51">51</xref>) demonstrated that point-by-point anonymization with DP, which is common, poses the risk of information leakage through temporal correlations between values. This was formalized as the temporal privacy leakage. The attack vector was viable in the performed experiments and thus posed the risk to one&#x2019;s privacy via spatio-temporal or continuous data.</p>
<p>Nergiz et al. (<xref ref-type="bibr" rid="B8">8</xref>) presented an approach for anonymization of trajectories, which preserves the spatio-temporal relation. It employs a group-and-link approach and anonymizes the data by releasing a representative constructed by choosing a random representation point in each group. While this approach preserves the spatio-temporal relation and order of the points by grouping non-overlapping areas of points, it does not preserve the data truth, as the representative might look nothing like any of the trajectories it represents because each representation point is generated independently and randomly. This can also cause large jumps in the representative, which threatens the continuity of the data.</p>
<p>Dankar and El Emam (<xref ref-type="bibr" rid="B10">10</xref>) reviewed the specific requirements for applying DP to health data, stipulating that it should be efficient, provide strong privacy guarantees, and exhibit adaptability. For the interactive approach, it should facilitate a wide range and a large number of queries. In the non-interactive case, special requirements are presented, as medical professionals and biomedical researchers like to &#x201C;look at the data&#x201D;; that is, explorative research is requested. As it is usually not their main task, many try to avoid changes in the processes and tools they have become used to, as they do not have the time to put much effort into changing and adapting them. A good utility for a wide range of queries is possible. This also holds for a non-interactive approach. However, besides the technical difficulties, a significant challenge lies in the social factor of a heterogeneous user group, which is reluctant to technical changes.</p>
<p>Olawoyin et al. (<xref ref-type="bibr" rid="B11">11</xref>) presented a novel approach for anonymizing spatio-temporal patient data. This approach aims to protect the temporal attributes, with a five-level temporal hierarchy and temporal representative points, while applying DP with Laplacian noise to the spatial attributes. The spatio-temporal relation is preserved by means of temporal representative points. Due to the very coarse generalization in the temporal hierarchy, this is not suitable for continuous trajectories. For instance, the order of data points becomes ambiguous.</p>
<p>A <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM255"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity-based mechanism with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM256"><mml:mi>t</mml:mi></mml:math></inline-formula>-closeness for time-continuous medical data was developed and successfully evaluated by Hammer et al. (<xref ref-type="bibr" rid="B30">30</xref>). This mechanism uses the Fr&#x00E9;chet distance to calculate the similarity between curves and the information loss as a utility metric. The data is optionally split along the time axis. This reduces the computational load drastically while having a positive impact on the utility of the evaluated datasets. It falls short of providing strong formal guarantees for privacy, as it is based on <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM257"><mml:mi>k</mml:mi></mml:math></inline-formula>-anonymity.</p>
<p>A novel method to provide privacy for ECGs employing generative adversarial networks (GANs) is proposed by Nolin-Lapalme et al. (<xref ref-type="bibr" rid="B41">41</xref>). This method aims to prevent the reidentification of biological sex based on the publication of the ECG. To achieve this, it is shown that a 12-electrode clinical ECG is indeed suitable to distinguish sex.</p>
<p>According to Becker (<xref ref-type="bibr" rid="B53">53</xref>), the R-wave is the most salient feature to classify an ECG according to sex. This morphological distinction of the heart is further explained by Moss (<xref ref-type="bibr" rid="B40">40</xref>).</p>
<p>Additionally, the rising interest in ECG data, especially for biometric authentication by Melzi et al. (<xref ref-type="bibr" rid="B54">54</xref>) and data sharing by Flanagin et al. (<xref ref-type="bibr" rid="B55">55</xref>), was highlighted by Nolin-Lapalme et al. (<xref ref-type="bibr" rid="B41">41</xref>). There already exist large ECG databases such as PhysioNet/CinC (<xref ref-type="bibr" rid="B56">56</xref>), PTB-XL (<xref ref-type="bibr" rid="B16">16</xref>), and large-scale medical databases like MIMIC IV (<xref ref-type="bibr" rid="B14">14</xref>), which enable the development and testing of such an approach without having to measure the records beforehand.</p>
<p>Especially, with the rise of artificial intelligence applications in recent years, the threats to one&#x2019;s privacy are also rising, especially in a medical environment (<xref ref-type="bibr" rid="B57">57</xref>). The viability of such an attack has already been described by Attia et al. (<xref ref-type="bibr" rid="B47">47</xref>), who used a CNN to determine the biological sex and age. It resulted in a 90.4&#x0025; accuracy for the sex classification and an average error of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM258"><mml:mn>6.9</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>5.6</mml:mn></mml:math></inline-formula> years for the age.</p>
<p>PrivECG, a CNN similar to that developed by Attia et al. (<xref ref-type="bibr" rid="B47">47</xref>), forms the basis of the evaluation of the approach of Nolin-Lapalme et al. (<xref ref-type="bibr" rid="B41">41</xref>). PrivECG aims to limit the classifyability of ECGs, thus increasing privacy. The original version of PrivECG resulted in a sex prediction accuracy of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM259"><mml:mn>0.686</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.012</mml:mn></mml:math></inline-formula> vs. the original <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM260"><mml:mn>0.882</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.022</mml:mn></mml:math></inline-formula>. This was further improved by PrivECG <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM261"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula>, resulting in a sex prediction accuracy of only <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM262"><mml:mn>0.529</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.014</mml:mn></mml:math></inline-formula> after sanitation, practically making the prediction impossible. There are some limitations to PrivECG. Nolin-Lapalme et al. (<xref ref-type="bibr" rid="B41">41</xref>) noted that an interaction between diseases and sex can exist, meaning a prediction of a disease could lead to an accurate prediction of the sex. Similarly, other attributes, which might be encoded in the ECG like age or ethnicity, remain unaffected by PrivECG (<xref ref-type="bibr" rid="B41">41</xref>). Additionally, it should be noted that no quantification of the achieved privacy level, like the case when using DP, can be given with PrivECG. The evaluation only compares the accuracy of one specific CNN on the original and sanitized databases. This might not generalize to other attack vectors and is certainly not based on any information-theoretic guarantees.</p>
<p>Furthermore, there are no remarks on general privacy, e.g., the risk of reidentification or linking attacks. PrivECG only aims at preventing sex classification of an ECG. Nolin-Lapalme et al. (<xref ref-type="bibr" rid="B41">41</xref>) described many metrics for the use in their algorithm and its evaluation. They proposed some general metrics, like the F1-score, i.e., the harmonic mean of the precision and recall, or the root mean square error. They also suggested using the Fr&#x00E9;chet distance (<xref ref-type="bibr" rid="B32">32</xref>) as a measure of the difference between ECGs before and after sanitation. However, they did not evaluate this idea further.</p>
<p>The Fr&#x00E9;chet distance was used as a similarity metric for different time-continuous medical datasets by Hammer et al. (<xref ref-type="bibr" rid="B30">30</xref>). Additionally, a few ECG-specific metrics of similarity were introduced, e.g., the average mean difference from the baseline or the average standard variation in R-wave amplitude for a single ECG cycle (<xref ref-type="bibr" rid="B41">41</xref>). However, further research on the viability of these metrics is needed to asses their usefulness.</p>
<p>Kaissis et al. (<xref ref-type="bibr" rid="B58">58</xref>) evaluated the application of privacy-preserving mechanisms to medical data, especially medical image data. They highlight that current approaches are often insufficient and risk patient privacy, making collaboration or sharing of data challenging. The increasing utilization of data, especially using AI in areas like medical image processing, can provide large benefits to medical professionals and patients. To achieve these benefits without sacrificing one&#x2019;s privacy, suitable privacy mechanisms are needed. DP is identified as a suitable candidate.</p>
<p>It is noted that the specifics of DP implementation for medical image data remain unclear (<xref ref-type="bibr" rid="B58">58</xref>). Qayyum et al. (<xref ref-type="bibr" rid="B59">59</xref>) provided an overview of the need for privacy and security in medical data settings and proposed some well-known solutions. They identified multiple approaches to ensure safety, clarify prediction causality, and reduce the risk of certain attack types. Additionally, DP is identified as one of the most promising candidates to provide privacy in medical data sharing and machine learning applications (<xref ref-type="bibr" rid="B59">59</xref>). While this paper does not address medical data-specific issues with DP, it does reference Beaulieu-Jones et al. (<xref ref-type="bibr" rid="B60">60</xref>) as an example of privacy-preserving machine learning in a medical environment.</p>
<p>Beaulieu-Jones et al. (<xref ref-type="bibr" rid="B60">60</xref>) presented an application of DP with cyclical weight transfer to the eICU database (<xref ref-type="bibr" rid="B15">15</xref>) and the Cancer Genome Atlas (TCGA) (<xref ref-type="bibr" rid="B61">61</xref>). From neither database, time-continuous data is used.</p>
<p>Beaulieu-Jones et al. (<xref ref-type="bibr" rid="B60">60</xref>) showed that integrating DP into the training of a machine learning model in a medical environment is feasible. However, it was limited to discrete data, meaning that their developed process cannot be applied directly to time-continuous data.</p>
<p>In summary, many anonymization approaches exist, most of which are not suitable for continuous or medical data. Some might be extendable to fit the specific needs, but none seem acceptable as is.</p>
</sec>
<sec id="s5" sec-type="conclusions"><label>5</label><title>Conclusion</title>
<p>The spatio-temporal structure of continuous medical data is essential for its utility, especially in diagnostic and predictive applications. However, existing anonymization approaches that aim to preserve this structure often fall short when applied to time-continuous data.</p>
<p>In this work, we discussed the precise requirements for anonymizing such data and evaluated different classes of mechanisms with respect to these requirements (<xref ref-type="sec" rid="s3b">Section 3.2</xref>). We also elaborated on a set of necessary properties for any anonymization technique in this domain (<xref ref-type="sec" rid="s3c">Section 3.3</xref>), highlighting the challenges of meeting strict privacy standards while maintaining clinical utility.</p>
<p>Medical data imposes particularly stringent demands on anonymization due to both the sensitivity of the information and the high risk to individuals in the event of a privacy breach. Unlike domains where damages can be mitigated post-breach, e.g., financial restitution in a banking scenario, medical data, once leaked, cannot be retracted. Consequently, anonymization mechanisms must rely on strong formal foundations capable of supporting provable and quantifiable privacy guarantees.</p>
<p>A key challenge in this context is the privacy&#x2013;utility tradeoff. While a higher level of privacy provides stronger protection against misuse and inference attacks, it often comes at the cost of reduced data utility. This is particularly problematic in medical domains where diagnostic accuracy or model performance can be critically dependent on subtle temporal patterns in the data. Moreover, utility metrics tend to be domain-specific; for instance, a metric suited to evaluating ECG data for sinoatrial node disorders might be inadequate for myocardial infarction detection. Therefore, any useful anonymization scheme must strike a careful balance between minimizing information loss and ensuring robust privacy guarantees.</p>
<p>Importantly, insufficient privacy is not only problematic in cases of outright data breaches&#x2014;it can also compromise individuals&#x2019; rights even when data is accessed by authorized parties. The trust patients place in medical professionals does not necessarily extend to insurers or third-party entities. Thus, any anonymization method must respect contextual integrity and consent-based data sharing.</p>
<p>To address these challenges, we advocate for a version of DP adapted specifically to time-continuous medical data. By incorporating temporal correlation handling, as highlighted in (<xref ref-type="bibr" rid="B51">51</xref>), and preserving order and continuity (<xref ref-type="statement" rid="th6">Theorems 3.1.c</xref>, <xref ref-type="statement" rid="th13">3.2.c</xref>), such a mechanism could offer a principled approach to managing the privacy&#x2013;utility tradeoff. Specifically, the noise addition process must be context-aware&#x2014;i.e., informed by preceding and subsequent data points&#x2014;to minimize the impact on utility while preserving privacy.</p>
<p>The development of such an adapted DP mechanism holds great promise. It could facilitate secure sharing of medical datasets beyond the originating institution, enable more effective training and validation of AI models, and ultimately lead to better clinical outcomes. A robust anonymization framework that respects both patient privacy and the needs of medical research has the potential to unlock significant progress&#x2014;ethically, legally, and scientifically.</p>
</sec>
</body>
<back>
<sec id="s6" sec-type="data-availability"><title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material; further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec id="s7" sec-type="author-contributions"><title>Author contributions</title>
<p>FH: Formal Analysis, Methodology, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing; TS: Methodology, Supervision, Writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec id="s8" sec-type="funding-information"><title>Funding</title>
<p>The author(s) declare that no financial support was received for the research and/or publication of this article.</p>
</sec>
<sec id="s9" sec-type="COI-statement"><title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="ai-statement"><title>Generative AI statement</title>
<p>The author(s) declare that no Generative AI was used in the creation of this manuscript.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<fn-group>
<fn id="FN0001"><p><sup>1</sup>The operators are needed for the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM263"><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> criterion.</p></fn>
<fn id="FN0002"><p><sup>2</sup>Removes, inserts, or changes the value.</p></fn>
</fn-group>
<sec id="s11" sec-type="disclaimer"><title>Publisher&#x0027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list><title>References</title>
<ref id="B1"><label>1.</label><citation citation-type="other"><collab>Wikimedia Commons</collab>. <article-title>File: a comparison of blood pressure and photoplethysmogram signals.svg</article-title> (<year>2023</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://commons.wikimedia.org/wiki/File:A_comparison_of_blood_pressure_ and_photoplethysmogram_signals.svg">https://commons.wikimedia.org/wiki/File:A&#x005F;comparison&#x005F;of&#x005F;blood&#x005F;pressure&#x005F; and&#x005F;photoplethysmogram&#x005F;signals.svg</ext-link> <comment>(accessed June 11, 2025)</comment>.</citation></ref>
<ref id="B2"><label>2.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Samarati</surname><given-names>P</given-names></name><name><surname>Sweeney</surname><given-names>L</given-names></name></person-group>. <article-title>Protecting privacy when disclosing information: k-anonymity and its enforcement through generalization and suppression</article-title>. <source>dataprivacylab.org</source> (<year>1998</year>).</citation></ref>
<ref id="B3"><label>3.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Li</surname><given-names>N</given-names></name><name><surname>Li</surname><given-names>T</given-names></name><name><surname>Venkatasubramanian</surname><given-names>S</given-names></name></person-group>. <article-title>t-closeness: privacy beyond k-anonymity and l-diversity</article-title>. In: <source>2007 IEEE 23rd International Conference on Data Engineering</source>. IEEE (<year>2006</year>). <comment>pp. 106&#x2013;15</comment>.</citation></ref>
<ref id="B4"><label>4.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Machanavajjhala</surname><given-names>A</given-names></name><name><surname>Kifer</surname><given-names>D</given-names></name><name><surname>Gehrke</surname><given-names>J</given-names></name><name><surname>Venkitasubramaniam</surname><given-names>M</given-names></name></person-group>. <article-title>l-Diversity: privacy beyond k-anonymity</article-title>. <source>ACM Trans Knowl Discov Data (TKDD)</source>. (<year>2007</year>) <volume>1</volume>(<issue>1</issue>):<fpage>3&#x2013;es</fpage>. <pub-id pub-id-type="doi">10.1145/1217299.1217302</pub-id></citation></ref>
<ref id="B5"><label>5.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Cao</surname><given-names>J</given-names></name><name><surname>Carminati</surname><given-names>B</given-names></name><name><surname>Ferrari</surname><given-names>E</given-names></name><name><surname>Tan</surname><given-names>KL</given-names></name></person-group>. <article-title>CASTLE: a delay-constrained scheme for <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM264"><mml:msub><mml:mi>k</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>-anonymizing data streams</article-title>. In: <source>2008 IEEE 24th International Conference on Data Engineering</source>. IEEE (<year>2008</year>). pp. <comment>1376&#x2013;8</comment>.</citation></ref>
<ref id="B6"><label>6.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cao</surname><given-names>J</given-names></name><name><surname>Carminati</surname><given-names>B</given-names></name><name><surname>Ferrari</surname><given-names>E</given-names></name><name><surname>Tan</surname><given-names>K-L</given-names></name></person-group>. <article-title>Castle: continuously anonymizing data streams</article-title>. <source>IEEE Trans Depend Secure Comput</source>. (<year>2010</year>) <volume>8</volume>(<issue>3</issue>):<fpage>337</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1109/TDSC.2009.47</pub-id></citation></ref>
<ref id="B7"><label>7.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cao</surname><given-names>J</given-names></name><name><surname>Karras</surname><given-names>P</given-names></name><name><surname>Kalnis</surname><given-names>P</given-names></name><name><surname>Tan</surname><given-names>K-L</given-names></name></person-group>. <article-title>Sabre: a sensitive attribute bucketization and redistribution framework for t-closeness</article-title>. <source>VLDB J</source>. (<year>2011</year>) <volume>20</volume>:<fpage>59</fpage>&#x2013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1007/s00778-010-0191-9</pub-id></citation></ref>
<ref id="B8"><label>8.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Nergiz</surname><given-names>ME</given-names></name><name><surname>Atzori</surname><given-names>M</given-names></name><name><surname>Saygin</surname><given-names>Y</given-names></name></person-group>. <article-title>Towards trajectory anonymization: a generalization-based approach</article-title>. In: <source>Proceedings of the SIGSPATIAL ACM GIS 2008 International Workshop on Security and Privacy in GIS and LBS</source> (<year>2008</year>). <comment>pp. 52&#x2013;61</comment>.</citation></ref>
<ref id="B9"><label>9.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Dwork</surname><given-names>C</given-names></name></person-group>. <article-title>Differential privacy</article-title>. In: <person-group person-group-type="author"><name><surname>Bugliesi</surname><given-names>M</given-names></name><name><surname>Preneel</surname><given-names>B</given-names></name><name><surname>Sassone</surname><given-names>V</given-names></name><name><surname>Wegener</surname><given-names>I</given-names></name></person-group>, editors, <source>International Colloquium on Automata, Languages, and Programming</source>. <publisher-loc>Springer</publisher-loc> (<year>2006</year>). pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>.</citation></ref>
<ref id="B10"><label>10.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Kamal Dankar</surname><given-names>F</given-names></name><name><surname>El Emam</surname><given-names>K</given-names></name></person-group>. <article-title>The application of differential privacy to health data</article-title>. In: <source>Proceedings of the 2012 Joint EDBT/ICDT Workshops</source> (<year>2012</year>). <comment>pp. 158&#x2013;66</comment>.</citation></ref>
<ref id="B11"><label>11.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Olawoyin</surname><given-names>AM</given-names></name><name><surname>Leung</surname><given-names>CK</given-names></name><name><surname>Choudhury</surname><given-names>R</given-names></name></person-group>. <article-title>Privacy-preserving spatio-temporal patient data publishing</article-title>. In: <source>Database and Expert Systems Applications: 31st International Conference, DEXA 2020; September 14&#x2013;17, 2020; Bratislava, Slovakia</source>. Springer (<year>2020</year>). <comment>pp. 407&#x2013;16</comment>.</citation></ref>
<ref id="B12"><label>12.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prasser</surname><given-names>F</given-names></name><name><surname>Eicher</surname><given-names>J</given-names></name><name><surname>Spengler</surname><given-names>H</given-names></name><name><surname>Bild</surname><given-names>R</given-names></name><name><surname>Kuhn</surname><given-names>KA</given-names></name></person-group>. <article-title>Flexible data anonymization using ARX &#x2013; current status and challenges ahead</article-title>. <source>Softw Pract Exp</source>. (<year>2020</year>) <volume>50</volume>(<issue>7</issue>):<fpage>1277</fpage>&#x2013;<lpage>304</lpage>. <pub-id pub-id-type="doi">10.1002/spe.2812</pub-id></citation></ref>
<ref id="B13"><label>13.</label><citation citation-type="other"><collab>Cardiovascular Medicine</collab>. <article-title>ECG interpretation: characteristics of the normal ECG (P-wave, QRS complex, ST segment, T-wave)</article-title> (<year>2021</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://ecgwaves.com/topic/ecg-normal-p-wave-qrs-complex-st-segment-t-w ave-j-point/">https://ecgwaves.com/topic/ecg-normal-p-wave-qrs-complex-st-segment-t-w ave-j-point/</ext-link> <comment>(accessed October 10, 2024)</comment>.</citation></ref>
<ref id="B14"><label>14.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Johnson</surname><given-names>A</given-names></name><name><surname>Bulgarelli</surname><given-names>L</given-names></name><name><surname>Pollard</surname><given-names>T</given-names></name><name><surname>Horng</surname><given-names>S</given-names></name><name><surname>Anthony Celi</surname><given-names>L</given-names></name><name><surname>Mark</surname><given-names>R</given-names></name></person-group>. <article-title>Mimic-iv</article-title>. <source>PhysioNet</source> (<year>2020</year>). pp. 49&#x2013;55. <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://physionet.org/content/mimiciv/3.1">https://physionet.org/content/mimiciv/3.1</ext-link> <comment>(accessed October 24, 2024)</comment>.</citation></ref>
<ref id="B15"><label>15.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pollard</surname><given-names>TJ</given-names></name><name><surname>Johnson</surname><given-names>AEW</given-names></name><name><surname>Raffa</surname><given-names>JD</given-names></name><name><surname>Celi</surname><given-names>LA</given-names></name><name><surname>Mark</surname><given-names>RG</given-names></name><name><surname>Badawi</surname><given-names>O</given-names></name></person-group>. <article-title>The eICU collaborative research database, a freely available multi-center database for critical care research</article-title>. <source>Sci Data</source>. (<year>2018</year>) <volume>5</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1038/sdata.2018.178</pub-id><pub-id pub-id-type="pmid">30482902</pub-id></citation></ref>
<ref id="B16"><label>16.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wagner</surname><given-names>P</given-names></name><name><surname>Strodthoff</surname><given-names>N</given-names></name><name><surname>Bousseljot</surname><given-names>R-D</given-names></name><name><surname>Kreiseler</surname><given-names>D</given-names></name><name><surname>Lunze</surname><given-names>FI</given-names></name><name><surname>Samek</surname><given-names>W</given-names></name><etal/></person-group>. <article-title>PTB-XL, a large publicly available electrocardiography dataset</article-title>. <source>Sci Data</source>. (<year>2020</year>) <volume>7</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1038/s41597-020-0495-6</pub-id><pub-id pub-id-type="pmid">31896794</pub-id></citation></ref>
<ref id="B17"><label>17.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wagner</surname><given-names>P</given-names></name><name><surname>Strodthoff</surname><given-names>N</given-names></name><name><surname>Bousseljot</surname><given-names>R-D</given-names></name><name><surname>Samek</surname><given-names>W</given-names></name><name><surname>Schaeffter</surname><given-names>T</given-names></name></person-group>. <article-title>PTB-XL, a large publicly available electrocardiography dataset</article-title> (<year>2022</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://physionet.org/content/ptb-xl/1.0.3/">https://physionet.org/content/ptb-xl/1.0.3/</ext-link> <comment>(accessed October 24, 2024)</comment>.</citation></ref>
<ref id="B18"><label>18.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pollard</surname><given-names>TJ</given-names></name><name><surname>W Johnson</surname><given-names>AE</given-names></name><name><surname>Raffa</surname><given-names>JD</given-names></name><name><surname>Celi</surname><given-names>LA</given-names></name><name><surname>Mark</surname><given-names>RG</given-names></name><name><surname>Badawi</surname><given-names>O</given-names></name></person-group>. <article-title>The eICU collaborative research database, a freely available multi-center database for critical care research</article-title>. <source>Sci Data</source>. (<year>2018</year>) <volume>5</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1038/sdata.2018.178</pub-id><pub-id pub-id-type="pmid">30482902</pub-id></citation></ref>
<ref id="B19"><label>19.</label><citation citation-type="other"><collab>Kardiovaskul&#x00E4;re Medizin Online</collab>. <article-title>EKG-Interpretation: Merkmale des normalen EKGs (P-Welle, QRS-Komplex, ST-Strecke, T-Welle)</article-title> (<year>2021</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://ekgecho.de/thema/normale-ekg-p-welle-qrs-komplex-st-strecke-t-w elle/">https://ekgecho.de/thema/normale-ekg-p-welle-qrs-komplex-st-strecke-t-w elle/</ext-link> <comment>(accessed June 5, 2024)</comment>.</citation></ref>
<ref id="B20"><label>20.</label><citation citation-type="other"><collab>Cardiovascular Medicine</collab>. <article-title>The ST segment: J-point, J-60 point, ST depression, ST elevation</article-title> (<year>2021</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://ecgwaves.com/topic/the-st-segment-j-point-j-60-point-st-depress ion-st-elevation/">https://ecgwaves.com/topic/the-st-segment-j-point-j-60-point-st-depress ion-st-elevation/</ext-link> <comment>(accessed September 30, 2024)</comment>.</citation></ref>
<ref id="B21"><label>21.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Oprea</surname><given-names>C</given-names></name><name><surname>Gr&#x00FC;ne</surname><given-names>M</given-names></name><name><surname>Buglowski</surname><given-names>M</given-names></name><name><surname>Olivier</surname><given-names>L</given-names></name><name><surname>Orlikowsky</surname><given-names>T</given-names></name><name><surname>Kowalewski</surname><given-names>S</given-names></name><etal/></person-group>. <article-title>Evaluating the explainable AI method grad-cam for breath classification on newborn time series data</article-title>. <source>arXiv</source> <comment>[Preprint] <italic>arXiv:2405.07590</italic> (2024)</comment>.</citation></ref>
<ref id="B22"><label>22.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sweeney</surname><given-names>L</given-names></name></person-group>. <article-title>k-anonymity: a model for protecting privacy</article-title>. <source>Int J Uncertain Fuzziness Knowl Based Syst</source>. (<year>2002</year>) <volume>10</volume>(<issue>05</issue>):<fpage>557</fpage>&#x2013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1142/S0218488502001648</pub-id></citation></ref>
<ref id="B23"><label>23.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Evfimievski</surname><given-names>A</given-names></name><name><surname>Gehrke</surname><given-names>J</given-names></name><name><surname>Srikant</surname><given-names>R</given-names></name></person-group>. <article-title>Limiting privacy breaches in privacy preserving data mining</article-title>. In: <source>Proceedings of the Twenty-Second ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems</source> <comment>(2003). pp. 211&#x2013;222</comment>.</citation></ref>
<ref id="B24"><label>24.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Dalenius</surname><given-names>T</given-names></name></person-group>. <article-title>Towards a methodology for statistical disclosure control</article-title>. <source>Statistics Sweden</source> (<year>1977</year>).</citation></ref>
<ref id="B25"><label>25.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Goldwasser</surname><given-names>S</given-names></name><name><surname>Micali</surname><given-names>S</given-names></name></person-group>. <article-title>Probabilistic encryption &#x0026; how to play mental poker keeping secret all partial information</article-title>. In: <source>STOC &#x2019;82: Proceedings of the Fourteenth Annual ACM Symposium on Theory of Computing</source>. <comment>New York, NY, USA: Association for Computing Machinery (1982) pp. 365&#x2013;77. ISBN 0897910702</comment>.</citation></ref>
<ref id="B26"><label>26.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Brickell</surname><given-names>J</given-names></name><name><surname>Shmatikov</surname><given-names>V</given-names></name></person-group>. <article-title>The cost of privacy: destruction of data-mining utility in anonymized data publishing</article-title>. In: <source>Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source> <comment>(2008). pp. 70&#x2013;8</comment>.</citation></ref>
<ref id="B27"><label>27.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Iyengar</surname><given-names>VS</given-names></name></person-group>. <article-title>Transforming data to satisfy privacy constraints</article-title>. In: <source>Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source> <comment>(2002). pp. 279&#x2013;88</comment>.</citation></ref>
<ref id="B28"><label>28.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>LeFevre</surname><given-names>K</given-names></name><name><surname>DeWitt</surname><given-names>DJ</given-names></name><name><surname>Ramakrishnan</surname><given-names>R</given-names></name></person-group>. <article-title>Workload-aware anonymization</article-title>. In: <source>Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source> <comment>(2006). pp. 277&#x2013;86</comment>.</citation></ref>
<ref id="B29"><label>29.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wang</surname><given-names>K</given-names></name><name><surname>Fung</surname><given-names>BCM</given-names></name><name><surname>Yu</surname><given-names>PS</given-names></name></person-group>. <article-title>Template-based privacy preservation in classification problems</article-title>. In: <source>Fifth IEEE International Conference on Data Mining (ICDM&#x2019;05)</source>. <comment>IEEE (2005). p. 8</comment>.</citation></ref>
<ref id="B30"><label>30.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hammer</surname><given-names>FGH</given-names></name><name><surname>Buglowski</surname><given-names>M</given-names></name><name><surname>Stollenwerk</surname><given-names>A</given-names></name></person-group>. <article-title>Semi-local time sensitive anonymization of clinical data</article-title>. <source>Sci Data</source>. (<year>2024</year>) <volume>11</volume>(<issue>1</issue>):<fpage>1412</fpage>. <pub-id pub-id-type="doi">10.1038/s41597-024-04192-1</pub-id><pub-id pub-id-type="pmid">39706828</pub-id></citation></ref>
<ref id="B31"><label>31.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>M&#x00E9;moli</surname><given-names>F</given-names></name></person-group>. <article-title>Gromov-Hausdorff distances in Euclidean spaces</article-title>. In: <source>2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops</source>. <comment>IEEE (2008). pp. 1&#x2013;8</comment>.</citation></ref>
<ref id="B32"><label>32.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alt</surname><given-names>H</given-names></name><name><surname>Godau</surname><given-names>M</given-names></name></person-group>. <article-title>Computing the Fr&#x00E9;chet distance between two polygonal curves</article-title>. <source>Int J Comput Geom Appl</source>. (<year>1995</year>) <volume>5</volume>(<issue>01n02</issue>):<fpage>75</fpage>&#x2013;<lpage>91</lpage>. <comment>doi: 10.1142/S0218195995000064</comment>. <pub-id pub-id-type="doi">10.1142/S0218195995000064</pub-id></citation></ref>
<ref id="B33"><label>33.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Eiter</surname><given-names>T</given-names></name><name><surname>Mannila</surname><given-names>H</given-names></name></person-group>. <article-title>Computing discrete Fr&#x00E9;chet distance</article-title>. <comment>Technical report, Citeseer (1994)</comment>.</citation></ref>
<ref id="B34"><label>34.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Andoni</surname><given-names>A</given-names></name><name><surname>Onak</surname><given-names>K</given-names></name></person-group>. <article-title>Approximating edit distance in near-linear time</article-title>. In: <source>Proceedings of the Forty-First Annual ACM Symposium on Theory of Computing</source> <comment>(2009). pp. 199&#x2013;204</comment>.</citation></ref>
<ref id="B35"><label>35.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Klein</surname><given-names>PN</given-names></name><name><surname>Sebastian</surname><given-names>TB</given-names></name><name><surname>Kimia</surname><given-names>BB</given-names></name></person-group>. <article-title>Shape matching using edit-distance: an implementation</article-title>. In: <source>Proceedings of the Twelfth Annual ACM-SIAM Symposium on Discrete Algorithms</source> <comment>(2001). pp. 781&#x2013;90</comment>.</citation></ref>
<ref id="B36"><label>36.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fox</surname><given-names>K</given-names></name><name><surname>Li</surname><given-names>X</given-names></name></person-group>. <article-title>Approximating the geometric edit distance</article-title>. <source>Algorithmica</source>. (<year>2022</year>) <volume>84</volume>(<issue>9</issue>):<fpage>2395</fpage>&#x2013;<lpage>413</lpage>. <pub-id pub-id-type="doi">10.1007/s00453-022-00966-4</pub-id></citation></ref>
<ref id="B37"><label>37.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Su</surname><given-names>H</given-names></name><name><surname>Liu</surname><given-names>S</given-names></name><name><surname>Zheng</surname><given-names>B</given-names></name><name><surname>Zhou</surname><given-names>X</given-names></name><name><surname>Zheng</surname><given-names>K</given-names></name></person-group>. <article-title>A survey of trajectory distance measures and performance evaluation</article-title>. <source>VLDB J</source>. (<year>2020</year>) <volume>29</volume>:<fpage>3</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1007/s00778-019-00574-9</pub-id></citation></ref>
<ref id="B38"><label>38.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nanni</surname><given-names>M</given-names></name><name><surname>Pedreschi</surname><given-names>D</given-names></name></person-group>. <article-title>Time-focused density-based clustering of trajectories of moving</article-title>. <source>J Intell Inf Syst</source>. (<year>2006</year>) <volume>27</volume>(<issue>3</issue>):<fpage>267</fpage>&#x2013;<lpage>89</lpage>. <pub-id pub-id-type="doi">10.1007/s10844-006-9953-7</pub-id></citation></ref>
<ref id="B39"><label>39.</label><citation citation-type="other"><collab>Wikimedia Commons</collab>. <article-title>File:ECG pacemaker syndrome.svg</article-title> (<year>2020</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://commons.wikimedia.org/w/index.php?title=File:ECG_pacemaker_synd rome.svg%26oldid=470717807">https://commons.wikimedia.org/w/index.php?title=File:ECG&#x005F;pacemaker&#x005F;synd rome.svg&#x0026;oldid=470717807</ext-link> <comment>(accessed June 5, 2024)</comment>.</citation></ref>
<ref id="B40"><label>40.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moss</surname><given-names>AJ</given-names></name></person-group>. <article-title>Gender differences in ECG parameters and their clinical implications</article-title>. <source>Ann Noninvasive Electrocardiol</source>. (<year>2010</year>) <volume>15</volume>(<issue>1</issue>):<fpage>1</fpage>. <pub-id pub-id-type="doi">10.1111/j.1542-474X.2009.00345.x</pub-id><pub-id pub-id-type="pmid">20146775</pub-id></citation></ref>
<ref id="B41"><label>41.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Nolin-Lapalme</surname><given-names>A</given-names></name><name><surname>Avram</surname><given-names>R</given-names></name><name><surname>Julie</surname><given-names>H</given-names></name></person-group>. <article-title>Privecg: generating private ECG for end-to-end anonymization</article-title>. In: <source>Machine Learning for Healthcare Conference</source>. <comment>PMLR (2023). pp. 509&#x2013;28</comment>.</citation></ref>
<ref id="B42"><label>42.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Yeom</surname><given-names>S</given-names></name><name><surname>Giacomelli</surname><given-names>I</given-names></name><name><surname>Fredrikson</surname><given-names>M</given-names></name><name><surname>Jha</surname><given-names>S</given-names></name></person-group>. <article-title>Privacy risk in machine learning: analyzing the connection to overfitting</article-title>. In: <source>2018 IEEE 31st Computer Security Foundations Symposium (CSF)</source> <comment>(2018). pp. 268&#x2013;82</comment>. <pub-id pub-id-type="doi">10.1109/CSF.2018.00027</pub-id></citation></ref>
<ref id="B43"><label>43.</label><citation citation-type="other"><collab>Wikimedia Commons</collab>. <article-title>File:ecg-rrinterval.svg &#x2013; wikimedia commons, the free media repository</article-title> (<year>2021</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://commons.wikimedia.org/w/index.php?title=File:ECG-RRinterval.svg %26oldid=569365855">https://commons.wikimedia.org/w/index.php?title=File:ECG-RRinterval.svg &#x0026;oldid=569365855</ext-link> <comment>(accessed June 5, 2024)</comment>.</citation></ref>
<ref id="B44"><label>44.</label><citation citation-type="other"><collab>Cardiovascular Medicine</collab>. <article-title>ECG signs of myocardial infarction: pathological Q-waves &#x0026; pathological R-waves</article-title> (<year>2021</year>). <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://ecgwaves.com/topic/ecg-criteria-myocardial-infarction-pathologi cal-q-waves-r-waves/">https://ecgwaves.com/topic/ecg-criteria-myocardial-infarction-pathologi cal-q-waves-r-waves/</ext-link> <comment>(accessed November 6, 2024)</comment>.</citation></ref>
<ref id="B45"><label>45.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Islam</surname><given-names>SMM</given-names></name><name><surname>Bori&#x0107;-Lubecke</surname><given-names>O</given-names></name><name><surname>Zheng</surname><given-names>Y</given-names></name><name><surname>Lubecke</surname><given-names>VM</given-names></name></person-group>. <article-title>Radar-based non-contact continuous identity authentication</article-title>. <source>Remote Sens (Basel)</source>. (<year>2020</year>) <volume>12</volume>(<issue>14</issue>):<fpage>2279</fpage>. <pub-id pub-id-type="doi">10.3390/rs12142279</pub-id></citation></ref>
<ref id="B46"><label>46.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Lin</surname><given-names>F</given-names></name><name><surname>Song</surname><given-names>C</given-names></name><name><surname>Zhuang</surname><given-names>Y</given-names></name><name><surname>Xu</surname><given-names>W</given-names></name><name><surname>Li</surname><given-names>C</given-names></name><name><surname>Ren</surname><given-names>K</given-names></name></person-group>. <article-title>Cardiac scan: a non-contact and continuous heart-based user authentication system</article-title>. In: <source>Proceedings of the 23rd Annual International Conference on Mobile Computing and Networking</source> <comment>(2017). pp. 315&#x2013;28</comment>.</citation></ref>
<ref id="B47"><label>47.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Attia</surname><given-names>ZI</given-names></name><name><surname>Friedman</surname><given-names>PA</given-names></name><name><surname>Noseworthy</surname><given-names>PA</given-names></name><name><surname>Lopez-Jimenez</surname><given-names>F</given-names></name><name><surname>Ladewig</surname><given-names>DJ</given-names></name><name><surname>Satam</surname><given-names>G</given-names></name><etal/></person-group>. <article-title>Age and sex estimation using artificial intelligence from standard 12-lead ECGs</article-title>. <source>Circ Arrhythm Electrophysiol</source>. (<year>2019</year>) <volume>12</volume>(<issue>9</issue>):<fpage>e007284</fpage>. <pub-id pub-id-type="doi">10.1161/CIRCEP.119.007284</pub-id><pub-id pub-id-type="pmid">31450977</pub-id></citation></ref>
<ref id="B48"><label>48.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Murthy</surname><given-names>S</given-names></name><name><surname>Bakar</surname><given-names>AA</given-names></name><name><surname>Rahim</surname><given-names>FA</given-names></name><name><surname>Ramli</surname><given-names>R</given-names></name></person-group>. <article-title>A comparative study of data anonymization techniques</article-title>. In: <source>2019 IEEE 5th Intl Conference on Big Data Security on Cloud (BigDataSecurity), IEEE Intl Conference on High Performance and Smart Computing (HPSC), and IEEE Intl Conference on Intelligent Data and Security (IDS)</source>. <comment>IEEE (2019). pp. 306&#x2013;9</comment>.</citation></ref>
<ref id="B49"><label>49.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coppola</surname><given-names>G</given-names></name><name><surname>Carit&#x00E0;</surname><given-names>P</given-names></name><name><surname>Corrado</surname><given-names>E</given-names></name><name><surname>Borrelli</surname><given-names>A</given-names></name><name><surname>Rotolo</surname><given-names>A</given-names></name><name><surname>Guglielmo</surname><given-names>M</given-names></name><etal/></person-group>. <article-title>St segment elevations: always a marker of acute myocardial infarction?</article-title> <source>Indian Heart J</source>. (<year>2013</year>) <volume>65</volume>(<issue>4</issue>):<fpage>412</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1016/j.ihj.2013.06.013</pub-id><pub-id pub-id-type="pmid">23993002</pub-id></citation></ref>
<ref id="B50"><label>50.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ibanez</surname><given-names>B</given-names></name><name><surname>James</surname><given-names>S</given-names></name><name><surname>Agewall</surname><given-names>S</given-names></name><name><surname>Antunes</surname><given-names>MJ</given-names></name><name><surname>Bucciarelli-Ducci</surname><given-names>C</given-names></name><name><surname>Bueno</surname><given-names>H</given-names></name><etal/></person-group>. <article-title>ESC guidelines for the management of acute myocardial infarction in patients presenting with st-segment elevation: the task force for the management of acute myocardial infarction in patients presenting with st-segment elevation of the European Society of Cardiology (ESC)</article-title>. <source>Eur Heart J</source>. (<year>2018</year>) <volume>39</volume>(<issue>2</issue>):<fpage>119</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1093/eurheartj/ehx393</pub-id><pub-id pub-id-type="pmid">28886621</pub-id></citation></ref>
<ref id="B51"><label>51.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cao</surname><given-names>Y</given-names></name><name><surname>Yoshikawa</surname><given-names>M</given-names></name><name><surname>Xiao</surname><given-names>Y</given-names></name><name><surname>Xiong</surname><given-names>L</given-names></name></person-group>. <article-title>Quantifying differential privacy in continuous data release under temporal correlations</article-title>. <source>IEEE Trans Knowl Data Eng</source>. (<year>2018</year>) <volume>31</volume>(<issue>7</issue>):<fpage>1281</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2018.2824328</pub-id><pub-id pub-id-type="pmid">31435181</pub-id></citation></ref>
<ref id="B52"><label>52.</label><citation citation-type="other"><collab>ExCard Research</collab>. <article-title>Myokardinfarkt &#x2013; fokus-ekg.de</article-title>. <comment>Available online at:</comment> <ext-link ext-link-type="uri" xlink:href="https://fokus-ekg.de/inhalt-von-a-z/ischaemie-und-infarkt/infarkt/">https://fokus-ekg.de/inhalt-von-a-z/ischaemie-und-infarkt/infarkt/</ext-link> <comment>(Accessed June, 25 2025)</comment>.</citation></ref>
<ref id="B53"><label>53.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Becker</surname><given-names>DE</given-names></name></person-group>. <article-title>Fundamentals of electrocardiography interpretation</article-title>. <source>Anesth Prog</source>. (<year>2006</year>) <volume>53</volume>(<issue>2</issue>):<fpage>53</fpage>&#x2013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.2344/0003-3006(2006)53[53:FOEI]2.0.CO;2</pub-id><pub-id pub-id-type="pmid">16863387</pub-id></citation></ref>
<ref id="B54"><label>54.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Melzi</surname><given-names>P</given-names></name><name><surname>Tolosana</surname><given-names>R</given-names></name><name><surname>Vera-Rodriguez</surname><given-names>R</given-names></name></person-group>. <article-title>ECG biometric recognition: review, system proposal, and benchmark evaluation</article-title>. <source>IEEE Access</source>. (<year>2023</year>) <volume>11</volume>:<fpage>15555</fpage>&#x2013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3244651</pub-id></citation></ref>
<ref id="B55"><label>55.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Flanagin</surname><given-names>A</given-names></name><name><surname>Curfman</surname><given-names>G</given-names></name><name><surname>Bibbins-Domingo</surname><given-names>K</given-names></name></person-group>. <article-title>Data sharing and the growth of medical knowledge</article-title>. <source>JAMA</source>. (<year>2022</year>) <volume>328</volume>(<issue>24</issue>):<fpage>2398</fpage>&#x2013;<lpage>9</lpage>. <comment>doi: 10.1001/jama.2022.22837</comment>. <pub-id pub-id-type="doi">10.1001/jama.2022.22837</pub-id><pub-id pub-id-type="pmid">36469324</pub-id></citation></ref>
<ref id="B56"><label>56.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberger</surname><given-names>AL</given-names></name><name><surname>Amaral</surname><given-names>LAN</given-names></name><name><surname>Glass</surname><given-names>L</given-names></name><name><surname>Hausdorff</surname><given-names>JM</given-names></name><name><surname>Ivanov</surname><given-names>PC</given-names></name><name><surname>Mark</surname><given-names>RG</given-names></name><etal/></person-group>. <article-title>Physiobank, physiotoolkit, and physionet</article-title>. <source>Circulation</source>. (<year>2000</year>) <volume>101</volume>(<issue>23</issue>):<fpage>e215</fpage>&#x2013;<lpage>20</lpage>. <comment>doi: 10.1161/01.CIR.101.23.e215</comment>.<pub-id pub-id-type="pmid">10851218</pub-id></citation></ref>
<ref id="B57"><label>57.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murdoch</surname><given-names>B</given-names></name></person-group>. <article-title>Privacy and artificial intelligence: challenges for protecting health information in a new era</article-title>. <source>BMC Med Ethics</source>. (<year>2021</year>) <volume>22</volume>:<fpage>1</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1186/s12910-021-00687-3</pub-id><pub-id pub-id-type="pmid">33388052</pub-id></citation></ref>
<ref id="B58"><label>58.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kaissis</surname><given-names>GA</given-names></name><name><surname>Makowski</surname><given-names>MR</given-names></name><name><surname>R&#x00FC;ckert</surname><given-names>D</given-names></name><name><surname>Braren</surname><given-names>RF</given-names></name></person-group>. <article-title>Secure, privacy-preserving and federated machine learning in medical imaging</article-title>. <source>Nat Mach Intell</source>. (<year>2020</year>) <volume>2</volume>(<issue>6</issue>):<fpage>305</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-020-0186-1</pub-id></citation></ref>
<ref id="B59"><label>59.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qayyum</surname><given-names>A</given-names></name><name><surname>Qadir</surname><given-names>J</given-names></name><name><surname>Bilal</surname><given-names>M</given-names></name><name><surname>Al-Fuqaha</surname><given-names>A</given-names></name></person-group>. <article-title>Secure and robust machine learning for healthcare: a survey</article-title>. <source>IEEE Rev Biomed Eng</source>. (<year>2020</year>) <volume>14</volume>:<fpage>156</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1109/RBME.2020.3013489</pub-id></citation></ref>
<ref id="B60"><label>60.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Beaulieu-Jones</surname><given-names>BK</given-names></name><name><surname>Yuan</surname><given-names>W</given-names></name><name><surname>Finlayson</surname><given-names>SG</given-names></name><name><surname>Steven Wu</surname><given-names>Z</given-names></name></person-group>. <article-title>Privacy-preserving distributed deep learning for clinical data</article-title>. <comment>arXiv [Preprint] <italic>arXiv:1812.01484</italic> (2018)</comment>.</citation></ref>
<ref id="B61"><label>61.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weinstein</surname><given-names>JN</given-names></name><name><surname>Collisson</surname><given-names>EA</given-names></name><name><surname>Mills</surname><given-names>GB</given-names></name><name><surname>Shaw</surname><given-names>KR</given-names></name><name><surname>Ozenberger</surname><given-names>BA</given-names></name><name><surname>Ellrott</surname><given-names>K</given-names></name><etal/></person-group>. <article-title>The cancer genome atlas pan-cancer analysis project</article-title>. <source>Nat Genet</source>. (<year>2013</year>) <volume>45</volume>(<issue>10</issue>):<fpage>1113</fpage>&#x2013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2764</pub-id><pub-id pub-id-type="pmid">24071849</pub-id></citation></ref></ref-list>
</back>
</article>