<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Sci.</journal-id>
<journal-title>Frontiers in Computer Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Sci.</abbrev-journal-title>
<issn pub-type="epub">2624-9898</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fcomp.2024.1379788</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Computer Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A matter of annotation: an empirical study on <italic>in situ</italic> and self-recall activity annotations from wearable sensors</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Hoelzemann</surname> <given-names>Alexander</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2349476/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Van Laerhoven</surname> <given-names>Kristof</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/192141/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
</contrib>
</contrib-group>
<aff><institution>Ubiquitous Computing, University of Siegen</institution>, <addr-line>Siegen</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jamie A. Ward, Goldsmiths University of London, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Mathias Ciliberto, University of Sussex, United Kingdom</p>
<p>Lei Zhang, Nanjing Normal University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Alexander Hoelzemann <email>alexander.hoelzemann&#x00040;uni-siegen.de</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>18</day>
<month>07</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>6</volume>
<elocation-id>1379788</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>01</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>02</day>
<month>07</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2024 Hoelzemann and Van Laerhoven.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Hoelzemann and Van Laerhoven</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Research into the detection of human activities from wearable sensors is a highly active field, benefiting numerous applications, from ambulatory monitoring of healthcare patients via fitness coaching to streamlining manual work processes. We present an empirical study that evaluates and contrasts four commonly employed annotation methods in user studies focused on in-the-wild data collection. For both the user-driven, <italic>in situ</italic> annotations, where participants annotate their activities during the actual recording process, and the recall methods, where participants retrospectively annotate their data at the end of each day, the participants had the flexibility to select their own set of activity classes and corresponding labels. Our study illustrates that different labeling methodologies directly impact the annotations&#x00027; quality, as well as the capabilities of a deep learning classifier trained with the data. We noticed that <italic>in situ</italic> methods produce less but more precise labels than recall methods. Furthermore, we combined an activity diary with a visualization tool that enables the participant to inspect and label their activity data. Due to the introduction of such a tool were able to decrease missing annotations and increase the annotation consistency, and therefore the F1-Score of the deep learning model by up to 8% (ranging between 82.1 and 90.4% F1-Score). Furthermore, we discuss the advantages and disadvantages of the methods compared in our study, the biases they could introduce, and the consequences of their usage on human activity recognition studies as well as possible solutions.</p></abstract>
<kwd-group>
<kwd>data annotation ambiguity</kwd>
<kwd>data labeling</kwd>
<kwd>deep learning</kwd>
<kwd>dataset</kwd>
<kwd>annotation method</kwd>
</kwd-group>
<contract-sponsor id="cn001">Deutsche Forschungsgemeinschaft<named-content content-type="fundref-id">10.13039/501100001659</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="5"/>
<equation-count count="0"/>
<ref-count count="63"/>
<page-count count="19"/>
<word-count count="14301"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Mobile and Ubiquitous Computing</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Sensor-based activity recognition is one of the research fields of Pervasive Computing developed with enormous speed and success by industry and science and influencing medicine, sports, industry, and therefore the daily lives of many people. However, current available smart devices are mostly capable of detecting periodic activities like simple locomotions. In order to recognize more complex activities a multimodal sensor input, such as Roggen et al. (<xref ref-type="bibr" rid="B40">2010</xref>), and more complex recognition models are needed. Many of the published datasets are made in controlled laboratory environments. Such data does not have the same characteristics and patterns as data recorded in-the-wild. Data that belongs to similar classes but is recorded in an uncontrolled vs. controlled environment can differ significantly since it contains more contextual information (Mekruksavanich and Jitpattanakul, <xref ref-type="bibr" rid="B30">2021</xref>). Furthermore, study participants tend to control their movements more while being monitored (Friesen et al., <xref ref-type="bibr" rid="B15">2020</xref>). The recording of long-term and real-world data is a tedious, time-consuming, and therefore a non-trivial task. Researchers have various motivations to record such datasets but the technical hurdles are still high and problems during the annotation process occur regularly. In Human Activity Recognition research, capturing long-term datasets presents a challenge: balancing precise labeling with minimal participant burden. Sole reliance on self-recall methods, like activity diaries (e.g., Zhao et al., <xref ref-type="bibr" rid="B63">2013</xref>), often leads to imprecise time indications that may not accurately reflect actual activity duration. Such incorrectly or noisy labeled data later on leads to a trained model that is less capable of detecting activities reliably (Natarajan et al., <xref ref-type="bibr" rid="B34">2013</xref>), due to unwanted temporal dependencies learned by wrongly annotated patterns (Bock et al., <xref ref-type="bibr" rid="B7">2022</xref>).</p>
<p>The field of HAR is witnessing a growing emphasis on real-world and long-term activity recognition. This focus stems from the need to address current limitations and achieve reliable recognition of complex daily activities. Existing long-term datasets often rely heavily on self-recall methods or additional tracking apps. These apps can set labels either automatically (Akbari et al., <xref ref-type="bibr" rid="B3">2021</xref>) or require manual selection (Cleland et al., <xref ref-type="bibr" rid="B13">2014</xref>). However, such approaches present challenges, leading many researchers to favor controlled environments for data collection. As a consequence, the number of publicly available &#x0201C;in-the-wild&#x0201D; datasets remains limited.</p>
<sec>
<title>1.1 Contribution</title>
<p>Our study focuses on the evaluation of 4 different annotation methods for labeling data in-the-wild: &#x02460; <italic>In situ (lat. on site or in position)</italic> with a button on a smartwatch, &#x02461; <italic>in situ</italic> with the app Strava<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> (an app that is available for iOS and Android smartphones), &#x02462; pure self-recall (writing an activity diary at the end of the day), and &#x002462; <italic>time-series</italic> assisted self-recall with the MAD-GUI (Ollenschl&#x000E4;ger et al., <xref ref-type="bibr" rid="B35">2022</xref>), which displays the sensor data visually and allows to annotate it interactively. Our study was conducted with 11 participants, 10 males, and one female, over 2 weeks. Participants wore a Bangle.js Version 1<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> smartwatch on their preferred hand, used Strava, and completed self-recall annotations every evening. In the first week, the participants were asked to write an activity diary at the end of the day without any helping material and additionally using two user-initiated methods (<italic>in situ button</italic> and <italic>in situ app</italic>) to manually set labels at the start and beginning of each activity. In the second week, the participants were given an additional visualization of the sensor data with an adapted version of the MAD-GUI annotation tool. With the help of this, participants then were instructed to label their data in hindsight with the activity diary as a mnemonic aid. Given labels from both weeks were compared to each other regarding the quality through visual inspection and statistical analysis with regard to the consistency and quantity of missing annotations across labeling methods. The participants in this study were given the freedom to self-report their activity classes based on the diverse range of pursuits encompassed within their daily lives. Consequently, the resulting dataset exhibits a heterogeneous composition, comprising both commonplace, routine activities such as <italic>walking, driving</italic>, and <italic>eating</italic>, as well as more specialized and niche activities like <italic>badminton, yoga, horse_riding</italic>, and gardening. Furthermore, we used a Shallow-DeepConv(LSTM) architecture (see Ord&#x000F3;&#x000F1;ez and Roggen, <xref ref-type="bibr" rid="B36">2016</xref>; Bock et al., <xref ref-type="bibr" rid="B8">2021</xref>), and trained models with a Leave-One-Day-Out cross-validation method of six previously selected subjects and each annotation method.</p>
</sec>
<sec>
<title>1.2 Impact</title>
<p>Annotating data, especially in real-world environments, is still very difficult and tedious. Labeling such data is always a trade-off between accuracy and workload for the study participants or annotators. We raise awareness among researchers to put more effort into exploring new annotation methods to overcome this issue. Our study shows that different labeling methodologies have a direct impact on the quality of annotations. With the deep learning analysis, we prove that this impacts the model capabilities directly. Therefore, we consider the evaluation of frequently used annotation methods for real-world and long-term studies to be crucial to give decision-makers of future studies a better base on which they can choose the annotation methodology for their study in a targeted way.</p></sec>
</sec>
<sec id="s2">
<title>2 Related work</title>
<p>A very limited number of datasets are currently publicly available which were recorded in the wild (e.g. Berlin and Van Laerhoven, <xref ref-type="bibr" rid="B6">2012</xref>; Thomaz et al., <xref ref-type="bibr" rid="B51">2015</xref>; Sztyler and Stuckenschmidt, <xref ref-type="bibr" rid="B49">2016</xref>; Gjoreski et al., <xref ref-type="bibr" rid="B19">2018</xref>; Vaizman et al., <xref ref-type="bibr" rid="B53">2018</xref>). Sztyler and Stuckenschmidt (<xref ref-type="bibr" rid="B49">2016</xref>) and Gjoreski et al. (<xref ref-type="bibr" rid="B19">2018</xref>) were captured in naturalistic settings, but the participants were equipped with multiple sensors on various body locations and were filmed by a third party during the exercises. Such visible equipment and the presence of an observer could potentially introduce a behavior bias (Yordanova et al., <xref ref-type="bibr" rid="B61">2018</xref>) in the data, as it may alter participants&#x00027; movement patterns due to the constant reminder that they are participating in a study (Friesen et al., <xref ref-type="bibr" rid="B15">2020</xref>). Furthermore, multimodal datasets recorded with multiple body-worn sensors, rather than a single Inertial Measurement Unit (IMU), have faced the challenge of proper inter-sensor synchronization (Hoelzemann et al., <xref ref-type="bibr" rid="B21">2019</xref>). A comprehensive dataset encompassing a diverse range of classes, accurate annotations, and recorded by a single device that is nearly unnoticeable to the participant (and therefore unlikely to influence their behavior or movement patterns), is not yet publicly available due to the aforementioned obstacles. According to Stikic et al. (<xref ref-type="bibr" rid="B47">2011</xref>) and later Cleland et al. (<xref ref-type="bibr" rid="B13">2014</xref>), we distinguish between 6 or 7, respectively, different methods and two environments (online/offline) of labeling data, the methods are (1) Indirect Observation, (2) Self-Recall, (3) Experience Sampling, (4) Video/Audio Recordings, (5) Time Diary, (6) Human Observer, (7) Prompted Labeling. Cruz-Sandoval et al. (<xref ref-type="bibr" rid="B14">2019</xref>) uses 4 different categories to classify data labeling approaches, these are <italic>(1) temporal (when)</italic>&#x02014;is the label conducted during or after the activity, <italic>(2) annotator (who)</italic>&#x02014;is the label given by the individual itself or by an observer, <italic>(3) scenario (where)</italic>&#x02014;is the activity labeled in a controlled (e.g laboratory) or uncontrolled (in-the-wild) environment, and <italic>(4) annotation mechanism (how)</italic>&#x02014;is the activity labeled manually, semi-automatically or fully-automatically. All labeling methods have their own benefits, and costs and come with a trade-off between required time and label accuracy. However, not every method is suitable for long-term and in-the-wild recording data. Reining et al. (<xref ref-type="bibr" rid="B38">2020</xref>), evaluated the annotation performance between six different human annotators of a MoCap (Motion Capturing) and IMU HAR Dataset for industrial deployment. They came to the conclusion that annotations were moderately consistent when subjects labeled the data for the first time. However, annotation quality improved after a revision by a domain expert. In the following, we would like to go into more detail on what we consider to be the most important labeling methods for the specific field of activity recognition.</p>
<sec>
<title>2.1 Annotation methods in activity recognition</title>
<sec>
<title>2.1.1 Self-recall</title>
<p>Self-recall methodologies are generally called methods in which study participants have to remember an event in the past. This methodology is used, for instance, in the medical field [e.g. in the diagnosis of injuries (Valuri et al., <xref ref-type="bibr" rid="B54">2005</xref>)], but also frequently in studies in the field of long-term activity recognition. Van Laerhoven et al. (<xref ref-type="bibr" rid="B55">2008</xref>) used this method during a study in which participants were asked to label their personal daily data at the end of the day. They noticed that the label quality depends heavily on the participant&#x00027;s recall and can therefore be very coarse. During a study conducted by Tapia et al. (<xref ref-type="bibr" rid="B50">2004</xref>), every 15 min a questionnaire was triggered in which participants needed the answer multiple choice questions about which of 35 predefined activities were recently performed.</p></sec>
<sec>
<title>2.1.2 App assisted labeling</title>
<p>Cleland et al. (<xref ref-type="bibr" rid="B13">2014</xref>) presented in 2014 the so called Prompted labeling. An approach that is already used by commercial smartwatches like the Apple Watch<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref>. In this study user&#x00027;s were asked to set a label for a time period which has been detected as an activity right after the activity stops. Akbari et al. (<xref ref-type="bibr" rid="B3">2021</xref>) leverages freely available Bluetooth Low Energy (BLE) information broadcasted by other nearby devices and combines this with wearable sensor data in order to detect context and direction changes. The participant is asked to set a new label whenever a change in the signal is detected. Gjoreski et al. (<xref ref-type="bibr" rid="B18">2017</xref>) published the SHL dataset which contains versatile labeled multimodal sensor data that has been labeled using an Android application that asked the user to set a label whenever they detected a position change via GPS. Tonkin et al. (<xref ref-type="bibr" rid="B52">2018</xref>) presented a smartphone app that was used in their experimental smart home environment with which study participants were able to either use voice-based labeling, select a label from a list of activities ordered by the corresponding location or scan NFC tags that were installed at locations in the smart house. Similar to Tonkin et al. (<xref ref-type="bibr" rid="B52">2018</xref>) and Vaizman et al. (<xref ref-type="bibr" rid="B53">2018</xref>) developed an open-source mobile app for recording sensor measurements in combination with a self-reported behavioral context (e.g. driving, eating, in class, showering). Sixty subjects participated in their study. The study found that most of the participants preferred to fill out their past behavior through a daily journal. Only some people preferred to set a label for an activity that they are about to do. Schr&#x000F6;der et al. (<xref ref-type="bibr" rid="B42">2016</xref>) developed a web-based GUI which can either be used on a smartphone, tablet, or a PC to label data recorded in a smart home environment. However, it is important to mention that, According to Cleland et al. (<xref ref-type="bibr" rid="B13">2014</xref>), the process of continually labeling data becomes laborious for participants and can result in a feeling of discomfort.</p></sec>
<sec>
<title>2.1.3 Unsupervised labeling</title>
<p>Unsupervised labeling is a methodology that uses clustering algorithms to first categorize new samples without deciding yet to which class a sample belongs. Leonardis et al. (<xref ref-type="bibr" rid="B27">2002</xref>) presented the concept of finding multiple subsets of eigenspaces where, according to Huynh (<xref ref-type="bibr" rid="B23">2008</xref>), each of them corresponds to an individual activity. Huynh uses this knowledge to develop the eigenspace growing algorithm, whereby, <italic>growing</italic> refers to an increasing set of samples as well as to increasing the so-called <italic>effective dimension</italic> of a corresponding eigenspace. Based on the reconstruction error (when a new sample is projected to an eigenspace), the algorithm tries to find the best-fitting representation of a sample with minimal redundancy. Hassan et al. (<xref ref-type="bibr" rid="B20">2021</xref>) recently published a methodology that uses the Pearson Correlation Coefficient to map very specific labels of a variety of datasets to 4 meta labels (inactive, active, walking, and driving) of the ExtraSensory Dataset (Vaizman et al., <xref ref-type="bibr" rid="B53">2018</xref>).</p></sec>
<sec>
<title>2.1.4 Human-in-the-Loop (Labeling)</title>
<p>Human-in-the-Loop (Labeling) is a collective term for methodologies that integrates human knowledge into their learning or labeling process. Besides of being applied in HAR research, such techniques are often used in Natural Language Processing (NLP) and according to Wu et al. (<xref ref-type="bibr" rid="B60">2022</xref>) the NLP community distinguishes between entity extraction (Gentile et al., <xref ref-type="bibr" rid="B16">2019</xref>; Zhang et al., <xref ref-type="bibr" rid="B62">2019</xref>), entity linking (Klie et al., <xref ref-type="bibr" rid="B26">2020</xref>), Q&#x00026;A tasks (Wallace et al., <xref ref-type="bibr" rid="B56">2019</xref>), and reading comprehension tasks (Bartolo et al., <xref ref-type="bibr" rid="B5">2020</xref>).</p></sec>
<sec>
<title>2.1.5 Active learning</title>
<p>Active learning is a machine learning strategy that currently receives a lot of attention in the HAR community. Such strategies involves a Human-in-the-Loop for labeling purposes. In the first step the learning algorithm automatically identifies relevant samples of a dataset which are posteriorly queued to be annotated by an expert. Incorporating a human guarantees high quality labels which directly leads to a better performing classifier. Whether a sample is determined to be relevant, and as well the decision to whom it may get presented for annotation purposes are the main focus of research in this field. Bota et al. (<xref ref-type="bibr" rid="B9">2019</xref>) presents a technique that relies on specific criteria defined by three different uncertainty-based selection functions to select samples that will be presented to an expert for labeling and then be propagated throughout the most similar samples. Adaimi and Thomaz (<xref ref-type="bibr" rid="B2">2019</xref>) benchmarks the performance of different Active Learning strategies and compared them, with regard to four different datasets with a fully-supervised approach. The authors came to the conclusion that Active Learning needs only 8%&#x02013;12% of the data to reach similar or even better results than a fully-supervised trained model. These results suggest that presenting pre-selected samples to a human for labeling purposes can reduce the amount of data needed to train a machine learning classifier significantly due to the increased quality of the labels. Miu et al. (<xref ref-type="bibr" rid="B32">2015</xref>) presented a system which used the Online Active Learning approach published by Sculley (<xref ref-type="bibr" rid="B44">2007</xref>) to bootstrap (Abney, <xref ref-type="bibr" rid="B1">2002</xref>) a machine learning classifier. The publication presented a smartphone app that asked the user right after finishing an activity, which activity has been performed. Afterwards a small subset of the labeled data was used to bootstrap a personalized machine learning classifier.</p></sec></sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3 Methodology</title>
<p>Our study is conducted with 11 participants, from which 10 are male and one is female. The participants are between 25 and 45 years old. Out of 11 participants, six are researchers in the field of signal processing and are used to read and work with sensor data. Participants were selected among acquaintances and colleagues.</p>
<sec>
<title>3.1 Study setup</title>
<p>The study was conducted over a period of 2 weeks, during which participants wore an open-source smartwatch on their chosen wrist. Throughout the two-week study period, the participants were instructed to use four different labeling methods in parallel, as illustrated in <xref ref-type="fig" rid="F1">Figure 1</xref>. In the first week they were asked to use the &#x002460; <italic>in situ button</italic>, &#x002461; <italic>in situ app</italic>, and &#x002462; <italic>pure self-recall</italic> methods. At the beginning of the 2nd week, we expanded the number of annotation methods with the &#x002463; <italic>time-series recall</italic>. This annotation method combines the activity diary with a graphical visualization of the participants&#x00027; daily data.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The study participants collected data for 14 days in total and annotated the data with 4 different methods: Labeling &#x002460; <italic>in situ</italic> with a mechanical button, &#x002461; <italic>in situ</italic> with an app, &#x002462; by writing a pure self-recall diary, and &#x002463; writing a self-recall diary assisted by visualization of their time-series data. The upper part of the figure is an artistic representation of our study, where the colored rectangles represent arbitrary annotations and the gray bars represent a full day of recorded activity data.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0001.tif"/>
</fig>
<p>&#x002460; The Bangle.js smartwatch has three mechanical buttons on the right side of the case. These buttons are programmed to record the number of consecutive button presses per minute. The button-press annotation method captures the total number of button presses along with their corresponding timestamps, enabling the delineation of the beginning and end of an activity within the time-series data. However, this approach does not inherently assign a specific label or description to the identified activity segment. To address this limitation, we employed an inference strategy that leveraged the temporal alignment of the button-press data with the annotations obtained from other methods. By identifying segments with similar timestamps across multiple annotation modalities, we could infer the appropriate label for the button-press annotations.</p>
<p>&#x002461; In addition, the participants were asked to track their activities with the smartphone app Strava. Strava is an activity tracker that is available for Android and iOS and freely downloadable from the app stores. The user can choose from a variety of predefined labels and start recording. Recording an activity starts a timer that runs until the user stops it. The time as well as the GPS position of the user during the activity is tracked and saved locally.</p>
<p>&#x002462; The <italic>pure self-recall</italic> methods consist of writing an activity diary on a daily basis at the end of the day. The participants were explicitly told that they should only write down the activities that they still remember 2 h after the measurement stopped.</p>
<p>&#x002463; The <italic>time-series recall</italic> method can be seen as a combination of an activity diary and a graphical representation of the raw sensor data. For visualization and labeling purposes, we provided the participants with an adapted version of the MAD-GUI. The GUI was published by Ollenschl&#x000E4;ger et al. (<xref ref-type="bibr" rid="B35">2022</xref>) and is a generic open-source Python package. Therefore, it can be integrated into one&#x00027;s project. Our adaptions to the package are available for download from a GitHub repository<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref>. It contains changes to the data loader, the definition of available labels, and color settings for displaying the 3D raw data.</p>
<sec>
<title>3.1.1 Annotation guidelines</title>
<p>The participants were provided with guidelines that instructed them to document recurring daily activities, encompassing both sports and activities of daily living, that exceeded a duration threshold of 10 min. However, the annotation process was deferred until &#x0007E;2 h after the cessation of the recording session. This temporal offset was implemented to allow for a reasonable time buffer, enabling participants to consolidate their experiences. For instance, if a daily recording concluded at 7 pm, the participant would typically annotate their data around 9 pm, allowing for a 2-h interim period. Each of these annotation methods represents a layer of annotation that is used for the visual, statistical, and deep learning evaluation. <xref ref-type="fig" rid="F1">Figure 1</xref> illustrates the overall concept.</p></sec>
<sec>
<title>3.1.2 Annotation process</title>
<p>To capture realistic daily data reflecting participants&#x00027; natural routines, we granted them complete autonomy in choosing their activity classes. Participants were not restricted to a specific activity protocol; instead, we left the decision of what to label entirely to their judgment. During the study&#x00027;s first week, participants employed methods &#x002460;&#x02014;&#x002462; concurrently. In the second week, method &#x002463; was introduced for them to utilize alongside the existing methods. The labels provided by the participants were later interpreted by the researchers and, when necessary, categorized into meta-classes. However, whenever a participant was specific about the activity performed, their label was not summed up into a meta-class. For example, activities such as <italic>yoga, badminton</italic>, or <italic>horse_riding</italic> were not combined under the meta-class <italic>sport</italic>.</p>
</sec>
</sec>
<sec>
<title>3.2 Hardware</title>
<p>Participants wore the commercial open-source smartwatch Bangle.js Version 1 with our open-source firmware<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref> installed. The device comes with a Nordic 64 MHz nRF52832 ARM Cortex-M4 processor with Bluetooth LE, 64 kB RAM, 512 kB on-chip flash, 4MB external flash, a heart rate monitor, a 3D accelerometer, and a 3D magnetometer. Our firmware only uses the 3D accelerometer and provides the user with the basic functions of a smartwatch, like displaying the time and counting steps. The data is recorded with 25 Hz, a sensitivity of &#x000B1;8g and saved on the devices&#x00027; memory with a delta compression algorithm. Therefore, we are able to save up to 8&#x02013;9 h (depending on how much of the data could be compressed) of data with the given parameters. The smartwatch stops recording as soon as the memory is full. At the end of the day, the participants need to upload their daily data and program the starting time for the next day using our upload web-tool<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref>.</p></sec></sec>
<sec id="s4">
<title>4 Statistical analysis</title>
<p>The labels were statistically analyzed based on their consistency using the Cohen &#x003BA; score as well as the number of missing annotations across all methods. The Cohen &#x003BA; score describes the agreement between two annotation methods, which is defined as follows &#x003BA; &#x0003D; (<italic>p</italic><sub>0</sub>&#x02212;<italic>p</italic><sub><italic>e</italic></sub>)/(1&#x02212;<italic>p</italic><sub><italic>e</italic></sub>), (see Artstein and Poesio, <xref ref-type="bibr" rid="B4">2008</xref>; Reining et al., <xref ref-type="bibr" rid="B38">2020</xref>). Where <italic>p</italic><sub>0</sub> is the observed agreement ratio and <italic>p</italic><sub><italic>e</italic></sub> is the expected agreement if both annotators assign labels randomly. The score shows how uniform two different annotators labeled the same data. For calculation purposes, an implementation provided by Scikit-Learn (<xref ref-type="bibr" rid="B43">2022</xref>), was used. Furthermore, missing annotations across methods are measured as the percentage of missing or incomplete annotations. The annotations of all methods were first compared with each other and matched based on the given time indications. Annotations that could not be assigned or were missing were marked accordingly and are the base for calculating this indicator, see Section 6.2 for more information. We used a similar representation as Brenner (<xref ref-type="bibr" rid="B10">1999</xref>) to visualize the matches among labeling methods. In this study, the authors compared genome annotations labeled by different annotators with regard to their error scores between different annotators.</p></sec>
<sec id="s5">
<title>5 Effects on deep learning performance</title>
<p>The deep learning analyses are performed using the DeepConvLSTM architecture (Ord&#x000F3;&#x000F1;ez and Roggen, <xref ref-type="bibr" rid="B36">2016</xref>) which is based on a Keras implementation of Hoelzemann and Van Laerhoven (<xref ref-type="bibr" rid="B22">2020</xref>). We did not perform hyperparameter tuning because it would involve a considerable amount of additional workload, since we trained 64 models independently during the evaluation. We therefore decided to opt out of the architecture with regards to efficiency rather than optimal classification results. Additionally, we don&#x00027;t expect that the actual experiment&#x02014;evaluating different annotation methods&#x02014;would benefit from hyperparameter tuning or gain any significant information and insights. Instead, we use the default hyperparameters provided by the authors. These are depicted in the <xref ref-type="fig" rid="F2">Figure 2</xref>. Furthermore, we reduce the number of LSTM layers to one and instead increase the number of hidden units of the only LSTM layer to 512. According to Bock et al. (<xref ref-type="bibr" rid="B8">2021</xref>), this modification decreases the runtime up to 48% compared to a two-layered DeepConvLSTM while significantly increasing the overall classification performance on 4/5 publicly available datasets: (Roggen et al., <xref ref-type="bibr" rid="B40">2010</xref>; Scholl et al., <xref ref-type="bibr" rid="B41">2015</xref>; Stisen et al., <xref ref-type="bibr" rid="B48">2015</xref>; Reyes-Ortiz et al., <xref ref-type="bibr" rid="B39">2016</xref>; Sztyler and Stuckenschmidt, <xref ref-type="bibr" rid="B49">2016</xref>). LSTM-Layers in general are important if the dataset contains sporadic activities (Bock et al., <xref ref-type="bibr" rid="B7">2022</xref>). However, our dataset does not and our evaluation aims to identify long periods of periodic activities, like walking or running. For this reason, we can conclude that additional LSTM layers are not needed. The implementation of Hoelzemann and Van Laerhoven (<xref ref-type="bibr" rid="B22">2020</xref>) incorporates BatchNormalization layers after each Convolutional layer, as well as MaxPooling for the transition between the final convolutional block and the LSTM layer, and a Dropout layer before classification. Each Convolutional layer employs a ReLU activation function. The inclusion of the BatchNormalization layers serves to accelerate training and mitigate the detrimental effects of internal covariate shift, as discussed further in Ioffe and Szegedy (<xref ref-type="bibr" rid="B24">2015</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The architecture consists of an <italic>Input</italic> Layer with the kernel-size 10 (window_size) &#x000D7; 10 (filter_length) &#x000D7; 3 (channels). The data is passed into three concatenated <italic>convolutional blocks</italic>, followed by a <italic>MaxPooling</italic> (kernel 2 &#x000D7; 1) where 50% of the data is filtered. The convolutional block consists of a convolutional layer with a variable kernel size of 5 &#x000D7; 1 &#x000D7; (<italic>n</italic>*64) following a ReLU activation function and a BatchNorm-Layer. We decided to use a single LSTM-Layer with the size of 512 units, as mentioned by Bock et al. (<xref ref-type="bibr" rid="B8">2021</xref>), which is followed by a Dropout-Layer that filters 30% of randomly selected samples of the window.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0002.tif"/>
</fig>
<sec>
<title>5.1 Preprocessing</title>
<p>To prepare the data for neural network training, we perform two preprocessing steps. First, we address minor inconsistencies in the device&#x00027;s sampling rate. The original data was collected at a rate of 12.5 Hz. However, for optimal performance with neural networks, a consistent and regular sampling rate is preferred. To achieve this, we upsample the data by a factor of two, resulting in a constant frequency of 25 Hz. This upsampling process essentially inserts additional data points between the existing ones, effectively increasing the resolution of the signal. The second preprocessing step involves rescaling the accelerometer data to a range between &#x02013;1 and 1.</p>
</sec>
<sec>
<title>5.2 Leave-one-day-out cross-validation</title>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> illustrates the train and test setting for the deep learning model. Instead of following the traditional Leave-One-Subject-Out strategy, we adapted it to our needs by using one day of the week for testing and training on the remaining days for each study participant and week. This approach was necessitated by the unique characteristics of our dataset. It consists of a predominant <italic>void</italic> class and a small number of samples per activity class and participant. To mitigate the issue of an disproportionately large <italic>void</italic> class, we trained our model with balanced class weights. By not limiting the participants in their choice of daily activities and not specifying predefined activity labels, we ended up with very unique sets of activities for each study participant. Given these circumstances, it is unrealistic to expect a model capable of generalizing across participants and days. The Leave-One-Day-Out strategy aims to maintain the consistency of class labels within each day&#x00027;s data, providing a more cohesive and reliable dataset for training and evaluation purposes. This strategy also mitigates the potential impact of participant-specific biases or variations in class labeling, leading to a more robust and accurate model. Furthermore, due to the in-the-wild recording setup, the intra-class differences (Bulling et al., <xref ref-type="bibr" rid="B11">2014</xref>) for comparatively simple activities, such as <italic>walking</italic> or <italic>running</italic> can be significant. Consequently, the impact of different labeling methods is expected to be more pronounced and visible in a personalized model compared to a generalized model.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Leave-One-Day-Out Cross Validation. The models are personally trained for every participant and are not intended to generalize across all study participants. Instead, a generalization across all days of 1 week is desired.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0003.tif"/>
</fig>
</sec>
<sec>
<title>5.3 Post-processing and classification</title>
<p>In the classification task, we initially segmented the data into fixed-length sliding windows of 2 s (50 samples). However, our objective extended beyond instantaneous classifications; we aimed to identify longer periods of recurring activities. To achieve this, we employed a post-processing technique involving a jumping window approach with a duration of 5 min. Within each 5-min window, a majority vote was applied to the individual 2-s window predictions. The activity class with the highest number of occurrences within the 5-min window was then assigned as the predominant activity for the entire window. This approach enabled us to capture sustained patterns of activities over extended periods, aligning with our goal of analyzing longer-term behavioral trends.</p></sec>
</sec>
<sec sec-type="results" id="s6">
<title>6 Results</title>
<p>Our participants were asked to annotate daily activities lasting more than 10 min. We did not limit them to a predefined set of classes; they independently decided on labels for their activities. After normalizing the labels (e.g., changing &#x0201C;going for a walk&#x0201D; to &#x0201C;walking&#x0201D;), the participants assigned 26 different labels: <italic>laying, sitting, walking, running, cycling, bus_driving, car_driving, cleaning, vacuum_cleaning, laundry, cooking, eating, shopping, showering, yoga, sport, playing_games, desk_work, guitar_playing, gardening, table_tennis, badminton, horse_riding, cleaning, reading, weightlifting, manual_work, dish_washing</italic>. Any unlabeled samples were classified as <italic>void</italic>. However, after excluding infrequent or non-standalone classes (e.g., <italic>shopping</italic> likely combines <italic>walking, standing, and sitting</italic>), we reduced the dataset to 23 labels (22 activities plus <italic>void</italic>): <italic>laying, sitting, walking, running, cycling, bus_driving, car_driving, vacuum_cleaning, laundry, cooking, eating, shopping, showering, yoga, sport, playing_games, desk_work, guitar_playing, gardening, table_tennis, badminton, horse_riding</italic>. Nevertheless, the graphical representation of the distribution and the table in Section 6.1 include the full scope of classes.</p>
<sec>
<title>6.1 Class distribution</title>
<p>The class distribution reflects a broad range of activity classes that represent the daily lives of our participants. These classes remain primarily participant-specific due to the absence of a predefined annotation protocol, which allowed participants the freedom to label activities according to their own interpretations. Consequently, we decided against employing a Leave-One-Subject-Out evaluation method, as it might introduce inconsistencies in the dataset due to the varying class labels assigned by different participants. The <italic>walking</italic> class is the most consistently annotated class across participants and annotation methods, although it may not represent the maximum amount of labeled data points. Notably, the <italic>void</italic> class, which is not visible in <xref ref-type="fig" rid="F4">Figure 4</xref>, accounts for a substantial portion of the labeled data, ranging from 80% to 96%, depending on the annotation method employed.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>This figure illustrates the relative prevalence of various activity classes within the dataset, excluding the <italic>void</italic> class, see <xref ref-type="table" rid="T1">Table 1</xref> for details. The class labeled as <italic>void</italic> represents the predominant category within the dataset, surpassing the frequency of the second most prevalent class, <italic>desk_work</italic>, by a substantial factor of 13. The figure illustrates a pronounced imbalance in the data distribution, both in terms of the annotation methodology employed and the distinct week-specific patterns observed in the annotation process.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0004.tif"/>
</fig>
<p><xref ref-type="table" rid="T1">Table 1</xref> shows that daily activities that do not require extensive planning and are inherent to most people&#x00027;s everyday lives tend to be the most consistently annotated classes. Among these are activities such as <italic>walking, cycling, car_driving</italic>, or <italic>cooking</italic>. On the other hand, activities like <italic>badminton, weightlifting, manual_work</italic>, and others are highly subject-dependent and occur only sporadically. The observation that the <italic>desk_work</italic> class exhibits the highest frequency is valid solely when employing the &#x002462; diary- or the &#x002463; GUI-methodology for data collection, suggesting a potential limitation or bias associated with this particular approach. The classes pertaining to physical activities such as <italic>running, bus_driving, yoga, badminton, weightlifting</italic>, and <italic>sport</italic> exhibit a distinct pattern of clustering within specific weeks, indicating a temporal dependency. This observation highlights the inherent bias introduced by the real-world recording environment in which the dataset was collected, potentially limiting the generalizability of the model to broader contexts. Furthermore, it is crucial to note that the size of the <italic>void</italic> class for the recall methods &#x002462; and &#x002463; is up to 16% smaller than the <italic>in situ</italic> methods &#x002460; and &#x002461;. While this disparity in class representation highlights the inherent complexities and challenges associated with the data collection and annotation procedures, it simultaneously presents an opportunity to address real-world imbalances and biases. By critically examining and accounting for these factors, the resulting models can potentially enhance their generalizability and applicability across diverse scenarios, ultimately contributing to a deeper understanding of the underlying phenomena under investigation.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>This table presents a comprehensive overview of the number of data points for each activity class, categorized according to the different annotation methods employed during two distinct weeks.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center" colspan="3"><bold>Week 1</bold></th>
<th valign="top" align="center" colspan="3"><bold>Week 2</bold></th>
<th valign="top" align="center" colspan="2"><bold>Instances</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>Button</bold></th>
<th valign="top" align="center"><bold>App</bold></th>
<th valign="top" align="center"><bold>Diary</bold></th>
<th valign="top" align="center"><bold>Button</bold></th>
<th valign="top" align="center"><bold>App</bold></th>
<th valign="top" align="center"><bold>Diary</bold></th>
<th valign="top" align="center"><bold>GUI</bold></th>
<th valign="top" align="center"><bold>Total</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Laying</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">135,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
</tr> <tr>
<td valign="top" align="left">Sitting</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">292,500</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">607,500</td>
<td valign="top" align="center">264,041</td>
<td valign="top" align="center">13</td>
</tr> <tr>
<td valign="top" align="left">Walking</td>
<td valign="top" align="center">1,409,177</td>
<td valign="top" align="center">1,380,112</td>
<td valign="top" align="center">2,462,923</td>
<td valign="top" align="center">1,446,000</td>
<td valign="top" align="center">1,168,525</td>
<td valign="top" align="center">2,160,712</td>
<td valign="top" align="center">2,228,337</td>
<td valign="top" align="center">245</td>
</tr> <tr>
<td valign="top" align="left">Running</td>
<td valign="top" align="center">16,500</td>
<td valign="top" align="center">74,600</td>
<td valign="top" align="center">84,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">4</td>
</tr> <tr>
<td valign="top" align="left">Cycling</td>
<td valign="top" align="center">573,000</td>
<td valign="top" align="center">574,100</td>
<td valign="top" align="center">647,880</td>
<td valign="top" align="center">616,500</td>
<td valign="top" align="center">330,000</td>
<td valign="top" align="center">610,500</td>
<td valign="top" align="center">784,224</td>
<td valign="top" align="center">92</td>
</tr> <tr>
<td valign="top" align="left">Bus_driving</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">67,500</td>
<td valign="top" align="center">15,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2</td>
</tr> <tr>
<td valign="top" align="left">Car_driving</td>
<td valign="top" align="center">708,000</td>
<td valign="top" align="center">659,950</td>
<td valign="top" align="center">1,683,176</td>
<td valign="top" align="center">340,500</td>
<td valign="top" align="center">155,450</td>
<td valign="top" align="center">1,025,985</td>
<td valign="top" align="center">954,882</td>
<td valign="top" align="center">111</td>
</tr> <tr>
<td valign="top" align="left">Vacuum_cleaning</td>
<td valign="top" align="center">49,500</td>
<td valign="top" align="center">66,000</td>
<td valign="top" align="center">67,500</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">42,000</td>
<td valign="top" align="center">60,000</td>
<td valign="top" align="center">47,955</td>
<td valign="top" align="center">10</td>
</tr> <tr>
<td valign="top" align="left">Laundry</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">27,000</td>
<td valign="top" align="center">30,000</td>
<td valign="top" align="center">33,000</td>
<td valign="top" align="center">76,500</td>
<td valign="top" align="center">112,500</td>
<td valign="top" align="center">106,316</td>
<td valign="top" align="center">19</td>
</tr> <tr>
<td valign="top" align="left">Cooking</td>
<td valign="top" align="center">229,500</td>
<td valign="top" align="center">45,632</td>
<td valign="top" align="center">269,132</td>
<td valign="top" align="center">214,500</td>
<td valign="top" align="center">96,000</td>
<td valign="top" align="center">382,500</td>
<td valign="top" align="center">419,848</td>
<td valign="top" align="center">44</td>
</tr> <tr>
<td valign="top" align="left">Eating</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">5,753</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">75,000</td>
<td valign="top" align="center">92,850</td>
<td valign="top" align="center">5</td>
</tr> <tr>
<td valign="top" align="left">Shopping</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">57,000</td>
<td valign="top" align="center">60,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">45,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
</tr> <tr>
<td valign="top" align="left">Showering</td>
<td valign="top" align="center">112,500</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">225,000</td>
<td valign="top" align="center">34,500</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">251,777</td>
<td valign="top" align="center">165,721</td>
<td valign="top" align="center">20</td>
</tr> <tr>
<td valign="top" align="left">Yoga</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">22,500</td>
<td valign="top" align="center">30,680</td>
<td valign="top" align="center">2</td>
</tr> <tr>
<td valign="top" align="left">Sport</td>
<td valign="top" align="center">39,000</td>
<td valign="top" align="center">46,088</td>
<td valign="top" align="center">45,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
</tr> <tr>
<td valign="top" align="left">Playing_games</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">648,926</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">45,000</td>
<td valign="top" align="center">100,771</td>
<td valign="top" align="center">6</td>
</tr> <tr>
<td valign="top" align="left">Desk_work</td>
<td valign="top" align="center">36,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3,646,985</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">4,537,832</td>
<td valign="top" align="center">1,137,214</td>
<td valign="top" align="center">21</td>
</tr> <tr>
<td valign="top" align="left">Guitar_playing</td>
<td valign="top" align="center">49,500</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">172,500</td>
<td valign="top" align="center">111,000</td>
<td valign="top" align="center">67,060</td>
<td valign="top" align="center">217,500</td>
<td valign="top" align="center">234,355</td>
<td valign="top" align="center">24</td>
</tr> <tr>
<td valign="top" align="left">Gardening</td>
<td valign="top" align="center">48,000</td>
<td valign="top" align="center">53,125</td>
<td valign="top" align="center">43,718</td>
<td valign="top" align="center">69,000</td>
<td valign="top" align="center">69,650</td>
<td valign="top" align="center">375,000</td>
<td valign="top" align="center">421,794</td>
<td valign="top" align="center">16</td>
</tr> <tr>
<td valign="top" align="left">Table_tennis</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">17,875</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">190,500</td>
<td valign="top" align="center">108,725</td>
<td valign="top" align="center">105,000</td>
<td valign="top" align="center">171,177</td>
<td valign="top" align="center">29</td>
</tr> <tr>
<td valign="top" align="left">Badminton</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">234,000</td>
<td valign="top" align="center">210,000</td>
<td valign="top" align="center">252,725</td>
<td valign="top" align="center">6</td>
</tr> <tr>
<td valign="top" align="left">Horse_riding</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">66,000</td>
<td valign="top" align="center">199,500</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">519,000</td>
<td valign="top" align="center">495,000</td>
<td valign="top" align="center">502,742</td>
<td valign="top" align="center">18</td>
</tr> <tr>
<td valign="top" align="left">Cleaning</td>
<td valign="top" align="center">70,500</td>
<td valign="top" align="center">100,500</td>
<td valign="top" align="center">240,000</td>
<td valign="top" align="center">48,000</td>
<td valign="top" align="center">40,500</td>
<td valign="top" align="center">502,500</td>
<td valign="top" align="center">215,034</td>
<td valign="top" align="center">25</td>
</tr> <tr>
<td valign="top" align="left">Reading</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">45,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
</tr> <tr>
<td valign="top" align="left">Weightlifting</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">84,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60,000</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2</td>
</tr> <tr>
<td valign="top" align="left">Manual_work</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">69,425</td>
<td valign="top" align="center">157,587</td>
<td valign="top" align="center">3</td>
</tr> <tr>
<td valign="top" align="left">Dish_washing</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">10,500</td>
<td valign="top" align="center">21,000</td>
<td valign="top" align="center">52,500</td>
<td valign="top" align="center">66,246</td>
<td valign="top" align="center">9</td>
</tr>
<tr>
<td valign="top" align="left">Void</td>
<td valign="top" align="center">61,581,367</td>
<td valign="top" align="center">61,685,562</td>
<td valign="top" align="center">53,948,051</td>
<td valign="top" align="center">60,240,422</td>
<td valign="top" align="center">60,510,012</td>
<td valign="top" align="center">51,369,691</td>
<td valign="top" align="center">55,083,923</td>
<td valign="top" align="center">539</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The columns are divided into two main sections, representing Week 1 and Week 2 of the data collection process. Within each week, the data points are further subdivided based on the annotation method used, labeled as button, app, diary, GUI (for Week 2 only).</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>6.2 Missing annotations and consistency across methods</title>
<p>Missing Annotations and the consistency of labels set over the course of one week varied greatly depending on the study participant. However, tendencies with regard to specific methods are observable. We computed missing annotations by merging all available annotations from the various methods used (button, app, diary, GUI) into an artificial global ground truth. We then compared each individual annotation layer against this consolidated ground truth to identify any missing annotations, leveraging the collective information as a reference point. Method &#x002460;, pressing the situ button on the smartwatch&#x00027;s case, was not consistently used by every participant. Furthermore, this method carries the risk that either setting one of the two markers (start or end) is forgotten. An annotation where one marker is missing becomes therefore obsolete. The app-assisted annotation method &#x002461;, for which we used the app Strava, is well accepted among the participants who agreed with using third-party software. However, 4 participants, namely <italic>74e4, 90a4, d8f2</italic>, and <italic>f30d</italic> did not use the app continuously or refused to use it completely due to concerns regarding their private data. Strava is a commercial app, that is freely available for download on the app stores, but it collects certain users&#x00027; metadata. To label a time period with Strava, the participant needs to (1) take the smartphone, (2) open the app, (3) start a timer, set a label, and (4) end the timer. This procedure contains significantly more steps than other methods. Therefore, the average value of missing annotations results in 46.40% (week 1) and 56.79% (week 2). One participant found the annotation process in general very tedious and therefore dropped out of the study. These data have been excluded from the dataset and the evaluation. Method &#x002462; <italic>pure self-recall</italic>, writing an activity diary, got well accepted by every participant. As <xref ref-type="fig" rid="F5">Figure 5</xref> shows and the results in <xref ref-type="table" rid="T2">Table 2</xref> proof, it is overall the most complete annotation method with an average amount of missing annotations of 4.30% for the first and 8.14% for the second week. By introducing the MAD-GUI, participants were able to inspect their daily data, get insights into what patterns of specific classes look like, and label them interactively. With an average amount of missing annotations of 7.67%, this method became the most complete during the second study week. <xref ref-type="table" rid="T3">Table 3</xref> shows the resulting Cohen &#x003BA; scores. Due to the constraint that only one labeling method can be compared to a second one and since, according to <xref ref-type="table" rid="T2">Table 2</xref>, the most consistent annotation methods are &#x02462; <italic>pure self-recall</italic> and &#x02463; <italic>time-series recall</italic>, we used these methods as our baseline and compared them with every other method used in the study. The second column indicates the comparison direction. The abbreviations used in this column are defined as follows: (&#x02462; C/W &#x02460;) <italic>pure self-recall</italic> compared with <italic>in situ button</italic>, (&#x02462; C/W &#x02461;) <italic>pure self-recall</italic> compared with <italic>in situ app</italic> and (&#x02462; C/W &#x02463;) <italic>pure self-recall</italic> compared with <italic>time-series recall</italic>. The direction (&#x02463; C/W &#x02462;) is not explicitly included since Cohens &#x003BA; is bidirectional and both directions result in the same score. The score indicates how similar two annotators, or in our study labeling methods, are to each other. The resulting score is a decimal value between &#x02013;1.0 and 1.0, where &#x02013;1.0 means that the two annotators differ at most and 1.0 means complete similarity. 0.0 denotes that the target method was not used on that specific day.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Missing annotations across all study participants and both weeks. The <italic>Y</italic>-axis shows the total number of annotations of one specific participant for the corresponding week. The color codes are as follows: <inline-graphic xlink:href="fcomp-06-1379788-i0001.tif"/> Annotation is missing, <inline-graphic xlink:href="fcomp-06-1379788-i0002.tif"/> Annotation is partially missing (start or stop time), <inline-graphic xlink:href="fcomp-06-1379788-i0003.tif"/> Annotation is complete. The figure is inspired by Brenner (<xref ref-type="bibr" rid="B10">1999</xref>), <xref ref-type="fig" rid="F1">Figure 1</xref>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0005.tif"/>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Missing annotations across all labeling methods (in %) of both weeks.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Subject</bold></th>
<th valign="top" align="center"><bold>2b88</bold></th>
<th valign="top" align="center"><bold>36fd</bold></th>
<th valign="top" align="center"><bold>74e4</bold></th>
<th valign="top" align="center"><bold>90a4</bold></th>
<th valign="top" align="center"><bold>834b</bold></th>
<th valign="top" align="center"><bold>4531</bold></th>
<th valign="top" align="center"><bold>a506</bold></th>
<th valign="top" align="center"><bold>d8f2</bold></th>
<th valign="top" align="center"><bold>eed7</bold></th>
<th valign="top" align="center"><bold>f30d</bold></th>
<th valign="top" align="center"><bold>fc25</bold></th>
<th valign="top" align="center"><bold>Avg</bold>.</th>
</tr>
</thead>
<tbody>
<tr style="background-color:#dee1e1">
<td valign="top" align="left" colspan="13"><bold>Week 1</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02460;<italic>in situ button</italic></td>
<td valign="top" align="center">40</td>
<td valign="top" align="center">70.59</td>
<td valign="top" align="center">79.41</td>
<td valign="top" align="center">52.18</td>
<td valign="top" align="center">36.37</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">26.32</td>
<td valign="top" align="center">96.15</td>
<td valign="top" align="center">45.46</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">26.81</td>
<td valign="top" align="center">40.95</td>
</tr> <tr>
<td valign="top" align="left">&#x02461;<italic>in situ app</italic></td>
<td valign="top" align="center">13.30</td>
<td valign="top" align="center">5.89</td>
<td valign="top" align="center">97.06</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">5.00</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">36.84</td>
<td valign="top" align="center">92.30</td>
<td valign="top" align="center">22.73</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">4.35</td>
<td valign="top" align="center">43.40</td>
</tr> <tr>
<td valign="top" align="left">&#x02462;<italic>pure self-recall</italic></td>
<td valign="top" align="center">6.67</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">4.55</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">31.58</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">4.55</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">4.30</td>
</tr> <tr style="background-color:#dee1e1">
<td valign="top" align="left" colspan="13"><bold>Week 2</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02460;<italic>in situ button</italic></td>
<td valign="top" align="center">23.08</td>
<td valign="top" align="center">73.33</td>
<td valign="top" align="center">92.00</td>
<td valign="top" align="center">82.14</td>
<td valign="top" align="center">8.33</td>
<td valign="top" align="center">76.47</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">95.33</td>
<td valign="top" align="center">61.11</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">27.78</td>
<td valign="top" align="center">49.05</td>
</tr> <tr>
<td valign="top" align="left">&#x02461;<italic>in situ app</italic></td>
<td valign="top" align="center">61.54</td>
<td valign="top" align="center">6.67</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">89.29</td>
<td valign="top" align="center">8.33</td>
<td valign="top" align="center">35.30</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">79.16</td>
<td valign="top" align="center">33.33</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">11.11</td>
<td valign="top" align="center">56.79</td>
</tr> <tr>
<td valign="top" align="left">&#x02462;<italic>pure self-recall</italic></td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">8.57</td>
<td valign="top" align="center">17.88</td>
<td valign="top" align="center">4.17</td>
<td valign="top" align="center">35.30</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">12.50</td>
<td valign="top" align="center">5.56</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">5.56</td>
<td valign="top" align="center">8.14</td>
</tr> <tr>
<td valign="top" align="left">&#x02463;<italic>time-series recall</italic></td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">22.88</td>
<td valign="top" align="center">39.29</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">16.67</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">5.56</td>
<td valign="top" align="center">7.67</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The columns contain the subject-ID of all participants. The last column shows the average percentage of missing annotations across every labeling method, for all participants.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Average similarity between annotation methods according to the Cohan &#x003BA; score for both study weeks.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Week, day</bold></th>
<th valign="top" align="center"><bold>Direction</bold></th>
<th valign="top" align="center"><bold>2b88</bold></th>
<th valign="top" align="center"><bold>36fd</bold></th>
<th valign="top" align="center"><bold>4531</bold></th>
<th valign="top" align="center"><bold>74e4</bold></th>
<th valign="top" align="center"><bold>834b</bold></th>
<th valign="top" align="center"><bold>90a4</bold></th>
<th valign="top" align="center"><bold>a506</bold></th>
<th valign="top" align="center"><bold>d8f2</bold></th>
<th valign="top" align="center"><bold>eed7</bold></th>
<th valign="top" align="center"><bold>f30d</bold></th>
<th valign="top" align="center"><bold>fc25</bold></th>
<th valign="top" align="center"><bold>Avg</bold>.</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" rowspan="2">{1,2}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.35</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.23</td>
<td valign="top" align="center">0.58</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.49</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="2">{1,2}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.05</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.74</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.05</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.73</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="2">{1,3}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.05</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.51</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.54</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="2">{1,4}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.0</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="2">{1,5}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.87</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">&#x02212;0.31</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.89</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="2">{1,6}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.07</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">0.84</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">&#x02212;0.14</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.15</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">&#x02212;0.07</td>
<td valign="top" align="center">0.41</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.84</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="2">{1,7}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.525</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.49</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.15</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.10</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.77</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{2,1}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.10</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.41</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.85</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.81</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.10</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center">0.11</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.46</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.58</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.18</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.55</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{2,2}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.21</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.05</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.36</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.81</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.46</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.11</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">1.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.45</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.46</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{2,3}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">0.68</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.82</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.54</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">&#x02212;0.18</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.58</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">1.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.41</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.65</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{2,4}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.05</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.23</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.78</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.14</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.26</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.87</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.92</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">1.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.76</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.85</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{2,5}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.95</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.41</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.54</td>
<td valign="top" align="center">0.54</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.92</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.16</td>
<td valign="top" align="center">0.26</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.40</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.47</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">0.5</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.46</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.44</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{2,6}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.20</td>
<td valign="top" align="center">0.86</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.87</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.86</td>
<td/>
</tr> <tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.83</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.82</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{2,7}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.40</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.86</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.41</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.14</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.86</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.24</td>
<td valign="top" align="center">&#x02212;0.01</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.84</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.78</td>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.78</td>
<td/>
</tr> <tr>
<td valign="top" align="left" rowspan="5">{Avg.Week2}</td>
<td valign="top" align="center">&#x02462; C/W &#x02460;</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.23</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.39</td>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02461;</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.36</td>
</tr>
<tr>
<td valign="top" align="center">&#x02462; C/W &#x02463;</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.12</td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.52</td>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02460;</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.15</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.07</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">0.48</td>
</tr>
<tr>
<td valign="top" align="center">&#x02463; C/W &#x02461;</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.21</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.0</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.41</td>
</tr></tbody>
</table>
</table-wrap>
<p>Comparing the &#x02462; <italic>pure self-recall</italic> method with the &#x02460; <italic>in situ button</italic> and &#x02461; <italic>in situ app</italic> method we can see that the final results for weeks 1 and 2 are proximate to one another. &#x02461; <italic>Pure self-recall</italic> compared with the &#x02463; <italic>time-series recall</italic> results in the highest similarity of 0.52. The comparison between the &#x02463; <italic>time-series recall</italic> and the &#x02460; <italic>in situ button</italic> as well as the &#x02461; <italic>in situ app</italic> assisted annotations result in higher similarity than the prior comparison of &#x02462; <italic>pure self-recall</italic> vs. both methods &#x02460; and &#x02461;. This means that subjects rather agree to the timestamps of the <italic>in situ</italic> methods than to a self-written activity diary as soon as they can visually inspect the accelerometer data.</p>
</sec>
<sec>
<title>6.3 Visual time-series analysis</title>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> shows exemplary the time-series of the sixth day of every participant&#x00027;s second week. The four bars that are visible above the accelerometer data are the labels set by the participants for every layer. The order is from bottom to top: &#x02460; <italic>in situ button</italic>, &#x02461; <italic>in situ app</italic>. &#x02461; <italic>pure self-recall</italic>, and &#x02463; <italic>time-series recall</italic>. Examples of labels that differ with regard to the applied labeling method are marked with red boxes. The <italic>x</italic>-axis of every subplot represents roughly 8&#x02013;9 h of data. Most of the day was not labeled and is therefore categorized as <italic>void</italic>. However, such long periods often contain shorter periods of other activities, like <italic>walking</italic>. This makes it difficult to define a distinguishable <italic>void</italic>-class, which results in false positive classifications of non-void samples. <xref ref-type="fig" rid="F6">Figure 6</xref> visually shows that each participant labels his or her data very subjectively. The long green-labeled periods of participant <italic>74e4</italic> represent the class <italic>desk_work</italic>. The only other participant that used this label is <italic>90a4</italic>. Since each of the study participants works in an office environment and thus conclusively works at a desk, we can assume that the same class is classified as <italic>void</italic> for all other study participants. This intra-class and inter-participant discrepancy becomes a problem whenever a model is trained that is supposed to generalize across individuals. To reduce these side effects and focus on the experiment itself, we decided to evaluate personalized models that take weekly data from participants into account.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Visualization of participants&#x00027; accelerometer data on the sixth day in the second week of the study, together with annotations set by them. The four layers in the upper part of every participant&#x00027;s daily data represent the four annotation methods. The order is from bottom to top: &#x02460; <italic>in situ button</italic>, &#x02461; <italic>in situ app</italic>, &#x02461; <italic>pure self-recall</italic> and &#x02463; <italic>time-series recall</italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0006.tif"/>
</fig>
<p>The <italic>in situ button</italic> annotation is empty for five participants: <italic>eed7, 36fd, 74e4, 90a4</italic>, and <italic>d8f2</italic>. Labels are only partially set or missing entirely for this annotation method and we therefore assume that participants tend to forget to press the button on the smartwatch. Both <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref>, support this assumption, as this labeling method shows a high percentage of missing annotations as well as a low Cohen &#x003BA; score of 0.36% (week 1) and 0.39% (week 2). The <italic>pure self-recall</italic> method &#x02462;, visible on the 2nd upper layer, is often misaligned compared to the <italic>in situ</italic> methods as well as the <italic>time-series recall</italic> method &#x02463;. Participants tend to round up or down the start- and stop-time in steps of 5 or 10 min. For example, the annotations in <xref ref-type="fig" rid="F6">Figure 6</xref> given by the subjects <italic>2b88, 834b</italic>, or <italic>f30d</italic>, show such incorrectly annotated data. The pink color represents the class <italic>walking</italic>. With a closer look at the corresponding time-series data, one can see that the <italic>in situ button</italic> annotation (bottom layer) and <italic>time-series recall</italic> annotation (top layer) belongs to the typical periodic pattern of walking than the period labeled by <italic>pure self-recall</italic>.</p>
<p>A consistent reliable performance in all labeling methods can only be observed at the participants <italic>4531</italic> and <italic>fc25</italic>. Other participants like <italic>eed7, 36fd, 74e4</italic>, or <italic>a506</italic> are very precise in their annotations across methods, but are missing at least one layer of labels. The complete collection of visualizations is available in our dataset repository<xref ref-type="fn" rid="fn0007"><sup>7</sup></xref>.</p>
</sec>
<sec>
<title>6.4 Effects on classification</title>
<p>The results of our deep learning evaluation<xref ref-type="fn" rid="fn0008"><sup>8</sup></xref> suggest that the annotation method chosen can have a crucial impact on the classification ability of a trained deep learning model. Depending on the chosen methodology, the average F1-Score results differ by up to 8%, as depicted in <xref ref-type="fig" rid="F7">Figure 7</xref>. In the first week, the <italic>in situ</italic> methodologies, button &#x02460; and app &#x02461;, generally perform better than the <italic>pure self-recall</italic> diary &#x02461;. Study participants mostly correctly estimated the duration of an activity, but tended to round up or down the start and end times. The <italic>in situ</italic> methods are up to 8% better than the <italic>pure self-recall</italic>, although the amount of annotated data available, due to missing annotations, is significantly lower than for other methods. Although, we work with a dataset recorded in-the-wild, the deep learning results generally show a high F1-Score. This is untypical for such datasets but can be explained by the fact that the majority of the daily data are assigned to the <italic>void</italic> class. This leaves proportionally only a few samples that are crucial for determining the classification performance.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>The overall mean F1-Scores for the Leave-One-Day-Out Cross Validation across all participants. In the first week, the participants used methods &#x02460; - &#x02461;. In the second week, we introduced method &#x02463;.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0007.tif"/>
</fig>
<p>Even though the number of available annotations that have been labeled by the study participants using the <italic>time-series recall</italic> method &#x02463; is significantly higher with 92.33%, the average F1-Score is 1.1% lower (89.00%) than the results reached with the App Assisted method (90.1%). To understand this result it is crucial to look at <xref ref-type="table" rid="T4">Table 4</xref> in detail and take meta-information about the participants into account. The participants mostly used their diary as a mnemonic aid for the graphical annotation method and tried to identify the corresponding periods in the acceleration data. The results of subjects <italic>2b88, a506</italic> and <italic>eed7</italic> show that the performance of the classifier could be increased with graphical assistance. However, the F1-Score of <italic>2b88</italic> is 0.01% below the F1-Score of the <italic>in situ app</italic> assisted annotation method &#x02461;. These subjects have in common that they are already trained in interpreting acceleration data due to their prior knowledge and thus assign samples to specific classes more precisely.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>In detail representation of the final F1-Scores for every annotation methodology and a week per study participant.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Subject</bold></th>
<th valign="top" align="center"><bold>2b88</bold></th>
<th valign="top" align="center"><bold>a506</bold></th>
<th valign="top" align="center"><bold>eed7</bold></th>
<th valign="top" align="center"><bold>fc25</bold></th>
<th valign="top" align="center"><bold>4531</bold></th>
<th valign="top" align="center"><bold>834b</bold></th>
<th valign="top" align="center"><bold>Average</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#dee1e1">
<td valign="top" align="left" colspan="8"><bold>Week 1</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02460;<italic>in situ button</italic></td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">90.4</td>
</tr> <tr>
<td valign="top" align="left">&#x02461;<italic>in situ app</italic></td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">85.5</td>
</tr> <tr>
<td valign="top" align="left">&#x02462;<italic>pure self-recall</italic></td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">83.0</td>
</tr> <tr style="background-color:#dee1e1">
<td valign="top" align="left" colspan="8"><bold>Week 2</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02460;<italic>in situ button</italic></td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">86.8</td>
</tr> <tr>
<td valign="top" align="left">&#x02461;<italic>in situ app</italic></td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">na</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">90.1</td>
</tr> <tr>
<td valign="top" align="left">&#x02462;<italic>pure self-recall</italic></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">82.1</td>
</tr> <tr>
<td valign="top" align="left">&#x02463;<italic>time-series recall</italic></td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">89.0</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The average F1-Scores are graphically visualized in <xref ref-type="fig" rid="F7">Figure 7</xref>.</p>
</table-wrap-foot>
</table-wrap>
<p>Subjects <italic>fc25, 4531</italic>, and <italic>834b</italic>, on the other hand, do not have prior knowledge. Apart from subject <italic>834b</italic>, the deep learning results show that presenting a visualization to an untrained participant rather harmed than helped the classifier. If one looks at the visualizations of day 1 and 6, week 2 of <italic>fc25</italic> (see <xref ref-type="fig" rid="F6">Figure 6</xref>, <xref ref-type="fig" rid="F8">8</xref>), the labels set by the subject with the help of the graphical interface, it is comprehensible that this study participant tended to be rather confused by the graphical representation and therefore labeled the data incorrectly.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Visualization of the 1st day in week 2 of subject <italic>fc25</italic>. Differences can be seen in the upper annotation layer (&#x02463; <italic>time-series recall</italic>), exhibiting larger differences regarding the annotated start- and stop times compared to other methods.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-06-1379788-g0008.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s7">
<title>7 Discussion</title>
<p>In our 2-week long-term study, we recorded the acceleration data of 11 participants using a smartwatch and analyzed it visually, statistically, and using deep learning. The findings of the visual and statistical analysis were confirmed by the deep learning result. They show that the underlying annotation procedure is crucial for the quality of the annotations and the success of the deep learning model.</p>
<p>The <italic>in situ button</italic> method &#x02460; offers accuracy but brings the risk that the setting of a label is forgotten entirely or incompletely set. However, this method can be combined with additional on-device feedback or a smartphone app, so that greater accuracy and consistency of the annotation can be achieved. This involves a considerable implementation effort, which many scientists avoid because such projects, although of their significant value to the community, attract little attention in the scientific world. The use of existing, but often commercial, software and hardware is all too often accompanied by a loss of privacy. As our research has shown in passing, many users therefore shy away from using such products.</p>
<p>Through our investigation of the consistency of annotations between methodologies, we were able to show that participants in our study seem to prefer to write an activity diary (<italic>pure self-recall</italic> method &#x02462;). This finding corresponds to what (Vaizman et al., <xref ref-type="bibr" rid="B53">2018</xref>) already points out. However, this method has the disadvantage that it can be imprecise, which is evident in the visualization of the data and annotations. Similarly, the activity diary methodology performed the least reliably among all methodologies, which has been confirmed by the deep learning model. Since the deep learning results using the <italic>in situ app</italic> annotations &#x02462; are almost similar to the results given by the <italic>time-series recall</italic> &#x02463;, even though the number of labeled samples is lower, it raises the question if a smaller set of high-quality annotations is more valuable for a classifier than a larger set of annotated data that comes with imprecise labels. This could mean that in future works we can reduce the amount of necessary training samples drastically if a certain annotation quality can be assured. However, this needs to be confirmed by further investigations.</p>
<p>Some participants reported that they found the support provided by the visual representation of the data helpful. The resulting Cohen &#x003BA; scores strengthen this impression since the F1-Scores are much higher when we compare the <italic>time-series recall</italic> with both <italic>in situ</italic> methods vs. the <italic>pure self-recall</italic>. This indicates that as soon as the participants received a visual inspection tool, they tended to annotate data at similar time periods as through the <italic>in situ</italic> methods since they can easily identify periods of activity that roughly correspond to the execution time they remember. Our participants reported similar preferences, which led us to the conclusion that a digital diary that includes data visualization could combine the benefits of both annotation methods.</p>
<p>However, the study also showed that participants can find it difficult to interpret the acceleration data correctly and thus set inaccurate annotations. As our trained models show, this also has a strong influence on the classification result. If such a tool is to be made available to study participants, it must be ensured that they have the necessary knowledge and tools to be able to interpret these data. Thus, to ensure the success of future long-term and real-world activity recognition projects, prior training of the study participants regarding data interpretation is of crucial importance if a data visualization is supposed to be used.</p>
<p>Apart from trying to solve annotation difficulties during the annotation phase itself, we can also partially counter wrong or noisy classified data by using machine learning techniques like Bootstrapping (see Miu et al., <xref ref-type="bibr" rid="B32">2015</xref>) or using a loss function that specifically tries to counteract this problem, such as Natarajan et al. (<xref ref-type="bibr" rid="B34">2013</xref>) and Ma et al. (<xref ref-type="bibr" rid="B29">2019</xref>). By using Bootstrapping, the machine or deep learning classifier is initially trained by a small subset of high-confident labels and further improved by using additional data. However, this technique comes with the trade-off that whenever wrong-labeled data is introduced as training data, the error will get propagated into the model. An effect that sooner or later occurs as long as the annotation methodologies themselves are not further researched. Other machine learning techniques that can work with noisy labels (see Song et al., <xref ref-type="bibr" rid="B46">2022</xref>), are already successfully tested for Computer Vision problems and can, in theory, be adopted for Human Activity Recognition. However, earlier research has shown that not every technique that is applicable in other fields is also applicable to sensor-based data (Hoelzemann and Van Laerhoven, <xref ref-type="bibr" rid="B22">2020</xref>).</p>
<p>Cause of missing annotations: We believe that specific activities are more likely to be forgotten during the labeling process than others. These activities are generally more spontaneous and require less dedicated preparation time. Examples might include classes like <italic>laying, sitting, walking, bus_driving, car_driving, eating</italic>, or <italic>desk_work</italic>. In contrast, other classes like <italic>shopping, yoga, playing_games, badminton, cooking</italic>, or <italic>horse_riding</italic> are often time-intensive, physically or mentally demanding, and frequently planned in advance or even take place at dedicated locations. Therefore, it is likely that participants find these activities easier to recall and label accurately. Obtaining separate annotations for each activity through distinct and dedicated annotation processes would have yielded valuable insights; however, this approach was deemed unfeasible for the participants involved in our study. The immense time commitment and laborious efforts required from our participants to annotate each activity individually would have imposed an unreasonable burden, rendering such a comprehensive annotation strategy impractical within the constraints of our study.</p>
<sec>
<title>7.1 Discussing different annotation biases</title>
<p>Directly quantifying the perceived workload of subjective tasks like data labeling is a complex challenge. This difficulty stems from several factors. Firstly, individual differences in mental stamina and task perception mean what one person finds laborious, another might find manageable (Smith et al., <xref ref-type="bibr" rid="B45">2019</xref>). Secondly, memory biases can lead to under- or overestimates of effort depending on the emotional context of the task or the participant&#x00027;s current state (Watkins, <xref ref-type="bibr" rid="B59">2002</xref>). Social desirability bias can also come into play, with participants potentially downplaying their workload to appear competent or exaggerating it to justify breaks (Chung and Monroe, <xref ref-type="bibr" rid="B12">2003</xref>). Therefore, accurately quantifying the workload associated with each of the four labeling methods presents a significant challenge. While surveys, like the NASA TLX (NASA, <xref ref-type="bibr" rid="B33">1986</xref>), asking participants about perceived effort hold some value, these results are inherently subjective and can be heavily influenced by individual experiences and biases. While aiming for a fully objective measure of workload is desirable, it might require collecting more personal data from participants. This additional data could include details like preferred wearable devices (e.g., smartwatches), smartphone usage patterns, or individual memory recall capabilities or the emotional state of a participant (Ghosh et al., <xref ref-type="bibr" rid="B17">2015</xref>). While recording and quantifying this type of personal data would have provided valuable insights, it would have also significantly increased the workload placed on participants. This additional workload fell outside the scope of the current study, which prioritized collecting data through the four predefined methods. However, we need to acknowledge that several biases could have been introduced due to the chosen annotation guidelines and tools. For example, the usage of <italic>in situ</italic> annotation methods during the day can have a positive effect on the self-recall capabilities of a participant at the end of the day. The comparison of consistencies across methods does not confirm that this effect indeed occurred. Every study participant showed an almost complete overall profile of self-recall annotations, even though the person has not used or has incomplete <italic>in situ</italic> annotations (see <xref ref-type="fig" rid="F5">Figure 5</xref>). However, deeper investigations are needed to be able to understand such effects better.</p>
<p>Yordanova et al. (<xref ref-type="bibr" rid="B61">2018</xref>) lists the following three biases for sensor-based human activity data: Self-Recall bias (Valuri et al., <xref ref-type="bibr" rid="B54">2005</xref>), Behavior bias (Friesen et al., <xref ref-type="bibr" rid="B15">2020</xref>) and the Self-Annotation bias (Yordanova et al., <xref ref-type="bibr" rid="B61">2018</xref>). We showed that indeed a time-deviation bias (which can be seen as a self-recall bias) has been introduced to annotations created with the <italic>pure self-recall</italic> method &#x02461;, and that such a bias affects the classifier negatively. However, visualizing the sensor data can counter this effect because it was easier for participants to detect active phases in hindsight.</p>
<p>A behavior bias can be neglected, because the participants were not monitored by a person or video camera during the day and the minimalistic setup of one wrist-worn smartwatch does not influence one&#x00027;s behavior since the wearing comfort of such a device is generally perceived as positive (Pal et al., <xref ref-type="bibr" rid="B37">2020</xref>). A self-annotation bias, a bias that occurs if the annotator labels their data in an isolated environment and cannot refer to an expert to verify an annotation, did occur as well. With the deep learning analysis, we were able to show that the classifier was less negatively impacted by this bias than by time-deviation bias.</p>
<p>A Parallel annotation bias can arise in two scenarios: when multiple annotators independently label the same data, or when a single annotator uses multiple labeling methods for the same data, where the application of one method influences the subsequent labeling decisions made with other methods. There are three main ways this bias can manifest:</p>
<list list-type="order">
<list-item><p>Anchoring bias (Lieder et al., <xref ref-type="bibr" rid="B28">2018</xref>): the initial labeling method might act as an anchor, subtly influencing the annotator&#x00027;s decisions when using subsequent methods, even if their initial assessment might differ.</p></list-item>
<list-item><p>Confirmation bias (Klayman, <xref ref-type="bibr" rid="B25">1995</xref>): the annotator might subconsciously favor interpretations that align with labels generated from previous methods, overlooking alternative possibilities.</p></list-item>
<list-item><p>Method bias (Min et al., <xref ref-type="bibr" rid="B31">2016</xref>): certain methods might inherently be easier or more difficult to use for specific types of data, potentially leading to systematic inconsistencies across the labeled data.</p></list-item>
</list>
<p>The presence of parallel annotation bias in this context suggests that the annotations might not be entirely independent between methods, potentially impacting the overall quality of the data. Anchoring and confirmation bias can lead to a lack of diversity in annotations and potentially perpetuate errors. Method bias can introduce inconsistencies that complicate data analysis. We recognize the possibility of parallel annotation bias in our dataset, where applying one labeling method might influence subsequent methods used by the same participant. However, prioritizing participant engagement, we opted for a parallel approach. This decision ensured the workload remained manageable and prevented participant dropout from the study.</p></sec>
</sec>
<sec sec-type="conclusions" id="s8">
<title>8 Conclusions</title>
<p>We argue that the annotation methodologies for benchmark datasets in Human Activity Recognition do not yet capture the attention it should. Data annotation is a laborious and time-consuming task that often cannot be performed accurately and conscientiously without the right tools. However, there is a very limited number of tools that can be used for this purpose and often they do not pass the prototype status.</p>
<p>Only a few scientific publications, such as Reining et al. (<xref ref-type="bibr" rid="B38">2020</xref>), focus on annotations and their quality. However, the use of properly annotated data drastically affects the final capacities of the trained machine or deep learning model. Therefore, we consider our study to be important for the HAR community, as it analyzes this topic in greater depth and thus provides important insights that go beyond the current state of science. <xref ref-type="table" rid="T5">Table 5</xref> summarizes the advantages and disadvantages of every method.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Comparison of advantages and disadvantages of all annotation methods used in this study.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Methodology</bold></th>
<th valign="top" align="left"><bold>Advantages</bold></th>
<th valign="top" align="left"><bold>Disadvantages</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">&#x02460;<italic>in situ button</italic></td>
<td valign="top" align="left">- Easy to implement and use - Can be improved with feedback mechanisms</td>
<td valign="top" align="left">- Participants tend to forget pressing a button<break/>- Many incomplete annotations that become unusable for the classifier</td>
</tr> <tr>
<td valign="top" align="left">&#x02461;<italic>in situ app</italic></td>
<td valign="top" align="left">- Tracking apps are already widely used and accepted, therefore low acceptance threshold<break/>- Can be improved with feedback mechanisms or additional smartphone functionalities<break/>- List of possible annotations can be expanded with minimum effort<break/>- Participants tend to set very precise annotations</td>
<td valign="top" align="left">- Data and privacy concerns if a commercial app is used<break/>- Participants often forgot to set an annotation, especially when they were unfamiliar with tracking apps<break/>- Implementation workload may be very high</td>
</tr> <tr>
<td valign="top" align="left">&#x02462;<italic>pure self-recall</italic></td>
<td valign="top" align="left">- Easy to use even without technical knowledge (a handwritten diary)<break/>- Most accepted method in our experiment<break/>- Annotations are very consistent</td>
<td valign="top" align="left">- Can be very imprecise<break/>- Only suitable for coarse activity labels and activities that were performed for long periods of time, like <italic>walking</italic> or <italic>running</italic></td>
</tr> <tr>
<td valign="top" align="left">&#x02463;<italic>time-series</italic> recall</td>
<td valign="top" align="left">- Visualization of data helps participants to set annotations more accurate than using the pure self-recall method &#x02462;</td>
<td valign="top" align="left">- Available tools are often in the state of a prototype and need additional developments and adjustments and are therefore not impromptu usable<break/>- Participants need to be trained to be able to interpret sensor data</td>
</tr></tbody>
</table>
</table-wrap>
<p>High-quality annotations are crucial for accurate activity recognition, especially in uncontrolled real-world settings where video recordings are unavailable for ground truth verification. To address this challenge, further research on activity data annotation methodologies is necessary. These methodologies should empower annotators to label data in a way that comprehensively captures the subtleties of everyday life. The annotations must not only be extensive but also complete and coherent, ensuring a consistent and well-defined understanding for the AI model to learn from. Furthermore, leveraging learning methodologies like Weakly Supervised Learning methods, exemplified by works such as Wang et al. (<xref ref-type="bibr" rid="B57">2019</xref>) and Wang et al. (<xref ref-type="bibr" rid="B58">2021</xref>), can potentially utilize datasets like ours. However, a more comprehensive evaluation is needed to determine their suitability for real-world application. The combination of a (handwritten) diary with a correction aided by a data visualization in hindsight shows the best results in terms of consistency and missing annotations and provides accurate start and end times. However, this combination results in additional work for the study participants and therefore, remains a trade-off between additional workload and annotation quality.</p>
<sec>
<title>8.1 Lesson learned</title>
<p>During this study, we gained insights about the effects of different annotation methods on the reliability and consistency of annotations and finally on the classifier itself, but also about training deep learning models on data recorded in-the-wild. In this chapter, we would like to share these insights to help other researchers perform their experiments more successfully. With regards to <xref ref-type="table" rid="T5">Table 5</xref>, we are able to narrow down specific study setups that either benefit more from self-recall or <italic>in situ</italic> annotation methods. As part of our annotation guidelines, we allowed our study participants to name their activities as they wished. Therefore, we were forced to simplify certain activities. To be able to create a real-world dataset that contains complex classes or even classes that consist of several subclasses, more elaborated annotation methods and tools must be developed. We believe that with the currently available resources, the hurdle lies very high for such datasets to be annotated accurately.</p>
<p>Our study includes people who cycle to work in their daily work routine and others who commute by public transport or work in a home office environment. Thus, each study participant has his or her set of daily repetitive activities. Due to the nature of our dataset as one recorded in a real-world and long-term scenario, the number of labeled samples is rather small, and given labels vary participant-dependent. This mix of factors creates a bias in the dataset and we concluded that a cross-participant train-/test-strategy is not appropriate for our study design and would not give meaningful insights, since every study participant has their own set of unique activities which are too different and hardly generalizable. Therefore, for certain studies, the commonly known and accepted Leave-One-Subject-Out Cross-Validation is not suitable.</p></sec>
</sec>
<sec sec-type="data-availability" id="s9">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/supplementary material.</p></sec>
<sec sec-type="ethics-statement" id="s10">
<title>Ethics statement</title>
<p>The studies involving humans were approved by Ethics Committee of the University of Siegen, ethics vote &#x00023;ER 12 2019. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.</p></sec>
<sec sec-type="author-contributions" id="s11">
<title>Author contributions</title>
<p>AH: Writing &#x02013; original draft, Visualization, Validation, Software, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. KV: Writing &#x02013; review &#x00026; editing, Supervision, Project administration, Methodology, Funding acquisition, Conceptualization.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s12">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This project was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) &#x02013; 425868829 and is part of Priority Program SPP2199 Scalable Interaction Paradigms for Pervasive Computing Environments.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The author(s) declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.</p>
</sec>
<sec sec-type="disclaimer" id="s13">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://www.strava.com/">https://www.strava.com/</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://www.espruino.com/banglejs">https://www.espruino.com/banglejs</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup><ext-link ext-link-type="uri" xlink:href="https://www.apple.com/watch/">https://www.apple.com/watch/</ext-link></p></fn>
<fn id="fn0004"><p><sup>4</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/ahoelzemann/mad-gui-adaptions/">https://github.com/ahoelzemann/mad-gui-adaptions/</ext-link></p></fn>
<fn id="fn0005"><p><sup>5</sup>Our smartwatch firmware is made publicly available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/kristofvl/BangleApps/tree/master/apps/activate">https://github.com/kristofvl/BangleApps/tree/master/apps/activate</ext-link>.</p></fn>
<fn id="fn0006"><p><sup>6</sup>Our web-tool is made publicly available at: <ext-link ext-link-type="uri" xlink:href="https://ubi29.informatik.uni-siegen.de/upload/">https://ubi29.informatik.uni-siegen.de/upload/</ext-link>.</p></fn>
<fn id="fn0007"><p><sup>7</sup><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.7654684">https://doi.org/10.5281/zenodo.7654684</ext-link></p></fn>
<fn id="fn0008"><p><sup>8</sup>Detailed results for every participant included in our deep learning evaluation can be accessed online on the Weights &#x00026; Biases platform: <ext-link ext-link-type="uri" xlink:href="https://tinyurl.com/4vxvfaed">https://tinyurl.com/4vxvfaed</ext-link>.</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abney</surname> <given-names>S.</given-names></name></person-group> (<year>2002</year>). <article-title>&#x0201C;Bootstrapping,&#x0201D;</article-title> in <source>Proceedings of the 40th annual meeting of the Association for Computational Linguistics</source> (<publisher-loc>Philadelphia, PA</publisher-loc>), <fpage>360</fpage>&#x02013;<lpage>367</lpage>. <pub-id pub-id-type="doi">10.3115/1073083.1073143</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Adaimi</surname> <given-names>R.</given-names></name> <name><surname>Thomaz</surname> <given-names>E.</given-names></name></person-group> (<year>2019</year>). <article-title>Leveraging active learning and conditional mutual information to minimize data annotation in human activity recognition</article-title>. <source>Proc. ACM Interact. Mob. Wearable Ubiquitous Technol</source>. <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1145/3351228</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Akbari</surname> <given-names>A.</given-names></name> <name><surname>Martinez</surname> <given-names>J.</given-names></name> <name><surname>Jafari</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>Facilitating human activity data annotation via context-aware change detection on smartwatches</article-title>. <source>ACM Trans. Embed. Comput. Syst</source>. <volume>20</volume>, <fpage>1</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1145/3431503</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Artstein</surname> <given-names>R.</given-names></name> <name><surname>Poesio</surname> <given-names>M.</given-names></name></person-group> (<year>2008</year>). <article-title>Inter-coder agreement for computational linguistics</article-title>. <source>Comput. Linguist</source>. <volume>34</volume>, <fpage>555</fpage>&#x02013;<lpage>596</lpage>. <pub-id pub-id-type="doi">10.1162/coli.07-034-R2</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bartolo</surname> <given-names>M.</given-names></name> <name><surname>Roberts</surname> <given-names>A.</given-names></name> <name><surname>Welbl</surname> <given-names>J.</given-names></name> <name><surname>Riedel</surname> <given-names>S.</given-names></name> <name><surname>Stenetorp</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Beat the AI: investigating adversarial human annotation for reading comprehension</article-title>. <source>Trans. Assoc. Comput. Linguist</source>. <volume>8</volume>, <fpage>662</fpage>&#x02013;<lpage>678</lpage>. <pub-id pub-id-type="doi">10.1162/tacl_a_00338</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Berlin</surname> <given-names>E.</given-names></name> <name><surname>Van Laerhoven</surname> <given-names>K.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Detecting leisure activities with dense motif discovery,&#x0201D;</article-title> in <source>Proceedings of the 2012 ACM Conference on Ubiquitous Computing</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>250</fpage>&#x02013;<lpage>259</lpage>. <pub-id pub-id-type="doi">10.1145/2370216.2370257</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bock</surname> <given-names>M.</given-names></name> <name><surname>Hoelzemann</surname> <given-names>A.</given-names></name> <name><surname>Moeller</surname> <given-names>M.</given-names></name> <name><surname>Van Laerhoven</surname> <given-names>K.</given-names></name></person-group> (<year>2022</year>). <article-title>Investigating (re) current state-of-the-art in human activity recognition datasets</article-title>. <source>Front. Comput. Sci</source>. <volume>4</volume>:<fpage>924954</fpage>. <pub-id pub-id-type="doi">10.3389/fcomp.2022.924954</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bock</surname> <given-names>M.</given-names></name> <name><surname>H&#x000F6;lzemann</surname> <given-names>A.</given-names></name> <name><surname>Moeller</surname> <given-names>M.</given-names></name> <name><surname>Van Laerhoven</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Improving deep learning for har with shallow lstms,&#x0201D;</article-title> in <source>2021 International Symposium on Wearable Computers</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>7</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1145/3460421.3480419</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bota</surname> <given-names>P.</given-names></name> <name><surname>Silva</surname> <given-names>J.</given-names></name> <name><surname>Folgado</surname> <given-names>D.</given-names></name> <name><surname>Gamboa</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>A semi-automatic annotation approach for human activity recognition</article-title>. <source>Sensors</source> <volume>19</volume>:<fpage>501</fpage>. <pub-id pub-id-type="doi">10.3390/s19030501</pub-id><pub-id pub-id-type="pmid">30691040</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brenner</surname> <given-names>S. E.</given-names></name></person-group> (<year>1999</year>). <article-title>Errors in genome annotation</article-title>. <source>Trends Genet</source>. <volume>15</volume>, <fpage>132</fpage>&#x02013;<lpage>133</lpage>. <pub-id pub-id-type="doi">10.1016/S0168-9525(99)01706-0</pub-id><pub-id pub-id-type="pmid">10203816</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bulling</surname> <given-names>A.</given-names></name> <name><surname>Blanke</surname> <given-names>U.</given-names></name> <name><surname>Schiele</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>A tutorial on human activity recognition using body-worn inertial sensors</article-title>. <source>ACM Comput. Surv</source>. <volume>46</volume>:<fpage>33</fpage>. <pub-id pub-id-type="doi">10.1145/2499621</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chung</surname> <given-names>J.</given-names></name> <name><surname>Monroe</surname> <given-names>G. S.</given-names></name></person-group> (<year>2003</year>). <article-title>Exploring social desirability bias</article-title>. <source>J. Bus. Ethics</source> <volume>44</volume>, <fpage>291</fpage>&#x02013;<lpage>302</lpage>. <pub-id pub-id-type="doi">10.1023/A:1023648703356</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cleland</surname> <given-names>I.</given-names></name> <name><surname>Han</surname> <given-names>M.</given-names></name> <name><surname>Nugent</surname> <given-names>C.</given-names></name> <name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>McClean</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Evaluation of prompted annotation of activity data recorded from a smart phone</article-title>. <source>Sensors</source> <volume>14</volume>, <fpage>15861</fpage>&#x02013;<lpage>15879</lpage>. <pub-id pub-id-type="doi">10.3390/s140915861</pub-id><pub-id pub-id-type="pmid">25166500</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cruz-Sandoval</surname> <given-names>D.</given-names></name> <name><surname>Beltran-Marquez</surname> <given-names>J.</given-names></name> <name><surname>Garcia-Constantino</surname> <given-names>M.</given-names></name> <name><surname>Gonzalez-Jasso</surname> <given-names>L. A.</given-names></name> <name><surname>Favela</surname> <given-names>J.</given-names></name> <name><surname>Lopez-Nava</surname> <given-names>I. H.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Semi-automated data labeling for activity recognition in pervasive healthcare</article-title>. <source>Sensors</source> <volume>19</volume>, <fpage>3035</fpage>. <pub-id pub-id-type="doi">10.3390/s19143035</pub-id><pub-id pub-id-type="pmid">31295850</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friesen</surname> <given-names>K. B.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Monaghan</surname> <given-names>P. G.</given-names></name> <name><surname>Oliver</surname> <given-names>G. D.</given-names></name> <name><surname>Roper</surname> <given-names>J. A.</given-names></name></person-group> (<year>2020</year>). <article-title>All eyes on you: how researcher presence changes the way you walk</article-title>. <source>Sci. Rep</source>. <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-020-73734-5</pub-id><pub-id pub-id-type="pmid">33051502</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gentile</surname> <given-names>A. L.</given-names></name> <name><surname>Gruhl</surname> <given-names>D.</given-names></name> <name><surname>Ristoski</surname> <given-names>P.</given-names></name> <name><surname>Welch</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Explore and exploit. dictionary expansion with human-in-the-loop,&#x0201D;</article-title> in <source>The Semantic Web: 16th International Conference, ESWC 2019, Portoro&#x0017E;, Slovenia, June 2-6, 2019, Proceedings 16</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>131</fpage>&#x02013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-21348-0_9</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ghosh</surname> <given-names>A.</given-names></name> <name><surname>Danieli</surname> <given-names>M.</given-names></name> <name><surname>Riccardi</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Annotation and prediction of stress and workload from physiological and inertial signals,&#x0201D;</article-title> in <source>2015 37th annual international conference of the IEEE engineering in medicine and biology society (EMBC)</source> (<publisher-loc>Milan</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1621</fpage>&#x02013;<lpage>1624</lpage>. <pub-id pub-id-type="doi">10.1109/EMBC.2015.7318685</pub-id><pub-id pub-id-type="pmid">26736585</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gjoreski</surname> <given-names>H.</given-names></name> <name><surname>Ciliberto</surname> <given-names>M.</given-names></name> <name><surname>Morales</surname> <given-names>F. J. O.</given-names></name> <name><surname>Roggen</surname> <given-names>D.</given-names></name> <name><surname>Mekki</surname> <given-names>S.</given-names></name> <name><surname>Valentin</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;A versatile annotated dataset for multimodal locomotion analytics with mobile devices,&#x0201D;</article-title> in <source>Proceedings of the 15th ACM Conference on Embedded Network Sensor Systems</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>2</lpage>. <pub-id pub-id-type="doi">10.1145/3131672.3136976</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gjoreski</surname> <given-names>H.</given-names></name> <name><surname>Ciliberto</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Morales</surname> <given-names>F. J. O.</given-names></name> <name><surname>Mekki</surname> <given-names>S.</given-names></name> <name><surname>Valentin</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>The university of sussex-huawei locomotion and transportation dataset for multimodal analytics with mobile devices</article-title>. <source>IEEE Access</source> <volume>6</volume>, <fpage>42592</fpage>&#x02013;<lpage>42604</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2018.2858933</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hassan</surname> <given-names>I.</given-names></name> <name><surname>Mursalin</surname> <given-names>A.</given-names></name> <name><surname>Salam</surname> <given-names>R. B.</given-names></name> <name><surname>Sakib</surname> <given-names>N.</given-names></name> <name><surname>Haque</surname> <given-names>H. Z.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Autoact: an auto labeling approach based on activities of daily living in the wild domain,&#x0201D;</article-title> in <source>2021 Joint 10th International Conference on Informatics, Electronics</source> &#x00026; <italic>Vision (ICIEV) and 2021 5th International Conference on Imaging, Vision</italic> &#x00026; <italic>Pattern Recognition (icIVPR)</italic> (Kitakyushu: IEEE), <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/ICIEVicIVPR52578.2021.9564211</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hoelzemann</surname> <given-names>A.</given-names></name> <name><surname>Odoemelem</surname> <given-names>H.</given-names></name> <name><surname>Van Laerhoven</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Using an in-ear wearable to annotate activity data across multiple inertial sensors,&#x0201D;</article-title> in <source>Proceedings of the 1st International Workshop on Earable Computing</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>14</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1145/3345615.3361136</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hoelzemann</surname> <given-names>A.</given-names></name> <name><surname>Van Laerhoven</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Digging deeper: towards a better understanding of transfer learning for human activity recognition,&#x0201D;</article-title> in <source>Proceedings of the 2020 International Symposium on Wearable Computers</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>50</fpage>&#x02013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1145/3410531.3414311</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huynh</surname> <given-names>D. T. G.</given-names></name></person-group> (<year>2008</year>). <source>Human actIvity Recognition with Wearable Sensors</source>. <publisher-loc>Darmstadt</publisher-loc>: <publisher-name>Technische Universit&#x000E4;t Darmstadt</publisher-name>, <fpage>59</fpage>&#x02013;<lpage>65</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ioffe</surname> <given-names>S.</given-names></name> <name><surname>Szegedy</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Batch normalization: accelerating deep network training by reducing internal covariate shift,&#x0201D;</article-title> in <source>International conference on machine learning</source> (<publisher-loc>Lille</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>448</fpage>&#x02013;<lpage>456</lpage>.<pub-id pub-id-type="pmid">35496726</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klayman</surname> <given-names>J.</given-names></name></person-group> (<year>1995</year>). <article-title>Varieties of confirmation bias</article-title>. <source>Psychol. Learn. Motiv</source>. <volume>32</volume>, <fpage>385</fpage>&#x02013;<lpage>418</lpage>. <pub-id pub-id-type="doi">10.1016/S0079-7421(08)60315-1</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Klie</surname> <given-names>J.-C.</given-names></name> <name><surname>de Castilho</surname> <given-names>R. E.</given-names></name> <name><surname>Gurevych</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;From zero to hero: human-in-the-loop entity linking in low resource domains,&#x0201D;</article-title> in <source>Proceedings of the 58th annual meeting of the association for computational linguistics</source> (<publisher-loc>Stroudsbourg, PA</publisher-loc>: <publisher-name>ACL</publisher-name>), <fpage>6982</fpage>&#x02013;<lpage>6993</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2020.acl-main.624</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leonardis</surname> <given-names>A.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name> <name><surname>Maver</surname> <given-names>J.</given-names></name></person-group> (<year>2002</year>). <article-title>Multiple eigenspaces</article-title>. <source>Pattern Recognit</source>. <volume>35</volume>, <fpage>2613</fpage>&#x02013;<lpage>2627</lpage>. <pub-id pub-id-type="doi">10.1016/S0031-3203(01)00198-4</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lieder</surname> <given-names>F.</given-names></name> <name><surname>Griffiths</surname> <given-names>T. L.</given-names></name> <name><surname>Huys</surname> <given-names>M.</given-names></name> <name><surname>Goodman</surname> <given-names>Q. J. N. D.</given-names></name></person-group> (<year>2018</year>). <article-title>The anchoring bias reflects rational use of cognitive resources</article-title>. <source>Psychon. Bull. Rev</source>. <volume>25</volume>, <fpage>322</fpage>&#x02013;<lpage>349</lpage>. <pub-id pub-id-type="doi">10.3758/s13423-017-1286-8</pub-id><pub-id pub-id-type="pmid">28484952</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>Z.</given-names></name> <name><surname>Wei</surname> <given-names>X.</given-names></name> <name><surname>Hong</surname> <given-names>X.</given-names></name> <name><surname>Gong</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Bayesian loss for crowd count estimation with point supervision,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF international conference on computer vision</source> (<publisher-loc>Seoul</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6142</fpage>&#x02013;<lpage>6151</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00624</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mekruksavanich</surname> <given-names>S.</given-names></name> <name><surname>Jitpattanakul</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Recognition of real-life activities with smartphone sensors using deep learning approaches,&#x0201D;</article-title> in <source>IEEE 12th International Conference on Software Engineering and Service Science (ICSESS)</source> (<publisher-loc>Beijing</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>243</fpage>&#x02013;<lpage>246</lpage>. <pub-id pub-id-type="doi">10.1109/ICSESS52187.2021.9522231</pub-id><pub-id pub-id-type="pmid">38339452</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Min</surname> <given-names>H.</given-names></name> <name><surname>Park</surname> <given-names>J.</given-names></name> <name><surname>Kim</surname> <given-names>H. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Common method bias in hospitality research: a critical review of literature and an empirical study</article-title>. <source>Int. J. Hosp. Manag</source>. <volume>56</volume>, <fpage>126</fpage>&#x02013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijhm.2016.04.010</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Miu</surname> <given-names>T.</given-names></name> <name><surname>Missier</surname> <given-names>P.</given-names></name> <name><surname>Pl&#x000F6;tz</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Bootstrapping personalised human activity recognition models using online active learning,&#x0201D;</article-title> in <source>2015 IEEE International Conference on Computer and Information Technology; Ubiquitous Computing and Communications; Dependable, Autonomic and Secure Computing; Pervasive Intelligence and Computing</source> (<publisher-loc>Liverpool</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1138</fpage>&#x02013;<lpage>1147</lpage>. <pub-id pub-id-type="doi">10.1109/CIT/IUCC/DASC/PICOM.2015.170</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><collab>NASA</collab></person-group> (<year>1986</year>). <source>NASA task load index (NASA-TLX), version 1.0: Paper and pencil package</source>. Moffett Field, CA.</citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Natarajan</surname> <given-names>N.</given-names></name> <name><surname>Dhillon</surname> <given-names>I. S.</given-names></name> <name><surname>Ravikumar</surname> <given-names>P. K.</given-names></name> <name><surname>Tewari</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Learning with noisy labels,&#x0201D;</article-title> in <source>Advances in neural information processing systems</source> (<publisher-loc>Lake Tahoe, NV</publisher-loc>), 26.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ollenschl&#x000E4;ger</surname> <given-names>M.</given-names></name> <name><surname>K&#x000FC;derle</surname> <given-names>A.</given-names></name> <name><surname>Mehringer</surname> <given-names>W.</given-names></name> <name><surname>Seifer</surname> <given-names>A.-K.</given-names></name> <name><surname>Winkler</surname> <given-names>J.</given-names></name> <name><surname>Ga&#x000DF;ner</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Mad gui: an open-source python package for annotation and analysis of time-series data</article-title>. <source>Sensors</source> <volume>22</volume>:<fpage>5849</fpage>. <pub-id pub-id-type="doi">10.3390/s22155849</pub-id><pub-id pub-id-type="pmid">35957406</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ord&#x000F3;&#x000F1;ez</surname> <given-names>F.</given-names></name> <name><surname>Roggen</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep convolutional and LSTM recurrent neural networks for multimodal wearable activity recognition</article-title>. <source>Sensors</source> <volume>16</volume>:<fpage>115</fpage>. <pub-id pub-id-type="doi">10.3390/s16010115</pub-id><pub-id pub-id-type="pmid">26797612</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pal</surname> <given-names>D.</given-names></name> <name><surname>Funilkul</surname> <given-names>S.</given-names></name> <name><surname>Vanijja</surname> <given-names>V.</given-names></name></person-group> (<year>2020</year>). <article-title>The future of smartwatches: assessing the end-users&#x00027; continuous usage using an extended expectation-confirmation model</article-title>. <source>Univers. Access. Inf. Soc</source>. <volume>19</volume>, <fpage>261</fpage>&#x02013;<lpage>281</lpage>. <pub-id pub-id-type="doi">10.1007/s10209-018-0639-z</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Reining</surname> <given-names>C.</given-names></name> <name><surname>Rueda</surname> <given-names>F. M.</given-names></name> <name><surname>Niemann</surname> <given-names>F.</given-names></name> <name><surname>Fink</surname> <given-names>G. A. ten Hompel, M.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Annotation performance for multi-channel time series har dataset in logistics,&#x0201D;</article-title> in <source>2020 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops)</source> (<publisher-loc>Austin, TX</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/PerComWorkshops48775.2020.9156170</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reyes-Ortiz</surname> <given-names>J.-L.</given-names></name> <name><surname>Oneto</surname> <given-names>L. Sam&#x000E0;, A.</given-names></name> <name><surname>Parra</surname> <given-names>X.</given-names></name> <name><surname>Anguita</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Transition-aware human activity recognition using smartphones</article-title>. <source>Neurocomputing</source> <volume>171</volume>, <fpage>754</fpage>&#x02013;<lpage>767</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2015.07.085</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Roggen</surname> <given-names>D.</given-names></name> <name><surname>Calatroni</surname> <given-names>A.</given-names></name> <name><surname>Rossi</surname> <given-names>M.</given-names></name> <name><surname>Holleczek</surname> <given-names>T.</given-names></name> <name><surname>F&#x000F6;rster</surname> <given-names>K.</given-names></name> <name><surname>Tr&#x000F6;ster</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>&#x0201C;Collecting complex activity datasets in highly rich networked sensor environments,&#x0201D;</article-title> in <source>2010 Seventh international conference on networked sensing systems (INSS)</source> (<publisher-loc>Kassel</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>233</fpage>&#x02013;<lpage>240</lpage>. <pub-id pub-id-type="doi">10.1109/INSS.2010.5573462</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Scholl</surname> <given-names>P. M.</given-names></name> <name><surname>Wille</surname> <given-names>M.</given-names></name> <name><surname>Van Laerhoven</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Wearables in the wet lab: a laboratory system for capturing and guiding experiments,&#x0201D;</article-title> in <source>Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>589</fpage>&#x02013;<lpage>599</lpage>. <pub-id pub-id-type="doi">10.1145/2750858.2807547</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Schr&#x000F6;der</surname> <given-names>M.</given-names></name> <name><surname>Yordanova</surname> <given-names>K.</given-names></name> <name><surname>Bader</surname> <given-names>S.</given-names></name> <name><surname>Kirste</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Tool support for the online annotation of sensor data,&#x0201D;</article-title> in <source>Proceedings of the 3rd International Workshop on Sensor-based Activity Recognition and Interaction</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1145/2948963.2948972</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="web"><person-group person-group-type="author"><collab>scikit-Learn</collab></person-group> (<year>2022</year>). <source>Cohen&#x00027;s kappa</source> - <italic>scikit-learn</italic>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://scikit-learn.org/stable/modules/generated/sklearn.metrics.cohen_kappa_score.html">https://scikit-learn.org/stable/modules/generated/sklearn.metrics.cohen_kappa_score.html</ext-link> (accessed Febuary 10, 2022).</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sculley</surname> <given-names>D.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Online active learning methods for fast label-efficient spam filtering,&#x0201D;</article-title> in <source>CEAS, Vol</source>. 7 (Mountain View, CA), <fpage>143</fpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>M. R.</given-names></name> <name><surname>Chai</surname> <given-names>R.</given-names></name> <name><surname>Nguyen</surname> <given-names>H. T.</given-names></name> <name><surname>Marcora</surname> <given-names>S. M.</given-names></name> <name><surname>Coutts</surname> <given-names>A. J.</given-names></name></person-group> (<year>2019</year>). <article-title>Comparing the effects of three cognitive tasks on indicators of mental fatigue</article-title>. <source>J. Psychol</source>. <volume>153</volume>, <fpage>759</fpage>&#x02013;<lpage>783</lpage>. <pub-id pub-id-type="doi">10.1080/00223980.2019.1611530</pub-id><pub-id pub-id-type="pmid">31188721</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>M.</given-names></name> <name><surname>Park</surname> <given-names>D.</given-names></name> <name><surname>Shin</surname> <given-names>Y.</given-names></name> <name><surname>Lee</surname> <given-names>J.-G.</given-names></name></person-group> (<year>2022</year>). <article-title>Learning from noisy labels with deep neural networks: a survey</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst</source>. <volume>34</volume>, <fpage>8135</fpage>&#x02013;<lpage>8153</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2022.3152527</pub-id><pub-id pub-id-type="pmid">35254993</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stikic</surname> <given-names>M.</given-names></name> <name><surname>Larlus</surname> <given-names>D.</given-names></name> <name><surname>Ebert</surname> <given-names>S.</given-names></name> <name><surname>Schiele</surname> <given-names>B.</given-names></name></person-group> (<year>2011</year>). <article-title>Weakly supervised recognition of daily life activities with wearable sensors</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>33</volume>, <fpage>2521</fpage>&#x02013;<lpage>2537</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2011.36</pub-id><pub-id pub-id-type="pmid">21339526</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Stisen</surname> <given-names>A.</given-names></name> <name><surname>Blunck</surname> <given-names>H.</given-names></name> <name><surname>Bhattacharya</surname> <given-names>S.</given-names></name> <name><surname>Prentow</surname> <given-names>T. S.</given-names></name> <name><surname>Kj&#x000E6;rgaard</surname> <given-names>M. B.</given-names></name> <name><surname>Dey</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>&#x0201C;Smart devices are different: assessing and mitigatingmobile sensing heterogeneities for activity recognition,&#x0201D;</article-title> in <source>Proceedings of the 13th ACM conference on embedded networked sensor systems</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>127</fpage>&#x02013;<lpage>140</lpage>. <pub-id pub-id-type="doi">10.1145/2809695.2809718</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sztyler</surname> <given-names>T.</given-names></name> <name><surname>Stuckenschmidt</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;On-body localization of wearable devices: an investigation of position-aware activity recognition,&#x0201D;</article-title> in <source>IEEE International Conference on Pervasive Computing and Communications</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1109/PERCOM.2016.7456521</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tapia</surname> <given-names>E. M.</given-names></name> <name><surname>Intille</surname> <given-names>S. S.</given-names></name> <name><surname>Larson</surname> <given-names>K.</given-names></name></person-group> (<year>2004</year>). <article-title>&#x0201C;Activity recognition in the home using simple and ubiquitous sensors,&#x0201D;</article-title> in <source>International conference on pervasive computing</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>158</fpage>&#x02013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-540-24646-6_10</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Thomaz</surname> <given-names>E.</given-names></name> <name><surname>Essa</surname> <given-names>I.</given-names></name> <name><surname>Abowd</surname> <given-names>G. D.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;A practical approach for recognizing eating moments with wrist-mounted inertial sensing,&#x0201D;</article-title> in <source>Proceedings of the 2015 ACM international joint conference on pervasive and ubiquitous computing</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>1029</fpage>&#x02013;<lpage>1040</lpage>. <pub-id pub-id-type="doi">10.1145/2750858.2807545</pub-id><pub-id pub-id-type="pmid">29520397</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tonkin</surname> <given-names>E. L.</given-names></name> <name><surname>Burrows</surname> <given-names>A.</given-names></name> <name><surname>Woznowski</surname> <given-names>P. R.</given-names></name> <name><surname>Laskowski</surname> <given-names>P.</given-names></name> <name><surname>Yordanova</surname> <given-names>K. Y.</given-names></name> <name><surname>Twomey</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Talk, text, tag? Understanding self-annotation of smart home data from a user&#x00027;s perspective</article-title>. <source>Sensors</source> <volume>18</volume>:<fpage>2365</fpage>. <pub-id pub-id-type="doi">10.3390/s18072365</pub-id><pub-id pub-id-type="pmid">30037046</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vaizman</surname> <given-names>Y.</given-names></name> <name><surname>Ellis</surname> <given-names>K.</given-names></name> <name><surname>Lanckriet</surname> <given-names>G.</given-names></name> <name><surname>Weibel</surname> <given-names>N.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Extrasensory app: data collection in-the-wild with rich user interface to self-report behavior,&#x0201D;</article-title> in <source>Proceedings of the 2018 CHI conference on human factors in computing systems</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1145/3173574.3174128</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Valuri</surname> <given-names>G.</given-names></name> <name><surname>Stevenson</surname> <given-names>M.</given-names></name> <name><surname>Finch</surname> <given-names>C.</given-names></name> <name><surname>Hamer</surname> <given-names>P.</given-names></name> <name><surname>Elliott</surname> <given-names>B.</given-names></name></person-group> (<year>2005</year>). <article-title>The validity of a four week self-recall of sports injuries</article-title>. <source>Inj. Prev</source>. <volume>11</volume>, <fpage>135</fpage>&#x02013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1136/ip.2003.004820</pub-id><pub-id pub-id-type="pmid">15933402</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Van Laerhoven</surname> <given-names>K.</given-names></name> <name><surname>Kilian</surname> <given-names>D.</given-names></name> <name><surname>Schiele</surname> <given-names>B.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Using rhythm awareness in long-term activity recognition,&#x0201D;</article-title> in <source>2008 12th IEEE International Symposium on Wearable Computers</source> (<publisher-loc>Pittsburgh, PA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>63</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1109/ISWC.2008.4911586</pub-id><pub-id pub-id-type="pmid">33628834</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wallace</surname> <given-names>E.</given-names></name> <name><surname>Rodriguez</surname> <given-names>P.</given-names></name> <name><surname>Feng</surname> <given-names>S.</given-names></name> <name><surname>Yamada</surname> <given-names>I.</given-names></name> <name><surname>Boyd-Graber</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Trick me if you can: human-in-the-loop generation of adversarial examples for question answering</article-title>. <source>Trans. Assoc. Comput. Linguist</source>. <volume>7</volume>, <fpage>387</fpage>&#x02013;<lpage>401</lpage>. <pub-id pub-id-type="doi">10.1162/tacl_a_00279</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>Attention-based convolutional neural network for weakly labeled human activities&#x00027; recognition with wearable sensors</article-title>. <source>IEEE Sens. J</source>. <volume>19</volume>, <fpage>7598</fpage>&#x02013;<lpage>7604</lpage>. <pub-id pub-id-type="doi">10.1109/JSEN.2019.2917225</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>Sequential weakly labeled multiactivity localization and recognition on wearable sensors using recurrent attention networks</article-title>. <source>IEEE Trans. Hum.-Mach. Syst</source>. <volume>51</volume>, <fpage>355</fpage>&#x02013;<lpage>364</lpage>. <pub-id pub-id-type="doi">10.1109/THMS.2021.3086008</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watkins</surname> <given-names>P. C.</given-names></name></person-group> (<year>2002</year>). <article-title>Implicit memory bias in depression</article-title>. <source>Cogn. Emot</source>. <volume>16</volume>, <fpage>381</fpage>&#x02013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1080/02699930143000536</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>X.</given-names></name> <name><surname>Xiao</surname> <given-names>L.</given-names></name> <name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Ma</surname> <given-names>T.</given-names></name> <name><surname>He</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>A survey of human-in-the-loop for machine learning</article-title>. <source>Future Gener. Comput. Syst</source>. <volume>134</volume>, <fpage>365</fpage>&#x02013;<lpage>381</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2022.05.014</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yordanova</surname> <given-names>K. Y.</given-names></name> <name><surname>Paiement</surname> <given-names>A.</given-names></name> <name><surname>Schr&#x000F6;der</surname> <given-names>M.</given-names></name> <name><surname>Tonkin</surname> <given-names>E.</given-names></name> <name><surname>Woznowski</surname> <given-names>P.</given-names></name> <name><surname>Olsson</surname> <given-names>C. M.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Challenges in annotation of user data for ubiquitous systems: results from the 1st ARDUOUS workshop</article-title>. <source>arXiv</source> [Preprint]. <pub-id pub-id-type="doi">10.48550/arXiv.1803.05843</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>He</surname> <given-names>L.</given-names></name> <name><surname>Dragut</surname> <given-names>E.</given-names></name> <name><surname>Vucetic</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;How to invest my time: lessons from human-in-the-loop entity extraction,&#x0201D;</article-title> in <source>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery</source> &#x00026; <italic>Data Mining</italic> (New York, NY: ACM), <fpage>2305</fpage>&#x02013;<lpage>2313</lpage>. <pub-id pub-id-type="doi">10.1145/3292500.3330773</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>K.</given-names></name> <name><surname>Du</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>&#x0201C;Healthy: a diary system based on activity recognition using smartphone,&#x0201D;</article-title> in <source>2013 IEEE 10th international conference on mobile Ad-Hoc and sensor systems</source> (<publisher-loc>Hangzhou</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>290</fpage>&#x02013;<lpage>294</lpage>. <pub-id pub-id-type="doi">10.1109/MASS.2013.14</pub-id><pub-id pub-id-type="pmid">28546137</pub-id></citation></ref>
</ref-list>
</back>
</article>