<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Surg.</journal-id>
<journal-title>Frontiers in Surgery</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Surg.</abbrev-journal-title>
<issn pub-type="epub">2296-875X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fsurg.2021.742160</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Surgery</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Continuous Feature-Based Tracking of the Inner Ear for Robot-Assisted Microsurgery</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Marzi</surname> <given-names>Christian</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1199604/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Prinzen</surname> <given-names>Tom</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Haag</surname> <given-names>Julia</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1510929/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Klenzner</surname> <given-names>Thomas</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Mathis-Ullrich</surname> <given-names>Franziska</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Health Robotics and Automation, Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology</institution>, <addr-line>Karlsruhe</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Otorhinolaryngology, Head &#x00026; Neck Surgery, University-Hospital D&#x000FC;sseldorf</institution>, <addr-line>D&#x000FC;sseldorf</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Vincent Van Rompaey, University of Antwerp, Belgium</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: AB Zulkiflee, University Malaya Medical Centre, Malaysia; Yann Nguyen, Sorbonne Universit&#x000E9;s, France</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Franziska Mathis-Ullrich <email>franziska.ullrich&#x00040;kit.edu</email></corresp>
<corresp id="c002">Thomas Klenzner <email>thomas.klenzner&#x00040;med.uni-duesseldorf.de</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Otorhinolaryngology - Head and Neck Surgery, a section of the journal Frontiers in Surgery</p></fn></author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>8</volume>
<elocation-id>742160</elocation-id>
<history>
<date date-type="received">
<day>15</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Marzi, Prinzen, Haag, Klenzner and Mathis-Ullrich.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Marzi, Prinzen, Haag, Klenzner and Mathis-Ullrich</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>Robotic systems for surgery of the inner ear must enable highly precise movement in relation to the patient. To allow for a suitable collaboration between surgeon and robot, these systems should not interrupt the surgical workflow and integrate well in existing processes. As the surgical microscope is a standard tool, present in almost every microsurgical intervention and due to it being in close proximity to the situs, it is predestined to be extended by assistive robotic systems. For instance, a microscope-mounted laser for ablation. As both, patient and microscope are subject to movements during surgery, a well-integrated robotic system must be able to comply with these movements. To solve the problem of on-line registration of an assistance system to the situs, the standard of care often utilizes marker-based technologies, which require markers being rigidly attached to the patient. This not only requires time for preparation but also increases invasiveness of the procedure and the line of sight of the tracking system may not be obstructed. This work aims at utilizing the existing imaging system for detection of relative movements between the surgical microscope and the patient. The resulting data allows for maintaining registration. Hereby, no artificial markers or landmarks are considered but an approach for feature-based tracking with respect to the surgical environment in otology is presented. The images for tracking are obtained by a two-dimensional RGB stream of a surgical microscope. Due to the bony structure of the surgical site, the recorded cochleostomy scene moves nearly rigidly. The goal of the tracking algorithm is to estimate motion only from the given image stream. After preprocessing, features are detected in two subsequent images and their affine transformation is computed by a random sample consensus (RANSAC) algorithm. The proposed method can provide movement feedback with up to 93.2 &#x003BC;m precision without the need for any additional hardware in the operating room or attachment of fiducials to the situs. In long term tracking, an accumulative error occurs.</p></abstract>
<kwd-group>
<kwd>tracking</kwd>
<kwd>feature-based</kwd>
<kwd>microscope</kwd>
<kwd>image-processing</kwd>
<kwd>inner ear</kwd>
<kwd>robotic surgery</kwd>
<kwd>cochlea implantation</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="2"/>
<equation-count count="8"/>
<ref-count count="18"/>
<page-count count="10"/>
<word-count count="5592"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Otologic microsurgery requires the surgeon to work at the limit of their visuo-tactile feedback and dexterity. The procedure of a cochlea implantation, for example, consists traditionally of a manually drilled, nearly cone-shaped access beginning on the outer surface of the skull with a diameter of around 30 mm and tapered to a 2 mm narrow opening to the middle-ear (posterior tympanotomy). After visualization of the round window, the cochlea can be opened through the round window or a cochleostomy, an artificial opening drilled by the surgeon. The surgeon then has to move a 0.3&#x02013;1 mm thin electrode array through the posterior tympanotomy in the even more narrow cochlea. Robotic systems can exceed human precision in order of multiple magnitudes. Therefore, it is obvious that otologic microsurgery can highly benefit from robotic assistance.</p>
<p>When introducing novel technological robotic aids into surgery, space is often a critical factor. The closer to the surgical situs, the more important it is to keep the spacial obstruction to a minimum. In microsurgical interventions, a surgical microscope is always present. Therefore, mounting an assistive robotic manipulator to a microscope&#x00027;s optic unit poses high potential for robotic support. This allows for bringing the robot into close proximity of the situs while maintaining obstruction to the surgeon on a similar level as in regular microsurgery.</p>
<p>While being widely established for ablation of soft tissue (for example in ophthalmology), robotic laser surgery is gaining increasing interest in ablation of bone. In otologic surgery different kind of handheld lasers are used to penetrate the footplate of the stapes and more recently robotic guided lasers for interventions in the inner ear are taken into clinical trials (<xref ref-type="bibr" rid="B1">1</xref>). Also ablation of larger volumes of bone tissue could be demonstrated to be ready for clinical applications as for example by AOT&#x00027;s recent certification of CARLO (<xref ref-type="bibr" rid="B2">2</xref>), a laser osteotome mounted on a collaborative robot arm. The latter is applied in craniofacial surgery and provides cleaner cuts as well as additional freedom in cut geometry. Also the research project MIRACLE (<xref ref-type="bibr" rid="B3">3</xref>) aims on ablation of bone. However, in this case a minimally invasive robotic approach is pursued to reduce trauma. In addition, interventions at the inner ear are in focus of laser ablation of bone (<xref ref-type="bibr" rid="B4">4</xref>&#x02013;<xref ref-type="bibr" rid="B6">6</xref>). In combination with sensory feedback about residual bone tissue, laser ablation provides a precise tool for opening of the cochlea. Such robotic systems in particular, could greatly benefit from integration into a surgical microscope toward clinical translation.</p>
<p>However, integration into a movable microscope will pose the challenge for the robotic system to maintain precise registration to the patient. Modern microscopes provide robotic support with position encoders as well as interfaces to marker based registration systems (<xref ref-type="bibr" rid="B7">7</xref>). Still, registration may be interrupted or become inaccurate by small, sudden movements, which can occur due to unintended contact with the microscope or movement of the patient. Compensating for such motions will be a necessary skill for any microscope-mounted robotic system manipulating tissue.</p>
<p>Modern surgical microscopes provide a magnified image of the surgical scene and integrate cameras or adapters for camera attachment. Often, the recorded images can be streamed to monitors in the operation room (OR) by standardized interfaces. Thus, the magnified image provides information available at no additional cost of hardware. Utilizing these images to derive movement information for a robotic system would thus be integrateable without increased efforts. In addition, such an image based tracking system would gain precision from the microscopes magnification.</p>
<p>State of the art for tracking the surgical microscope (and other tools or the patient) remain retro-reflective markers detected optically by infrared (IR) cameras in combination with IR-LED (<xref ref-type="bibr" rid="B7">7</xref>). Recent works have focused on using features based tracking in microscopes images for augmentation and registration of preoperative data. For example in (<xref ref-type="bibr" rid="B8">8</xref>), the pose of the cochlea is augmented for navigation support. Here, <italic>Speeded Up Robust Features</italic> (SURF) were used for maintaining the augmented images registered. In (<xref ref-type="bibr" rid="B9">9</xref>), the tips of the instruments for microsurgical intervention had to be colored green to allow for pose estimation through the microscope&#x00027;s image.</p>
<p>Extending the modification of tools or introduction of fiducials, this work aims on processing the microscope images based on natural features to gather information of the relative movement between the microscope and the patient. These tracking information can then be made available to enable robotic assistance. Cochleostomy is used as an example for a common and standardized intervention with high potential for automation.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<sec>
<title>2.1. Imaging Setup</title>
<p>The investigated method aims on interfering as little as possible with the existing surgical workflow. This also means no additional hardware should be introduced into the operating room or, in particular, in proximity to the patient. Therefore, the existing imaging capabilities of commercial surgical microscopes should be utilized. To record microscope images, most conventional microscopes are equipped with standardized flanges to attach a camera as it is often used for documentation in current practice. Here, a computer with a frame grabber (DeckLink Recorder Mini 4K, Blackmagic Design Pty Ltd, Victoria, Australia) is used to gain access to the image frames. The processing computer is equipped with an Intel(R) Core(TM) i7-8086K CPU and GeForce GTX 1080 Ti GPU. These components are the only additional hardware, which could be easily positioned outside of the OR.</p>
</sec>
<sec>
<title>2.2. Image Processing Pipeline</title>
<sec>
<title>2.2.1. Framework</title>
<p>To facilitate data exchange and enable a connection to a future robotic system the Robot Operation System (ROS, Distribution <italic>Noetic</italic>) is used as a software framework. A ROS driver for the frame grabber was developed to provide the images from the microscope to ROS. The raw frames are submitted to the image processing node on a ROS topic. For representation of pose information, ROS&#x00027; dedicated data structure called <italic>TF-Three</italic> is used. It represents pose information in a hierarchical structure and is easily expandable and accessible in a network.</p>
</sec>
<sec>
<title>2.2.2. Scene Tracking</title>
<p>Due to the bony structure of the surgical site, the cochleostomy scene is assumed to move rigidly and tissue deformations can be neglected. Movement is tracked in 2D in the microscope image plane, as illustrated in <xref ref-type="fig" rid="F1">Figure 1</xref>. The proposed tracking algorithm provides an estimate of the relative motion between the surgical situs and microscope, given only the microscope&#x00027;s image stream and no further information. Motivated by microscope mounted robotic systems, this information would be sufficient to allow for compensation of unintended motion of either patient or microscope. In the proposed method, two subsequent images are compared and their affine transformation</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>00</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>01</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Overview of the proposed method. Two consecutive microscope image frames <italic>F</italic><sub><italic>i</italic></sub> and <italic>F</italic><sub><italic>i</italic>&#x0002B;1</sub> are processed to identify features, which are used to estimate the transformation <italic>T</italic><sub><italic>i</italic></sub> and maintain an initial registration by updating position <italic>x</italic><sub><italic>i</italic></sub> to <italic>x</italic><sub><italic>i</italic>&#x0002B;1</sub>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fsurg-08-742160-g0001.tif"/>
</fig>
<p>is estimated. The algorithm consists of three main steps. The flowchart in <xref ref-type="fig" rid="F2">Figure 2</xref> outlines the algorithm. First, a feature detection algorithm (see section 2.2.4) identifies distinct natural landmarks. Second, the identified features are matched. An example of these identified features is displayed in <xref ref-type="fig" rid="F3">Figure 3</xref>. Third, a transformation model between the established matches is estimated. Additional preprocessing to detect reflection artifacts in the images can increase tracking robustness for some surgical scenes. Here, thresholding is used to confine illumination artifacts in the field of view. The full image processing pipeline is illustrated in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>UML diagram of the proposed algorithm.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fsurg-08-742160-g0002.tif"/>
</fig>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Two microscope images of a moving situs. In each frame, features detected by the proposed algorithm are marked.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fsurg-08-742160-g0003.tif"/>
</fig>
</sec>
<sec>
<title>2.2.3. Image Preprocessing</title>
<p>Lighting-dependent artifacts appear as pixels with distinctively high color values in the microscope image. This delimits the affected points from their neighboring points. Accordingly, thresholding is a reasonable approach for reflection detection (<xref ref-type="bibr" rid="B10">10</xref>). As reflections are often prone to wrongly serving as detected features, thresholding is conducted before feature detection. It is conducted for each pixel comprising a saturation <italic>S</italic> and intensity <italic>I</italic>. If the statement in Equation (2) holds true, the pixel is added to the mask.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">max</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x0222A;</mml:mo><mml:mi>S</mml:mi><mml:mo>&#x0003E;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">max</mml:mtext></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here, <italic>I</italic><sub>max</sub> is the image&#x00027;s maximum intensity and <italic>S</italic><sub>max</sub> the image&#x00027;s maximum saturation. Parameters &#x003C4;<sub>1</sub> and &#x003C4;<sub>2</sub> are the respective thresholds, which were iteratively identified and evaluated. Sufficiently suitable values are given by &#x003C4;<sub>1</sub> &#x0003D; 0.8 and &#x003C4;<sub>2</sub> &#x0003D; 0.2. Preprocessing generates a mask, which excludes part of the images from further processing.</p>
</sec>
<sec>
<title>2.2.4. Feature Detection</title>
<p>The masked image is used to detect features utilizing the <italic>Oriented FAST and Rotated BRIEF</italic> (ORB) algorithm first presented by (<xref ref-type="bibr" rid="B11">11</xref>). It was developed as an alternative to the patented <italic>Scale Invariant Feature Transform</italic> (SIFT) algorithm (<xref ref-type="bibr" rid="B12">12</xref>). ORB is faster than SIFT and other alternatives like SURF, while being more sensitive to movements and more robust (<xref ref-type="bibr" rid="B13">13</xref>). The ORB feature detector is invariant to translation, rotation and scaling of the image, as well as robust against illumination changes and noise. The first step of the ORB algorithm is the detection of keypoints. These are generated by the <italic>Features from Accelerated Segment Test</italic> (FAST), which are combined with an orientation measure. For all keypoints found, a Binary Robust Independent Elementary Features (BRIEF) descriptor is computed. The number <italic>k</italic> of desired keypoints depends on the size of the obtained images. For 1080p images, <italic>k</italic> is suggested to be set to 2,000 according to the results by (<xref ref-type="bibr" rid="B14">14</xref>). Here, the ORB algorithm is implemented in Python using the image processing library OpenCV (<xref ref-type="bibr" rid="B15">15</xref>).</p>
</sec>
<sec>
<title>2.2.5. Transformation Model</title>
<p>Natural landmark detection results in a set of keypoints and their descriptors. Given two such sets obtained from images that share image features, the next step in our tracking algorithm is to find the corresponding matches between two images based on the detected features. The found matches are then used to estimate the affine transformation between these scenes. Since surgical scenes do not vary significantly in color or features it is likely that many keypoints are matched incorrectly despite the computed descriptors. Thus, a model estimation algorithm that is robust against a high ratio of mismatches (outliers) is required. The Random Sample Consensus (RANSAC) algorithm (<xref ref-type="bibr" rid="B16">16</xref>) estimates a model&#x00027;s parameters based on a set of data <italic>D</italic> which contains more points than are required for model description.</p>
<p>The desired model is the affine transformation <italic>T</italic> (see Equation 1). The set <italic>D</italic> is formed by tuples of ORB features with matching BRIEF descriptor in two subsequent images. The model is estimated to approximate the best affine transformation with respect to the translations of the features. For the developed image processing software, the implementation of RANSAC from the Python library <italic>scikit-images</italic> (<xref ref-type="bibr" rid="B17">17</xref>) was used. The affine transformations <italic>T</italic><sub><italic>n</italic></sub> of each iteration <italic>n</italic> can be cascaded to form an accumulated position <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the measured trajectory (formed by all <italic>x</italic><sub><italic>i</italic></sub> &#x02208; {1, &#x02026;, <italic>n</italic>}) of the relative movement of situs and microscope.</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:mover><mml:mo>&#x0220F;</mml:mo><mml:mi>n</mml:mi></mml:mover></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The evaluated position is passed to the TF-tree in ROS to easily be accessible by any connected robotic system.</p>
</sec>
</sec>
<sec>
<title>2.3. Experimental Evaluation</title>
<p>For evaluation of the proposed algorithm, a robot is used to create a precise reference movements of a specimen. The trajectories are captured through the microscope and the image processing pipeline estimates the movement. Comparing estimated movement and reference movement allows for determination of a tracking error.</p>
<p>The setup for evaluation consist of a commercial surgical microscope (OPMI Pro Magis/S8, Carl Zeiss AG, Oberkochen, Germany). A camera (Canon EOS 100D) is attached to the side port of the microscope, recording a video stream. The video stream is captured by the frame grabber card in the processing computer. Below the microscope the surgical scene is set up on a Stewart platform (M-850, Physik Instrumente GmbH, Karlsruhe, Germany) that allows for defined control of precise reference movement with a repeatability of 2&#x003BC;m. The robot is controlled by the processing computer using ROS. The complete experimental setup is depicted in <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Overview of the evaluation setup. A camera is attached to the side port of a surgical microscope. Below, the phantom is attached to a Stewart platform (covered by drapes). The robot is used for generating precise reference movement data.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fsurg-08-742160-g0004.tif"/>
</fig>
<p>To evaluate the presented tracking pipeline on several levels of realism and allow for comparison between different domains, three specimens are evaluated:</p>
<list list-type="order">
<list-item><p>A temporal bone phantom <bold>(TBP)</bold> (PHACON GmbH, Leipzig, Germany) comprising only of bone-like material (see <xref ref-type="fig" rid="F5">Figure 5A</xref>)</p></list-item>
<list-item><p>A temporal bone phantom comprising of bone-like material covered with multilayered skin-like material <bold>(TBPs)</bold> (PHACON GmbH, Leipzig, Germany). The skin incision is held apart by self-retaining retractors to facilitate good visualization (see <xref ref-type="fig" rid="F5">Figure 5B</xref>)</p></list-item>
<list-item><p>A cadaveric temporal bone <bold>(CTP)</bold>. The skin incision is held apart by self-retaining retractors to facilitate good visualization (see <xref ref-type="fig" rid="F5">Figure 5C</xref>).</p></list-item>
</list>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The three specimens evaluated, representing different levels of realism: <bold>(A)</bold> temporal bone phantom (TBP), <bold>(B)</bold> temporal bone phantom with skin-like material (TBPs), <bold>(C)</bold> cadaveric temporal bone (CTP).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fsurg-08-742160-g0005.tif"/>
</fig>
<p>All models and phantoms have been prepared to represent the last surgical phase before opening the cochlea. Therefore, a skin incision, mastoidectomy and posterior tympanotomy have been previously performed. The microscope is set up to provide a view similar to visualization during a surgical intervention.</p>
<p>The Stewart platform provides 6 degrees of freedom motion, however only translation movement along its x- and y-axes are used for reference motion (compare <xref ref-type="fig" rid="F4">Figure 4</xref>). For data recording, the x-axis and y-axis of the robot are aligned manually to the image axes.</p>
<p>Motion of the patient is then simulated by driving the robot along a predefined trajectory. First, linear translational movement in x- and y-directions are evaluated. To also cover combinations of x- and y-motion in the 2D image space, we additionally evaluated spiral motion of the robot. The processing of the image data, as well as the control of the robot and sampling of reference data were conducted on the same computer to allow for data synchronization. The data was recorded for later evaluation as <italic>rosbag</italic>, ROS&#x00027; data recording format. Equations (4) and (5) define the waypoints for the chosen trajectories. As soon as one waypoint was reached by the robot, the next one was passed to the robot&#x00027;s controller. In between the waypoints, the used controller interpolates a linear trajectory. The robot conducted the movement with its maximum speed of 2 mm/s. Linear translations in x- and y-directions (i.e., cross-shape) are defined by the waypoints</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo stretchy="true">{</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>10</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>-</mml:mo><mml:mn>10</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>10</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mo>-</mml:mo><mml:mn>10</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo stretchy="true">}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The spiral trajectory is defined by the waypoints</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M6"><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>c</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mn>10</mml:mn><mml:mfrac><mml:mi>n</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:mfrac><mml:mi>sin</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>4</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mfrac><mml:mi>n</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>10</mml:mn><mml:mfrac><mml:mi>n</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:mfrac><mml:mi>cos</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>4</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mfrac><mml:mi>n</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x02003;</mml:mtext><mml:msub><mml:mo>&#x0007C;</mml:mo><mml:mo>&#x02200;</mml:mo></mml:msub><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy='false'>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>25</mml:mn><mml:mo stretchy='false'>]</mml:mo><mml:mo>&#x02208;</mml:mo><mml:mi>&#x02115;</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Frame to Frame Precision</title>
<p>To evaluate the precision of the algorithm, for two consecutive frames the estimated affine transformations is compared to the reference transformation of the robot. For the evaluated trajectories the translation error <inline-formula><mml:math id="M8"><mml:mover accent="true"><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:math></inline-formula> is calculated by Equation (6) from the translations given from the algorithm <inline-formula><mml:math id="M9"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">ORB</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:math></inline-formula> and reference from robot <inline-formula><mml:math id="M10"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">robot</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">ORB</mml:mtext></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">robot</mml:mtext></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>For each trajectory (linear and spiral), <inline-formula><mml:math id="M12"><mml:mover accent="true"><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:math></inline-formula> is calculated for all two consecutive frames. This results in an error distribution, which is plotted for each inner ear model. Error distributions are presented for the linear trajectories (<xref ref-type="fig" rid="F6">Figures 6A&#x02013;C</xref>), and spiral trajectories (<xref ref-type="fig" rid="F6">Figures 6D,E</xref>).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Error (<inline-formula><mml:math id="M7"><mml:mover accent="true"><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:math></inline-formula>) distribution for the evaluated scenes. The top row displays 2D-errors for the linear trajectories on TBP <bold>(A)</bold>, TBPs <bold>(B)</bold>, and CTP <bold>(C)</bold>. The bottom row displays 2D-errors for the spiral trajectories on TBP <bold>(D)</bold>, TBPs <bold>(E)</bold>, and CTP <bold>(F)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fsurg-08-742160-g0006.tif"/>
</fig>
<p>The mean absolute error distance &#x003BC; is derived from the <italic>n</italic> sets of consecutive frames by</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003BC;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:mover class="msup"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:mover></mml:mstyle><mml:mo>||</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>||</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><xref ref-type="table" rid="T1">Table 1</xref> summarizes the mean errors of the tracked motion and their standard deviations for each specimen.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary of the tracking precision results.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>TBP</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>TBPs</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>CTP</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Std</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Std</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Std</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Linear</td>
<td valign="top" align="center">93.9</td>
<td valign="top" align="center">118.4</td>
<td valign="top" align="center">135.8</td>
<td valign="top" align="center">114.3</td>
<td valign="top" align="center">110.1</td>
<td valign="top" align="center">112.6</td>
</tr>
<tr>
<td valign="top" align="left">Spiral</td>
<td valign="top" align="center">98.7</td>
<td valign="top" align="center">79.0</td>
<td valign="top" align="center">97.2</td>
<td valign="top" align="center">85.9</td>
<td valign="top" align="center">93.2</td>
<td valign="top" align="center">80.0</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>For each sample and each tested trajectory the mean error for the tracking error of two consecutive frames and their standard deviation are presented (all values in &#x003BC;m)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>Error distributions in <xref ref-type="fig" rid="F6">Figure 6</xref> show that the x and y locations of the error correlate with the number of trajectory sections with a constant orientation. Execution of linear trajectories in x- and y-directions results in error aggregation around <italic>x</italic> &#x0003D; 0 and <italic>y</italic> &#x0003D; 0. The spiral trajectories result in error aggregation along distinct angles.</p>
</sec>
<sec>
<title>3.2. Trajectories</title>
<p>A set of affine transformation is estimated from the image stream. Cascading these transformations and applying them to the initial pose, results in an estimation of the current pose. The translational information of these poses can be displayed as the scenes full trajectory. This trajectory allows for comparison to the reference trajectories as executed by the Stewart platform. <xref ref-type="fig" rid="F7">Figures 7A&#x02013;C</xref> show reference trajectories (blue) and the image based trace of motion (red) for linear trajectories for each inner ear model (i.e., TBP, TBPs, and CTP). <xref ref-type="fig" rid="F7">Figures 7D&#x02013;F</xref> show reference trajectories and the image based trace of motion for the spiral trajectories for each scene.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Linear <bold>(A&#x02013;C)</bold> and spiral <bold>(D&#x02013;F)</bold> trajectories for TBP <bold>(A,D)</bold>, TBPs <bold>(B,E)</bold>, and CTP <bold>(C,F)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fsurg-08-742160-g0007.tif"/>
</fig>
<p>The tracked linear trajectories (i.e., cross-shape) display an offset to the reference path but returns to its original starting point for all three models. For the spiral trajectories the accumulating pose error results in a total error of the final position.</p>
</sec>
<sec>
<title>3.3. Performance</title>
<p>The average duration of each algorithmic step in the process is listed in <xref ref-type="table" rid="T2">Table 2</xref>. These values refer to the runtime per image for images of size 1,920 &#x000D7; 1,080 and 960 &#x000D7; 540 px. The runtime is measured using 30 random images of the surgical site. The total runtime is listed for the implemented algorithm using scikit&#x00027;s RANSAC implementation, which was used for model estimation in this work.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Results of performance evaluation.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Step</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Runtime mean &#x000B1; standard deviation in ms</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>1,920 &#x000D7; 1,080</bold></th>
<th valign="top" align="center"><bold>960 &#x000D7; 540</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Feature detection</td>
<td valign="top" align="center">74 &#x000B1; 9</td>
<td valign="top" align="center">18 &#x000B1; 1</td>
</tr>
<tr>
<td valign="top" align="left">Matching</td>
<td valign="top" align="center">16 &#x000B1; 3</td>
<td valign="top" align="center">12 &#x000B1; 8</td>
</tr>
<tr>
<td valign="top" align="left">Model estimation</td>
<td valign="top" align="center">246 &#x000B1; 77</td>
<td valign="top" align="center">234 &#x000B1; 11</td>
</tr>
<tr>
<td valign="top" align="left">Preprocessing</td>
<td valign="top" align="center">36 &#x000B1; 9</td>
<td valign="top" align="center">17 &#x000B1; 5</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Approximate total time</td>
<td valign="top" align="center">327</td>
<td valign="top" align="center">281</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Runtime for 1,920 &#x000D7; 1,080 (Full HD) and 960 &#x000D7; 540 images are compared</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>The precision result exhibit few deviations between the phantoms (TBP, TBPs) and the human model (CTP). This leads to the conclusion, that the proposed method is well-suited for application in surgery independent of the specific domain.</p>
<p>The smallest tissue manipulation necessary for the intervention of cochlea implantation is the 2,000 &#x003BC;m opening to the middle-ear. The average translational error for all trajectories and scenarios (93.2&#x02013;135.8&#x003BC;m) is more then one magnitudes below. Therefore, the frame to frame tracking proves suitable for supporting the localization of an assistive robotic system.</p>
<p>In the error distribution diagrams in <xref ref-type="fig" rid="F6">Figure 6</xref> a strong correlation between the error and the direction of movement can be observed. For the linear trajectories erroneous motion only occurred along the x- and y-axes. The respective error distributions exhibit errors along the x- and y-axes. This leads to the conclusion, that the presented algorithm can determine the direction of a translation with significantly higher precision than the magnitude of the same translation. For the spiral trajectories the errors are distributed more evenly. The observation of aggregation along distinct angles (i.e., creating the star-like error distribution), can also be explained by the conclusion of higher angular precision. As the spiral trajectory is interpolated by linear sections, translation occurs section-wise linearly and for each section errors aggregate in the respective direction of translation.</p>
<p>The presented algorithm for reconstruction of the trajectory incrementally traces the current relative pose of microscope to patient. Position information only relies on the last increment of the pose as it is derived from the last two consecutive images. Therefore, it suffers from typical loss in precision over time as errors accumulate. For linear trajectories along x- and y-directions this effect is sufficiently small. However, when combining translation in multiple directions in the spiral trajectories, the accumulated error increases over time. The latter displays an accumulating overall position error in all three inner ear models. Presumably, this observation is caused by the aforementioned uncertainty in distance exceeding the angular uncertainty.</p>
<p>Despite the relatively high pose error after conduction of the spiral trajectories, the proposed method is suitable for extension by initial registration of the scene, which may be marker based or manually conducted. The trajectories evaluated in this work have a longer duration (44 s for the linear, 30 s for the spiral) compared to a shock caused by unintended motion in the OR. Thus, more iterations evaluated, which increase the accumulative error, in contrast to a real surgical scenario. In a robotic intervention such as motivated in section 1, global registration is likely to be considered a necessary prerequisite anyway. As an example, we envision a manual registration of an ablation laser spot before the ablation process. This could for example be conducted by manual input through a joystick. Extending this with the presented method represents a reliable safety measure against short unintended motion of patient or microscope.</p>
<p>The current algorithm and used hardware allows for processing of the microscope images in real-time with approximately 3 Hz. The major limitation is given by the RANSAC algorithm. Here scikit&#x00027;s implementation was used as it offers greater flexibility in implementation however in preliminary studies also an openCV implementation&#x00027;s runtime was evaluated and resulted in a significant reduction in the exection of RANSAC from an average of 246 ms down to 4 ms. This demonstrates the high potential software as well as hardware optimization offers for increasing the frame rates.</p>
<p>The presented method is limited to tracking an initially conducted registration and compensate for small errors occurring over short periods of time. The initial registration is outside the scope of this work as several methods have previously been presented. Initial registration methodologies strongly depend on the intervention and the applied robotic system. The presented method is prone to long term drifts of the pose due to accumulation of errors. As the scene can be expected to display only small and fast changes in pose. A suggested improvement may be to compare the current frame not only to the most recent one but also to past image data like a user defined initial frame or images captured multiple iterations earlier. The estimated transform from these frames can be used to correct a global drift of the tracked pose.</p>
<p>For appropriate integration to a robotic system, a frame rate suitable to the robotics kinematics needs to be reached by optimizing hardware, image resolution and implementation. For high speed (short term) tracking of relative pose changes, the here presented method could be extended by the use of inertial measurement units. However, these would require integration into the robotic system as well as attachment to the patient.</p>
<p>The presented method has been evaluated for feature-based tracking of inner ear models in two dimensions only. Here, we assume planar motion of surgical situs in the microscope image. To extend this method to covering full 6D pose estimation, i.e., three translations and three rotations, the estimated model needs to be expanded to a 3D-Transformation, as in</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>00</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>01</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>02</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>20</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>For application in clinical intervention the surgical scene might become less rigid for example due to moving instruments (robotic or manual). The same issue is likely to occur for manipulations of the surgical field, obstructions by blood or residual tissue from drilling. If these artifacts only cover small areas of the field of view they are likely to be filtered by the RANSAC algorithm. Future work could investigate the robustness of the presented algorithm against such artifacts. Further approaches could research the masking of instruments and residual tissue in the image before feature detection to avoid falsely using features on the tools instead of the situs for tracking. This challenge could be solved by semantic segmentation of the instruments prior to executing the tracking algorithm, as demonstrated in Bodenstedt et al. (<xref ref-type="bibr" rid="B18">18</xref>) for laparoscopic scenes. With sufficient training data, typical instruments are masked from the scene and only the situs&#x00027; image information are utilized for tracking.</p>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>A method for feature-based tracking of the inner ear for compensation of unintended motion was proposed. It is motivated by its use as safety feature enabling microscope mounted medical robotic assistance. Aiming for application in various fields of microsurgery, the application in cochlea implantation was regarded exemplary. Images from a surgical microscope are processed to derive pose changes between patient and microscope. These information can serve as input for compensating motion of a microscope mounted robotic system. Two consecutive images are analyzed for ORB features, which are matched and an affine transformation is estimated by a RANSAC algorithm. The transform is published in the Robot Operating System for integration into robotic systems. Making use of existing hardware in the OR during microsurgery, the microscope image stream is available for processing without introduction of additional hardware. This potentially allows for simple clinical translation of the proposed method. Evaluation showed sub-millimeter accuracy for frame to frame pose changes but revealed increasing offset in absolute pose due to accumulating errors. Application as shock countermeasure seems promising, however, clinical translation will require extension to 3D tracking and optimized performance.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec id="s7">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by Ethikkommission an der Med. Fakult&#x000E4;t der HHU D&#x000FC;sseldorf Moorenstr. 5 D-40225 D&#x000FC;sseldorf FWA-Nr.: 00000829 HHS IRB Registration Nr.: IRB00001579. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>CM and JH conceived and implemented the algorithm. CM, TP, and JH conceived, designed, and executed the experimental study. CM, TP, TK, and FM-U analyzed and involved in interpretation of data and made final approval of the version to be published. CM, TP, and FM-U drafted the article. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>

<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vittoria</surname> <given-names>S</given-names></name> <name><surname>Lahlou</surname> <given-names>G</given-names></name> <name><surname>Torres</surname> <given-names>R</given-names></name> <name><surname>Daoudi</surname> <given-names>H</given-names></name> <name><surname>Mosnier</surname> <given-names>I</given-names></name> <name><surname>Mazalaigue</surname> <given-names>S</given-names></name> <etal/></person-group>. <article-title>Robot-based assistance in middle ear surgery and cochlear implantation: first clinical report</article-title>. <source>Eur Arch Otorhinolaryngol</source>. (<year>2021</year>) <volume>278</volume>:<fpage>77</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1007/s00405-020-06070-z</pub-id><pub-id pub-id-type="pmid">32458123</pub-id></citation></ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ureel</surname> <given-names>M</given-names></name> <name><surname>Augello</surname> <given-names>M</given-names></name> <name><surname>Holzinger</surname> <given-names>D</given-names></name> <name><surname>Wilken</surname> <given-names>T</given-names></name> <name><surname>Berg</surname> <given-names>BI</given-names></name> <name><surname>Zeilhofer</surname> <given-names>HF</given-names></name> <etal/></person-group>. <article-title>Cold ablation robot-guided laser osteotome (Carloa&#x000AE;): from bench to bedside</article-title>. <source>J Clin Med</source>. (<year>2021</year>) <volume>10</volume>:<fpage>450</fpage>. <pub-id pub-id-type="doi">10.3390/jcm10030450</pub-id><pub-id pub-id-type="pmid">33498921</pub-id></citation></ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rauter</surname> <given-names>G</given-names></name></person-group>. <article-title>The miracle.</article-title> In: <person-group person-group-type="editor"><name><surname>Stobinger</surname> <given-names>S</given-names></name> <name><surname>Klompfl</surname> <given-names>F</given-names></name> <name><surname>Schmidt</surname> <given-names>M</given-names></name> <name><surname>Zeilhofer</surname> <given-names>HF</given-names></name></person-group> editors. <source>Lasers in Oral and Maxillofacial Surgery</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature</publisher-name> (<year>2020</year>). p. <fpage>247</fpage>&#x02013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-29604-9_19</pub-id></citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y</given-names></name> <name><surname>Pfeiffer</surname> <given-names>T</given-names></name> <name><surname>Weller</surname> <given-names>M</given-names></name> <name><surname>Wieser</surname> <given-names>W</given-names></name> <name><surname>Huber</surname> <given-names>R</given-names></name> <name><surname>Raczkowsky</surname> <given-names>J</given-names></name> <etal/></person-group>. <article-title>Optical coherence tomography guided laser cochleostomy: towards the accuracy on tens of micrometer scale</article-title>. <source>Biomed Res Int</source>. (<year>2014</year>) <volume>2014</volume>:<fpage>251814</fpage>. <pub-id pub-id-type="doi">10.1155/2014/251814</pub-id><pub-id pub-id-type="pmid">25295253</pub-id></citation></ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y</given-names></name></person-group>. <source>Optical Coherence Tomography Guided Laser-Cochleostomy</source>. <publisher-loc>Chichester</publisher-loc>: <publisher-name>Scientific Publishing</publisher-name> (<year>2015</year>).<pub-id pub-id-type="pmid">25295253</pub-id></citation></ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kahrs</surname> <given-names>LA</given-names></name> <name><surname>Burgner</surname> <given-names>J</given-names></name> <name><surname>Klenzner</surname> <given-names>T</given-names></name> <name><surname>Raczkowsky</surname> <given-names>J</given-names></name> <name><surname>Schipper</surname> <given-names>J</given-names></name> <name><surname>W&#x000F6;rn</surname> <given-names>H</given-names></name></person-group>. <article-title>Planning and simulation of microsurgical laser bone ablation</article-title>. <source>Int J Comput Assist Radiol Surg</source>. (<year>2010</year>) <volume>5</volume>:<fpage>155</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1007/s11548-009-0303-4</pub-id><pub-id pub-id-type="pmid">20033520</pub-id></citation></ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>L</given-names></name> <name><surname>Fei</surname> <given-names>B</given-names></name></person-group>. <article-title>Comprehensive review of surgical microscopes: technology development and medical applications</article-title>. <source>J BiomedOptics</source>. (<year>2021</year>) <volume>26</volume>:<fpage>1</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1117/1.JBO.26.1.010901</pub-id><pub-id pub-id-type="pmid">33398948</pub-id></citation></ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hussain</surname> <given-names>R</given-names></name> <name><surname>Lalande</surname> <given-names>A</given-names></name> <name><surname>Berihu Girum</surname> <given-names>K</given-names></name> <name><surname>Guigou</surname> <given-names>C</given-names></name> <name><surname>Grayeli</surname> <given-names>AB</given-names></name></person-group>. <article-title>Augmented reality for inner ear procedures: visualization of the cochlear central axis in microscopic videos</article-title>. <source>Int J Comput Assist Radiol Surg</source>. (<year>2020</year>) <volume>15</volume>:<fpage>1703</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1007/s11548-020-02240-w</pub-id><pub-id pub-id-type="pmid">32737858</pub-id></citation></ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Giraldez</surname> <given-names>JG</given-names></name> <name><surname>Talib</surname> <given-names>H</given-names></name> <name><surname>Caversaccio</surname> <given-names>M</given-names></name> <name><surname>Ballester</surname> <given-names>MAG</given-names></name></person-group>. <article-title>Multimodal augmented reality system for surgical microscopy.</article-title> In: <person-group person-group-type="editor"><name><surname>Cleary</surname> <given-names>KR</given-names></name> <name><surname>Robert</surname> <given-names>L</given-names></name> <name><surname>Galloway</surname> <given-names>J</given-names></name></person-group> editors. <source>Medical Imaging 2006: Visualization, Image-Guided Procedures, and Display</source>. <publisher-loc>San Diego, CA</publisher-loc>: <publisher-name>International Society for Optics and Photonics</publisher-name>. SPIE (<year>2006</year>). p. <fpage>537</fpage>&#x02013;<lpage>44</lpage>.<pub-id pub-id-type="pmid">16279086</pub-id></citation></ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Groeger</surname> <given-names>M</given-names></name> <name><surname>Sepp</surname> <given-names>W</given-names></name> <name><surname>Ortmaier</surname> <given-names>T</given-names></name> <name><surname>Hirzinger</surname> <given-names>G</given-names></name></person-group>. <article-title>Reconstruction of image structure in presence of specular reflections.</article-title> In: <person-group person-group-type="editor"><name><surname>Radig</surname> <given-names>B</given-names></name> <name><surname>Florczyk</surname> <given-names>S</given-names></name></person-group> editors. <source>Pattern Recognition</source>. <publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2001</year>). p. <fpage>53</fpage>&#x02013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1007/3-540-45404-7_8</pub-id></citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rublee</surname> <given-names>E</given-names></name> <name><surname>Rabaud</surname> <given-names>V</given-names></name> <name><surname>Konolige</surname> <given-names>K</given-names></name> <name><surname>Bradski</surname> <given-names>G</given-names></name></person-group>. <article-title>ORB: an efficient alternative to SIFT or SURF</article-title>. In: <source>2011 International Conference on Computer Vision</source>. <publisher-loc>Barcelona</publisher-loc> (<year>2011</year>). p. <fpage>2564</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2011.6126544</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lowe</surname> <given-names>DG</given-names></name></person-group>. <article-title>Distinctive image features from scale-invariant keypoints</article-title>. <source>Int J Comput Vision</source>. (<year>2004</year>) <volume>60</volume>:<fpage>91</fpage>&#x02013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1023/B:VISI.0000029664.99615.94</pub-id></citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Karami</surname> <given-names>E</given-names></name> <name><surname>Prasad</surname> <given-names>S</given-names></name> <name><surname>Shehata</surname> <given-names>MS</given-names></name></person-group>. <article-title>Image matching using SIFT, SURF, BRIEF and ORB: performance comparison for distorted images</article-title>. <source>arXiv [Preprint]. arXiv:1710.02726.</source> (<year>2017</year>).</citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mur-Artal</surname> <given-names>R</given-names></name> <name><surname>Montiel</surname> <given-names>JMM</given-names></name> <name><surname>Tardas</surname> <given-names>JD</given-names></name></person-group>. <article-title>ORB-SLAM: a versatile and accurate monocular SLAM system</article-title>. <source>IEEE Trans Robot</source>. (<year>2015</year>) <volume>31</volume>:<fpage>1147</fpage>&#x02013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1109/TRO.2015.2463671</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bradski</surname> <given-names>G</given-names></name></person-group>. <article-title>The OpenCV library</article-title>. <source>Dobbs J Softw Tools</source>. (<year>2000</year>) <volume>120</volume>:<fpage>122</fpage>&#x02013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fischler</surname> <given-names>MA</given-names></name> <name><surname>Bolles</surname> <given-names>RC</given-names></name></person-group>. <article-title>Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography</article-title>. <source>Commun ACM</source>. (<year>1981</year>) <volume>24</volume>:<fpage>381</fpage>&#x02013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1145/358669.358692</pub-id></citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van der Walt</surname> <given-names>S</given-names></name> <name><surname>Sch&#x000F6;nberger</surname> <given-names>JL</given-names></name> <name><surname>Nunez-Iglesias</surname> <given-names>J</given-names></name> <name><surname>Boulogne</surname> <given-names>F</given-names></name> <name><surname>Warner</surname> <given-names>JD</given-names></name> <name><surname>Yager</surname> <given-names>N</given-names></name> <etal/></person-group>. <article-title>scikit-image: image processing in Python</article-title>. <source>PeerJ</source>. (<year>2014</year>) <volume>2</volume>:<fpage>e453</fpage>. <pub-id pub-id-type="doi">10.7717/peerj.453</pub-id><pub-id pub-id-type="pmid">25024921</pub-id></citation></ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bodenstedt</surname> <given-names>S</given-names></name> <name><surname>Allan</surname> <given-names>M</given-names></name> <name><surname>Agustinos</surname> <given-names>A</given-names></name> <name><surname>Du</surname> <given-names>X</given-names></name> <name><surname>Garcia-Peraza-Herrera</surname> <given-names>L</given-names></name> <name><surname>Kenngott</surname> <given-names>H</given-names></name> <etal/></person-group>. <article-title>Comparative evaluation of instrument segmentation and tracking methods in minimally invasive surgery</article-title>. <source>arXiv [Preprint]. arXiv:1805.02475.</source> (<year>2018</year>).</citation>
</ref>
</ref-list> 
</back>
</article>