<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Digit. Humanit.</journal-id>
<journal-title>Frontiers in Digital Humanities</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Digit. Humanit.</abbrev-journal-title>
<issn pub-type="epub">2297-2668</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdigh.2018.00001</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Digital Humanities</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Bridging the Gap: Enriching YouTube Videos with Jazz Music Annotations</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Balke</surname> <given-names>Stefan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<uri xlink:href="http://frontiersin.org/people/u/440985"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Dittmar</surname> <given-names>Christian</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Abe&#x000DF;er</surname> <given-names>Jakob</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/520664"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Frieler</surname> <given-names>Klaus</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/485345"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Pfleiderer</surname> <given-names>Martin</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/240444"/>
</contrib>
<contrib contrib-type="author">
<name><surname>M&#x000FC;ller</surname> <given-names>Meinard</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/520984"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>International Audio Laboratories Erlangen</institution>, <addr-line>Erlangen</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Semantic Music Technologies Group, Fraunhofer IDMT</institution>, <addr-line>Ilmenau</addr-line>, <country>Germany</country></aff>
<aff id="aff3"><sup>3</sup><institution>Jazzomat Research Project, University of Music Franz Liszt</institution>, <addr-line>Weimar</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Mark Brian Sandler, Queen Mary University of London, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Anna Wolf, Hanover University of Music Drama and Media, Germany; Michael Scott Cuthbert, Massachusetts Institute of Technology, United States</p></fn>
<corresp content-type="corresp" id="cor1">&#x0002A;Correspondence: Stefan Balke, <email>stefan.balke&#x00040;audiolabserlangen.de</email></corresp>
<fn fn-type="other" id="fn001"><p>Specialty section: This article was submitted to Digital Musicology, a section of the journal Frontiers in Digital Humanities</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>02</month>
<year>2018</year>
</pub-date>
<pub-date pub-type="collection">
<year>2018</year>
</pub-date>
<volume>5</volume>
<elocation-id>1</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>10</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>16</day>
<month>01</month>
<year>2018</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2018 Balke, Dittmar, Abe&#x000DF;er, Frieler, Pfleiderer and M&#x000FC;ller.</copyright-statement>
<copyright-year>2018</copyright-year>
<copyright-holder>Balke, Dittmar, Abe&#x000DF;er, Frieler, Pfleiderer and M&#x000FC;ller</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Web services allow permanent access to music from all over the world. Especially in the case of web services with user-supplied content, e.g., YouTube&#x02122;, the available metadata is often incomplete or erroneous. On the other hand, a vast amount of high-quality and musically relevant metadata has been annotated in research areas such as Music Information Retrieval (MIR). Although they have great potential, these musical annotations are often inaccessible to users outside the academic world. With our contribution, we want to bridge this gap by enriching publicly available multimedia content with musical annotations available in research corpora, while maintaining easy access to the underlying data. Our web-based tools offer researchers and music lovers novel possibilities to interact with and navigate through the content. In this paper, we consider a research corpus called the Weimar Jazz Database (WJD) as an illustrating example scenario. The WJD contains various annotations related to famous jazz solos. First, we establish a link between the WJD annotations and corresponding YouTube videos employing existing retrieval techniques. With these techniques, we were able to identify 988 corresponding YouTube videos for 329 solos out of 456 solos contained in the WJD. We then embed the retrieved videos in a recently developed web-based platform and enrich the videos with solo transcriptions that are part of the WJD. Furthermore, we integrate publicly available data resources from the Semantic Web in order to extend the presented information, for example, with a detailed discography or artists-related information. Our contribution illustrates the potential of modern web-based technologies for the digital humanities, and novel ways for improving access and interaction with digitized multimedia content.</p>
</abstract>
<kwd-group>
<kwd>music information retrieval</kwd>
<kwd>digital humanities</kwd>
<kwd>audio processing</kwd>
<kwd>semantic web</kwd>
<kwd>multimedia</kwd>
</kwd-group>
<contract-num rid="cn01">MU 2686/6-1, MU 2686/12-1, MU 2686/10-1, MU 2686/11-1, PF 669/7-1</contract-num>
<contract-sponsor id="cn01">Deutsche Forschungsgemeinschaft<named-content content-type="fundref-id">10.13039/501100001659</named-content></contract-sponsor>
<counts>
<fig-count count="4"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="48"/>
<page-count count="11"/>
<word-count count="7884"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<label>1</label> <title>Introduction</title>
<p>Online video platforms, such as YouTube, make billions of videos available to users from all over the world. Many of these videos contain recordings of music performances. Often, these performances are tagged with basic metadata&#x02014;mainly the artist and the title of the song. However, since this metadata is not curated, it might be incomplete or incorrect. The lack of reliable metadata makes it hard to identify particular recordings, especially for music genres where many renditions of the same musical work exist (e.g., symphonies in Western classical music, ragas in Indian music, or standards in jazz music). Imagine a jazz student who is practicing a jazz solo played by a famous musician and is now interested in the original recording. In the case that the student searches for a musician whose name is not mentioned in the metadata (e.g., because the musician was &#x0201C;only&#x0201D; a sideman in the band), a textual search may not be successful or may result in too many irrelevant results. Assuming that the student has already a partial or even a complete transcription of the solo available, content-based retrieval techniques could help to resolve this problem. Here, <italic>content-based</italic> means that, in the comparison of music data, the system makes use of the raw music data itself (e.g., from the music recording or the YouTube video), rather than relying on manually generated keywords referring to the artists&#x02019; names, the song&#x02019;s title or lyrics (M&#x000FC;ller, <xref ref-type="bibr" rid="B28">2007</xref>).</p>
<p>Jazz musicians, musicologists, and publishers have made many jazz solo transcriptions publicly available during the last decades, e.g., Hal Leonard&#x02019;s <italic>Omnibook</italic> series.<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> One comprehensive corpus of solo transcriptions is the Weimar Jazz Database (WJD), which consists of 456 (as of May 2017) transcriptions of instrumental solos in jazz recordings performed by a wide range of renowned musicians (Pfleiderer et al., <xref ref-type="bibr" rid="B34">2017</xref>). The solos have been manually transcribed by musicology and jazz students. In addition, the database offers various music-related annotations such as chord sequences or beat positions. We believe that these annotations are a great resource that could help musicians and other researchers in gaining a deeper understanding of jazz music. However, these annotations and the underlying audio material are not directly accessible, mainly for two reasons. First, the audio files originate from commercial music recordings which are protected by copyright and ancillary copyright laws. Therefore, they cannot be made publicly available by scientific institutions. This restricts the usefulness of the dataset for scientific research, where both the annotations and the corresponding audio material are required. Second, the annotations are encoded in a database format which is not easily accessible for users without technical skills. Both problems apply to many scientific datasets which offer musical annotations for commercial music recordings. Simply switching to music recordings that are released under public domain licenses is not an option for research questions which rely on specific music recordings. In our approach, we try to bypass some of these copyright restrictions by using music recordings that are publicly available via YouTube. However, there is no doubt that both musicians and composers should be gratified financially for the music they create according to national and international copyright and ancillary copyright laws. YouTube seems to guarantee this financial entitlement through agreements with national copyright collecting societies. By contrast, for scientific institutions offering music databases, it is very difficult or impossible to handle these legal claims. As a case study, we focus on the recordings that have corresponding annotations in the WJD.</p>
<p>As the main contribution of this paper, we introduce various retrieval methods based on metadata and content-based descriptors and show how these techniques can be applied for identifying and enriching YouTube videos. In the following, we sketch a typical two-stage retrieval scenario which is then described in more detail in the subsequent sections (Figure <xref ref-type="fig" rid="F1">1</xref> provides an overview). In this example, we are interested in the song <italic>Jordu</italic>, recorded by Clifford Brown in 1954 (Figure <xref ref-type="fig" rid="F1">1</xref>A). In the first step, we use the title and the name of the soloist as provided by the WJD to perform a metadata-based search on YouTube (Figure <xref ref-type="fig" rid="F1">1</xref>B). This search results in a list of candidates. Besides relevant music recordings, this list may also contain other recordings by the same artist or cover versions by other artists. Using the recording associated to the WJD&#x02019;s annotations, we apply an audio-based retrieval approach to identify the relevant music recordings in this list of candidates (Figure <xref ref-type="fig" rid="F1">1</xref>C). The result of this matching procedure is a list of relevant documents that can be used to link the WJD&#x02019;s annotations to the YouTube videos. The retrieved video is then embedded in a web-based application (Figure <xref ref-type="fig" rid="F1">1</xref>D). In addition, we use the annotations provided by the WJD to further enrich the video, e.g., by offering new navigation possibilities based on the song structure or transcriptions of the song&#x02019;s solo. As a result, the user is able to follow the soloist&#x02019;s improvisation in a piano-roll-like representation. For intuition and hands-on experience with this concept, our web-based application can be accessed under the following address: <uri xlink:href="http://mir.audiolabs.uni-erlangen.de/jazztube">http://mir.audiolabs.uni-erlangen.de/jazztube</uri>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Illustration of the two-stage retrieval scenario applied to retrieve videos from YouTube. <bold>(A)</bold> Overview of the various annotations for Clifford Brown&#x02019;s solo <italic>Jordu</italic> contained in the Weimar Jazz Database. <bold>(B)</bold> First retrieval stage: text-based retrieval on YouTube resulting in a list of candidates. <bold>(C)</bold> Second retrieval stage: content-based retrieval using the solo recording from the WJD as query. <bold>(D)</bold> Identified video embedded in a web-based demonstrator and enriched with the annotations obtained from the WJD. Panel <bold>(D)</bold> has been created by the authors and, therefore, no permission is required for its use in this manuscript.</p></caption>
<graphic xlink:href="fdigh-05-00001-g001.tif"/>
</fig>
<p>The remainder of this paper is structured as follows. We start by giving a brief overview of the literature and related projects (Section <xref ref-type="sec" rid="S2">2</xref>). Then, we introduce the different data resources used in this study (Section <xref ref-type="sec" rid="S3">3</xref>). Subsequently, we describe the various retrieval procedures that are used to link the WJD to the YouTube videos (Section <xref ref-type="sec" rid="S4">4</xref>). Finally, we present a web-based service that integrates the introduced data resources in a unifying user interface (Section <xref ref-type="sec" rid="S5">5</xref>).</p>
</sec>
<sec id="S2">
<label>2</label> <title>Related Work</title>
<p>Similar web-based services that aim to enhance the listening experience have been proposed in the past. <italic>Songle</italic>,<xref ref-type="fn" rid="fn2"><sup>2</sup></xref> for instance, lets users explore music from different perspectives (Goto et al., <xref ref-type="bibr" rid="B19">2011</xref>). In this web-based service, computational approaches are used to annotate music recordings (including beats, melodic lines, or chords). Afterward, these generated annotations are presented in a web-based interface. Since the automatically generated annotations may contain errors, the users can correct them or add new ones. The annotations contained in <italic>Songle</italic> can then be used in third-party applications or research projects (e.g., for singing-voice analysis). Another service called <italic>Songrium</italic>,<xref ref-type="fn" rid="fn3"><sup>3</sup></xref> allows users to add lyrics to publicly available videos (e.g., obtained from YouTube). In addition, the lyrics can be visualized and played back along with the linked video similar to karaoke applications. For an overview of other systems by Goto and colleagues, we refer to the literature, see Goto (<xref ref-type="bibr" rid="B17">2011</xref>, <xref ref-type="bibr" rid="B18">2014</xref>).</p>
<p>Another project that aims at enhancing the listening experience, especially for classical music, is called <italic>PHENICX</italic> (Performances as Highly Enriched aNd Interactive Concert eXperiences) (Gasser et al., <xref ref-type="bibr" rid="B15">2015</xref>; Liem et al., <xref ref-type="bibr" rid="B24">2015</xref>, <xref ref-type="bibr" rid="B25">2017</xref>; Melenhorst et al., <xref ref-type="bibr" rid="B27">2015</xref>). As one main functionality, suitable visualizations are generated in real-time and displayed during the live performance of an orchestra. Such visualizations may be a rendition of a musical score (score-following applications) or an animation controlled by the baton movements of the orchestra&#x02019;s conductor. Furthermore, as in our scenario, the project offers a web-based service, which allows the playback of enriched videos.<xref ref-type="fn" rid="fn4"><sup>4</sup></xref> In the research project <italic>Freisch&#x000FC;tz Digital</italic>,<xref ref-type="fn" rid="fn5"><sup>5</sup></xref> user interfaces for dealing with critical editions in an opera scenario were developed (Pr&#x000E4;tzlich et al., <xref ref-type="bibr" rid="B35">2015</xref>; R&#x000F6;wenstrunk et al., <xref ref-type="bibr" rid="B40">2015</xref>). In this scenario, an essential step is to link the different sheet music editions with the various existing music recordings. These alignments are then used in special user interfaces that may support musicologists in their work on critical editions.</p>
<p>Besides publicly available music recordings or videos, the internet offers additional information (metadata or textual annotations) for music recordings. Many services offer metadata in a structured way, often following standardized data formats as defined in the <italic>Semantic Web</italic> (Berners-Lee et al., <xref ref-type="bibr" rid="B6">2001</xref>). The Semantic Web contains standardized schemas, called <italic>ontologies</italic>, for exchanging different kinds of data. A way to exchange musical annotations is defined in the <italic>Music Ontology</italic> (Raimond et al., <xref ref-type="bibr" rid="B37">2007</xref>). One of the most frequently used services in the Semantic Web is <italic>DBpedia</italic><xref ref-type="fn" rid="fn6"><sup>6</sup></xref> which offers information from Wikipedia in a structured data format. Popular services for music metadata in general are <italic>MusicBrainz</italic><xref ref-type="fn" rid="fn7"><sup>7</sup></xref> or <italic>Discogs</italic>.<xref ref-type="fn" rid="fn8"><sup>8</sup></xref> In particular for jazz music, the <italic>JDISC</italic><xref ref-type="fn" rid="fn9"><sup>9</sup></xref> project aims to provide complete discographies for a number of selected artists. Another related project is called <italic>Linked Jazz</italic>,<xref ref-type="fn" rid="fn10"><sup>10</sup></xref> which offers relationships between jazz musicians in a structured way (Pattuelli, <xref ref-type="bibr" rid="B32">2012</xref>). Besides sharing metadata, researchers have used YouTube as a way of specifying datasets that were used in their experiments (Schoeffler and Herre, <xref ref-type="bibr" rid="B44">2014</xref>). In particular, for audio applications, Google released <italic>AudioSet</italic>, a dataset consisting of over two million 10-s sound clips obtained from YouTube which have then been labeled by human annotators (Gemmeke et al., <xref ref-type="bibr" rid="B16">2017</xref>).</p>
<p>This work follows similar concepts as used in the <italic>SyncPlayer</italic> (Kurth et al., <xref ref-type="bibr" rid="B23">2005</xref>; Thomas et al., <xref ref-type="bibr" rid="B47">2009</xref>; Damm et al., <xref ref-type="bibr" rid="B10">2012</xref>). The SyncPlayer offers various ways of interacting and navigating with a large, multimodal corpus of music recordings, sheet music, and lyrics. Furthermore, users are able to search within this corpus by specifying a short melodic phrase or an excerpt from the lyrics. The results are then presented in an interactive graphical user interface that allows auditioning the results. In previous works, we studied the use of interfaces for two different music scenarios. In Balke et al. (<xref ref-type="bibr" rid="B3">2017a</xref>), a web-based user interface motivated by applications in jazz-piano education is presented. In particular, a video recording, a piano-roll representation, and additional annotations are incorporated in a unifying interface that allows the user to simultaneously play back the different media objects. A related approach focusses on the opera <italic>Die Walk&#x000FC;re (The Valkyrie)</italic> from Richard Wagner&#x02019;s cycle <italic>Der Ring des Nibelungen (The Ring of the Nibelung)</italic>. The goal of the interface is to supply intuitive functions that allow a user to easily access and explore all available data (including different recordings, videos, lyrics, sheet music) associated with a large-scale work such as an opera.</p>
</sec>
<sec id="S3">
<label>3</label> <title>Data Resources</title>
<p>In this paper, we consider jazz-related data of different modality stemming from different resources. We now introduce the Weimar Jazz Database (WJD), the relevant jazz recordings, the streaming platform YouTube from which we obtain videos, and the used web resources for additional metadata.</p>
<sec id="S3-1">
<label>3.1</label> <title>Weimar Jazz Database (WJD)</title>
<p>The WJD is part of the Jazzomat Research Project,<xref ref-type="fn" rid="fn11"><sup>11</sup></xref> which aims at a better understanding of creative processes in improvisations using computational methods (Pfleiderer et al., <xref ref-type="bibr" rid="B34">2017</xref>). The WJD comprises 456 (as of July 2017) high-quality solo transcriptions (similar to a piano-roll representation), extracted from 343 tracks taken from 197 different records. The solos are performed by a wide range of renowned jazz musicians in the period from 1925 to 2009 (e.g., Louis Armstrong, Don Byas, or Chris Potter). All solos were manually annotated by musicology and jazz students at the University of Music Franz Liszt Weimar using the <italic>SonicVisualiser</italic> (Cannam et al., <xref ref-type="bibr" rid="B8">2006</xref>). The annotators had different musical backgrounds but a general familiarity with jazz music, mostly through listening and playing. The produced transcriptions were then inspected with an automated verification procedure that primarily searched for syntactical errors and suspicious annotations, such as beat outliers. In a final step, the transcription was cross-checked by an experienced supervisor and added to the database. Table <xref ref-type="table" rid="T1">1</xref> lists the number of solo transcriptions grouped by the 13 different occurring solo instruments. As one might expect for jazz music, the database is biased toward tenor saxophone and trumpet solos, which represent about 56% of the currently available solo transcriptions.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Solo instruments occurring in the WJD.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Abbr.</th>
<th valign="top" align="left">Instrument</th>
<th valign="top" align="center">&#x00023;Solos</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">cl</td>
<td align="left" valign="top">Clarinet</td>
<td align="center" valign="top">15</td>
</tr>
<tr>
<td align="left" valign="top">bcl</td>
<td align="left" valign="top">Bass clarinet</td>
<td align="center" valign="top">2</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">ss</td>
<td align="left" valign="top">Soprano saxophone</td>
<td align="center" valign="top">23</td>
</tr>
<tr>
<td align="left" valign="top">as</td>
<td align="left" valign="top">Alto saxophone</td>
<td align="center" valign="top">80</td>
</tr>
<tr>
<td align="left" valign="top">ts</td>
<td align="left" valign="top">Tenor saxophone</td>
<td align="center" valign="top">157</td>
</tr>
<tr>
<td align="left" valign="top">ts-c</td>
<td align="left" valign="top">Tenor saxophone in C</td>
<td align="center" valign="top">1</td>
</tr>
<tr>
<td align="left" valign="top">bs</td>
<td align="left" valign="top">Baritone saxophone</td>
<td align="center" valign="top">11</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">tp</td>
<td align="left" valign="top">Trumpet</td>
<td align="center" valign="top">102</td>
</tr>
<tr>
<td align="left" valign="top">cor</td>
<td align="left" valign="top">Cornet</td>
<td align="center" valign="top">15</td>
</tr>
<tr>
<td align="left" valign="top">tb</td>
<td align="left" valign="top">Trombone</td>
<td align="center" valign="top">26</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">g</td>
<td align="left" valign="top">Guitar</td>
<td align="center" valign="top">6</td>
</tr>
<tr>
<td align="left" valign="top">p</td>
<td align="left" valign="top">Piano</td>
<td align="center" valign="top">6</td>
</tr>
<tr>
<td align="left" valign="top">vib</td>
<td align="left" valign="top">Vibraphone</td>
<td align="center" valign="top">12</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">13</td>
<td align="left" valign="top"/>
<td align="center" valign="top">&#x02211; 456</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>The first column introduces an abbreviation, whereas the last column indicates the number of solos of the respective instrument</italic>.</p></table-wrap-foot></table-wrap>
<p>Figure <xref ref-type="fig" rid="F2">2</xref>A shows the distribution of the solos with respect to their durations and recording years. The solos have a minimum duration of 19&#x02009;s (Steve Coleman&#x02019;s second solo on <italic>Cross-Fade</italic>), a maximum duration of 818&#x02009;s (John Coltrane&#x02019;s solo on <italic>Impressions</italic>), and an average duration of 107&#x02009;s. Similarly, Figure <xref ref-type="fig" rid="F2">2</xref>B indicates the distribution of the whole tracks (which usually contain more than a single solo part), with a minimum duration of 128&#x02009;s, a maximum duration of 1620&#x02009;s, and an average duration of 354&#x02009;s. From all 343 tracks, there are 247 tracks with one annotated solo part, 80 tracks with two, 15 with three, and a single track with four annotated solo parts. Summing over the number of annotated note events in all solo transcriptions results in over 200,000 elements.<xref ref-type="fn" rid="fn12"><sup>12</sup></xref></p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>(A)</bold> 456 solos and <bold>(B)</bold> 343 tracks considered in the WJD, represented according to their duration and respective recording years. The dashed lines indicate the limits for the maximum duration of the considered YouTube videos.</p></caption>
<graphic xlink:href="fdigh-05-00001-g002.tif"/>
</fig>
<p>Figure <xref ref-type="fig" rid="F3">3</xref> displays the beginning of <italic>Clifford Brown&#x02019;s</italic> solo on <italic>Jordu</italic> as an example for the data contained in the WJD. Figure <xref ref-type="fig" rid="F3">3</xref>A shows a time-frequency representation (see Section <xref ref-type="sec" rid="S3-2">3.2</xref> for details) of this excerpt superimposed by the available solo transcriptions (each note is represented by a red rectangle) and measure positions (represented as blue vertical lines). Figure <xref ref-type="fig" rid="F3">3</xref>B shows a sheet music representation derived from the solo annotations. Note that deriving sheet music from the transcriptions requires algorithms that are able to quantize the onsets and durations of the annotated note events into musically meaningful notes, see Frieler and Pfleiderer (<xref ref-type="bibr" rid="B14">2017</xref>).<xref ref-type="fn" rid="fn13"><sup>13</sup></xref></p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Beginning of <italic>Clifford Brown&#x02019;s</italic> solo on <italic>Jordu</italic>. <bold>(A)</bold> Log-frequency spectrogram, where rectangles (red) indicate the solo transcription and vertical lines (blue) the annotated measure positions. <bold>(B)</bold> Sheet music representation of the transcribed solo.</p></caption>
<graphic xlink:href="fdigh-05-00001-g003.tif"/>
</fig>
</sec>
<sec id="S3-2">
<label>3.2</label> <title>Jazz Recordings</title>
<p>A typical jazz recording consists of a soloist who is accompanied by a rhythm section (e.g., double bass, piano, and drums). From an engineering perspective, such a recording is a sequence of amplitude values sampled from a microphone signal (or a mixture of multiple signals). By applying digital signal processing methods, one can analyze and manipulate such signals. A common way to analyze music signals is to transform them into a time-frequency representation, e.g., a spectrogram. For example, in Figure <xref ref-type="fig" rid="F2">2</xref>A, we show an excerpt of a spectrogram from <italic>Clifford Brown&#x02019;s</italic> solo on <italic>Jordu</italic>. There exist different approaches to obtain such a time-frequency representation.<xref ref-type="fn" rid="fn14"><sup>14</sup></xref> In particular, we use a logarithmically spaced frequency axis with a bandwidth of a single semitone per frequency band (row in the spectrogram)&#x02014;motivated by human&#x02019;s logarithmic perception of frequency and the equal-tempered scale underlying the music. In this representation, one can locate note onsets and durations, as well as harmonic partials generated by the sounding instruments. For an overview of computational approaches and music processing in general, we refer to the literature, e.g., M&#x000FC;ller (<xref ref-type="bibr" rid="B29">2015</xref>), Knees and Schedl (<xref ref-type="bibr" rid="B22">2016</xref>), and Weihs et al. (<xref ref-type="bibr" rid="B48">2016</xref>).</p>
</sec>
<sec id="S3-3">
<label>3.3</label> <title>Videos</title>
<p>There exist many different web services that offer users to publish videos. Among these services, YouTube<xref ref-type="fn" rid="fn15"><sup>15</sup></xref> is without doubt the largest and most famous platform for video sharing. For our scenario, we are particularly interested in YouTube videos that contain music&#x02014;especially the music that underlies the WJD. Some of the offered music videos are official releases by record labels, but the majority are videos uploaded by private platform users. Especially the music videos uploaded by the private users often contain only a static image (cover art) or a slideshow while the audio track is a digitized version of the commercially available record. By embedding YouTube videos in a web service, one relies on the availability of these videos. Due to user deletions, copyright infringements, or legal constraints in some countries, videos may not be available. However, YouTube has a lot of redundancy, i.e., the same music recording may be available in more than one version.</p>
</sec>
<sec id="S3-4">
<label>3.4</label> <title>Additional Metadata</title>
<p>In addition to the solo transcriptions, the WJD contains basic metadata for the music recordings (e.g., artist and record name), as well as a special identifier for the MusicBrainz<xref ref-type="fn" rid="fn16"><sup>16</sup></xref> platform. MusicBrainz is a community-driven platform, which collects music metadata and makes it publicly available. With the identifier available in the WJD, one is able to request a comprehensive list of available metadata from the MusicBrainz platform (e.g., participating musicians, producer&#x02019;s name, and so on). Furthermore, MusicBrainz can serve as a gateway to other web services that offer different kinds of metadata or even other multimedia objects (e.g., pictures of the artist). This &#x0201C;web of data&#x0201D; is often referred to as the Semantic Web (Berners-Lee et al., <xref ref-type="bibr" rid="B6">2001</xref>). By using the web service DBpedia,<xref ref-type="fn" rid="fn17"><sup>17</sup></xref> we can furthermore obtain and integrate content published on Wikipedia (e.g., bibliographic information about the artist).</p>
</sec>
</sec>
<sec id="S4">
<label>4</label> <title>Retrieval and Linking Strategies</title>
<p>In this section, we report on experiments where we systematically created links between the annotations contained in the WJD and corresponding YouTube videos. The retrieval task is as follows: given a specific music recording or a solo annotation provided by the WJD as query, identify the relevant videos in the pool of YouTube videos. Since the number of YouTube videos is very large, we follow a two-step retrieval strategy that we describe in the following.</p>
<sec id="S4-1">
<label>4.1</label> <title>Retrieval Scenario</title>
<p>We started by formalizing our retrieval task following (M&#x000FC;ller, <xref ref-type="bibr" rid="B29">2015</xref>). Let &#x1D49F; be the set of all documents available on YouTube. A YouTube document <italic>D</italic>&#x02009;<italic>&#x02208;</italic>&#x02009;&#x1D49F; consists of the video and the available metadata. Let &#x1D4AC; be a collection of documents available within the WJD. A WJD document <italic>Q</italic>&#x02009;<italic>&#x02208;</italic>&#x02009;&#x1D4AC; consists of a solo annotation, the underlying music excerpts, as well as metadata. In our scenario, the document <italic>Q</italic> served as query, whereas &#x1D49F; was the database to search in. Given a query <italic>Q</italic>, the retrieval task was to identify the corresponding documents <italic>D</italic>. In our scenario, we followed a two-step retrieval strategy. First, we performed a metadata-based retrieval using the YouTube search engine. For a query <italic>Q</italic>, the result of the text-based retrieval is denoted as <inline-formula><mml:math id="M1"><mml:mrow><mml:msubsup><mml:mi mathvariant="script">D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02282;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi mathvariant='script'>D</mml:mi></mml:mrow></mml:math></inline-formula>. In the second step, we performed content-based retrieval only based on <inline-formula><mml:math id="M2"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> to identify the relevant documents, denoted as <inline-formula><mml:math id="M3"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Rel</mml:mtext></mml:mrow></mml:msubsup><mml:mtext>&#x02009;</mml:mtext><mml:mo>&#x02286;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec id="S4-2">
<label>4.2</label> <title>Text-Based Retrieval</title>
<p>In the first step of our retrieval strategy, we extracted a subset of possible video candidates from YouTube. These were retrieved by performing two text-based queries (per solo) using the standard YouTube search engine (using YouTube&#x02019;s default settings). The first text-based query term consisted of the name of the soloist and the song title (e.g., <monospace>John Coltrane Kind of Blue</monospace>). Since the soloist is not always the artist who released the record, we performed a second text-based query that consisted of the artist&#x02019;s name under which the record was released, followed by the song title (in our example: <monospace>Miles Davis Kind of Blue</monospace>). From each retrieval result, we took the top 20 candidates (or less, depending on the number of YouTube search results). Furthermore, we only considered videos that are shorter or equal to 1000&#x02009;s to avoid videos where users uploaded, for instance, complete records to YouTube (rather than individual songs).</p>
<p>In our experiments, using the first text-based query terms for all 456 solos considered in the WJD led to a pool of 4,114 video candidates. The second text-based query resulted in a pool of 4,069. In a next step, we fused the two candidate pools together, where we removed duplicates by using the video identifiers attached to every YouTube video. Our final candidate pool comprised 5,199 video candidates&#x02014;resulting in approximately 12 candidates per query (solo). Note that <inline-formula><mml:math id="M4"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> may still contain cover versions and other irrelevant documents. The following audio-based retrieval step is intended to resolve this issue.</p>
</sec>
<sec id="S4-3">
<label>4.3</label> <title>Audio-Based Retrieval</title>
<p>In a second step, we used the audio recordings from the WJD to refine the list of candidates <inline-formula><mml:math id="M5"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> obtained from the text-based retrieval. This task is also known as <italic>audio identification</italic> and can be approached in many different ways, see, e.g., Cano et al. (<xref ref-type="bibr" rid="B9">2005</xref>), M&#x000FC;ller (<xref ref-type="bibr" rid="B29">2015</xref>), and Arzt (<xref ref-type="bibr" rid="B1">2016</xref>). Our method is based on chroma features and diagonal matching which is easily extendable to retrieval scenarios with different query and database modalities (e.g., matching solo transcription against audio recordings or matching audio excerpt against sheet music representations). Furthermore, since we performed our retrieval only on the small subsets <inline-formula><mml:math id="M6"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, we do not consider efficiency issues here. In particular, we used a chroma variant called CENS with a feature rate of 5&#x02009;Hz (M&#x000FC;ller et al., <xref ref-type="bibr" rid="B31">2005</xref>; M&#x000FC;ller and Ewert, <xref ref-type="bibr" rid="B30">2011</xref>).<xref ref-type="fn" rid="fn18"><sup>18</sup></xref> We compared a query <italic>Q</italic> with each of the documents <inline-formula><mml:math id="M7"><mml:mrow><mml:mi>D</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> by using diagonal matching. This comparison yields a distance value <inline-formula><mml:math id="M8"><mml:mrow><mml:msub><mml:mi>&#x003B4;</mml:mi><mml:mrow><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mo stretchy='false'>[</mml:mo><mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>]</mml:mo></mml:mrow></mml:math></inline-formula> for each pair (<italic>Q,D</italic>), where &#x003B4;<italic><sub>Q,D</sub></italic>&#x02009;&#x0003D;&#x02009;0 refers to a perfect match and &#x003B4;<italic><sub>Q,D</sub></italic>&#x02009;&#x0003D;&#x02009;1.0 to a poor match. By sorting the documents <inline-formula><mml:math id="M9"><mml:mrow><mml:mi>D</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> by &#x003B4;<italic><sub>Q,D</sub></italic> in an ascending order, one receives a ranked list. In this ranked list, the most similar documents (w.r.t. to the used distance function) are listed on top. In the case of extracting the relevant documents, one has to further process this ranked list. For instance, one may mark a document as relevant if &#x003B4;<italic><sub>Q,D</sub></italic> is smaller than a threshold &#x003C4; &#x02208;[0:1]. All relevant documents that fulfill this condition are then collected in the subset <inline-formula><mml:math id="M10"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Rel</mml:mtext></mml:mrow></mml:msubsup><mml:mo>&#x02286;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>.</p>
<p>Using this retrieval approach with a threshold &#x003C4;&#x02009;&#x0003D;&#x02009;0.1, we were able to identify 988 relevant videos for 329 solos on YouTube (on average 3 relevant videos per solo, min&#x02009;&#x0003D;&#x02009;1, max&#x02009;&#x0003D;&#x02009;9). For 92 queries, we retrieved 1 relevant document, for 67 queries 2, for 60 queries 3, and for 110 queries more than 3 documents. However, for 124 queries, we were not able to find any relevant videos. We found different reasons for this from manually inspecting some candidate lists. One obvious reason is that the metadata-based retrieval step did not return any relevant documents in <inline-formula><mml:math id="M11"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> (e.g., for the textual <monospace>David Murray Ask me Now</monospace>). Sometimes, only other versions of the same song are available on YouTube, for instance, the textual query <monospace>Art Pepper Anthropology</monospace> yields mainly results for the version of this song from the record <italic>Art Pepper</italic>&#x02009;&#x0002B;&#x02009;<italic>Eleven: Modern Jazz Classics</italic> instead of the relevant version from the record <italic>The Intimate Art Pepper</italic>. Furthermore, in many instances, we found that relevant documents were present in <inline-formula><mml:math id="M12"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, but not recognized since the distance value &#x003B4;<italic><sub>Q,D</sub></italic> surpassed the chosen threshold &#x003C4;&#x02009;&#x0003D;&#x02009;0.1 by a small margin.</p>
</sec>
<sec id="S4-4">
<label>4.4</label> <title>Solo-Based Retrieval</title>
<p>The previous experiments were based on the assumption that we have access to the music recordings underlying the WJD annotations. However, in certain scenarios this might not be the case, for instance, when only a score representation of the piece or the solo is available. In this case, audio identification is no longer possible and one needs more general retrieval strategies. In the following experiment, we simulate a retrieval scenario by using the WJD&#x02019;s solo transcriptions (see Figure <xref ref-type="fig" rid="F3">3</xref>A) as query and convert them to chroma features. This constitutes a challenging retrieval task, where one needs to compare monophonic queries (the solo transcriptions) against polyphonic audio mixtures (music recordings contained in the YouTube videos).</p>
<p>In a first experiment for this advanced retrieval task, we took the same list of candidates <inline-formula><mml:math id="M13"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Text</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and parameters as used in Section <xref ref-type="sec" rid="S4-3">4.3</xref> and only exchanged the audio-based queries against solo-based queries. In order to evaluate the results, we took the results <inline-formula><mml:math id="M14"><mml:mrow><mml:msubsup><mml:mi mathvariant='script'>D</mml:mi><mml:mi>Q</mml:mi><mml:mrow><mml:mtext>Rel</mml:mtext></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> from the audio-based retrieval as reference. The solo-based retrieval is considered as correct if among the top-<italic>K</italic> documents in the ranked list, there is at least one relevant document. The results for this Top-K evaluation measure are shown in Table <xref ref-type="table" rid="T2">2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Top-K matching rate for the solo-based retrieval.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td align="left" valign="top">K</td>
<td align="center" valign="top">1</td>
<td align="center" valign="top">3</td>
<td align="center" valign="top">5</td>
<td align="center" valign="top">10</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">20</td>
</tr>
<tr>
<td align="left" valign="top">Top-K</td>
<td align="center" valign="top">0.85</td>
<td align="center" valign="top">0.97</td>
<td align="center" valign="top">0.99</td>
<td align="center" valign="top">1.00</td>
<td align="center" valign="top">1.00</td>
<td align="center" valign="top">1.00</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>The Top-K matching rate is calculated by dividing the Top-K matches by the 329 retrieved solos from the audio-based retrieval</italic>.</p></table-wrap-foot></table-wrap>
<p>We retrieve for 85% of the queries a relevant document at rank 1 (Top-1). For 99% of the queries, the first relevant document is within the Top-5 matches. The mean reciprocal rank for the first matches for all queries is 0.91 (&#x003C3;&#x02009;&#x0003D;&#x02009;0.22). Although the solos and audio recordings vary in their degree of polyphony, we reach respectable results. The main reason is that the solo transcriptions are relatively long and perfectly aligned, leading to a high &#x0201C;discriminative power.&#x0201D; Furthermore, the queries are very unique, since they stem from an improvisation. When the queries get shorter, usually the discriminative power decreases rapidly, as they may represent more frequently used patterns.</p>
</sec>
<sec id="S4-5">
<label>4.5</label> <title>Perspectives</title>
<p>So far, our approach for retrieving videos from YouTube relies on either audio recordings or very clean solo transcriptions taken as queries. This is exactly the situation we had in our WJD scenario. In other scenarios, one may have to deal with imperfect or less specific queries. For instance, a query might be a person humming the solo, which then requires an extra step for extracting the fundamental frequency from the hummed melody (query-by-humming), see, e.g., Pauws (<xref ref-type="bibr" rid="B33">2002</xref>), Ryyn&#x000E4;nen and Klapuri (<xref ref-type="bibr" rid="B41">2008</xref>), and Salamon et al. (<xref ref-type="bibr" rid="B43">2013</xref>). Furthermore, the query may have a different tuning, may be transposed to another key, or played with rhythmic variations. A possible solution could be to use multiple queries, e.g., transposing the query in all possible 12 keys and performing a separate retrieval for each resulting version. Another way of handling some of these issues is to use different feature representations, e.g., features that are robust against temporal deformations, see Arzt (<xref ref-type="bibr" rid="B1">2016</xref>), Sonnleitner and Widmer (<xref ref-type="bibr" rid="B46">2016</xref>), and Sonnleitner et al. (<xref ref-type="bibr" rid="B45">2016</xref>).</p>
<p>In a related retrieval scenario, described in Balke et al. (<xref ref-type="bibr" rid="B2">2016</xref>), audio recordings were retrieved from a database containing Western classical music recordings by using monophonic queries with a duration of only a few measures. Besides the discrepancy in the degree of polyphony between query and database documents, tuning, key, and tempo deviations, which frequently occur in Western classical music performances, make this retrieval task very challenging. A common preprocessing step, which targets the &#x0201C;polyphony gap&#x0201D; between query and database document, is to enhance the predominant melody in audio recordings. In Salamon et al. (<xref ref-type="bibr" rid="B43">2013</xref>), the authors used a so-called <italic>salience representation</italic> in a query-by-humming system which led to a substantial increase in performance (Salamon and G&#x000F3;mez, <xref ref-type="bibr" rid="B42">2012</xref>). In Balke et al. (<xref ref-type="bibr" rid="B4">2017b</xref>), a data-driven approach is used to estimate a salience representation for jazz music recordings which showed a similar performance as the aforementioned, salience-based method.</p>
<p>Another untapped resource for jazz music retrieval is the many publicly available solo transcriptions. However, these transcriptions are typically not available in a machine-readable format. In this case, one could use Optical Music Recognition (OMR) systems to convert sheet music images to symbolic music representations. This conversion may introduce errors, such as missing notes, wrongly detected clefs, key signatures, or accidentals, see Byrd and Schindele (<xref ref-type="bibr" rid="B7">2006</xref>), Bellini et al. (<xref ref-type="bibr" rid="B5">2007</xref>), Fremerey et al. (<xref ref-type="bibr" rid="B13">2009</xref>), Raphael and Wang (<xref ref-type="bibr" rid="B38">2011</xref>), Rebelo et al. (<xref ref-type="bibr" rid="B39">2012</xref>), Balke et al. (<xref ref-type="bibr" rid="B2">2016</xref>). A recently proposed approach for score-following tries to circumvent the difficult OMR step by directly working on the scanned images of the sheet music (Dorfer et al., <xref ref-type="bibr" rid="B11">2016</xref>). Two Convolutional Neural Networks (CNN)&#x02014;one applied to the sheet music and a second one to the audio recordings&#x02014;are used for feature extraction. In an extra layer, these features are then combined to retrieve temporal relationships between the two modalities, for instance, with a learned embedding space (Raffel and Ellis, <xref ref-type="bibr" rid="B36">2016</xref>; Dorfer et al., <xref ref-type="bibr" rid="B12">2017</xref>). Currently, new OMR approaches based on deep neural networks show promising results and may lead to a significant increase in conversion quality (Haji&#x0010D; and Dorfer, <xref ref-type="bibr" rid="B20">2017</xref>).</p>
</sec>
</sec>
<sec id="S5">
<label>5</label> <title>Application</title>
<p>In this section, we present the functionalities of our web-based application, called <italic>JazzTube</italic>, which allows users to easily access the WJD&#x02019;s annotations, as well as the corresponding YouTube videos, in different interactive ways. The application offers various ways to access the WJD. First, tables of the compositions, soloists, and transcribed solos contained in the WJD are given in the form of suitable tables. Furthermore, one can access the information on the record, the track, and at the solo level.</p>
<sec id="S5-1">
<label>5.1</label> <title>Solo View</title>
<p>Figure <xref ref-type="fig" rid="F4">4</xref> shows a screenshot of the core functionality of our interactive, web-based user interface. In the top panel, some general information about the solo (Figure <xref ref-type="fig" rid="F4">4</xref>A) is shown. Many of these entries are hyperlinks and lead to the artist&#x02019;s overview page or the corresponding track. Furthermore, several possibilities of exporting the solo transcription, either as comma-separated values (CSV) or as sheet music, are offered. The conversion from the annotations to the sheet music is obtained by using the algorithm described in Pfleiderer et al. (<xref ref-type="bibr" rid="B34">2017</xref>). Below this basic information, all available YouTube videos are listed (Figure <xref ref-type="fig" rid="F4">4</xref>B). Having more than one match gives alternatives to the user. Note that YouTube videos may have different recording qualities or may disappear from YouTube. After pressing the play button, the corresponding YouTube video is automatically retrieved and embedded in the website (Figure <xref ref-type="fig" rid="F4">4</xref>C). Below the YouTube player, a piano-roll representation of the solo transcription is presented running synchronously with the video playback (Figure <xref ref-type="fig" rid="F4">4</xref>D). Finally, at the bottom, additional statistics about the solo (e.g., pitch histograms) are provided (Figure <xref ref-type="fig" rid="F4">4</xref>E).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Screenshot of our web-based interface called <italic>JazzTube</italic>. <bold>(A)</bold> Metadata and export functionalities. <bold>(B)</bold> List of linked YouTube videos. <bold>(C)</bold> Embedded YouTube video. <bold>(D)</bold> Piano-roll representation of the solo transcription synchronized with the YouTube video. <bold>(E)</bold> Additional statistics. Panel <bold>(C)</bold> has been created by the authors and, therefore, no permission is required for its use in this manuscript.</p></caption>
<graphic xlink:href="fdigh-05-00001-g004.tif"/>
</fig>
</sec>
<sec id="S5-2">
<label>5.2</label> <title>Soloist View</title>
<p>Starting from an overview table of available soloists, the user can navigate to the soloist view containing additional details about the artist. Here, one can also find the available solo transcriptions for the given soloist. Furthermore, <italic>Semantic Web</italic> technologies are used to perform a search query on <italic>DBpedia</italic> to retrieve further details. Usually the received response is very rich in information. Currently, a short biography and a link to the corresponding <italic>Wikipedia</italic> entry for further reading are included. In addition to the biographical data, further relationships to other artists, obtained from the <italic>LinkedJazz</italic> project, are embedded.</p>
</sec>
<sec id="S5-3">
<label>5.3</label> <title>Technical Details</title>
<p>Our web-based demonstrator is a typical client-server application. The client uses the Hypertext Transfer Protocol (HTTP) to perform requests to the server (e.g., by entering an URL through a web browser). These requests are then processed by the server and the response is displayed in the user&#x02019;s web browser. For setting the layout, we use the open-source framework <italic>Bootstrap</italic>.<xref ref-type="fn" rid="fn19"><sup>19</sup></xref> This framework allows for designing a website for different devices (e.g., laptops, tablets, or smartphones). Interactions and animations within the client are realized with <italic>JavaScript</italic>.<xref ref-type="fn" rid="fn20"><sup>20</sup></xref> In particular, a framework called <italic>D3 (Data-Driven Documents)</italic><xref ref-type="fn" rid="fn21"><sup>21</sup></xref> for visualizing the piano-roll is employed. For the server backend, the Python framework <italic>Flask</italic> is used.<xref ref-type="fn" rid="fn22"><sup>22</sup></xref></p>
</sec>
<sec id="S5-4">
<label>5.4</label> <title>Possible Advancements for JazzTube</title>
<p>In the case of the Weimar Jazz Database, looking at the scrolling piano-roll visualization of a jazz improvisation while simultaneously listening to the recording could be both of high educational value and a great pleasure. To relate sounding music to moving pitch contours, rhythms, changing event densities, and recurring or contrasting motifs and patterns, which are easily recognizable from a piano-roll visualization, can enrich and deepen the understanding of the tonal, rhythmical, and formal dimensions of the music in an inimitable way. Moreover, recognizing musical passages visually immediately before listening to the sounds can contribute to the play with musical expectancies (or &#x0201C;sweet anticipation&#x0201D; (Huron, <xref ref-type="bibr" rid="B21">2006</xref>)), which lie at the heart of the pleasures of listening to music.</p>
<p>For the future, several extensions to the current form of <italic>JazzTube</italic> are desirable. The piano-roll representation could be extended with different layers of annotations, such as phrases, midlevel units, chords, choruses, form part, or tone formation, which are already available in the WJD. Coloring or annotating events with respect to different functions, e.g., roots of underlying chords, passing tones, or melodic accents, would give an even deeper insight into the inner structure of an improvisation. In the case of jazz, automated identification and annotation of patterns and licks would provide options for analysis not easily achievable with traditional paper and pencil tools. Furthermore, retrieving patterns or motifs from the database would be of great value, for example, by selecting a few tones in a solo and finding and displaying all cross-references in the corpus. Finally, adding options of score-following would be of great help, since music notation is still the standard communication and representation tools of musicians and musicologists. On a different footing, the vast educational implications could be further exploited by adding specialized display options or specifically designed course materials and tutorials (e.g., on jazz history) based on the contents and possibilities of <italic>JazzTube</italic>.</p>
</sec>
</sec>
<sec id="S6">
<label>6</label> <title>Conclusion</title>
<p>With <italic>JazzTube</italic>, we offer researchers and music lovers novel possibilities to interact with and navigate through the content of the WJD. With JazzTube&#x02019;s innovative approach to link scientific music databases, including metadata, transcriptions, and further annotations to the corresponding audio recordings that are publicly available via YouTube, copyright restrictions can be bypassed in an elegant way. The approach of <italic>JazzTube</italic> could open up a way for music projects to connect metadata and annotations with audio recordings that cannot be freely provided on the internet but can be used for searching for the corresponding audio recordings at YouTube. This could be a way to easily link, e.g., the recording metadata provided within the JDISC<xref ref-type="fn" rid="fn23"><sup>23</sup></xref> project with YouTube recordings. Furthermore, we envision that <italic>JazzTube</italic> is a source for inspiration and fosters the necessary dialog between musicologists and computer scientists to further advance the field of <italic>Digital Humanities</italic>.</p>
</sec>
<sec id="S7" sec-type="author-contributor">
<title>Author Contributions</title>
<p>Many people have contributed to this paper in various ways. In collaboration with all authors, SB and MM developed the presented concepts and wrote the manuscript. SB carried out the retrieval experiments and realized the web-based interface. CD was involved in technical discussions and writing. MP, JA, and KF were responsible for the data generation and curation of the annotations used in this study. All authors contributed to revisions and additions of the manuscript.</p>
</sec>
<sec id="S8">
<title>Conflict of Interest Statement</title>
<p>This research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>We would like to thank all student annotators of the Jazzomat research project and Patricio L&#x000F3;pez-Serrano for proof-reading the manuscript. SB wants to thank Brian McFee for giving inspiration on web-based systems with his work on the JDISC project together with Dan P. W. Ellis (Columbia University, New York City).</p>
</ack>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work was supported by the German Research Foundation (DFG) under grant numbers: MU 2686/6-1 (SB and MM), MU 2686/11-1 (SB and MM), MU 2686/12-1 (SB and MM), MU 2686/10-1 (CD and MM), and PF 669/7-1 (KF, JA, and MP).</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Arzt</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <source>Flexible and Robust Music Tracking</source>. Ph.D. thesis, <publisher-name>Universit&#x000E4;t Linz</publisher-name>, <publisher-loc>Linz</publisher-loc>.</citation></ref>
<ref id="B2"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Balke</surname> <given-names>S.</given-names></name> <name><surname>Arifi-M&#x000FC;ller</surname> <given-names>V.</given-names></name> <name><surname>Lamprecht</surname> <given-names>L.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Retrieving audio recordings using musical themes</article-title>. In <conf-name>Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)</conf-name>, <fpage>281</fpage>&#x02013;<lpage>285</lpage>. <conf-loc>Shanghai, China</conf-loc>.</citation></ref>
<ref id="B3"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Balke</surname> <given-names>S.</given-names></name> <name><surname>Bie&#x000DF;mann</surname> <given-names>P.</given-names></name> <name><surname>Trump</surname> <given-names>S.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name></person-group> (<year>2017a</year>). <article-title>Konzeption und Umsetzung webbasierter Werkzeuge f&#x000FC;r das Erlernen von Jazz-Piano</article-title>. In <conf-name>Proceedings of the GI Jahrestagung</conf-name>, <fpage>61</fpage>&#x02013;<lpage>73</lpage>. <conf-loc>Chemnitz, Germany</conf-loc>.</citation></ref>
<ref id="B4"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Balke</surname> <given-names>S.</given-names></name> <name><surname>Dittmar</surname> <given-names>C.</given-names></name> <name><surname>Abe&#x000DF;er</surname> <given-names>J.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name></person-group> (<year>2017b</year>). <article-title>Data-driven solo voice enhancement for Jazz music retrieval</article-title>. In <conf-name>Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)</conf-name>, <fpage>196</fpage>&#x02013;<lpage>200</lpage>. <conf-loc>New Orleans, USA</conf-loc>.</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bellini</surname> <given-names>P.</given-names></name> <name><surname>Bruno</surname> <given-names>I.</given-names></name> <name><surname>Nesi</surname> <given-names>P.</given-names></name></person-group> (<year>2007</year>). <article-title>Assessing optical music recognition tools</article-title>. <source>Comput. Music J.</source> <volume>31</volume>: <fpage>68</fpage>&#x02013;<lpage>93</lpage>.<pub-id pub-id-type="doi">10.1162/comj.2007.31.1.68</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berners-Lee</surname> <given-names>T.</given-names></name> <name><surname>Hendler</surname> <given-names>J.</given-names></name> <name><surname>Lassila</surname> <given-names>O.</given-names></name></person-group> (<year>2001</year>). <article-title>The semantic web</article-title>. <source>Scientific American</source> <volume>284</volume>: <fpage>28</fpage>&#x02013;<lpage>37</lpage>.<pub-id pub-id-type="doi">10.1038/scientificamerican0501-34</pub-id></citation></ref>
<ref id="B7"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Byrd</surname> <given-names>D.</given-names></name> <name><surname>Schindele</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>Prospects for improving OMR with multiple recognizers</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>41</fpage>&#x02013;<lpage>46</lpage>. <conf-loc>Victoria, Canada</conf-loc>.</citation></ref>
<ref id="B8"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Cannam</surname> <given-names>C.</given-names></name> <name><surname>Landone</surname> <given-names>C.</given-names></name> <name><surname>Sandler</surname> <given-names>M.</given-names></name> <name><surname>Bello</surname> <given-names>J.P.</given-names></name></person-group> (<year>2006</year>). <article-title>The sonic visualiser: a visualisation platform for semantic descriptors from musical signals</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>324</fpage>&#x02013;<lpage>327</lpage>. <conf-loc>Victoria, Canada</conf-loc>.</citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cano</surname> <given-names>P.</given-names></name> <name><surname>Batlle</surname> <given-names>E.</given-names></name> <name><surname>Kalker</surname> <given-names>T.</given-names></name> <name><surname>Haitsma</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <article-title>A review of audio fingerprinting</article-title>. <source>J. VLSI Signal Process.</source> <volume>41</volume>: <fpage>271</fpage>&#x02013;<lpage>84</lpage>.<pub-id pub-id-type="doi">10.1007/s11265-005-4151-3</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Damm</surname> <given-names>D.</given-names></name> <name><surname>Fremerey</surname> <given-names>C.</given-names></name> <name><surname>Thomas</surname> <given-names>V.</given-names></name> <name><surname>Clausen</surname> <given-names>M.</given-names></name> <name><surname>Kurth</surname> <given-names>F.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>A digital library framework for heterogeneous music collections: from document acquisition to cross-modal interaction</article-title>. <source>Int. J. Digit. Libr.</source> <volume>12</volume>: <fpage>53</fpage>&#x02013;<lpage>71</lpage>.<pub-id pub-id-type="doi">10.1007/s00799-012-0087-y</pub-id></citation></ref>
<ref id="B11"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Dorfer</surname> <given-names>M.</given-names></name> <name><surname>Arzt</surname> <given-names>A.</given-names></name> <name><surname>Widmer</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Towards score following in sheet music images</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>789</fpage>&#x02013;<lpage>795</lpage>. <conf-loc>New York, USA</conf-loc>.</citation></ref>
<ref id="B12"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Dorfer</surname> <given-names>M.</given-names></name> <name><surname>Arzt</surname> <given-names>A.</given-names></name> <name><surname>Widmer</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Learning audio-sheet music correspondences for score identification and offline alignment</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>115</fpage>&#x02013;<lpage>122</lpage>. <conf-loc>Suzhou, China</conf-loc>.</citation></ref>
<ref id="B13"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Fremerey</surname> <given-names>C.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name> <name><surname>Clausen</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Towards bridging the gap between sheet music and audio</article-title>. In <conf-name>Knowledge Representation for Intelligent Music Processing</conf-name>, Edited by <person-group person-group-type="editor"><name><surname>Selfridge-Field</surname> <given-names>E.</given-names></name> <name><surname>Wiering</surname> <given-names>F.</given-names></name> <name><surname>Wiggins</surname> <given-names>G.A.</given-names></name></person-group>, <fpage>9051</fpage>. <conf-loc>Dagstuhl, Germany</conf-loc>: <conf-sponsor>Schloss Dagstuhl &#x02013; Leibniz-Zentrum fuer Informatik</conf-sponsor>.</citation></ref>
<ref id="B14"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Frieler</surname> <given-names>K.</given-names></name> <name><surname>Pfleiderer</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Onbeat oder offbeat? &#x000DC;berlegungen zur symbolischen Darstellung von Musik am Beispiel der metrischen Quantisierung</article-title>. In <conf-name>Proceedings of the GI Jahrestagung</conf-name>, <fpage>111</fpage>&#x02013;<lpage>125</lpage>. <conf-loc>Mainz, Germany</conf-loc>.</citation></ref>
<ref id="B15"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Gasser</surname> <given-names>M.</given-names></name> <name><surname>Arzt</surname> <given-names>A.</given-names></name> <name><surname>Gadermaier</surname> <given-names>T.</given-names></name> <name><surname>Grachten</surname> <given-names>M.</given-names></name> <name><surname>Widmer</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Classical music on the web &#x02013; user interfaces and data representations</article-title>. In <conf-name>Proceedings of the International Conference on Music Information Retrieval (ISMIR)</conf-name>, <fpage>571</fpage>&#x02013;<lpage>577</lpage>. <conf-loc>M&#x000E1;laga, Spain</conf-loc></citation></ref>
<ref id="B16"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Gemmeke</surname> <given-names>J.F.</given-names></name> <name><surname>Ellis</surname> <given-names>D.P.W.</given-names></name> <name><surname>Freedman</surname> <given-names>D.</given-names></name> <name><surname>Jansen</surname> <given-names>A.</given-names></name> <name><surname>Lawrence</surname> <given-names>W.</given-names></name> <name><surname>Moore</surname> <given-names>R.C.</given-names></name> <etal/></person-group> (<year>2017</year>). <article-title>Audio set: an ontology and human-labeled dataset for audio events</article-title>. In <conf-name>Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)</conf-name>, <fpage>776</fpage>&#x02013;<lpage>780</lpage>. <conf-loc>New Orleans, USA</conf-loc>.</citation></ref>
<ref id="B17"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Goto</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Music listening in the future: augmented music-understanding interfaces and crowd music listening</article-title>. In <conf-name>Proceedings of the Audio Engineering Society (AES) Conference on Semantic Audio</conf-name>, <conf-loc>Ilmenau, Germany</conf-loc>.</citation></ref>
<ref id="B18"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Goto</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Frontiers of music information research based on signal processing</article-title>. In <conf-name>Proceedings of the International Conference on Signal Processing (ICSP)</conf-name>, <fpage>7</fpage>&#x02013;<lpage>14</lpage>. <conf-loc>Hangzhou, China</conf-loc>.</citation></ref>
<ref id="B19"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Goto</surname> <given-names>M.</given-names></name> <name><surname>Yoshii</surname> <given-names>K.</given-names></name> <name><surname>Fujihara</surname> <given-names>H.</given-names></name> <name><surname>Mauch</surname> <given-names>M.</given-names></name> <name><surname>Nakano</surname> <given-names>T.</given-names></name></person-group> (<year>2011</year>). <article-title>Songle: a web service for active music listening improved by user contributions</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>311</fpage>&#x02013;<lpage>316</lpage>. <conf-loc>Miami, Florida</conf-loc>.</citation></ref>
<ref id="B20"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Haji&#x0010D;</surname> <given-names>J.</given-names> <suffix>Jr.</suffix></name> <name><surname>Dorfer</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Prototyping full-pipeline optical music recognition with musicmarker</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR): Late Breaking Session</conf-name>, <conf-loc>Suzhou, China</conf-loc>.</citation></ref>
<ref id="B21"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Huron</surname> <given-names>D.B.</given-names></name></person-group> (<year>2006</year>). <source>Sweet Anticipation: Music and the Psychology of Expectation</source>. <publisher-name>The MIT Press</publisher-name>, <publisher-loc>Cambridge, Massachusetts</publisher-loc>.</citation></ref>
<ref id="B22"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Knees</surname> <given-names>P.</given-names></name> <name><surname>Schedl</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <source>Music Similarity and Retrieval</source>. <publisher-name>Springer Verlag</publisher-name>, <publisher-loc>Berlin, Heidelberg</publisher-loc>.</citation></ref>
<ref id="B23"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Kurth</surname> <given-names>F.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name> <name><surname>Damm</surname> <given-names>D.</given-names></name> <name><surname>Fremerey</surname> <given-names>C.</given-names></name> <name><surname>Ribbrock</surname> <given-names>A.</given-names></name> <name><surname>Clausen</surname> <given-names>M.</given-names></name></person-group> (<year>2005</year>). <article-title>SyncPlayer &#x02013; an advanced system for multimodal music access</article-title>. In <conf-name>Proceedings of the International Conference on Music Information Retrieval (ISMIR)</conf-name>, <fpage>381</fpage>&#x02013;<lpage>388</lpage>. <conf-loc>London, UK</conf-loc>.</citation></ref>
<ref id="B24"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Liem</surname> <given-names>C.C.S.</given-names></name> <name><surname>G&#x000F3;mez</surname> <given-names>E.</given-names></name> <name><surname>Schedl</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>PHENICX: innovating the classical music experience</article-title>. In <conf-name>Proceedings of the IEEE International Conference on Multimedia and Expo Workshops (ICMEW)</conf-name>, <fpage>1</fpage>&#x02013;<lpage>4</lpage>. <conf-loc>Torino, Italy</conf-loc>.</citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liem</surname> <given-names>C.C.S.</given-names></name> <name><surname>G&#x000F3;mez</surname> <given-names>E.</given-names></name> <name><surname>Tzanetakis</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Multimedia technologies for enriched music performance, production, and consumption</article-title>. <source>IEEE MultiMedia</source> <volume>24</volume>: <fpage>20</fpage>&#x02013;<lpage>3</lpage>.<pub-id pub-id-type="doi">10.1109/MMUL.2017.20</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McFee</surname> <given-names>B.</given-names></name> <name><surname>McVicar</surname> <given-names>M.</given-names></name> <name><surname>Nieto</surname> <given-names>O.</given-names></name> <name><surname>Balke</surname> <given-names>S.</given-names></name> <name><surname>Thom&#x000E9;</surname> <given-names>C.</given-names></name> <name><surname>Liang</surname> <given-names>D.</given-names></name> <etal/></person-group> (<year>2017</year>). <source>Librosa 0.5.0</source>. Zenodo.<pub-id pub-id-type="doi">10.5281/zenodo.293021</pub-id></citation></ref>
<ref id="B27"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Melenhorst</surname> <given-names>M.S.</given-names></name> <name><surname>van der Sterren</surname> <given-names>R.</given-names></name> <name><surname>Arzt</surname> <given-names>A.</given-names></name> <name><surname>Martorell</surname> <given-names>A.</given-names></name> <name><surname>Liem</surname> <given-names>C.C.</given-names></name></person-group> (<year>2015</year>). <article-title>A tablet app to enrich the live and post-live experience of classical concerts</article-title>. In <conf-name>Proceedings of the International Workshop on Interactive Content Consumption (WSICC)</conf-name>, <conf-loc>Brussels, Belgium</conf-loc>.</citation></ref>
<ref id="B28"><citation citation-type="book"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). <source>Information Retrieval for Music and Motion</source>. <publisher-name>Springer Verlag</publisher-name>, <publisher-loc>Berlin, Heidelberg</publisher-loc>.</citation></ref>
<ref id="B29"><citation citation-type="book"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <source>Fundamentals of Music Processing</source>. <publisher-name>Springer Verlag</publisher-name>, <publisher-loc>Berlin, Heidelberg</publisher-loc>.</citation></ref>
<ref id="B30"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name> <name><surname>Ewert</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>Chroma toolbox: MATLAB implementations for extracting variants of chroma-based audio features</article-title>. In <conf-name>Proceedings of the International Conference on Music Information Retrieval (ISMIR)</conf-name>, <fpage>215</fpage>&#x02013;<lpage>220</lpage>. <conf-loc>Miami, Florida</conf-loc>.</citation></ref>
<ref id="B31"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name> <name><surname>Kurth</surname> <given-names>F.</given-names></name> <name><surname>Clausen</surname> <given-names>M.</given-names></name></person-group> (<year>2005</year>). <article-title>Chroma-based statistical audio features for audio matching</article-title>. In <conf-name>Proceedings of the IEEE Workshop on Applications of Signal Processing (WASPAA)</conf-name>, <fpage>275</fpage>&#x02013;<lpage>278</lpage>. <conf-loc>New Paltz, NY</conf-loc>.</citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pattuelli</surname> <given-names>M.C.</given-names></name></person-group> (<year>2012</year>). <article-title>Personal name vocabularies as linked open data: a case study of Jazz artist names</article-title>. <source>J. Info. Sci.</source> <volume>38</volume>: <fpage>558</fpage>&#x02013;<lpage>65</lpage>.<pub-id pub-id-type="doi">10.1177/0165551512455989</pub-id></citation></ref>
<ref id="B33"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Pauws</surname> <given-names>S.</given-names></name></person-group> (<year>2002</year>). <article-title>CubyHum: a fully operational query by humming system</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>187</fpage>&#x02013;<lpage>196</lpage>. <conf-loc>Paris, France</conf-loc>.</citation></ref>
<ref id="B34"><citation citation-type="book"><person-group person-group-type="editor"><name><surname>Pfleiderer</surname> <given-names>M.</given-names></name> <name><surname>Frieler</surname> <given-names>K.</given-names></name> <name><surname>Abe&#x000DF;er</surname> <given-names>J.</given-names></name> <name><surname>Zaddach</surname> <given-names>W.-G.</given-names></name> <name><surname>Burkhart</surname> <given-names>B.</given-names></name></person-group> eds. (<year>2017</year>). <source>Inside the Jazzomat. New Perspectives for Jazz Research</source>. <publisher-loc>Mainz, Germany</publisher-loc>: <publisher-name>Schott Campus</publisher-name>.</citation></ref>
<ref id="B35"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Pr&#x000E4;tzlich</surname> <given-names>T.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name> <name><surname>Bohl</surname> <given-names>B.W.</given-names></name> <name><surname>Veit</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Freisch&#x000FC;tz digital: demos of audio-related contributions</article-title>. In <source>Demos and Late Breaking News of the International Society for Music Information Retrieval Conference (ISMIR)</source>, <publisher-loc>Mal&#x000E1;ga, Spain</publisher-loc>.</citation></ref>
<ref id="B36"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Raffel</surname> <given-names>C.</given-names></name> <name><surname>Ellis</surname> <given-names>D.P.W.</given-names></name></person-group> (<year>2016</year>). <article-title>Pruning subsequence search with attention-based embedding</article-title>. In <conf-name>Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)</conf-name>, <fpage>554</fpage>&#x02013;<lpage>558</lpage>. <conf-loc>Shanghai, China</conf-loc>.</citation></ref>
<ref id="B37"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Raimond</surname> <given-names>Y.</given-names></name> <name><surname>Abdallah</surname> <given-names>S.</given-names></name> <name><surname>Sandler</surname> <given-names>M.</given-names></name> <name><surname>Giasson</surname> <given-names>F.</given-names></name></person-group> (<year>2007</year>). <article-title>The music ontology</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>417</fpage>&#x02013;<lpage>422</lpage>. <conf-loc>Vienna, Austria</conf-loc>.</citation></ref>
<ref id="B38"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Raphael</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>New approaches to optical music recognition</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>305</fpage>&#x02013;<lpage>310</lpage>. <conf-loc>Miami, FL</conf-loc>.</citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rebelo</surname> <given-names>A.</given-names></name> <name><surname>Fujinaga</surname> <given-names>I.</given-names></name> <name><surname>Paszkiewicz</surname> <given-names>F.</given-names></name> <name><surname>Marcal</surname> <given-names>A.R.S.</given-names></name> <name><surname>Guedes</surname> <given-names>C.</given-names></name> <name><surname>Cardoso</surname> <given-names>J.S.</given-names></name></person-group> (<year>2012</year>). <article-title>Optical music recognition: state-of-the-art and open issues</article-title>. <source>Int. J. Multimedia Information Retr.</source> <volume>1</volume>: <fpage>173</fpage>&#x02013;<lpage>90</lpage>.<pub-id pub-id-type="doi">10.1007/s13735-012-0004-6</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>R&#x000F6;wenstrunk</surname> <given-names>D.</given-names></name> <name><surname>Pr&#x000E4;tzlich</surname> <given-names>T.</given-names></name> <name><surname>Betzwieser</surname> <given-names>T.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>M.</given-names></name> <name><surname>Szwillus</surname> <given-names>G.</given-names></name> <name><surname>Veit</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Das Gesamtkunstwerk Oper aus Datensicht &#x02013; Aspekte des Umgangs mit einer heterogenen Datenlage im BMBF-Projekt &#x0201C;Freisch&#x000FC;tz Digital&#x0201D;</article-title>. <source>Datenbank Spektrum</source> <volume>15</volume>: <fpage>65</fpage>&#x02013;<lpage>72</lpage>.<pub-id pub-id-type="doi">10.1007/s13222-015-0179-0</pub-id></citation></ref>
<ref id="B41"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ryyn&#x000E4;nen</surname> <given-names>M.</given-names></name> <name><surname>Klapuri</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>Query by humming of MIDI and audio using locality sensitive hashing</article-title>. In <source>IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>, <fpage>2249</fpage>&#x02013;<lpage>2252</lpage>. <publisher-loc>Las Vegas, NV</publisher-loc>.</citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salamon</surname> <given-names>J.</given-names></name> <name><surname>G&#x000F3;mez</surname> <given-names>E.</given-names></name></person-group> (<year>2012</year>). <article-title>Melody extraction from polyphonic music signals using pitch contour characteristics</article-title>. <source>IEEE Trans. Audio Speech Lang. Process.</source> <volume>20</volume>: <fpage>1759</fpage>&#x02013;<lpage>70</lpage>.<pub-id pub-id-type="doi">10.1109/TASL.2012.2188515</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salamon</surname> <given-names>J.</given-names></name> <name><surname>Serr&#x000E0;</surname> <given-names>J.</given-names></name> <name><surname>G&#x000F3;mez</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). <article-title>Tonal representations for music retrieval: from version identification to query-by-humming</article-title>. <source>Int. J. Multimedia Info. Retr.</source> <volume>2</volume>: <fpage>45</fpage>&#x02013;<lpage>58</lpage>.<pub-id pub-id-type="doi">10.1007/s13735-012-0026-0</pub-id></citation></ref>
<ref id="B44"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Schoeffler</surname> <given-names>M.</given-names></name> <name><surname>Herre</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>The influence of audio quality on the popularity of music videos: a YouTube case study</article-title>. In <conf-name>Proceedings of the International Workshop on Internet-Scale Multimedia Management</conf-name>, <fpage>35</fpage>&#x02013;<lpage>38</lpage>. <conf-loc>Orlando, Florida, USA</conf-loc>.</citation></ref>
<ref id="B45"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Sonnleitner</surname> <given-names>R.</given-names></name> <name><surname>Arzt</surname> <given-names>A.</given-names></name> <name><surname>Widmer</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Landmark-based audio fingerprinting for DJ mix monitoring</article-title>. In <conf-name>Proceedings of the International Conference on Music Information Retrieval (ISMIR)</conf-name>, <fpage>185</fpage>&#x02013;<lpage>191</lpage>. <conf-loc>New York City, NY</conf-loc>.</citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sonnleitner</surname> <given-names>R.</given-names></name> <name><surname>Widmer</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Robust quad-based audio fingerprinting</article-title>. <source>IEEE Trans. Audio Speech Lang. Process.</source> <volume>24</volume>: <fpage>409</fpage>&#x02013;<lpage>21</lpage>.<pub-id pub-id-type="doi">10.1109/TASLP.2015.2509248</pub-id></citation></ref>
<ref id="B47"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Thomas</surname> <given-names>V.</given-names></name> <name><surname>Fremerey</surname> <given-names>C.</given-names></name> <name><surname>Damm</surname> <given-names>D.</given-names></name> <name><surname>Clausen</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>SLAVE: a score-lyrics-audio-video-explorer</article-title>. In <conf-name>Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)</conf-name>, <fpage>717</fpage>&#x02013;<lpage>722</lpage>. <conf-loc>Kobe, Japan</conf-loc>.</citation></ref>
<ref id="B48"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Weihs</surname> <given-names>C.</given-names></name> <name><surname>Jannach</surname> <given-names>D.</given-names></name> <name><surname>Vatolkin</surname> <given-names>I.</given-names></name> <name><surname>Rudolph</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <source>Music Data Analysis: Foundations and Applications</source>. <publisher-name>CRC Press</publisher-name>, <publisher-loc>Abingdon, UK</publisher-loc>.</citation></ref>
</ref-list>
<fn-group>
<fn id="fn1"><p><sup>1</sup><uri xlink:href="https://www.halleonard.com/search/search.action?seriesfeature&#x0003D;OMNIBK">https://www.halleonard.com/search/search.action?seriesfeature&#x0003D;OMNIBK</uri>.</p></fn>
<fn id="fn2"><p><sup>2</sup><uri xlink:href="http://songle.jp">http://songle.jp</uri>.</p></fn>
<fn id="fn3"><p><sup>3</sup><uri xlink:href="http://songrium.jp">http://songrium.jp</uri>.</p></fn>
<fn id="fn4"><p><sup>4</sup><uri xlink:href="http://phenicx.prototype.videodock.com">http://phenicx.prototype.videodock.com</uri>.</p></fn>
<fn id="fn5"><p><sup>5</sup><uri xlink:href="http://www.freischuetz-digital.de">http://www.freischuetz-digital.de</uri>.</p></fn>
<fn id="fn6"><p><sup>6</sup><uri xlink:href="http://www.dbpedia.org">http://www.dbpedia.org</uri>.</p></fn>
<fn id="fn7"><p><sup>7</sup><uri xlink:href="https://www.musicbrainz.org">https://www.musicbrainz.org</uri>.</p></fn>
<fn id="fn8"><p><sup>8</sup><uri xlink:href="https://www.discogs.com">https://www.discogs.com</uri>.</p></fn>
<fn id="fn9"><p><sup>9</sup><uri xlink:href="http://jdisc.columbia.edu">http://jdisc.columbia.edu</uri>.</p></fn>
<fn id="fn10"><p><sup>10</sup><uri xlink:href="https://www.linkedjazz.org">https://www.linkedjazz.org</uri>.</p></fn>
<fn id="fn11"><p><sup>11</sup><uri xlink:href="http://jazzomat.hfm-weimar.de">http://jazzomat.hfm-weimar.de</uri>.</p></fn>
<fn id="fn12"><p><sup>12</sup>Additional statistics:&#x02009;<uri xlink:href="http://mir.audiolabs.uni-erlangen.de/jazztube/statistics/">http://mir.audiolabs.uni-erlangen.de/jazztube/statistics/</uri>.</p></fn>
<fn id="fn13"><p><sup>13</sup>The sheet music representation was generated by using the <italic>LilyPond</italic> (<uri xlink:href="http://www.lilypond.org/">http://www.lilypond.org/</uri>) export which can be obtained from the WJD by using the <italic>MeloSpyGUI</italic> (<uri xlink:href="http://jazzomat.hfm-weimar.de/download/download.html&#x00023;download-melospygui">http://jazzomat.hfm-weimar.de/download/download.html&#x00023;download-melospygui</uri>).</p></fn>
<fn id="fn14"><p><sup>14</sup>For the example in Figure 2a, we use the semitone filterbank described in (M&#x000FC;ller, <xref ref-type="bibr" rid="B28">2007</xref>; M&#x000FC;ller and Ewert, <xref ref-type="bibr" rid="B30">2011</xref>).</p></fn>
<fn id="fn15"><p><sup>15</sup><uri xlink:href="http://www.youtube.com">http://www.youtube.com</uri>.</p></fn>
<fn id="fn16"><p><sup>16</sup><uri xlink:href="http://www.musicbrainz.org/">http://www.musicbrainz.org/</uri>.</p></fn>
<fn id="fn17"><p><sup>17</sup><uri xlink:href="http://wiki.dbpedia.org">http://wiki.dbpedia.org</uri>.</p></fn>
<fn id="fn18"><p><sup>18</sup>All computations can be done by using the implementations provided by the Python library <italic>librosa</italic> (McFee et al., <xref ref-type="bibr" rid="B26">2017</xref>).</p></fn>
<fn id="fn19"><p><sup>19</sup><uri xlink:href="http://www.getbootstrap.com">http://www.getbootstrap.com</uri>.</p></fn>
<fn id="fn20"><p><sup>20</sup><uri xlink:href="https://www.javascript.com">https://www.javascript.com</uri>.</p></fn>
<fn id="fn21"><p><sup>21</sup><uri xlink:href="https://www.d3js.org">https://www.d3js.org</uri>.</p></fn>
<fn id="fn22"><p><sup>22</sup><uri xlink:href="http://flask.pocoo.org">http://flask.pocoo.org</uri>.</p></fn>
<fn id="fn23"><p><sup>23</sup><uri xlink:href="http://jdisc.columbia.edu">http://jdisc.columbia.edu</uri>.</p></fn>
</fn-group>
</back>
</article>
