<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Vet. Sci.</journal-id>
<journal-title>Frontiers in Veterinary Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Vet. Sci.</abbrev-journal-title>
<issn pub-type="epub">2297-1769</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fvets.2017.00002</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Veterinary Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Using Machine Learning to Predict Swine Movements within a Regional Program to Improve Control of Infectious Diseases in the US</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Valdes-Donoso</surname> <given-names>Pablo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<uri xlink:href="http://frontiersin.org/people/u/255234"/>
</contrib>
<contrib contrib-type="author">
<name><surname>VanderWaal</surname> <given-names>Kimberly</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/254628"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Jarvis</surname> <given-names>Lovell S.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/303642"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wayne</surname> <given-names>Spencer R.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Perez</surname> <given-names>Andres M.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/198846"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Veterinary Population Medicine, College of Veterinary Medicine, University of Minnesota</institution>, <addr-line>St. Paul, MN</addr-line>, <country>USA</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Agricultural and Resource Economics, University of California Davis</institution>, <addr-line>Davis, CA</addr-line>, <country>USA</country></aff>
<aff id="aff3"><sup>3</sup><institution>Veterinary Services Pipestone</institution>, <addr-line>Pipestone, MN</addr-line>, <country>USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Salome D&#x000FC;rr, University of Bern, Switzerland</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Marco De Nardi, Safoso, Switzerland; Hartmut H. K. Lentz, Friedrich Loeffler Institute, Germany; Vitaly Belik, Freie Universit&#x000E4;t Berlin, Germany</p></fn>
<corresp content-type="corresp" id="cor1">&#x0002A;Correspondence: Pablo Valdes-Donoso, <email>pablovd&#x00040;umn.edu</email></corresp>
<fn fn-type="other" id="fn002"><p>Specialty section: This article was submitted to Veterinary Epidemiology and Economics, a section of the journal Frontiers in Veterinary Science</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>01</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>4</volume>
<elocation-id>2</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>09</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>01</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Valdes-Donoso, VanderWaal, Jarvis, Wayne and Perez.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Valdes-Donoso, VanderWaal, Jarvis, Wayne and Perez</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Between-farm animal movement is one of the most important factors influencing the spread of infectious diseases in food animals, including in the US swine industry. Understanding the structural network of contacts in a food animal industry is prerequisite to planning for efficient production strategies and for effective disease control measures. Unfortunately, data regarding between-farm animal movements in the US are not systematically collected and thus, such information is often unavailable. In this paper, we develop a procedure to replicate the structure of a network, making use of partial data available, and subsequently use the model developed to predict animal movements among sites in 34 Minnesota counties. First, we summarized two networks of swine producing facilities in Minnesota, then we used a machine learning technique referred to as random forest, an ensemble of independent classification trees, to estimate the probability of pig movements between farms and/or markets sites located in two counties in Minnesota. The model was calibrated and tested by comparing predicted data and observed data in those two counties for which data were available. Finally, the model was used to predict animal movements in sites located across 34 Minnesota counties. Variables that were important in predicting pig movements included between-site distance, ownership, and production type of the sending and receiving farms and/or markets. Using a weighted-kernel approach to describe spatial variation in the centrality measures of the predicted network, we showed that the south-central region of the study area exhibited high aggregation of predicted pig movements. Our results show an overlap with the distribution of outbreaks of porcine reproductive and respiratory syndrome, which is believed to be transmitted, at least in part, though animal movements. While the correspondence of movements and disease is not a causal test, it suggests that the predicted network may approximate actual movements. Accordingly, the predictions provided here might help to design and implement control strategies in the region. Additionally, the methodology here may be used to estimate contact networks for other livestock systems when only incomplete information regarding animal movements is available.</p>
</abstract>
<kwd-group>
<kwd>swine industry</kwd>
<kwd>pig movements</kwd>
<kwd>regional control programs</kwd>
<kwd>Minnesota</kwd>
<kwd>random forest</kwd>
<kwd>social network analysis</kwd>
</kwd-group>
<counts>
<fig-count count="10"/>
<table-count count="6"/>
<equation-count count="1"/>
<ref-count count="53"/>
<page-count count="13"/>
<word-count count="8872"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<title>Introduction</title>
<p>Between-farm direct or indirect contact <italic>via</italic> movement of animals or biological materials (e.g., semen), or cross-contamination through inputs such as machinery or human workers, is among the most important factors contributing to disease spread in food animals (<xref ref-type="bibr" rid="B1">1</xref>). Farm-to-farm contacts spread diseases that affect the US swine industry, including porcine reproductive and respiratory syndrome (PRRS) and porcine epidemic diarrhea (PED). For both PRRS and PED, animal movements (e.g., gilts, boars, weaned pigs, feeder pigs, and cull animals) represent one of the most important disease transmission routes between farms (<xref ref-type="bibr" rid="B2">2</xref>&#x02013;<xref ref-type="bibr" rid="B6">6</xref>).</p>
<p>Understanding the network structure of food animal industries is critical for efficient production and disease control. For example, the sharing of information among agents (e.g., farmers, suppliers, and brokers) within a network may result in an increase of economic efficiency due to the selection of strategies that can decrease production and/or transaction costs (<xref ref-type="bibr" rid="B7">7</xref>). Indeed, social network analysis (SNA) is an analytical tool that has been widely used in the field of veterinary medicine to design disease control plans (<xref ref-type="bibr" rid="B8">8</xref>). SNA has been used to quantify the nature of connections (referred to as <italic>edges</italic> or <italic>contacts</italic>) among elements (<italic>nodes</italic> or <italic>vertices</italic>) in a population (<xref ref-type="bibr" rid="B9">9</xref>). Nodes may be farms or other facilities (e.g., slaughter houses, truck wash disinfection stations, or feed plants) from, to, or through which, animal populations are connected, and contacts among nodes may be categorized as direct or indirect (<xref ref-type="bibr" rid="B8">8</xref>, <xref ref-type="bibr" rid="B10">10</xref>). SNA enables researchers to better understand animal movement patterns and, consequently, provide insights on how diseases diffuse in a given industry (<xref ref-type="bibr" rid="B11">11</xref>&#x02013;<xref ref-type="bibr" rid="B13">13</xref>). For example, in many livestock industries, a minority of farms typically account for the majority of animal movements (<xref ref-type="bibr" rid="B14">14</xref>&#x02013;<xref ref-type="bibr" rid="B16">16</xref>). Identification of those few farms, often referred to as &#x0201C;hotspots&#x0201D; or &#x0201C;super-spreaders&#x0201D; for disease transmission, may help formulating contingency plans to control high impact diseases, as timely intervention to targeted farms may enhance the probability of such plans being successful (<xref ref-type="bibr" rid="B13">13</xref>&#x02013;<xref ref-type="bibr" rid="B15">15</xref>, <xref ref-type="bibr" rid="B17">17</xref>). Similarly, efforts to improve animal management and biosecurity in super-spreaders may also contribute to reducing disease risk and prevalence (<xref ref-type="bibr" rid="B17">17</xref>, <xref ref-type="bibr" rid="B18">18</xref>).</p>
<p>The US swine industry is characterized by large numbers of documented pig movements within and between states and regions. From 1970 to 2001, the number of pigs moved from one state to another (or from Canada to the US) increased from 30 to 50&#x02009;million (<xref ref-type="bibr" rid="B19">19</xref>). Increases in the number and distance of movements reflect growth in the number of farms specializing in specific phases of the production cycle. Indeed, a growing proportion of feeder and finishing swine farms are located in the Midwest in close proximity to the grain used to feed pigs (<xref ref-type="bibr" rid="B20">20</xref>). In contrast, breeding populations tend to be located in areas distant from the major growing pig regions, such as the southeastern US, where grain inputs are not as critical (<xref ref-type="bibr" rid="B20">20</xref>&#x02013;<xref ref-type="bibr" rid="B22">22</xref>). Whereas the regional specialization of different industry components has undoubtedly improved efficiency, the necessary movement of animals between the two regions alters the risk of long-distance disease spread (<xref ref-type="bibr" rid="B1">1</xref>, <xref ref-type="bibr" rid="B22">22</xref>&#x02013;<xref ref-type="bibr" rid="B24">24</xref>).</p>
<p>Animal movements are only partially regulated in the US, and no source provides complete information on such movements. For example, the United State Department of Agriculture, through the animal disease traceability program, collects information on movements of cattle, bison, equines, sheep and goats, swine, and poultry, only when movements cross state boundaries, except when livestock are moved to slaughter facilities or chicks moved from hatcheries (<xref ref-type="bibr" rid="B25">25</xref>). The lack of movement data creates a particular problem for the control of diseases, such as PRRS. In that context, regional control programs (RCPs), voluntarily organized and coordinated by producers, serve as means to share sanitary status information among farmers located in a given area. Sharing information within an RCP in Minnesota (RCP-N212) has been correlated with a decrease in PRRS incidence (<xref ref-type="bibr" rid="B26">26</xref>), and thus, one may hypothesize that sharing additional information about pig movements would further improve control program effectiveness. Unfortunately, lack of information about between-farm movements hinders attempts to describe network structure, hence impairing ability to prevent and control disease.</p>
<p>To elucidate the role of network structure in the spread of swine diseases, the relation between PRRS manifestation and animal movements between farms (and other related sites, such as buyer stations or market sites) was assessed in two counties in Minnesota (<xref ref-type="bibr" rid="B27">27</xref>). A positive association between positive PRRS status and the number of direct and indirect suppliers (in-reach degree) was observed in one county, but no additional network measures were significantly correlated with positive PRRS status (<xref ref-type="bibr" rid="B27">27</xref>). Although that early study provided valuable insights about pig movements between sites and their potential contribution to disease spread, a more complete assessment of the structure of contacts is required to understand disease spread. We use data from Wayne (<xref ref-type="bibr" rid="B27">27</xref>) and more recent data collected by the RCP-N212 to build a predictive movement model between sites, which is then used to estimate a complete movement network for the RCP-N212 in Minnesota. The results may be incorporated into a disease-spread model to help explain disease dynamics and support disease prevention and control activities within the RCP-N212.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2-1">
<title>Data Sources</title>
<p>We used two complementary sets of data to construct our model. The first dataset, referred to as the network building data set, included information on pig movements related to two counties being used to fit the model, whereas the second dataset included information on sites located within the broader RCP-N212 area, which was used for prediction purposes. The first dataset included information collected in two counties that were geographically located within the boundaries of the second dataset; however, the two datasets were collected separately. The first data set was based on surveys conducted with owners, managers, and veterinarians on farms and at market sites located in Stevens and Rice counties, Minnesota in 2006 (<xref ref-type="bibr" rid="B27">27</xref>). Animal movement data included origin and destination of sites in and out of Stevens and Rice, geographic locations, and the production type of sites and owner. Production types included boar stud (BS), farrowing (Fa), nursery (N), finishing (Fi) farms, and market sites (M). This last type encompasses buying stations or slaughter plants. Two networks were described in the building dataset, one for each county, i.e., a Stevens network (SN) and a Rice network (RN). Each network contained data on directional animal movements between any given site located within the county and a number of sites located either inside or outside the county.</p>
<p>The second data set, referred to as RCP-N212, contained information on geographical location, owner, and type of site for premises enrolled in the RCP-N212. This data set contains roughly 38% of total swine premises with 100 or more animals located in Minnesota (<xref ref-type="bibr" rid="B28">28</xref>). Data were collected between 2012 and 2015. The RCP-N212 comprised 34 counties in Minnesota, including Stevens and Rice counties (<xref ref-type="bibr" rid="B26">26</xref>). The University of Minnesota manages the RCP-N212 data under the terms of an agreement with swine producers that protects the confidentiality of the data.</p>
</sec>
<sec id="S2-2">
<title>Network Description</title>
<p>The structures of SN and RN were described using SNA representing directional flows of animal movements between sites. The site-level connectivity of each network was described using <italic>in-</italic> and <italic>out-degree</italic>, calculated as the number of pig movements received or sent by a specific site to or from other sites. <italic>Betweenness</italic>, defined as the number of directed paths that pass through a given site, when the shortest paths between other pairs of sites are traced (<xref ref-type="bibr" rid="B9">9</xref>), was also estimated. Metrics were stratified by site type (e.g., BS, Fa, N, Fi, and M) for SN and RN, and differences in centrality measures between types were analyzed using Kruskal&#x02013;Wallis tests. To assess the correlation between types of sites in each network, the assortativity coefficient (<italic>r</italic>) for a mixing matrix was used, as defined by elsewhere (<xref ref-type="bibr" rid="B29">29</xref>), so that
<disp-formula id="E1"><mml:math id="M1"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mstyle><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
where <italic>e<sub>ij</sub></italic> is the fraction of animal movements in the network that connects sites of type <italic>i</italic> to type <italic>j</italic>, and <italic>a<sub>i</sub></italic> and <italic>b<sub>i</sub></italic> are fractions of destination-or-origin, respectively, of a movement that is attached to site type <italic>i</italic>. A value of <italic>r</italic>&#x02009;&#x0003D;&#x02009;0 indicates no assortative mixing or a random network, while a value of <italic>r</italic>&#x02009;&#x0003D;&#x02009;1 indicates complete assortativity, i.e., that all movements are between sites of the same production type. Alternatively, if <italic>r</italic>&#x02009;&#x0003C;&#x02009;0 links are more likely to connect two different types of nodes, which is closer to having a randomly mixed network where links often connect unlike nodes (e.g., different type of sites) (<xref ref-type="bibr" rid="B29">29</xref>).</p>
<p>Four metrics at the network level were estimated, namely, (1) <italic>network density</italic>, calculated as the fraction of movements that are present in the network relative to the total number possible; (2) <italic>clustering coefficient</italic>, used as a measure of cohesiveness and defined as the probability that two sites that are linked to a common site are also linked to each other; (3) <italic>diameter</italic>, calculated as the largest distance between two sites in the network, where distance is the shortest path between two sites; and (4) the <italic>mean path length</italic>, calculated as the mean length of the shortest paths connecting two sites (<xref ref-type="bibr" rid="B9">9</xref>, <xref ref-type="bibr" rid="B30">30</xref>).</p>
<p>Figures and statistical computations were preformed using R V.3.1.1 (<xref ref-type="bibr" rid="B31">31</xref>), including the packages <italic>ggplot2</italic> (<xref ref-type="bibr" rid="B32">32</xref>), <italic>maps</italic> (<xref ref-type="bibr" rid="B33">33</xref>), <italic>MASS</italic> (<xref ref-type="bibr" rid="B34">34</xref>), and <italic>igraph</italic> (<xref ref-type="bibr" rid="B35">35</xref>).</p>
</sec>
<sec id="S2-3">
<title>Network Prediction</title>
<p>Due to the inherent attributes of nodes, their dimensional distribution, connection features, etc. predicting networks can be challenging (<xref ref-type="bibr" rid="B36">36</xref>, <xref ref-type="bibr" rid="B37">37</xref>). Unsupervised and supervised methods have been used to try to elucidate network structures. Unsupervised approaches seek to assign scores to possible links between nodes mainly based on the neighborhood characteristics of each node and path-distances between nodes. While the first estimates the likelihood of a link between two nodes based on the degree of overlap of their neighbors, the second searches for the shortest path-distance among all possible combinations between nodes (<xref ref-type="bibr" rid="B36">36</xref>). For example, the preferential attachment prediction has been used to estimate potential connections of a node given the proportional number of neighbors that it has (<xref ref-type="bibr" rid="B38">38</xref>), or the Katz coefficient scores the possible links between two nodes subject to a given length paths (<xref ref-type="bibr" rid="B39">39</xref>). While unsupervised methods have been popular in network prediction, they fail to handle network dynamics, the mutual dependence of components, and other features inherent of the network structure (e.g., an unbalanced number of links), thus often leading to unstable performance (<xref ref-type="bibr" rid="B37">37</xref>). Among supervised methods, random forest (RF) has shown high levels of classification accuracy compared to others techniques such as bagging (<xref ref-type="bibr" rid="B37">37</xref>, <xref ref-type="bibr" rid="B40">40</xref>), so it is the approach used here. Here, information provided by SN and RN was used in a RF model to predict animal movements between sites. After predictions were obtained, parameters were extrapolated to predict movements for the entire number of farms within the RCP-N212.</p>
<sec id="S2-3-1">
<title>RF Model</title>
<p>Models based on classification trees are built using a single rule or a set of rules for a number of variables that split data to predict possible outcomes. A RF is an ensemble of independent classification trees created from bootstrap samples chosen with replacement from a training data set, in which aggregated estimates from each ensemble generates a final prediction of the probability that a given outcome occurs (<xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B41">41</xref>), e.g., a link between two sites. The samples that are not selected as bootstrap samples are called &#x0201C;out-of-bag&#x0201D; (OOB) samples and are used to estimate the error rate. The OOB error rate is reduced by ranking predictors and subsequently removing those considered less important. Calculating the difference in accuracy between models in which predictors are present or removed is used to assess predictor importance. Differences are normalized across all trees generated and then ranked based on accuracy of prediction (<xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B41">41</xref>).</p>
<p>Using all sites from SN and RN, we created a new dataset (referred to as RF-data) that contained all possible origin-destination pairs of sites within each network. Per each possible pair of sites, we assigned a dichotomous outcome (yes, no) variable (also referred to as a class variable) indicating whether or not the animal movement has occurred between that pair. We used the geographical location of each site to estimate the pairwise Euclidean distance (kilometers) between farms, and generated a dichotomous variable (yes, no) indicating whether or not each pair had a common owner. Additionally, we generated 25 dummy variables, each denoting a possible pair of site types (e.g., Fa&#x02013;Fi, Fi&#x02013;M, BS&#x02013;Fa, etc.), being 1 if the pair site type combination was true and 0 otherwise.</p>
<p>The effectiveness of model prediction is determined using a portion of the data that has not been used to build and tune the model (<xref ref-type="bibr" rid="B40">40</xref>). Thus, we split the RF-data randomly, using 75% of observations to build and tune the model (referred to as the training dataset), and the remaining 25% to test or validate our model (referred to as the testing dataset or validation set). In other words, we used the training dataset to create (i.e., train and tune) the RF model and then used the testing dataset to qualify its performance through a confusion matrix: a two by two table displaying the number of observed and predicted movements reported from the model (Table <xref ref-type="table" rid="T1">1</xref>). While there is no widely accepted rule-of-thumb for splitting the data, it is preferred to use a larger amount of information for the training set in order to reduce the variance of the parameter estimates (<xref ref-type="bibr" rid="B40">40</xref>). Also, we insured that the training dataset contained the same proportion of class variables (yes and no) as the original RF-data by using a data partition function executed by the <italic>caret</italic> package (<xref ref-type="bibr" rid="B42">42</xref>) in R (<xref ref-type="bibr" rid="B31">31</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Confusion matrix for the class variable (i.e., animal movements&#x02009;&#x0003D;&#x02009;yes or no)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left" rowspan="2">Predicted</th>
<th valign="top" align="center" colspan="2">Observed<hr/></th>
<th valign="top" align="center" rowspan="2">Total</th>
</tr><tr>
<th valign="top" align="center">Yes</th>
<th valign="top" align="center">No</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Yes</td>
<td align="center" valign="top">a</td>
<td align="center" valign="top">b</td>
<td align="center" valign="top">a&#x02009;&#x0002B;&#x02009;b</td>
</tr>
<tr>
<td align="left" valign="top">No</td>
<td align="center" valign="top">c</td>
<td align="center" valign="top">d</td>
<td align="center" valign="top">c&#x02009;&#x0002B;&#x02009;d</td>
</tr>
<tr>
<td align="left" valign="top">Total</td>
<td align="center" valign="top">a&#x02009;&#x0002B;&#x02009;c</td>
<td align="center" valign="top">b&#x02009;&#x0002B;&#x02009;d</td>
<td align="center" valign="top"><italic>N</italic></td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>Cells indicate the number of true positives (a), false positives (b), true negatives (d), and false negatives (c)</italic>.</p></table-wrap-foot></table-wrap>
<p>On the other hand, because we anticipated that RF-data would be unbalanced (i.e., only a small fraction of observations were class variable &#x0201C;yes&#x0201D;), a <italic>post hoc</italic> down-sampling approach was implemented to balance the data, i.e., we used a sample that has roughly the same proportion of each outcome class. The down-sampling technique is an efficient way to improve predictions, particularly when using bootstrap samples, given that no information is lost during the process (<xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B43">43</xref>). We used a wrapper provided by the <italic>train</italic> function in the <italic>caret</italic> package (<xref ref-type="bibr" rid="B42">42</xref>) in R (<xref ref-type="bibr" rid="B31">31</xref>) to improve model consistency and to determine the desired standard resampling and performance testing (<xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B44">44</xref>). We ran and tuned the RF model using 1,500 trees for each training dataset (unbalanced and balanced), and we implemented 10-fold cross-validations to estimate and rank the most important predictors.</p>
<p>Subsequently, we compared performance comparing predictive (or expected) versus observed movements for both, unbalanced and balanced testing datasets by using their confusion matrixes. As result, we compared the accuracy, Kappa statistic, specificity, sensitivity, and the area under the receiver-operating characteristic (ROC) curve. The ROC curve is a graphical method to test predictive performance by contrasting true positive and negative values. The accuracy rate <inline-formula><mml:math id="M2"><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>AR</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mi>d</mml:mi></mml:mrow><mml:mi>N</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> was used to measure the agreement between the predicted and the observed classes, although AR does not provide any information on the type of error the model is producing. The Kappa statistic (&#x003BA;) was used to quantify the relation between observed <inline-formula><mml:math id="M3"><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>O</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mi>d</mml:mi></mml:mrow><mml:mi>N</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> and expected accuracy <inline-formula><mml:math id="M4"><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>*</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>+</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>*</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>+</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>, so that <inline-formula><mml:math id="M5"><mml:mrow><mml:mo>&#x003BA;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>O</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>E</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:math></inline-formula> serves as a proxy for model performance (<xref ref-type="bibr" rid="B40">40</xref>). Sensitivity <inline-formula><mml:math id="M6"><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>Se</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mi>a</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> and specificity <inline-formula><mml:math id="M7"><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>Sp</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mi>d</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>+</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> were used to measure the capability of the model to predict true movements (i.e., &#x0201C;yes&#x0201D;) and non-movements (i.e., &#x0201C;no&#x0201D;), respectively, whereas the area under the ROC (AUC) was used to assess the trade-off between increasing sensitivity and decreasing specificity or vice versa. With RF, the final prediction as to whether movement occurs between a pair of sites (i.e., class variable&#x02009;&#x0003D;&#x02009;yes) is based on a given probability threshold (i.e., 0.5). We tested varying threshold probabilities (i.e., &#x02265;0.5) to maximize the &#x003BA; value.</p>
<p>Our RF model utilized a data set of 14,307 observations (75% of all observations) and 28 variables, thus complexity of the algorithm is given by <italic>O</italic>(<italic>v</italic>&#x02009;&#x000D7;&#x02009;<italic>n</italic>log(<italic>n</italic>)), where <italic>v</italic> is the number of variables and <italic>n</italic> is the number of observations. This analysis took around 45&#x02009;min to complete on a standard MacBook Pro<sup>&#x000AE;</sup>, though other packages such as ranger and random jungle may achieve faster performances for larger data sets and down-sampling can further optimize run times (<xref ref-type="bibr" rid="B45">45</xref>).</p>
<p>Finally, using the observed and predicted animal movements, we conducted an SNA to contrast centrality measurements between the observed and predicted network using the non-parametric Kruskal&#x02013;Wallis test.</p>
</sec>
<sec id="S2-3-2">
<title>Prediction of the RCP-N212 Network</title>
<p>We used our final model to predict animal movements among sites in the RCP-N212. Using data from RCP-N212, we generated all possible pair combinations among sites located within that RCP area. Similar to the analyses performed with the RF-data, we estimated Euclidean distances (kilometers) between each pair of sites and generated a dichotomous (yes, no) variable for ownership and 25 dummy variables for possible combinations of types of sites. Acknowledging that movements of animals must also occur to and from sites located out of the RCP-N212 and to avoid overestimations in the number of movements occurring within sites in the RCP, we restricted the number of predicted movements among sites in the RCP-N212 using the maximum values of <italic>in-</italic> and <italic>out-degree</italic> per each type of site observed in SN and RN.</p>
<p>We summarized the distributions of centrality measures of the predicted network for the entire RCP-N212 and for each of its 34 counties. We used the same metrics as described in Section &#x0201C;<xref ref-type="sec" rid="S2-2">Network Description</xref>&#x0201D; at site and network levels. We performed a spatial analysis of centrality measures using a 2D-kernel density estimation. This allowed us to evaluate the intensity of pig movements (e.g., to, through, and from other sites) in a given unit of space by approximating its probability density function (<xref ref-type="bibr" rid="B34">34</xref>, <xref ref-type="bibr" rid="B46">46</xref>, <xref ref-type="bibr" rid="B47">47</xref>).</p>
</sec>
</sec>
</sec>
<sec id="S3">
<title>Results</title>
<sec id="S3-1">
<title>Network Description</title>
<p>The network building data set included 237 sites (220 farms and 17 market sites), of which 33 and 19% were located within Stevens County and Rice County, respectively. The remaining 48% of sites were not located within those counties, but involved animal movements to or from Stevens and Rice (Table <xref ref-type="table" rid="T2">2</xref>), some covering long distances (Figure <xref ref-type="fig" rid="F1">1</xref>). We identified 474 animal movements (286 movements in SN and 215 movements in RN), some of which connected the two networks (Table <xref ref-type="table" rid="T2">2</xref>). Only 14% of all site types located in Stevens or Rice had movements with sites located in the other county, and all these were movements from finishing farms to market sites.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p><bold>Number of sites by production type and network</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Network</th>
<th valign="top" align="center" colspan="2">BS</th>
<th valign="top" align="center" colspan="2">Fa</th>
<th valign="top" align="center" colspan="2">N</th>
<th valign="top" align="center" colspan="2">Fi</th>
<th valign="top" align="center" colspan="2">M</th>
<th valign="top" align="center" colspan="2">Total</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Rice network</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">(0)</td>
<td align="center" valign="top">21</td>
<td align="center" valign="top">(8)</td>
<td align="center" valign="top">18</td>
<td align="center" valign="top">(7)</td>
<td align="center" valign="top">52</td>
<td align="center" valign="top">(28)</td>
<td align="center" valign="top">6</td>
<td align="center" valign="top">(1)</td>
<td align="center" valign="top">97</td>
<td align="center" valign="top">(44)</td>
</tr>
<tr>
<td align="left" valign="top">Stevens network</td>
<td align="center" valign="top">3</td>
<td align="center" valign="top">(1)</td>
<td align="center" valign="top">43</td>
<td align="center" valign="top">(20)</td>
<td align="center" valign="top">12</td>
<td align="center" valign="top">(7)</td>
<td align="center" valign="top">43</td>
<td align="center" valign="top">(23)</td>
<td align="center" valign="top">6</td>
<td align="center" valign="top">(1)</td>
<td align="center" valign="top">107</td>
<td align="center" valign="top">(52)</td>
</tr>
<tr>
<td align="left" valign="top">Both</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">(0)</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">(0)</td>
<td align="center" valign="top">0</td>
<td align="center" valign="top">(0)</td>
<td align="center" valign="top">28</td>
<td align="center" valign="top">(27)</td>
<td align="center" valign="top">5</td>
<td align="center" valign="top">(2)</td>
<td align="center" valign="top">33</td>
<td align="center" valign="top">(29)</td>
</tr>
<tr>
<td align="left" valign="top">Total</td>
<td align="center" valign="top">3</td>
<td align="center" valign="top">(1)</td>
<td align="center" valign="top">64</td>
<td align="center" valign="top">(28)</td>
<td align="center" valign="top">30</td>
<td align="center" valign="top">(14)</td>
<td align="center" valign="top">123</td>
<td align="center" valign="top">(78)</td>
<td align="center" valign="top">17</td>
<td align="center" valign="top">(4)</td>
<td align="center" valign="top">237</td>
<td align="center" valign="top">(125)</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>Values in parentheses indicate number of sites inside the noted county</italic>.</p>
<p><italic>BS, boar stud; Fa, farrowing; N, nursery; Fi, finishing; M, market sites</italic>.</p></table-wrap-foot></table-wrap>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Geographical representation of the Stevens network and Rice network of animal movements between swine farms and/or market sites</bold>. Dots represent geographical location of sites and straight gray lines represent animal movements. [Source: Wayne (<xref ref-type="bibr" rid="B27">27</xref>)].</p></caption>
<graphic xlink:href="fvets-04-00002-g001.tif"/>
</fig>
<p>Graphical representations of the networks indicate a confluence of paths toward finishing farms and then to market sites (Figure <xref ref-type="fig" rid="F2">2</xref>). Thus, markets and finishing farms served as hubs in the network. As expected, the most likely movements occurred between sites of different types (<italic>r</italic>&#x02009;&#x0003D;&#x02009;&#x02212;0.13 and <italic>r</italic>&#x02009;&#x0003D;&#x02009;&#x02212;0.16 for SN and RN, respectively) that followed downstream flows, i.e., a vertical structure (Table <xref ref-type="table" rid="T3">3</xref>). For example, movements from farrowing (Fa) or nursery (N) farms to finishing farms (Fi) were more frequent compared to other possible types of destinations, i.e., market sites (M), boar studs (BS), farrowing (Fa), or nursery (N) farms (Table <xref ref-type="table" rid="T3">3</xref>). Markets were the most likely destinations for finishing farms (<italic>e</italic><sub>FiM</sub> &#x0003D; 0.40 and <italic>e</italic><sub>FiM</sub> &#x0003D; 0.41 for SN and RN, respectively), although finishers also sent pigs into upstream destinations, including nurseries and farrowing farms (e.g., N, Fa, etc.), probably to provide replacement animals (Table <xref ref-type="table" rid="T3">3</xref>). The most likely destination for farrowing farms was finishers, followed by nurseries, consistent with the industry trend to eliminate nurseries as midpoint stations (<xref ref-type="bibr" rid="B20">20</xref>) (Figure <xref ref-type="fig" rid="F2">2</xref>; Table <xref ref-type="table" rid="T3">3</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Graphical representation of Rice (A) and Stevens networks (B)</bold>. Circles and squares represent either swine farms (BS, boar stud; Fa, farrowing; N, nursery; and Fi, finishing) or market sites (M).</p></caption>
<graphic xlink:href="fvets-04-00002-g002.tif"/>
</fig>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p><bold>Mixing matrix <italic>e<sub>ij</sub></italic> for type of farm in Rice network (RN) and Stevens network (SN)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left" colspan="2">Network<hr/></th>
<th valign="top" align="center" colspan="6">SN<hr/></th>
<th valign="top" align="center" colspan="6">RN<hr/></th>
</tr><tr>
<th valign="top" align="left" colspan="2">Destination</th>
<th valign="top" align="center">BS</th>
<th valign="top" align="center">Fa</th>
<th valign="top" align="center">Fi</th>
<th valign="top" align="center">M</th>
<th valign="top" align="center">N</th>
<th valign="top" align="center"><italic>a<sub>i</sub></italic></th>
<th valign="top" align="center">BS</th>
<th valign="top" align="center">Fa</th>
<th valign="top" align="center">Fi</th>
<th valign="top" align="center">M</th>
<th valign="top" align="center">N</th>
<th valign="top" align="center"><italic>a<sub>i</sub></italic></th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" rowspan="5">Origin</td>
<td align="left" valign="top">BS</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">&#x02013;</td>
</tr>
<tr>
<td align="left" valign="top">Fa</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">0.09</td>
<td align="center" valign="top">0.10</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.26</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">0.06</td>
<td align="center" valign="top">0.12</td>
<td align="center" valign="top">0.05</td>
<td align="center" valign="top">0.05</td>
<td align="center" valign="top">0.27</td>
</tr>
<tr>
<td align="left" valign="top">Fi</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.10</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.40</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.58</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.41</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.45</td>
</tr>
<tr>
<td align="left" valign="top">M</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.02</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.02</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.02</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.02</td>
</tr>
<tr>
<td align="left" valign="top">N</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.13</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.13</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.24</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.26</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2"><italic>b<sub>i</sub></italic></td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.14</td>
<td align="center" valign="top">0.29</td>
<td align="center" valign="top">0.53</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">1.00</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">&#x02013;</td>
<td align="center" valign="top">0.36</td>
<td align="center" valign="top">0.50</td>
<td align="center" valign="top">0.05</td>
<td align="center" valign="top">1.00</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>BS, boar stud; Fa, farrowing; N, nursery; Fi, finishing; M, market sites</italic>.</p>
<p><italic>r<sub>SN</sub>&#x02009;&#x0003D;&#x02009;&#x02212;0.13 and r<sub>RN</sub>&#x02009;&#x0003D;&#x02009;&#x02212;0.16</italic>.</p></table-wrap-foot></table-wrap>
<p>Whereas <italic>in-degree</italic> and <italic>betweenness</italic> were slightly higher in SN than RN (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.05 and <italic>P</italic>&#x02009;&#x0003D;&#x02009;0.04, respectively), there was no statistical difference between the two networks in out-degree (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.31). In contrast, <italic>in-degree</italic> varied across production types for both networks (<italic>P</italic>&#x02009;&#x0003C;&#x02009;0.01 for both), with markets having a higher <italic>in-degree</italic> (mean&#x02009;&#x0003D;&#x02009;11.7, SD&#x02009;&#x0003D;&#x02009;16.8, min&#x02009;&#x0003D;&#x02009;0 and max&#x02009;&#x0003D;&#x02009;57) (Figure <xref ref-type="fig" rid="F3">3</xref>). Nurseries exhibited significantly higher out-degree than other production types (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.01 for SN and <italic>P</italic>&#x02009;&#x0003C;&#x02009;0.01 for RN), each shipping animals to three different sites on average, with a maximum of 12. <italic>Betweenness</italic> did not significantly differ across production types within RN (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.17) but was statistically different among different types of sites in SN (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.02) (Table <xref ref-type="table" rid="T4">4</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Boxplot of betweenness, in- and out-degree in Rice network and Stevens network by production type (BS, boar stud; Fa, farrowing; N, nursery; Fi, finishing; M, market sites)</bold>. Boxes indicate the first and third percentile; middle bars represent the median and diamonds represent the mean.</p></caption>
<graphic xlink:href="fvets-04-00002-g003.tif"/>
</fig>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p><bold>Summary of centrality measures at site-level using both Stevens network and Rice network</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Centrality measure</th>
<th valign="top" align="center">BS</th>
<th valign="top" align="center">Fa</th>
<th valign="top" align="center">Fi</th>
<th valign="top" align="center">M</th>
<th valign="top" align="center">N</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" rowspan="3">Betweenness</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">4.20</td>
<td align="center" valign="top">2.07</td>
<td align="center" valign="top">11.56</td>
<td align="center" valign="top">5.61</td>
</tr>
<tr>
<td align="center" valign="top">(0)</td>
<td align="center" valign="top">(1.43)</td>
<td align="center" valign="top">(0.51)</td>
<td align="center" valign="top">(8.67)</td>
<td align="center" valign="top">(1.66)</td>
</tr>
<tr>
<td align="center" valign="top">0<sup>a</sup></td>
<td align="center" valign="top">68<sup>a</sup></td>
<td align="center" valign="top">46.21<sup>a</sup></td>
<td align="center" valign="top">185.14<sup>a</sup></td>
<td align="center" valign="top">31<sup>a</sup></td>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">In-degree</td>
<td align="center" valign="top">0.67</td>
<td align="center" valign="top">0.92</td>
<td align="center" valign="top">1.05</td>
<td align="center" valign="top">11.73</td>
<td align="center" valign="top">0.77</td>
</tr>
<tr>
<td align="center" valign="top">(0.67)</td>
<td align="center" valign="top">(0.14)</td>
<td align="center" valign="top">(0.07)</td>
<td align="center" valign="top">(3.59)</td>
<td align="center" valign="top">(0.1)</td>
</tr>
<tr>
<td align="center" valign="top">2<sup>a</sup></td>
<td align="center" valign="top">5<sup>a</sup></td>
<td align="center" valign="top">5<sup>a</sup></td>
<td align="center" valign="top">57<sup>a</sup></td>
<td align="center" valign="top">2<sup>a</sup></td>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">Out-degree</td>
<td align="center" valign="top">1.00</td>
<td align="center" valign="top">2.08</td>
<td align="center" valign="top">1.74</td>
<td align="center" valign="top">0.46</td>
<td align="center" valign="top">3.07</td>
</tr>
<tr>
<td align="center" valign="top">(0)</td>
<td align="center" valign="top">(0.26)</td>
<td align="center" valign="top">(0.15)</td>
<td align="center" valign="top">(0.18)</td>
<td align="center" valign="top">(0.62)</td>
</tr>
<tr>
<td align="center" valign="top">1<sup>a</sup></td>
<td align="center" valign="top">8<sup>a</sup></td>
<td align="center" valign="top">12<sup>a</sup></td>
<td align="center" valign="top">3<sup>a</sup></td>
<td align="center" valign="top">12<sup>a</sup></td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>Values denote means, SEs in parenthesis, and maximums with superscript &#x0201C;a.&#x0201D;</italic></p>
<p><italic>BS, boar stud; Fa, farrowing; N, nursery; Fi, finishing; M, market sites</italic>.</p></table-wrap-foot></table-wrap>
<p>Both networks exhibited similar densities, 0.014 for SN and 0.013 for RN. However, RN was relatively more cohesive than SN, as shown by a higher clustering coefficient, revealing a 0.007 and a 0.056 probability, respectively, that two sites moving animals to a common site were also connected to each other. As a result, RN also had a smaller diameter (4) than SN (5), though mean path lengths were relatively similar (1.82 for RN and 1.85 for SN). In turn, the distances between sites varied considerably, from less than 1 to more than 1,000&#x02009;km. The overall mean distance between sites was 111&#x02009;km, with nurseries and farrowing farms receiving animals from longer distances and boar studs shipping animals to sites located more than 500&#x02009;km away (Table <xref ref-type="table" rid="T5">5</xref>).</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p><bold>Summary of distances (km) between origin and destination by type of site</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left" colspan="2">Type destination</th>
<th valign="top" align="center">BS</th>
<th valign="top" align="center">Fa</th>
<th valign="top" align="center">Fi</th>
<th valign="top" align="center">M</th>
<th valign="top" align="center">N</th>
<th valign="top" align="center">Mean</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" rowspan="15">Type origin</td>
<td align="left" valign="top" rowspan="3">BS</td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top">1,894.81</td>
<td align="center" valign="top">20.36</td>
<td align="center" valign="top"/>
<td align="center" valign="top" rowspan="3">645.18</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top">(&#x02013;)</td>
<td align="center" valign="top">(13.68)</td>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top">1,894.81<sup>a</sup></td>
<td align="center" valign="top">34.04<sup>a</sup></td>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">Fa</td>
<td align="center" valign="top"/>
<td align="center" valign="top">113.71</td>
<td align="center" valign="top">122.50</td>
<td align="center" valign="top">32.95</td>
<td align="center" valign="top">170.83</td>
<td align="center" valign="top" rowspan="3">102.98</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top">(46.72)</td>
<td align="center" valign="top">(38.56)</td>
<td align="center" valign="top">(7.78)</td>
<td align="center" valign="top">(61.61)</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top">700.64<sup>a</sup></td>
<td align="center" valign="top">1,582.98<sup>a</sup></td>
<td align="center" valign="top">214.84<sup>a</sup></td>
<td align="center" valign="top">1,066.67<sup>a</sup></td>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">Fi</td>
<td align="center" valign="top">9.94</td>
<td align="center" valign="top">181.90</td>
<td align="center" valign="top">98.16</td>
<td align="center" valign="top">122.97</td>
<td align="center" valign="top">21.51</td>
<td align="center" valign="top" rowspan="3">128.49</td>
</tr>
<tr>
<td align="center" valign="top">7.65</td>
<td align="center" valign="top">(43.15)</td>
<td align="center" valign="top">(22.28)</td>
<td align="center" valign="top">(8.75)</td>
<td align="center" valign="top">(&#x02013;)</td>
</tr>
<tr>
<td align="center" valign="top">17.58<sup>a</sup></td>
<td align="center" valign="top">1285.83<sup>a</sup></td>
<td align="center" valign="top">271.21<sup>a</sup></td>
<td align="center" valign="top">354.65<sup>a</sup></td>
<td align="center" valign="top">21.51<sup>a</sup></td>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">M</td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top">267.49</td>
<td align="center" valign="top"/>
<td align="center" valign="top" rowspan="3">267.49</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top">(49.57)</td>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top">432.79<sup>a</sup></td>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">N</td>
<td align="center" valign="top"/>
<td align="center" valign="top">126.73</td>
<td align="center" valign="top">48.98</td>
<td align="center" valign="top">42.12</td>
<td align="center" valign="top">22.08</td>
<td align="center" valign="top" rowspan="3">50.21</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top">(101.77)</td>
<td align="center" valign="top">(7.73)</td>
<td align="center" valign="top">(30.68)</td>
<td align="center" valign="top">(&#x02013;)</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="center" valign="top">228.50<sup>a</sup></td>
<td align="center" valign="top">270.83<sup>a</sup></td>
<td align="center" valign="top">72.81<sup>a</sup></td>
<td align="center" valign="top">22.08<sup>a</sup></td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">Mean</td>
<td align="center" valign="top">9.94</td>
<td align="center" valign="top">155.76</td>
<td align="center" valign="top">90.30</td>
<td align="center" valign="top">109.95</td>
<td align="center" valign="top">158.91</td>
<td align="center" valign="top">111.14</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>Values denote means, SEs in parenthesis, and maximums with superscript &#x0201C;a.&#x0201D;</italic></p>
<p><italic>SEs could not be calculated in all cases due to small sample size</italic>.</p>
<p><italic>BS, boar stud; Fa, farrowing; N, nursery; Fi, finishing; M, market sites</italic>.</p></table-wrap-foot></table-wrap>
</sec>
<sec id="S3-2">
<title>Network Prediction</title>
<sec id="S3-2-1">
<title>Random Forest</title>
<p>There were 19,075 possible pairs for RN and SN. Among them, a minority (2.6%) corresponded to true movements (i.e., class&#x02009;&#x0003D;&#x02009;yes). RF models for both balanced (similar proportion of class variable &#x0201C;yes&#x0201D; and &#x0201C;no&#x0201D;) and unbalanced datasets used 1,500 trees, and the optimal number of predictors (<italic>m</italic><sub>try</sub>) estimated was 27 and 20, respectively. We observed a higher &#x003BA; for the unbalanced dataset, indicating a higher accuracy (Table <xref ref-type="table" rid="T6">6</xref>). However, use of the unbalanced datasets resulted in predictions that were strongly biased toward the majority class, with the class variable &#x0201C;no&#x0201D; accounting for 97.4% of total pairs.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p><bold>Results of the random forest (RF) analyses for balanced and unbalanced datasets</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Model</th>
<th valign="top" align="center">Accuracy</th>
<th valign="top" align="center">Kappa</th>
<th valign="top" align="center">Sensitivity</th>
<th valign="top" align="center">Specificity</th>
<th valign="top" align="center">Area under receiver-operating characteristic</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Balanced RF model</td>
<td align="center" valign="top">0.88</td>
<td align="center" valign="top">0.23</td>
<td align="center" valign="top">0.808</td>
<td align="center" valign="top">0.883</td>
<td align="center" valign="top">0.93</td>
</tr>
<tr>
<td align="left" valign="top">Unbalanced RF</td>
<td align="center" valign="top">0.98</td>
<td align="center" valign="top">0.34</td>
<td align="center" valign="top">0.232</td>
<td align="center" valign="top">0.997</td>
<td align="center" valign="top">0.85</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The balanced dataset optimized sensitivity with a low penalty to specificity. Moreover, inspection of the AUC indicated that false positives and negatives were minimized with the balanced dataset (Figure <xref ref-type="fig" rid="F4">4</xref>). However, the 0.5 default probability threshold used by the RF to predict animal movement between a pair of sites (i.e., class variable&#x02009;&#x0003D;&#x02009;yes) resulted in low agreement (i.e., when &#x003BA;&#x02009;&#x0003C;&#x02009;0.3) between observed versus predicted movements (<xref ref-type="bibr" rid="B40">40</xref>) (Table <xref ref-type="table" rid="T6">6</xref>). Increasing the threshold from 0.5 to 0.85 resulted in an increase in agreement (&#x003BA;&#x02009;&#x0003D;&#x02009;0.5, Figure <xref ref-type="fig" rid="F5">5</xref>) between observed and predicted movements. The most important variables predicting movements were farm type (downstream combinations from finishers and farrowing farms to market sites), sharing the same owner, and distance (Figure <xref ref-type="fig" rid="F6">6</xref>).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Receiver-operating characteristic curves for the RF analyses using balanced and imbalanced datasets</bold>.</p></caption>
<graphic xlink:href="fvets-04-00002-g004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>Kappa statistic accompanying the threshold probability for the proportion of predicting movements out of the total pairs using balanced data</bold>. Sensitivity (Se) and specificity (Sp) are also reported through different thresholds.</p></caption>
<graphic xlink:href="fvets-04-00002-g005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Rank of the 15 most important variables for network prediction using the balanced dataset</bold>.</p></caption>
<graphic xlink:href="fvets-04-00002-g006.tif"/>
</fig>
<p>Comparing the observed (<italic>O</italic>) and predicted (or expected, <italic>E</italic>) networks based on observed and predicted animal movements from use of the <italic>testing dataset</italic>, model predictions provided a reasonable approximation of real movements (Figure <xref ref-type="fig" rid="F7">7</xref>). Overall, there were no statistical differences between both networks in <italic>betweenness</italic> (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.38) and <italic>in-degree</italic> (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.97), whereas values for the <italic>out-degree</italic> were significantly different (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.02). For the latter, we predicted that, on average, a site would deliver animals to 1.4 sites (SD&#x02009;&#x0003D;&#x02009;1.3), compared to 1 site (SD&#x02009;&#x0003D;&#x02009;1.0) observed in the real network.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Boxplot of betweenness, in- and out-degrees in the testing dataset for observed (<italic>O</italic>) and predicted (<italic>E</italic>) movements by type of site</bold>. Boxes indicate the first and third percentile; middle bar represents the median and diamonds represent the mean values.</p></caption>
<graphic xlink:href="fvets-04-00002-g007.tif"/>
</fig>
<p>Furthermore, the patterns of connectivity across farm types were qualitatively similar (Figure <xref ref-type="fig" rid="F7">7</xref>). Whereas comparisons across production types within observed and predicted networks did not show significant differences in <italic>betweenness</italic> (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.32 and <italic>P</italic>&#x02009;&#x0003D;&#x02009;0.35, respectively), <italic>in-</italic> and <italic>out-degree</italic> were statistically different across production types in both observed (<italic>P</italic>&#x02009;&#x0003C;&#x02009;0.01 for both) and predicted networks (<italic>P</italic>&#x02009;&#x0003C;&#x02009;0.01, and <italic>P</italic>&#x02009;&#x0003D;&#x02009;0.01, respectively). For example, market sites in the observed data received animals from 7.4 different sites, whereas the model predicted receptions from 8.4 different sites. Similarly, whereas the model predicted that a nursery would ship animals into 2.1 farms, observed values indicated 1.7 different farms (Figure <xref ref-type="fig" rid="F7">7</xref>). On the other hand, there were no significant differences when comparing the observed to predicted centrality metrics by production type (<italic>P</italic>&#x02009;&#x0003E;&#x02009;0.05) for all, except <italic>betweenness</italic> of market (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.02) and <italic>out-degree</italic> of finishing farms (<italic>P</italic>&#x02009;&#x0003D;&#x02009;0.002).</p>
<p>Additionally, average distances between observed and predicted movements did not significantly vary across types of sites (N, Fa, Fi, and M with <italic>P</italic>-values of 0.70, 0.05, 0.29, 0.94, respectively) (Figure <xref ref-type="fig" rid="F8">8</xref>). Among farms, we noticed that finishing farms shipped animals the longest distances (observed and predicted averages 128.4 and 130.7&#x02009;km, respectively), whereas markets on average received animals from 107.3&#x02009;km away versus a prediction of 115.7&#x02009;km.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>Distribution of distances (km) of shipments (out) and entries (in) for observed (<italic>O</italic>) and predicted (<italic>E</italic>) movements by type of location</bold>.</p></caption>
<graphic xlink:href="fvets-04-00002-g008.tif"/>
</fig>
</sec>
<sec id="S3-2-2">
<title>Prediction of the RCP-N212 Network</title>
<p>The RC-N212 dataset contained 830 sites, 65.1% of which specialized in the last stage of production (e.g., growing, finisher, wean-to-finish), and 32% characterized as farrowing farms or nurseries. Only 1.2% of total sites recorded in RCP-N212 were market sites, which were located in only 4 out of the 34 counties in which RCP-N212 sites were located. We generated 688,070 possible origin-destination pairs, and our model predicted that 0.9% of those pairs were likely to move animals between them, using a probability threshold &#x0003E;0.85. However, if the number of likely links for a given farm exceeded the maximum observed <italic>in-</italic> or <italic>out-degree</italic> for its production type (Table <xref ref-type="table" rid="T4">4</xref>), the number of contacts was restricted to the maximum degree by randomly selecting from the highly probable links. This process resulted in a network where 0.4% of the total pairs were likely to move pigs between them.</p>
<p>Unsurprisingly, market sites reached the maximum allowable <italic>in-degree</italic>, receiving pigs from 57 sites, whereas farrowing farms, nurseries, and finishers were expected to receive animals (perhaps replacements), on average, from 4, 1, and 2 sites, respectively. On the other hand, the model predicted that nurseries and farrowing farms would ship pigs (i.e., <italic>out-degree</italic>) to 12 and 10 farms, respectively (Figure <xref ref-type="fig" rid="F9">9</xref>A). <italic>Betweenness</italic> was highest in farrowing and nursery farms, followed by finishing farms (Figure <xref ref-type="fig" rid="F9">9</xref>A).</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p><bold>Boxplot of betweenness, in- and out-degrees distributions (A), and distance (km) distribution of shipments (out) and entries (in) (B) for the predicted network in RCP-N212</bold>.</p></caption>
<graphic xlink:href="fvets-04-00002-g009.tif"/>
</fig>
<p>The model using RCP-N212 data predicted animal movement distances that were slightly different from those predicted using the testing dataset presented in the previous section. Finishing farms were expected to ship (i.e., <italic>out-degree</italic>) animals through, on average, 141.4&#x02009;km (SD&#x02009;&#x0003D;&#x02009;75.5&#x02009;km), whereas nurseries, on average, 113.6&#x02009;km away (SD&#x02009;&#x0003D;&#x02009;67.7&#x02009;km) to sites within the RCP-N212. In turn, farrowing farms and market sites were expected to receive animals from longer distances (mean&#x02009;&#x0003D;&#x02009;168.4&#x02009;km, SD&#x02009;&#x0003D;&#x02009;51.0&#x02009;km, and mean&#x02009;&#x0003D;&#x02009;104.3&#x02009;km, SD&#x02009;&#x0003D;&#x02009;98.5&#x02009;km, respectively) (Figure <xref ref-type="fig" rid="F9">9</xref>). The density of the predicted network in RCP-N212 was 0.004, with a clustering coefficient of 2.8%, and a mean path length of 7.26.</p>
<p>Predicted pig movements in the RCP-N212 covered large spatial areas, and only 14% were within the same county. In general, most predicted pig movements passed through several counties, with a maximum of 11 counties. Finally, the predicted network for the RCP-N212 suggested a major aggregation of movements to and from sites located in areas toward the southern part of the regional program (Figure <xref ref-type="fig" rid="F10">10</xref>).</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p><bold>Plot of geographical density for expected out-degree (A), in-degree (B), and betweenness (C) of sites in the RCP-N212</bold>. Counties delineated compose the area of the RCP-N212.</p></caption>
<graphic xlink:href="fvets-04-00002-g010.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="S4" sec-type="discussion">
<title>Discussion</title>
<p>The aim of this research was to predict animal movements among sites located within a given RCP. Unfortunately, data of movement networks are often incomplete or unavailable for food animal industries characterized by a large number of animal movements between sites, such as the US swine industry (<xref ref-type="bibr" rid="B19">19</xref>, <xref ref-type="bibr" rid="B25">25</xref>). Therefore, we employed machine-learning techniques to illustrate how models may be fitted by using a subset of the data to increase their completeness and accuracy. Specifically, using information available in only two counties, we studied the likelihood of possible movements among sites in a larger-scale swine disease RCP in Minnesota, referred to as RCP-N212. In general, networks predicted by the RF model were consistent with the observed data used for model training and testing in terms of both spatial and production-type connectivity patterns.</p>
<p>The SN and RN networks exhibited relatively similar centrality patterns and a marked flow of animal movements from upstream to downstream sites in the production chain. However, the network structure of both counties also indicated that finishing (and even market) sites might provide animal replacement (e.g., gilts and boars) to upstream sites. Because outbreaks of diseases, such as PRRS, are also common in downstream sites (<xref ref-type="bibr" rid="B26">26</xref>), movements from those sites to upstream sites could perpetuate disease in those areas. Indeed, previous research has shown that despite an overall decrease in the occurrence of PRRS, spatial and temporal aggregations of that disease allowed for continued hotspots throughout the period of study (<xref ref-type="bibr" rid="B26">26</xref>). These interactions merit further analysis to explain swine disease dynamics, especially for industry-persistent diseases such as PRRS (<xref ref-type="bibr" rid="B2">2</xref>, <xref ref-type="bibr" rid="B6">6</xref>, <xref ref-type="bibr" rid="B48">48</xref>).</p>
<p>While model results suggest that ownership and distance are strongly related to the probability of pig movement between sites, the production type of the origin and destination sites also influenced the probability of pig movements from one location to another. Moreover, if we consider that different types of farms might share transportation services, whereby farms may ship or receive different type of animals (e.g., feeder pigs and finishing pigs), such mixing might facilitate spread disease <italic>via</italic> contaminated vehicles (<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B49">49</xref>). Thus, we suggest that a complete evaluation of disease risks associated with transportation of pigs between facilities should take into account factors such as the commercial relationship between sites, including contractual agreements in the US swine industry (<xref ref-type="bibr" rid="B50">50</xref>), and site production type (<xref ref-type="bibr" rid="B6">6</xref>).</p>
<p>As mentioned previously, we used information available from two small, county-based networks to estimate parameters for predictions of animal movements that closely fit observed movements as judged by standard statistical tests. We were able to validate our predictions within SN and RN. The parameters generated by our model were used to predict animal movements between sites over a larger area, i.e., RCP-N212. This is essentially an out-of-sample prediction. As data on actual animal movements were not available for RCP-N212, we cannot directly test the accuracy of our predictions for the larger network. The results for larger network appear reasonable in that they are consistent with the topology of the SN and RN networks, and these results may be useful in helping understand actual (but unobservable) animal movements in Minnesota. We believe that such out-of-sample prediction is warranted for the scale of RCP-N212, given that farms within this program are similar to the farms in SN and RN in terms of geography, demography, and management. However, predictions that would rely on more extensive extrapolation (such as at the scale of multiple states) would not be appropriate given the scope of our sampling. In addition, the value of being able to assemble a full network may not be to target individual farms, but rather to capture possible regional patterns in connectivity.</p>
<p>Based on comparisons between observed, county-level data and regional-level model predictions, such as the distribution of movement distances and general attributes of the network, we believe that our predicted full network for RCP-N212 has structural features that are similar to the partial data used to estimate the probability of movement between farms. Thus, our findings appear reasonable and provide insight to better understand animal movement patterns within the RCP-N212. This, in turn, may help farmers design private strategies for sanitary management, as well as aid policy-makers in structure-based decisions. However, given inherent limitations to predictive modeling, we acknowledge that our results may provide only general insights about movement patterns, thus additional work must be done before strong conclusions can be made regarding the utility of the predictions achieved in this way.</p>
<p>The RCP-N212 covers 34 out of 87 counties in Minnesota, accounting for 38% of the total swine facilities with 100 or more heads in the state. Because farm distribution is heterogeneous within Minnesota, with a greater number of farms toward the south (<xref ref-type="bibr" rid="B28">28</xref>), it is reasonable to infer that the distribution and type of sites across the RCP-N212 should influence our network predictions. Furthermore, given that 65.1% of the sites are dedicated to the last stage of production, we acknowledge that a fraction of sites within the RCP-N212 must trade animals with sites located in neighboring states, such as Iowa or Illinois, or even more distant states such as North Carolina, where a high number of sow farms are located (<xref ref-type="bibr" rid="B20">20</xref>, <xref ref-type="bibr" rid="B51">51</xref>, <xref ref-type="bibr" rid="B52">52</xref>). This could have led to an overestimation of movements between some sites in our model, especially for those underrepresented in the RCP. To tackle this issue, we applied two constraints to our model: (1) we restricted prediction of movements by increasing the probability threshold to 85% when assigning a possible link between two sites and (2) we constrained the maximum number of links per site using maximum values of <italic>in-</italic> and <italic>out-degree</italic>. As a result our movement predictions within the RCP-N212 were conservative, showing a lower density in the expected network (0.4%) than in the two observed networks (1.4% and 1.3% for SN and RN, respectively), although some metrics might have been overestimated, such as <italic>out-degree</italic> for nurseries. This is because there are very few nursery farms in this region, and many of the finishing farms are actually sourcing for pigs from other states outside of RCP-N212. However, our algorithm restricted their choice of nurseries to those within RCP-N212, perhaps leading to a false inflation of their <italic>out-degree</italic>. Future work should expand to larger geographic regions that encompass all stages of production, capturing movements between states. We anticipate that we could improve model predictions by obtaining information regarding contract relationships and animal movements among Minnesota sites and suppliers located outside RCP-N212.</p>
<p>In the predicted RCP-N212 network, we found that sites with higher <italic>in-</italic> and <italic>out-degrees</italic> overlap with areas where spatial and temporal aggregations of PRRS have occurred (<xref ref-type="bibr" rid="B26">26</xref>). Therefore, it is possible that movements, in addition to farm density, might play an important role in the persistent circulation of disease in the area. The co-aggregation of animal movements and PRRS, a disease believed to be transmitted, at least in part, by animal movements (<xref ref-type="bibr" rid="B2">2</xref>, <xref ref-type="bibr" rid="B6">6</xref>, <xref ref-type="bibr" rid="B49">49</xref>), suggests that our predicted network might be capturing important features of the underlying industry structure, which indirectly supports the validity of our network predictions.</p>
<p>The characterization of network structures often may help for planning production and designing strategies to control animal disease (<xref ref-type="bibr" rid="B7">7</xref>, <xref ref-type="bibr" rid="B8">8</xref>, <xref ref-type="bibr" rid="B11">11</xref>, <xref ref-type="bibr" rid="B18">18</xref>). The approach developed here is an early step for helping in design strategies to control swine diseases regionally. For example, the spread of swine pathogens within the full network can be simulated using computational models, which would be valuable for both predicting patterns of between-farm spread and for evaluating alternate intervention and control strategies. Among them, for instance, vaccination strategies that maximize the collective good could be quantitatively explored, including minimum levels of coverage that may prevent disease circulation in the network. Additionally, the approach developed here may reduce time and cost for data collection, as collection of movements among a partial set of sites might be sufficient to predict movements among a larger set of sites.</p>
<p>Among classification techniques, there are several approaches that might be used to predict possible outcomes, such as links between sites. While the focus of this paper is not to provide an exhaustive review of these techniques, here we offer some ground for further discussion and perhaps comparative studies. The RF approach has high accuracy without overfitting, it is also fairly stable to the presence of outliers and noise, and it may handle the correlation between predictors (<xref ref-type="bibr" rid="B40">40</xref>, <xref ref-type="bibr" rid="B41">41</xref>, <xref ref-type="bibr" rid="B53">53</xref>). This may be important in the context of this study, as some atypical movements between sites may occur, predictors may be correlated, and the probability of animal movement between two or more sites may often occur in a non-linear fashion. Alternatively, other supervised techniques might be used. For example, support vector machines, a vector function based technique that splits the data for classification purposes, might resolve non-linearity in the data by using a non-linear kernel function, though its performance sometimes might be compromised (<xref ref-type="bibr" rid="B40">40</xref>).</p>
<p>In conclusion, we present an approach to predict the network structure of contacts between and among farms in a region by using partial data. Our results, combined with information on the occurrence of disease in the area (i.e., outbreaks of PRRS within the RCP-N212), may be incorporated into a disease transmission model that will help to evaluate the effectiveness of prevention and control strategies in a region, with the ultimate objective of mitigating the impact of endemic disease and hypothetical epidemic incursions. The approach here may also be applied to other regions and production systems, where information on animal movements is only partially regulated, thus improving decision-makers&#x02019; ability to plan and implement disease surveillance and control activities.</p>
</sec>
<sec id="S5" sec-type="author-contributor">
<title>Author Contributions</title>
<p>All the authors have met the four criteria described at the guidelines: PVD designed and data interpretation, revised and approved the version to be published, and agreed to be accountable for all aspects of the work. KV and LJ designed, revised and approved the version to be published, and agreed to be accountable for all aspects of the work. SW data interpretation, revised and approved the version to be published, and agreed to be accountable for all aspects of the work. AP designed, revised and approved the version to be published, and agreed to be accountable for all aspects of the work.</p>
</sec>
<sec id="S6">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>The Becas-Chile program of the National Commission for Scientific and Technological Research (CONICYT) has supported PV-D to develop his Master and PhD in the US. Additional support has been provided by the University of Minnesota MnDrive program and by the National Pork Board. Authors thank Jaclyn R. Aliperti, PhD(c), UCD, for providing valuable editorial advice. LS acknowledges support from the National Institute of Food and Agriculture (NIFA).</p>
</ack>
<sec id="S7">
<title>Funding</title>
<p>No funding was provided for the development of this article.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><label>1</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>F&#x000E8;vre</surname> <given-names>EM</given-names></name> <name><surname>Bronsvoort</surname> <given-names>BMDC</given-names></name> <name><surname>Hamilton</surname> <given-names>KA</given-names></name> <name><surname>Cleaveland</surname> <given-names>S</given-names></name></person-group>. <article-title>Animal movements and the spread of infectious diseases</article-title>. <source>Trends Microbiol</source> (<year>2006</year>) <volume>14</volume>:<fpage>125</fpage>&#x02013;<lpage>31</lpage>.<pub-id pub-id-type="doi">10.1016/j.tim.2006.01.004</pub-id></citation></ref>
<ref id="B2"><label>2</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Albina</surname> <given-names>E</given-names></name></person-group>. <article-title>Epidemiology of porcine reproductive and respiratory syndrome (PRRS): an overview</article-title>. <source>Vet Microbiol</source> (<year>1997</year>) <volume>55</volume>:<fpage>309</fpage>&#x02013;<lpage>16</lpage>.<pub-id pub-id-type="doi">10.1016/S0378-1135(96)01322-3</pub-id><pub-id pub-id-type="pmid">9220627</pub-id></citation></ref>
<ref id="B3"><label>3</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dee</surname> <given-names>S</given-names></name> <name><surname>Deen</surname> <given-names>J</given-names></name> <name><surname>Rossow</surname> <given-names>K</given-names></name> <name><surname>Weise</surname> <given-names>C</given-names></name> <name><surname>Eliason</surname> <given-names>R</given-names></name> <name><surname>Otake</surname> <given-names>S</given-names></name> <etal/></person-group> <article-title>Mechanical transmission of porcine reproductive and respiratory syndrome virus throughout a coordinated sequence of events during warm weather</article-title>. <source>Can J Vet Res</source> (<year>2003</year>) <volume>67</volume>:<fpage>12</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="pmid">12528824</pub-id></citation></ref>
<ref id="B4"><label>4</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dee</surname> <given-names>SA</given-names></name> <name><surname>Deen</surname> <given-names>J</given-names></name> <name><surname>Otake</surname> <given-names>S</given-names></name> <name><surname>Pijoan</surname> <given-names>C</given-names></name></person-group>. <article-title>An experimental model to evaluate the role of transport vehicles as a source of transmission of porcine reproductive and respiratory syndrome virus to susceptible pigs</article-title>. <source>Can J Vet Res</source> (<year>2004</year>) <volume>68</volume>:<fpage>128</fpage>&#x02013;<lpage>33</lpage>.</citation></ref>
<ref id="B5"><label>5</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dee</surname> <given-names>S</given-names></name> <name><surname>Clement</surname> <given-names>T</given-names></name> <name><surname>Schelkopf</surname> <given-names>A</given-names></name> <name><surname>Nerem</surname> <given-names>J</given-names></name> <name><surname>Knudsen</surname> <given-names>D</given-names></name> <name><surname>Christopher-Hennings</surname> <given-names>J</given-names></name> <etal/></person-group> <article-title>An evaluation of contaminated complete feed as a vehicle for porcine epidemic diarrhea virus infection of na&#x000EF;ve pigs following consumption via natural feeding behavior: proof of concept</article-title>. <source>BMC Vet Res</source> (<year>2014</year>) <volume>10</volume>:<fpage>1</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1186/s12917-014-0220-9</pub-id></citation></ref>
<ref id="B6"><label>6</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Perez</surname> <given-names>AM</given-names></name> <name><surname>Davies</surname> <given-names>PR</given-names></name> <name><surname>Goodell</surname> <given-names>CK</given-names></name> <name><surname>Holtkamp</surname> <given-names>DJ</given-names></name> <name><surname>Mondaca-Fern&#x000E1;ndez</surname> <given-names>E</given-names></name> <name><surname>Poljak</surname> <given-names>Z</given-names></name> <etal/></person-group> <article-title>Lessons learned and knowledge gaps about the epidemiology and control of porcine reproductive and respiratory syndrome virus in North America</article-title>. <source>J Am Vet Med Assoc</source> (<year>2015</year>) <volume>246</volume>:<fpage>1304</fpage>&#x02013;<lpage>17</lpage>.<pub-id pub-id-type="doi">10.2460/javma.246.12.1304</pub-id></citation></ref>
<ref id="B7"><label>7</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sydow</surname> <given-names>J</given-names></name> <name><surname>Windeler</surname> <given-names>A</given-names></name></person-group>. <article-title>Organizing and evaluating interfirm networks: a structurationist perspective on network processes and effectiveness</article-title>. <source>Organ Sci</source> (<year>1998</year>) <volume>9</volume>:<fpage>265</fpage>&#x02013;<lpage>84</lpage>.<pub-id pub-id-type="doi">10.1287/orsc.9.3.265</pub-id></citation></ref>
<ref id="B8"><label>8</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mart&#x000ED;nez-L&#x000F3;pez</surname> <given-names>B</given-names></name> <name><surname>Perez</surname> <given-names>AM</given-names></name> <name><surname>S&#x000E1;nchez-Vizca&#x000ED;no</surname> <given-names>JM</given-names></name></person-group>. <article-title>Social network analysis. Review of general concepts and use in preventive veterinary medicine</article-title>. <source>Transbound Emerg Dis</source> (<year>2009</year>) <volume>56</volume>:<fpage>109</fpage>&#x02013;<lpage>20</lpage>.<pub-id pub-id-type="doi">10.1111/j.1865-1682.2009.01073.x</pub-id><pub-id pub-id-type="pmid">19341388</pub-id></citation></ref>
<ref id="B9"><label>9</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Jackson</surname> <given-names>MO</given-names></name></person-group>. <source>Social and Economic Networks</source>. <publisher-loc>Princeton</publisher-loc>: <publisher-name>Princeton University Press</publisher-name> (<year>2008</year>).</citation></ref>
<ref id="B10"><label>10</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hagerman</surname> <given-names>AD</given-names></name> <name><surname>Mccarl</surname> <given-names>BA</given-names></name> <name><surname>Carpenter</surname> <given-names>TE</given-names></name> <name><surname>Ward</surname> <given-names>MP</given-names></name> <name><surname>O&#x02019;Brien</surname> <given-names>J</given-names></name></person-group>. <article-title>Emergency vaccination to control foot-and-mouth disease: implications of its inclusion as a U.S. policy option</article-title>. <source>Appl Econ Perspect Policy</source> (<year>2011</year>) <volume>34</volume>:<fpage>119</fpage>&#x02013;<lpage>46</lpage>.<pub-id pub-id-type="doi">10.1093/aepp/ppr039</pub-id></citation></ref>
<ref id="B11"><label>11</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>N&#x000F6;remark</surname> <given-names>M</given-names></name> <name><surname>H&#x000E5;kansson</surname> <given-names>N</given-names></name> <name><surname>Lewerin</surname> <given-names>SS</given-names></name> <name><surname>Lindberg</surname> <given-names>A</given-names></name> <name><surname>Jonsson</surname> <given-names>A</given-names></name></person-group>. <article-title>Network analysis of cattle and pig movements in Sweden: measures relevant for disease control and risk based surveillance</article-title>. <source>Prev Vet Med</source> (<year>2011</year>) <volume>99</volume>:<fpage>78</fpage>&#x02013;<lpage>90</lpage>.<pub-id pub-id-type="doi">10.1016/j.prevetmed.2010.12.009</pub-id><pub-id pub-id-type="pmid">21288583</pub-id></citation></ref>
<ref id="B12"><label>12</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rautureau</surname> <given-names>S</given-names></name> <name><surname>Dufour</surname> <given-names>B</given-names></name> <name><surname>Durand</surname> <given-names>B</given-names></name></person-group>. <article-title>Structural vulnerability of the French swine industry trade network to the spread of infectious diseases</article-title>. <source>Animal</source> (<year>2012</year>) <volume>6</volume>:<fpage>1152</fpage>&#x02013;<lpage>62</lpage>.<pub-id pub-id-type="doi">10.1017/s1751731111002631</pub-id></citation></ref>
<ref id="B13"><label>13</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>B&#x000FC;ttner</surname> <given-names>K</given-names></name> <name><surname>Krieter</surname> <given-names>J</given-names></name> <name><surname>Traulsen</surname> <given-names>A</given-names></name> <name><surname>Traulsen</surname> <given-names>I</given-names></name></person-group>. <article-title>Epidemic spreading in an animal trade network &#x02013; comparison of distance-based and network-based control measures</article-title>. <source>Transbound Emerg Dis</source> (<year>2016</year>) <volume>63</volume>:<fpage>e122</fpage>&#x02013;<lpage>34</lpage>.<pub-id pub-id-type="doi">10.1111/tbed.12245</pub-id><pub-id pub-id-type="pmid">25056832</pub-id></citation></ref>
<ref id="B14"><label>14</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Natale</surname> <given-names>F</given-names></name> <name><surname>Giovannini</surname> <given-names>A</given-names></name> <name><surname>Savini</surname> <given-names>L</given-names></name> <name><surname>Palma</surname> <given-names>D</given-names></name> <name><surname>Possenti</surname> <given-names>L</given-names></name> <name><surname>Fiore</surname> <given-names>G</given-names></name> <etal/></person-group> <article-title>Network analysis of Italian cattle trade patterns and evaluation of risks for potential disease spread</article-title>. <source>Prev Vet Med</source> (<year>2009</year>) <volume>92</volume>:<fpage>341</fpage>&#x02013;<lpage>50</lpage>.<pub-id pub-id-type="doi">10.1016/j.prevetmed.2009.08.026</pub-id><pub-id pub-id-type="pmid">19775765</pub-id></citation></ref>
<ref id="B15"><label>15</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bajardi</surname> <given-names>P</given-names></name> <name><surname>Barrat</surname> <given-names>A</given-names></name> <name><surname>Savini</surname> <given-names>L</given-names></name> <name><surname>Colizza</surname> <given-names>V</given-names></name></person-group>. <article-title>Optimizing surveillance for livestock disease spreading through animal movements</article-title>. <source>J R Soc Interface</source> (<year>2012</year>) <volume>9</volume>:<fpage>2814</fpage>&#x02013;<lpage>25</lpage>.<pub-id pub-id-type="doi">10.1098/rsif.2012.0289</pub-id><pub-id pub-id-type="pmid">22728387</pub-id></citation></ref>
<ref id="B16"><label>16</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>VanderWaal</surname> <given-names>KL</given-names></name> <name><surname>Picasso</surname> <given-names>C</given-names></name> <name><surname>Enns</surname> <given-names>EA</given-names></name> <name><surname>Craft</surname> <given-names>ME</given-names></name> <name><surname>Alvarez</surname> <given-names>J</given-names></name> <name><surname>Fernandez</surname> <given-names>F</given-names></name> <etal/></person-group> <article-title>Network analysis of cattle movements in Uruguay: quantifying heterogeneity for risk-based disease surveillance and control</article-title>. <source>Prev Vet Med</source> (<year>2016</year>) <volume>123</volume>:<fpage>12</fpage>&#x02013;<lpage>22</lpage>.<pub-id pub-id-type="doi">10.1016/j.prevetmed.2015.12.003</pub-id><pub-id pub-id-type="pmid">26708252</pub-id></citation></ref>
<ref id="B17"><label>17</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mardones</surname> <given-names>FO</given-names></name> <name><surname>Martinez-Lopez</surname> <given-names>B</given-names></name> <name><surname>Valdes-Donoso</surname> <given-names>P</given-names></name> <name><surname>Carpenter</surname> <given-names>TE</given-names></name> <name><surname>Perez</surname> <given-names>AM</given-names></name></person-group>. <article-title>The role of fish movements and the spread of infectious salmon anemia virus (ISAV) in Chile, 2007&#x02013;2009</article-title>. <source>Prev Vet Med</source> (<year>2014</year>) <volume>114</volume>:<fpage>37</fpage>&#x02013;<lpage>46</lpage>.<pub-id pub-id-type="doi">10.1016/j.prevetmed.2014.01.012</pub-id></citation></ref>
<ref id="B18"><label>18</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lentz</surname> <given-names>HHK</given-names></name> <name><surname>Koher</surname> <given-names>A</given-names></name> <name><surname>H&#x000F6;vel</surname> <given-names>P</given-names></name> <name><surname>Gethmann</surname> <given-names>J</given-names></name> <name><surname>Sauter-Louis</surname> <given-names>C</given-names></name> <name><surname>Selhorst</surname> <given-names>T</given-names></name> <etal/></person-group> <article-title>Disease spread through animal movements: a static and temporal network analysis of pig trade in Germany</article-title>. <source>PLoS One</source> (<year>2016</year>) <volume>11</volume>:<fpage>e0155196</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pone.0155196</pub-id><pub-id pub-id-type="pmid">27152712</pub-id></citation></ref>
<ref id="B19"><label>19</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Shields</surname> <given-names>DA</given-names></name> <name><surname>Mathews</surname> <given-names>K</given-names></name></person-group>. <source>Interstate Livestock Movements</source>. <publisher-name>Economic Research Service</publisher-name> (<year>2003</year>). Available from: <uri xlink:href="http://www.ers.usda.gov">http://www.ers.usda.gov</uri></citation></ref>
<ref id="B20"><label>20</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>McBride</surname> <given-names>WD</given-names></name> <name><surname>Key</surname> <given-names>N</given-names></name></person-group>. <source>U.S. Hog Production from 1992 to 2009: Technology, Restructuring, and Productivity Growth</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>US Department of Agriculture, Economic Research Service</publisher-name> (<year>2013</year>).</citation></ref>
<ref id="B21"><label>21</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hurt</surname> <given-names>C</given-names></name></person-group>. <article-title>Industrialization in the pork industry</article-title>. <source>Choices</source> (<year>1994</year>) <volume>9</volume>:<fpage>9</fpage>&#x02013;<lpage>13</lpage>.</citation></ref>
<ref id="B22"><label>22</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Key</surname> <given-names>N</given-names></name> <name><surname>McBride</surname> <given-names>W</given-names></name></person-group>. <source>The Changing Economics of U.S. Hog Production, ERR-52</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>United States Department of Agriculture, Economic Research Service</publisher-name> (<year>2007</year>).</citation></ref>
<ref id="B23"><label>23</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tilman</surname> <given-names>D</given-names></name> <name><surname>Cassman</surname> <given-names>KG</given-names></name> <name><surname>Matson</surname> <given-names>PA</given-names></name> <name><surname>Naylor</surname> <given-names>R</given-names></name> <name><surname>Polasky</surname> <given-names>S</given-names></name></person-group>. <article-title>Agricultural sustainability and intensive production practices</article-title>. <source>Nature</source> (<year>2002</year>) <volume>418</volume>:<fpage>671</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="doi">10.1038/nature01014</pub-id><pub-id pub-id-type="pmid">12167873</pub-id></citation></ref>
<ref id="B24"><label>24</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>MacDonald</surname> <given-names>JM</given-names></name> <name><surname>McBride</surname> <given-names>WD</given-names></name></person-group>. <source>The Transformation of US Livestock Agriculture: Scale, Efficiency, and Risks</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>US Department of Agriculture, Economic Research Service</publisher-name> (<year>2009</year>).</citation></ref>
<ref id="B25"><label>25</label><citation citation-type="book"><collab>USDA</collab>. <source>Animal Disease Traceabililty</source>. <publisher-loc>Washington DC</publisher-loc>: (<year>2016</year>). Available: <uri xlink:href="https://www.aphis.usda.gov/aphis/ourfocus/animalhealth/SA_Traceability">https://www.aphis.usda.gov/aphis/ourfocus/animalhealth/SA_Traceability</uri></citation></ref>
<ref id="B26"><label>26</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Valdes-Donoso</surname> <given-names>P</given-names></name> <name><surname>Jarvis</surname> <given-names>LS</given-names></name> <name><surname>Wright</surname> <given-names>D</given-names></name> <name><surname>Alvarez</surname> <given-names>J</given-names></name> <name><surname>Perez</surname> <given-names>AM</given-names></name></person-group>. <article-title>Measuring progress on the control of porcine reproductive and respiratory syndrome (PRRS) at a regional level: the Minnesota N212 regional control project (Rcp) as a working example</article-title>. <source>PLoS One</source> (<year>2016</year>) <volume>11</volume>:<fpage>e0149498</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pone.0149498</pub-id><pub-id pub-id-type="pmid">26895148</pub-id></citation></ref>
<ref id="B27"><label>27</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Wayne</surname> <given-names>SR</given-names></name></person-group>. <source>Assessment of Demographics and Network Structure of Swine Populations in Relation to Regional Disease Transmission and Control</source>. <publisher-loc>St Paul, MN</publisher-loc>: <publisher-name>University of Minnesota</publisher-name> (<year>2011</year>).</citation></ref>
<ref id="B28"><label>28</label><citation citation-type="book"><collab>USDA</collab>. <source>Census of Agriculture 2012. Minnesota: State and County Data</source>. <publisher-name>N.a.S. Service</publisher-name> (<year>2014</year>). Available from: <uri xlink:href="http://www.agcensus.usda.gov/Publications/2012/Full_Report/Census_by_State/Minnesota/index.asp">http://www.agcensus.usda.gov/Publications/2012/Full_Report/Census_by_State/Minnesota/index.asp</uri></citation></ref>
<ref id="B29"><label>29</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Newman</surname> <given-names>MEJ</given-names></name></person-group>. <article-title>Mixing patterns in networks</article-title>. <source>Phys Rev E</source> (<year>2003</year>) <volume>67</volume>:<fpage>1</fpage>&#x02013;<lpage>13</lpage>.<pub-id pub-id-type="doi">10.1103/PhysRevE.67.026126</pub-id></citation></ref>
<ref id="B30"><label>30</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Wasserman</surname> <given-names>S</given-names></name> <name><surname>Faust</surname> <given-names>K</given-names></name></person-group>. <source>Social Network Analysis: Methods and applications</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name> (<year>1994</year>).</citation></ref>
<ref id="B31"><label>31</label><citation citation-type="book"><collab>R Development Core Team</collab>. <source>R: A Language and Environment for Statistical Computing</source>. <publisher-loc>Vienna, Austria</publisher-loc>: <publisher-name>R.F.F.S. Computing</publisher-name> (<year>2015</year>).</citation></ref>
<ref id="B32"><label>32</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Wickham</surname> <given-names>H</given-names></name></person-group>. <source>ggplot2: Elegant Graphics for Data Analysis</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2009</year>).</citation></ref>
<ref id="B33"><label>33</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Becker</surname> <given-names>RA</given-names></name> <name><surname>Wilks</surname> <given-names>AR</given-names></name></person-group>. <source>maps: Draw Geographical Maps</source>. (<year>2014</year>).</citation></ref>
<ref id="B34"><label>34</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Venables</surname> <given-names>WN</given-names></name> <name><surname>Ripley</surname> <given-names>BD</given-names></name></person-group>. <source>Modern Applied Statistics with S</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2002</year>).</citation></ref>
<ref id="B35"><label>35</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Csardi</surname> <given-names>G</given-names></name> <name><surname>Nepusz</surname> <given-names>T</given-names></name></person-group>. <article-title>The igraph software package for complex network research</article-title>. <source>InterJournal Complex Syst</source> (<year>2006</year>).</citation></ref>
<ref id="B36"><label>36</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liben-Nowell</surname> <given-names>D</given-names></name> <name><surname>Kleinberg</surname> <given-names>J</given-names></name></person-group>. <article-title>The link-prediction problem for social networks</article-title>. <source>J Am Soc Inf Sci Technol</source> (<year>2007</year>) <volume>58</volume>:<fpage>1019</fpage>&#x02013;<lpage>31</lpage>.<pub-id pub-id-type="doi">10.1002/asi.20591</pub-id></citation></ref>
<ref id="B37"><label>37</label><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Lichtenwalter</surname> <given-names>RN</given-names></name> <name><surname>Lussier</surname> <given-names>JT</given-names></name> <name><surname>Chawla</surname> <given-names>NV</given-names></name></person-group>. <article-title>New perspectives and methods in link prediction</article-title>. <conf-name>16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</conf-name>. <conf-loc>New York, USA</conf-loc> (<year>2010</year>).</citation></ref>
<ref id="B38"><label>38</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Newman</surname> <given-names>MEJ</given-names></name></person-group>. <article-title>Clustering and preferential attachment in growing networks</article-title>. <source>Phys Rev E</source> (<year>2001</year>) <volume>64</volume>:<fpage>025102</fpage>.<pub-id pub-id-type="doi">10.1103/PhysRevE.64.025102</pub-id><pub-id pub-id-type="pmid">11497639</pub-id></citation></ref>
<ref id="B39"><label>39</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katz</surname> <given-names>L</given-names></name></person-group>. <article-title>A new status index derived from sociometric analysis</article-title>. <source>Psychometrika</source> (<year>1953</year>) <volume>18</volume>:<fpage>39</fpage>&#x02013;<lpage>43</lpage>.<pub-id pub-id-type="doi">10.1007/BF02289026</pub-id></citation></ref>
<ref id="B40"><label>40</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Kuhn</surname> <given-names>M</given-names></name> <name><surname>Johnson</surname> <given-names>K</given-names></name></person-group>. <source>Applied Predictive Modeling</source>. <publisher-loc>New York, U S</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2013</year>).</citation></ref>
<ref id="B41"><label>41</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L</given-names></name></person-group>. <article-title>Random forests</article-title>. <source>Mach Learn</source> (<year>2001</year>) <volume>45</volume>:<fpage>5</fpage>&#x02013;<lpage>32</lpage>.<pub-id pub-id-type="doi">10.1023/A:1017934522171</pub-id></citation></ref>
<ref id="B42"><label>42</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuhn</surname> <given-names>M</given-names></name></person-group>. <source>Caret: Classification and Regression Training</source>. (<year>2015</year>).</citation></ref>
<ref id="B43"><label>43</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>C</given-names></name> <name><surname>Liaw</surname> <given-names>A</given-names></name> <name><surname>Breiman</surname> <given-names>L</given-names></name></person-group>. <source>Using Random Forest to Learn Imbalanced Data</source>. <publisher-loc>Berkeley, CA</publisher-loc>: <publisher-name>University of California</publisher-name> (<year>2004</year>).</citation></ref>
<ref id="B44"><label>44</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuhn</surname> <given-names>M</given-names></name></person-group>. <article-title>Building predictive models in R using the caret package</article-title>. <source>J Stat Softw</source> (<year>2008</year>) <volume>28</volume>:<fpage>1</fpage>&#x02013;<lpage>26</lpage>.<pub-id pub-id-type="doi">10.18637/jss.v028.i05</pub-id></citation></ref>
<ref id="B45"><label>45</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wright</surname> <given-names>MN</given-names></name> <name><surname>Ziegler</surname> <given-names>A</given-names></name></person-group>. <article-title>ranger: a fast implementation of random forests for high dimensional data in C&#x0002B;&#x0002B; and R</article-title>. <fpage>arXiv</fpage> preprint arXiv:1508.04409 [Online] (<year>2015</year>). Available from: <uri xlink:href="https://arxiv.org/abs/1508.04409">https://arxiv.org/abs/1508.04409</uri> (accessed January 11, 2016).</citation></ref>
<ref id="B46"><label>46</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Silverman</surname> <given-names>BW</given-names></name></person-group>. <source>Density Estimation for Statistics and Data Analysis</source>. <publisher-loc>London</publisher-loc>: <publisher-name>CRC Press</publisher-name> (<year>1986</year>).</citation></ref>
<ref id="B47"><label>47</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duong</surname> <given-names>T</given-names></name></person-group>. <article-title>ks: kernel density estimation and kernel discriminant analysis for multivariate data in R</article-title>. <source>J Stat Softw</source> (<year>2007</year>) <volume>21</volume>:<fpage>1</fpage>&#x02013;<lpage>16</lpage>.<pub-id pub-id-type="doi">10.18637/jss.v021.i07</pub-id></citation></ref>
<ref id="B48"><label>48</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Corzo</surname> <given-names>CA</given-names></name> <name><surname>Mondaca</surname> <given-names>E</given-names></name> <name><surname>Wayne</surname> <given-names>S</given-names></name> <name><surname>Torremorell</surname> <given-names>M</given-names></name> <name><surname>Dee</surname> <given-names>S</given-names></name> <name><surname>Davies</surname> <given-names>P</given-names></name> <etal/></person-group> <article-title>Control and elimination of porcine reproductive and respiratory syndrome virus</article-title>. <source>Virus Res</source> (<year>2010</year>) <volume>154</volume>:<fpage>185</fpage>&#x02013;<lpage>92</lpage>.<pub-id pub-id-type="doi">10.1016/j.virusres.2010.08.016</pub-id><pub-id pub-id-type="pmid">20837071</pub-id></citation></ref>
<ref id="B49"><label>49</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dee</surname> <given-names>S</given-names></name> <name><surname>Deen</surname> <given-names>J</given-names></name> <name><surname>Burns</surname> <given-names>D</given-names></name> <name><surname>Douthit</surname> <given-names>G</given-names></name> <name><surname>Pijoan</surname> <given-names>C</given-names></name></person-group>. <article-title>An evaluation of disinfectants for the sanitation of porcine reproductive and respiratory syndrome virus-contaminated transport vehicles at cold temperatures</article-title>. <source>Can J Vet Res</source> (<year>2005</year>) <volume>69</volume>:<fpage>64</fpage>&#x02013;<lpage>70</lpage>.<pub-id pub-id-type="pmid">15745225</pub-id></citation></ref>
<ref id="B50"><label>50</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Giamalva</surname> <given-names>J</given-names></name></person-group>. <source>Pork and Swine. Industry and Trade Summary</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>International Trade Commission</publisher-name> (<year>2014</year>).</citation></ref>
<ref id="B51"><label>51</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>McBride</surname> <given-names>WD</given-names></name> <name><surname>Key</surname> <given-names>N</given-names></name></person-group>. <source>Economic and Structural Relationships in U.S. Hog Production</source>. <publisher-name>USDA-ERS Agricultural Economic Report &#x02013; SSRN Electronic Journal</publisher-name> (<year>2003</year>).</citation></ref>
<ref id="B52"><label>52</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alvarez</surname> <given-names>J</given-names></name> <name><surname>Valdes-Donoso</surname> <given-names>P</given-names></name> <name><surname>Tousignant</surname> <given-names>S</given-names></name> <name><surname>Alkhamis</surname> <given-names>M</given-names></name> <name><surname>Morrison</surname> <given-names>R</given-names></name> <name><surname>Perez</surname> <given-names>A</given-names></name></person-group>. <article-title>Novel analytic tools for the study of porcine reproductive and respiratory syndrome virus (PRRSv) in endemic settings: lessons learned in the U.S</article-title>. <source>Porcine Health Manag</source> (<year>2016</year>) <volume>2</volume>:<fpage>1</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="doi">10.1186/s40813-016-0019-0</pub-id></citation></ref>
<ref id="B53"><label>53</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liaw</surname> <given-names>A</given-names></name> <name><surname>Wiener</surname> <given-names>M</given-names></name></person-group>. <article-title>Classification and regression by randomForest</article-title>. <source>R News</source> (<year>2002</year>) <volume>2</volume>:<fpage>18</fpage>&#x02013;<lpage>22</lpage>.<pub-id pub-id-type="doi">10.1057/9780230509993</pub-id></citation></ref>
</ref-list>
</back>
</article>
