<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.0 20120330//EN" "JATS-journalpublishing1.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">JDS</journal-id>
<journal-title-group><journal-title>Journal of Data Science</journal-title></journal-title-group>
<issn pub-type="epub">1683-8602</issn><issn pub-type="ppub">1680-743X</issn><issn-l>1680-743X</issn-l>
<publisher>
<publisher-name>School of Statistics, Renmin University of China</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">JDS1136</article-id>
<article-id pub-id-type="doi">10.6339/24-JDS1136</article-id>
<article-categories><subj-group subj-group-type="heading">
<subject>Statistical Data Science</subject></subj-group></article-categories>
<title-group>
<article-title>Identifying Anomalous Data Entries in Repeated Surveys<xref ref-type="fn" rid="j_jds1136_fn_001"><sup>✩</sup></xref></article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-0446-1328</contrib-id>
<name><surname>Sartore</surname><given-names>Luca</given-names></name><email xlink:href="mailto:luca.sartore@usda.gov">luca.sartore@usda.gov</email><email xlink:href="mailto:lsartore@niss.org">lsartore@niss.org</email><xref ref-type="aff" rid="j_jds1136_aff_001">1</xref><xref ref-type="aff" rid="j_jds1136_aff_002">2</xref><xref ref-type="corresp" rid="cor2">∗</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-3387-3484</contrib-id>
<name><surname>Chen</surname><given-names>Lu</given-names></name><xref ref-type="aff" rid="j_jds1136_aff_001">1</xref><xref ref-type="aff" rid="j_jds1136_aff_002">2</xref>
</contrib>
<contrib contrib-type="author">
<name><surname>van Wart</surname><given-names>Justin</given-names></name><xref ref-type="aff" rid="j_jds1136_aff_002">2</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0009-0008-9482-5316</contrib-id>
<name><surname>Dau</surname><given-names>Andrew</given-names></name><xref ref-type="aff" rid="j_jds1136_aff_002">2</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-9828-968X</contrib-id>
<name><surname>Bejleri</surname><given-names>Valbona</given-names></name><xref ref-type="aff" rid="j_jds1136_aff_002">2</xref>
</contrib>
<aff id="j_jds1136_aff_001"><label>1</label><institution>National Institute of Statistical Sciences</institution>, 1750 K Street NW Suite 1100, Washington DC, 20006, <country>USA</country></aff>
<aff id="j_jds1136_aff_002"><label>2</label><institution>United States Department of Agriculture, National Agriculture Statistics Service</institution>, 1400 Independence Avenue SW, Washington DC, 20250, <country>USA</country></aff>
</contrib-group>
<author-notes>
<fn id="j_jds1136_fn_001"><label>✩</label>
<p>The findings and conclusions in this article are those of the authors and should not be construed to represent any official USDA or US Government determination or policy. This research was supported in part by the intramural research program of the US Department of Agriculture, National Agriculture Statistics Service.</p></fn><corresp id="cor2"><label>∗</label>Corresponding author. Email: <ext-link ext-link-type="uri" xlink:href="mailto:luca.sartore@usda.gov">luca.sartore@usda.gov</ext-link> or <ext-link ext-link-type="uri" xlink:href="mailto:lsartore@niss.org">lsartore@niss.org</ext-link>.</corresp>
</author-notes>
<pub-date pub-type="ppub"><year>2024</year></pub-date><pub-date pub-type="epub"><day>8</day><month>8</month><year>2024</year></pub-date><volume>22</volume><issue>3</issue><fpage>436</fpage><lpage>455</lpage><supplementary-material id="S1" content-type="archive" xlink:href="jds1136_s001.zip" mimetype="application" mime-subtype="x-zip-compressed">
<caption>
<title>Supplementary Material</title>
</caption>
</supplementary-material><history><date date-type="received"><day>30</day><month>11</month><year>2023</year></date><date date-type="accepted"><day>16</day><month>4</month><year>2024</year></date></history>
<permissions><copyright-statement>2024 The Author(s). Published by the School of Statistics and the Center for Applied Statistics, Renmin University of China.</copyright-statement><copyright-year>2024</copyright-year>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>Open access article under the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">CC BY</ext-link> license.</license-p></license></permissions>
<abstract>
<p>The presence of outliers in a dataset can substantially bias the results of statistical analyses. In general, micro edits are often performed manually on all records to correct for outliers. A set of constraints and decision rules is used to simplify the editing process. However, agricultural data collected through repeated surveys are characterized by complex relationships that make revision and vetting challenging. Therefore, maintaining high data-quality standards is not sustainable in short timeframes. The United States Department of Agriculture’s (USDA’s) National Agricultural Statistics Service (NASS) has partially automated its editing process to improve the accuracy of final estimates. NASS has investigated several methods to modernize its anomaly detection system because simple decision rules may not detect anomalies that break linear relationships. In this article, a computationally efficient method that identifies format-inconsistent, historical, tail, and relational anomalies at the data-entry level is introduced. Four separate scores (i.e., one for each anomaly type) are computed for all nonmissing values in a dataset. A distribution-free method motivated by the Bienaymé-Chebyshev’s inequality is used for scoring the data entries. Fuzzy logic is then considered for combining four individual scores into one final score to determine the outliers. The performance of the proposed approach is illustrated with an application to NASS survey data.</p>
</abstract>
<kwd-group>
<label>Keywords</label>
<kwd>agricultural data</kwd>
<kwd>Bienaymé-Chebyshev’s inequality</kwd>
<kwd>cellwise outliers</kwd>
<kwd>fuzzy logic</kwd>
<kwd>outlier detection</kwd>
<kwd>statistical analysis</kwd>
</kwd-group>
<funding-group><funding-statement>This research was supported by the intramural research program of the US Department of Agriculture, National Agricultural Statistics Service (NASS).</funding-statement></funding-group>
</article-meta>
</front>
<back>
<ref-list id="j_jds1136_reflist_001">
<title>References</title>
<ref id="j_jds1136_ref_001">
<mixed-citation publication-type="journal"> <string-name><surname>Agostinelli</surname> <given-names>C</given-names></string-name>, <string-name><surname>Leung</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yohai</surname> <given-names>VJ</given-names></string-name>, <string-name><surname>Zamar</surname> <given-names>RH</given-names></string-name> (<year>2015</year>). <article-title>Robust estimation of multivariate location and scatter in the presence of cellwise and casewise contamination</article-title>. <source><italic>Test</italic></source>, <volume>24</volume>(<issue>3</issue>): <fpage>441</fpage>–<lpage>461</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1007/s11749-015-0450-6" xlink:type="simple">https://doi.org/10.1007/s11749-015-0450-6</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_002">
<mixed-citation publication-type="journal"> <string-name><surname>Alqallaf</surname> <given-names>F</given-names></string-name>, <string-name><surname>Van Aelst</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yohai</surname> <given-names>VJ</given-names></string-name>, <string-name><surname>Zamar</surname> <given-names>RH</given-names></string-name> (<year>2009</year>). <article-title>Propagation of outliers in multivariate data</article-title>. <source><italic>The Annals of Statistics</italic></source>, <volume>37</volume>(<issue>1</issue>): <fpage>311</fpage>–<lpage>331</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1214/07-AOS588" xlink:type="simple">https://doi.org/10.1214/07-AOS588</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_003">
<mixed-citation publication-type="journal"> <string-name><surname>Bienaymé</surname> <given-names>IJ</given-names></string-name> (<year>1867</year>). <article-title>Considérations à l’appui de la découverte de Laplace sur la loi de probabilité dans la méthode des moindres carrés</article-title>. <source><italic>Journal de Mathématiques Pures et Appliquées</italic></source>, <volume>2</volume>(<issue>12</issue>): <fpage>158</fpage>–<lpage>176</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_004">
<mixed-citation publication-type="journal"> <string-name><surname>Chepulis</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Shevlyakov</surname> <given-names>G</given-names></string-name> (<year>2020</year>). <article-title>On outlier detection with the Chebyshev type inequalities</article-title>. <source><italic>Journal of the Belarusian State University. Mathematics and Informatics</italic></source>, <volume>3</volume>: <fpage>28</fpage>–<lpage>35</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.33581/2520-6508-2020-3-28-35" xlink:type="simple">https://doi.org/10.33581/2520-6508-2020-3-28-35</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_005">
<mixed-citation publication-type="journal"> <string-name><surname>Dagum</surname> <given-names>L</given-names></string-name>, <string-name><surname>Menon</surname> <given-names>R</given-names></string-name> (<year>1998</year>). <article-title>OpenMP: An industry standard API for shared-memory programming</article-title>. <source><italic>IEEE Computational Science and Engineering</italic></source>, <volume>5</volume>(<issue>1</issue>): <fpage>46</fpage>–<lpage>55</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1109/99.660313" xlink:type="simple">https://doi.org/10.1109/99.660313</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_006">
<mixed-citation publication-type="journal"> <string-name><surname>Daniel</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Gragg</surname> <given-names>WB</given-names></string-name>, <string-name><surname>Kaufman</surname> <given-names>L</given-names></string-name>, <string-name><surname>Stewart</surname> <given-names>GW</given-names></string-name> (<year>1976</year>). <article-title>Reorthogonalization and stable algorithms for updating the Gram-Schmidt <inline-formula id="j_jds1136_ineq_001"><alternatives><mml:math>
<mml:mi mathvariant="italic">Q</mml:mi>
<mml:mi mathvariant="italic">R</mml:mi></mml:math><tex-math><![CDATA[$QR$]]></tex-math></alternatives></inline-formula> factorization</article-title>. <source><italic>Mathematics of Computation</italic></source>, <volume>30</volume>(<issue>136</issue>): <fpage>772</fpage>–<lpage>795</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1090/S0025-5718-1976-0431641-8" xlink:type="simple">https://doi.org/10.1090/S0025-5718-1976-0431641-8</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_007">
<mixed-citation publication-type="book"> <string-name><surname>De Waal</surname> <given-names>T</given-names></string-name>, <string-name><surname>Pannekoek</surname> <given-names>J</given-names></string-name>, <string-name><surname>Scholtus</surname> <given-names>S</given-names></string-name> (<year>2011</year>). <source><italic>Handbook of Statistical Data Editing and Imputation</italic></source>, volume <volume>563</volume>. <publisher-name>John Wiley &amp; Sons</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_008">
<mixed-citation publication-type="journal"> <string-name><surname>Filzmoser</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gregorich</surname> <given-names>M</given-names></string-name> (<year>2020</year>). <article-title>Multivariate outlier detection in applied data analysis: Global, local, compositional and cellwise outliers</article-title>. <source><italic>Mathematical Geosciences</italic></source>, <volume>52</volume>(<issue>8</issue>): <fpage>1049</fpage>–<lpage>1066</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1007/s11004-020-09861-6" xlink:type="simple">https://doi.org/10.1007/s11004-020-09861-6</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_009">
<mixed-citation publication-type="journal"> <string-name><surname>Flynn</surname> <given-names>M</given-names></string-name> (<year>1966</year>). <article-title>Very high-speed computing systems</article-title>. <source><italic>Proceedings of the IEEE</italic></source>, <volume>54</volume>(<issue>12</issue>): <fpage>1901</fpage>–<lpage>1909</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1109/PROC.1966.5273" xlink:type="simple">https://doi.org/10.1109/PROC.1966.5273</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_010">
<mixed-citation publication-type="journal"> <string-name><surname>Gupta</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>J</given-names></string-name> (<year>1991</year>). <article-title>Theory of t-norms and fuzzy inference methods</article-title>. <source><italic>Fuzzy Sets and Systems</italic></source>, <volume>40</volume>(<issue>3</issue>): <fpage>431</fpage>–<lpage>450</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1016/0165-0114(91)90171-L" xlink:type="simple">https://doi.org/10.1016/0165-0114(91)90171-L</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_011">
<mixed-citation publication-type="journal"> <string-name><surname>Hampel</surname> <given-names>FR</given-names></string-name> (<year>1974</year>). <article-title>The influence curve and its role in robust estimation</article-title>. <source><italic>Journal of the American Statistical Association</italic></source>, <volume>69</volume>(<issue>346</issue>): <fpage>383</fpage>–<lpage>393</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1080/01621459.1974.10482962" xlink:type="simple">https://doi.org/10.1080/01621459.1974.10482962</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_012">
<mixed-citation publication-type="journal"> <string-name><surname>Heydarian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Doyle</surname> <given-names>TE</given-names></string-name>, <string-name><surname>Samavi</surname> <given-names>R</given-names></string-name> (<year>2022</year>). <article-title>MLCM: Multi-label confusion matrix</article-title>. <source><italic>IEEE Access</italic></source>, <volume>10</volume>: <fpage>19083</fpage>–<lpage>19095</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1109/ACCESS.2022.3151048" xlink:type="simple">https://doi.org/10.1109/ACCESS.2022.3151048</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_013">
<mixed-citation publication-type="journal"> <string-name><surname>Hidiroglou</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Berthelot</surname> <given-names>JM</given-names></string-name> (<year>1986</year>). <article-title>Statistical editing and imputation for periodic business surveys</article-title>. <source><italic>Survey Methodology</italic></source>, <volume>12</volume>(<issue>1</issue>): <fpage>73</fpage>–<lpage>83</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_014">
<mixed-citation publication-type="book"> <string-name><surname>Huber</surname> <given-names>PJ</given-names></string-name>, <string-name><surname>Ronchetti</surname> <given-names>EM</given-names></string-name> (<year>1981</year>). <source><italic>Robust Statistics</italic></source>. <publisher-name>John Wiley &amp; Sons</publisher-name>, <publisher-loc>New York</publisher-loc>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_015">
<mixed-citation publication-type="other"> <string-name><surname>Miller</surname> <given-names>D</given-names></string-name>, <string-name><surname>Robbins</surname> <given-names>M</given-names></string-name>, <string-name><surname>Habiger</surname> <given-names>J</given-names></string-name> (<year>2010</year>). Examining the challenges of missing data analysis in phase three of the agricultural resource management survey. <italic>JSM Proceedings. American Statistical Association Section on Survey Research Methods.</italic></mixed-citation>
</ref>
<ref id="j_jds1136_ref_016">
<mixed-citation publication-type="journal"> <string-name><surname>O’Gorman</surname> <given-names>TJ</given-names></string-name> (<year>1994</year>). <article-title>The effect of cosmic rays on the soft error rate of a DRAM at ground level</article-title>. <source><italic>IEEE Transactions on Electron Devices</italic></source>, <volume>41</volume>(<issue>4</issue>): <fpage>553</fpage>–<lpage>557</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1109/16.278509" xlink:type="simple">https://doi.org/10.1109/16.278509</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_017">
<mixed-citation publication-type="other"> <string-name><surname>Raymaekers</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rousseeuw</surname> <given-names>PJ</given-names></string-name> (<year>2019</year>). Handling cellwise outliers by sparse regression and robust covariance. arXiv preprint: <uri>https://arxiv.org/abs/1912.12446</uri>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_018">
<mixed-citation publication-type="other"> <string-name><surname>Raymaekers</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rousseeuw</surname> <given-names>PJ</given-names></string-name>, <string-name><surname>Van den Bossche</surname> <given-names>W</given-names></string-name>, <string-name><surname>Hubert</surname> <given-names>M</given-names></string-name> (<year>2023</year>). <monospace>cellWise</monospace><italic>: Analyzing data with cellwise outliers</italic>. CRAN, R package version 2.5.2.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_019">
<mixed-citation publication-type="journal"> <string-name><surname>Rousseeuw</surname> <given-names>PJ</given-names></string-name>, <string-name><surname>Van den Bossche</surname> <given-names>W</given-names></string-name> (<year>2018</year>). <article-title>Detecting deviating data cells</article-title>. <source><italic>Technometrics</italic></source>, <volume>60</volume>(<issue>2</issue>): <fpage>135</fpage>–<lpage>145</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1080/00401706.2017.1340909" xlink:type="simple">https://doi.org/10.1080/00401706.2017.1340909</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_020">
<mixed-citation publication-type="journal"> <string-name><surname>Rubin</surname> <given-names>DB</given-names></string-name> (<year>1976</year>). <article-title>Inference and missing data</article-title>. <source><italic>Biometrika</italic></source>, <volume>63</volume>(<issue>3</issue>): <fpage>581</fpage>–<lpage>592</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1093/biomet/63.3.581" xlink:type="simple">https://doi.org/10.1093/biomet/63.3.581</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_021">
<mixed-citation publication-type="journal"> <string-name><surname>Sandqvist</surname> <given-names>AP</given-names></string-name> (<year>2016</year>). <article-title>Identifizierung von Ausreissern in eindimensionalen gewichteten Umfragedaten</article-title>. <source><italic>KOF Analysen</italic></source>, <volume>2016</volume>(<issue>2</issue>): <fpage>45</fpage>–<lpage>56</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_022">
<mixed-citation publication-type="journal"> <string-name><surname>Sedgewick</surname> <given-names>R</given-names></string-name> (<year>1978</year>). <article-title>Implementing quicksort programs</article-title>. <source><italic>Communications of the ACM</italic></source>, <volume>21</volume>(<issue>10</issue>): <fpage>847</fpage>–<lpage>857</lpage>. <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.1145/359619.359631" xlink:type="simple">https://doi.org/10.1145/359619.359631</ext-link></mixed-citation>
</ref>
<ref id="j_jds1136_ref_023">
<mixed-citation publication-type="journal"> <string-name><surname>Stigler</surname> <given-names>SM</given-names></string-name> (<year>1973</year>). <article-title>The asymptotic distribution of the trimmed mean</article-title>. <source><italic>The Annals of Statistics</italic></source>, <volume>1</volume>(<issue>3</issue>): <fpage>472</fpage>–<lpage>477</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_024">
<mixed-citation publication-type="journal"> <string-name><surname>Tchebichef</surname> <given-names>P</given-names></string-name> (<year>1867</year>). <article-title>Des valeurs moyennes</article-title>. <source><italic>Journal de Mathématiques Pures et Appliquées</italic></source>, <volume>2</volume>(<issue>12</issue>): <fpage>177</fpage>–<lpage>184</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1136_ref_025">
<mixed-citation publication-type="book"> <string-name><surname>Zwillinger</surname> <given-names>D</given-names></string-name> (<year>2018</year>). <source><italic>Standard Mathematical Tables and Formulas</italic></source>. <publisher-name>CRC Press</publisher-name>.</mixed-citation>
</ref>
</ref-list>
</back>
</article>
