<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.0 20120330//EN" "JATS-journalpublishing1.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">JDS</journal-id>
<journal-title-group><journal-title>Journal of Data Science</journal-title></journal-title-group>
<issn pub-type="epub">1683-8602</issn><issn pub-type="ppub">1680-743X</issn><issn-l>1680-743X</issn-l>
<publisher>
<publisher-name>School of Statistics, Renmin University of China</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">JDS1035</article-id>
<article-id pub-id-type="doi">10.6339/22-JDS1035</article-id>
<article-categories><subj-group subj-group-type="heading">
<subject>Data Science Reviews</subject></subj-group></article-categories>
<title-group>
<article-title>Econometrics at Scale: Spark up Big Data in Economics<xref ref-type="fn" rid="j_jds1035_fn_001"><sup>✩</sup></xref></article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Bluhm</surname><given-names>Benjamin</given-names></name><xref ref-type="aff" rid="j_jds1035_aff_001">1</xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Cutura</surname><given-names>Jannic Alexander</given-names></name><email xlink:href="mailto:jannic.cutura@ecb.europa.eu">jannic.cutura@ecb.europa.eu</email><xref ref-type="aff" rid="j_jds1035_aff_002">2</xref><xref ref-type="corresp" rid="cor2">∗</xref>
</contrib>
<aff id="j_jds1035_aff_001"><label>1</label>Senior Data Scientist</aff>
<aff id="j_jds1035_aff_002"><label>2</label><institution>European Central Bank</institution>, Sonnemannstraße 20, 60314 Frankfurt am Main, <country>Germany</country></aff>
</contrib-group>
<author-notes>
<fn id="j_jds1035_fn_001"><label>✩</label>
<p>The views expressed in this paper are those of the authors alone and do not represent the view of the European Central Bank (ECB).</p></fn><corresp id="cor2"><label>∗</label>Corresponding author. Email: <ext-link ext-link-type="uri" xlink:href="mailto:jannic.cutura@ecb.europa.eu">jannic.cutura@ecb.europa.eu</ext-link>.</corresp>
</author-notes>
<pub-date pub-type="ppub"><year>2022</year></pub-date><pub-date pub-type="epub"><day>7</day><month>4</month><year>2022</year></pub-date><volume>20</volume><issue>3</issue><fpage>413</fpage><lpage>436</lpage><supplementary-material id="S1" content-type="archive" xlink:href="jds1035_s001.zip" mimetype="application" mime-subtype="x-zip-compressed">
<caption>
<title>Supplementary Material</title>
<p>Supplementary material is available on our github page, containing all codes to replicate the results along links to the data. Additional instructions are also available, detailing how to setup the AWS infrastructure: <uri>https://github.com/benjaminbluhm/econometrics_at_scale</uri>.</p>
</caption>
</supplementary-material><history><date date-type="received"><day>9</day><month>11</month><year>2021</year></date><date date-type="accepted"><day>12</day><month>1</month><year>2022</year></date></history>
<permissions><copyright-statement>2022 The Author(s). Published by the School of Statistics and the Center for Applied Statistics, Renmin University of China.</copyright-statement><copyright-year>2022</copyright-year>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>Open access article under the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">CC BY</ext-link> license.</license-p></license></permissions>
<abstract>
<p>This paper provides an overview of how to use “big data” for social science research (with an emphasis on economics and finance). We investigate the performance and ease of use of different Spark applications running on a distributed file system to enable the handling and analysis of data sets which were previously not usable due to their size. More specifically, we explain how to use Spark to (i) explore big data sets which exceed retail grade computers memory size and (ii) run typical statistical/econometric tasks including cross sectional, panel data and time series regression models which are prohibitively expensive to evaluate on stand-alone machines. By bridging the gap between the abstract concept of Spark and ready-to-use examples which can easily be altered to suite the researchers need, we provide economists and social scientists more generally with the theory and practice to handle the ever growing datasets available. The ease of reproducing the examples in this paper makes this guide a useful reference for researchers with a limited background in data handling and distributed computing.</p>
</abstract>
<kwd-group>
<label>Keywords</label>
<kwd>Apache Spark</kwd>
<kwd>distributed computing</kwd>
<kwd>econometrics</kwd>
</kwd-group>
<funding-group><award-group><funding-source xlink:href="https://doi.org/10.13039/501100000533">Bank of England</funding-source></award-group><funding-statement>We gratefully acknowledge a travel grant sponsored by the Bank of England. We gratefully acknowledge research support from the Leibniz Institute for Financial Research SAFE. </funding-statement></funding-group>
</article-meta>
</front>
<body/>
<back>
<ref-list id="j_jds1035_reflist_001">
<title>References</title>
<ref id="j_jds1035_ref_001">
<mixed-citation publication-type="journal"> <string-name><surname>Arellano</surname> <given-names>M</given-names></string-name> (<year>1987</year>). <article-title>Practitioners’ corner: computing robust standard errors for within-groups estimators</article-title>. <source>Oxford Bulletin of Economics and Statistics</source>, <volume>49</volume>(<issue>4</issue>): <fpage>431</fpage>–<lpage>434</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_002">
<mixed-citation publication-type="other"> <string-name><surname>Aruoba</surname> <given-names>SB</given-names></string-name>, <string-name><surname>Fernandez-Villaverde</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rubio-Ramirez</surname> <given-names>JF</given-names></string-name> (2003). <italic>Comparing Solution Methods for Dynamic Equilibrium Economies</italic>. PIER Working Paper Archive 04-003. Penn Institute for Economic Research, Department of Economics, University of Pennsylvania.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_003">
<mixed-citation publication-type="journal"> <string-name><surname>Athey</surname> <given-names>S</given-names></string-name>, <string-name><surname>Imbens</surname> <given-names>GW</given-names></string-name> (<year>2017</year>). <article-title>The state of applied econometrics: causality and policy evaluation</article-title>. <source>The Journal of Economic Perspectives</source>, <volume>31</volume>(<issue>2</issue>): <fpage>3</fpage>–<lpage>32</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_004">
<mixed-citation publication-type="book"> <string-name><surname>Baltagi</surname> <given-names>B</given-names></string-name> (<year>2008</year>). <source>Econometric Analysis of Panel Data</source>. <publisher-name>John Wiley &amp; Sons</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_005">
<mixed-citation publication-type="other"> <string-name><surname>Boneva</surname> <given-names>L</given-names></string-name>, <string-name><surname>Böninghausen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Fache Rousová</surname> <given-names>L</given-names></string-name>, <string-name><surname>Letizia</surname> <given-names>E</given-names></string-name>, et al. (2019). Derivatives transactions data and their use in central bank analysis. <italic>Economic Bulletin Articles</italic>, 6.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_006">
<mixed-citation publication-type="journal"> <string-name><surname>Böse</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Flunkert</surname> <given-names>V</given-names></string-name>, <string-name><surname>Gasthaus</surname> <given-names>J</given-names></string-name>, <string-name><surname>Januschowski</surname> <given-names>T</given-names></string-name>, <string-name><surname>Lange</surname> <given-names>D</given-names></string-name>, <string-name><surname>Salinas</surname> <given-names>D</given-names></string-name>, et al. (<year>2017</year>). <article-title>Probabilistic demand forecasting at scale</article-title>. <source>Proceedings of the VLDB Endowment</source>, <volume>10</volume>: <fpage>1694</fpage>–<lpage>1705</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_007">
<mixed-citation publication-type="book"> <string-name><surname>Cameron</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Trivedi</surname> <given-names>PK</given-names></string-name> (<year>2005</year>). <source>Microeconometrics: methods and applications</source>. <publisher-name>Cambridge University Press</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_008">
<mixed-citation publication-type="journal"> <string-name><surname>Cavallo</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rigobon</surname> <given-names>R</given-names></string-name> (<year>2016</year>). <article-title>The billion prices project: using online prices for measurement and research</article-title>. <source>The Journal of Economic Perspectives</source>, <volume>30</volume>(<issue>2</issue>): <fpage>151</fpage>–<lpage>78</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_009">
<mixed-citation publication-type="book"> <string-name><surname>Chambers</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zaharia</surname> <given-names>M</given-names></string-name> (<year>2018</year>). <source>Spark – The Definitive Guide: Big Data Processing Made Simple</source>. <publisher-name>O’Reilly Media, Incorporated</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_010">
<mixed-citation publication-type="chapter"> <string-name><surname>Chun</surname> <given-names>BG</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>B</given-names></string-name>, <string-name><surname>Jeon</surname> <given-names>B</given-names></string-name>, <string-name><surname>Jeong</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>G</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JY</given-names></string-name>, <etal>et al.</etal> (<year>2016</year>). <chapter-title>Dolphin: runtime optimization for distributed machine learning</chapter-title>. In: <source>The ML Systems Workshop at ICML</source>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_011">
<mixed-citation publication-type="journal"> <string-name><surname>Clemen</surname> <given-names>RT</given-names></string-name> (<year>1989</year>). <article-title>Combining forecasts: a review and annotated bibliography</article-title>. <source>International Journal of Forecasting</source>, <volume>5</volume>(<issue>4</issue>): <fpage>559</fpage>–<lpage>583</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_012">
<mixed-citation publication-type="other"> <string-name><surname>Correia</surname> <given-names>S</given-names></string-name> (2016). <italic>Linear Models with High-Dimensional Fixed Effects: An Efficient and Feasible Estimator</italic>. Technical Report. Working Paper.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_013">
<mixed-citation publication-type="chapter"> <string-name><surname>Dean</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ghemawat</surname> <given-names>S</given-names></string-name> (<year>2004</year>). <chapter-title>Mapreduce: simplified data processing on large clusters</chapter-title>. In: <source>OSDI’04: Sixth Symposium on Operating System Design and Implementation</source>, <fpage>137</fpage>–<lpage>150</lpage>. <publisher-loc>San Francisco, CA</publisher-loc>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_014">
<mixed-citation publication-type="journal"> <string-name><surname>Dick-Nielsen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Feldhütter</surname> <given-names>P</given-names></string-name>, <string-name><surname>Lando</surname> <given-names>D</given-names></string-name> (<year>2012</year>). <article-title>Corporate bond liquidity before and after the onset of the subprime crisis</article-title>. <source>Journal of Financial Economics</source>, <volume>103</volume>(<issue>3</issue>): <fpage>471</fpage>–<lpage>492</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_015">
<mixed-citation publication-type="journal"> <string-name><surname>Duchin</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sosyura</surname> <given-names>D</given-names></string-name> (<year>2014</year>). <article-title>Safer ratios, riskier portfolios: banks response to government aid</article-title>. <source>Journal of Financial Economics</source>, <volume>113</volume>(<issue>1</issue>): <fpage>1</fpage>–<lpage>28</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_016">
<mixed-citation publication-type="journal"> <string-name><surname>Edwards</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Harris</surname> <given-names>LE</given-names></string-name>, <string-name><surname>Piwowar</surname> <given-names>MS</given-names></string-name> (<year>2007</year>). <article-title>Corporate bond market transaction costs and transparency</article-title>. <source>The Journal of Finance</source>, <volume>62</volume>(<issue>3</issue>): <fpage>1421</fpage>–<lpage>1451</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_017">
<mixed-citation publication-type="journal"> <string-name><surname>Einav</surname> <given-names>L</given-names></string-name>, <string-name><surname>Levin</surname> <given-names>J</given-names></string-name> (<year>2014</year>). <article-title>Economics in the age of big data</article-title>. <source>Science (New York, N. Y.)</source>, <volume>346</volume>(<issue>6210</issue>): <fpage>1243089</fpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_018">
<mixed-citation publication-type="book"> <string-name><surname>Ferguson</surname> <given-names>TS</given-names></string-name> (<year>2017</year>). <source>A Course in Large Sample Theory</source>. <publisher-name>Routledge</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_019">
<mixed-citation publication-type="book"> <string-name><surname>Fernández-Villaverde</surname> <given-names>J</given-names></string-name>, <string-name><surname>Valencia</surname> <given-names>DZ</given-names></string-name> (<year>2018</year>). <source>A Practical Guide to Parallelization in Economics</source>. <publisher-name>National Bureau of Economic Research</publisher-name>, <publisher-loc>Cambridge, MA</publisher-loc>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_020">
<mixed-citation publication-type="other"> <string-name><surname>Flom</surname> <given-names>P</given-names></string-name> (2013). Hypothesis testing with big data. Cross Validated. (Version: 2013-08-13).</mixed-citation>
</ref>
<ref id="j_jds1035_ref_021">
<mixed-citation publication-type="book"> <string-name><surname>Foster</surname> <given-names>I</given-names></string-name>, <string-name><surname>Ghani</surname> <given-names>R</given-names></string-name>, <string-name><surname>Jarmin</surname> <given-names>RS</given-names></string-name>, <string-name><surname>Kreuter</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lane</surname> <given-names>J</given-names></string-name> (<year>2016</year>). <source>Big Data and Social Science: A Practical Guide to Methods and Tools</source>. <publisher-name>Chapman and Hall/CRC</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_022">
<mixed-citation publication-type="other"> <string-name><surname>Gao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ru</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name> (<year>2019</year>). <source>What do a Billion Observations Say About Distance and Relationship Lending?</source> <comment>Working Paper. Technical Report</comment>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_023">
<mixed-citation publication-type="other"> <string-name><surname>Gaure</surname> <given-names>S</given-names></string-name> (2019). lfe: Linear Group Fixed Effects. 2.8-7.1 edition.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_024">
<mixed-citation publication-type="journal"> <string-name><surname>Gentzkow</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kelly</surname> <given-names>BT</given-names></string-name>, <string-name><surname>Taddy</surname> <given-names>M</given-names></string-name> (<year>2019</year>). <article-title>Text as data</article-title>. <source>Journal of Economic Literature</source>. <volume>57</volume>(<issue>3</issue>): <fpage>535</fpage>–<lpage>374</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_025">
<mixed-citation publication-type="chapter"> <string-name><surname>Ghemawat</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gobioff</surname> <given-names>H</given-names></string-name>, <string-name><surname>Leung</surname> <given-names>ST</given-names></string-name> (<year>2003</year>). <chapter-title>The Google file system</chapter-title>. In: <source>Proceedings of the 19th ACM Symposium on Operating Systems Principles</source>, <fpage>20</fpage>–<lpage>43</lpage>, <publisher-loc>Bolton Landing, NY</publisher-loc>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_026">
<mixed-citation publication-type="journal"> <string-name><surname>Gilje</surname> <given-names>EP</given-names></string-name>, <string-name><surname>Loutskina</surname> <given-names>E</given-names></string-name>, <string-name><surname>Strahan</surname> <given-names>PE</given-names></string-name> (<year>2016</year>). <article-title>Exporting liquidity: branch banking and financial integration</article-title>. <source>The Journal of Finance</source>, <volume>71</volume>(<issue>3</issue>): <fpage>1159</fpage>–<lpage>1184</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_027">
<mixed-citation publication-type="journal"> <string-name><surname>Greenwald</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khanna</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal> (<year>2001</year>). <article-title>Space-efficient online computation of quantile summaries</article-title>. <source>ACM SIGMOD Record</source>, <volume>30</volume>(<issue>2</issue>): <fpage>58</fpage>–<lpage>66</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_028">
<mixed-citation publication-type="journal"> <string-name><surname>Grimmer</surname> <given-names>J</given-names></string-name>, <string-name><surname>Stewart</surname> <given-names>BM</given-names></string-name> (<year>2013</year>). <article-title>Text as data: the promise and pitfalls of automatic content analysis methods for political texts</article-title>. <source>Political Analysis</source>, <volume>21</volume>(<issue>03</issue>): <fpage>267</fpage>–<lpage>297</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_029">
<mixed-citation publication-type="journal"> <string-name><surname>Hamermesh</surname> <given-names>DS</given-names></string-name> (<year>2013</year>). <article-title>Six decades of top economics publishing: who and how?</article-title> <source>Journal of Economic Literature</source>, <volume>51</volume>(<issue>1</issue>): <fpage>162</fpage>–<lpage>172</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_030">
<mixed-citation publication-type="book"> <string-name><surname>Hamilton</surname> <given-names>JD</given-names></string-name> (<year>1994</year>). <source>Time Series Analysis</source>. <publisher-name>Princeton Univ. Press</publisher-name>, <publisher-loc>Princeton, NJ</publisher-loc>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_031">
<mixed-citation publication-type="journal"> <string-name><surname>Hansen</surname> <given-names>C</given-names></string-name> (<year>2007</year>). <article-title>Asymptotic properties of a robust variance matrix estimator for panel data when t is large</article-title>. <source>Journal of Econometrics</source>, <volume>141</volume>: <fpage>597</fpage>–<lpage>620</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_032">
<mixed-citation publication-type="other"> <collab>Irving-Fisher-Committee</collab> (2020). Irving Fisher Committee on Central Bank Statistics. 2019 ifc Annual Report. (accessed 15/01/2020).</mixed-citation>
</ref>
<ref id="j_jds1035_ref_033">
<mixed-citation publication-type="journal"> <string-name><surname>Jankowitsch</surname> <given-names>R</given-names></string-name>, <string-name><surname>Nagler</surname> <given-names>F</given-names></string-name>, <string-name><surname>Subrahmanyam</surname> <given-names>MG</given-names></string-name> (<year>2014</year>). <article-title>The determinants of recovery rates in the us corporate bond market</article-title>. <source>Journal of Financial Economics</source>, <volume>114</volume>(<issue>1</issue>): <fpage>155</fpage>–<lpage>177</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_034">
<mixed-citation publication-type="book"> <string-name><surname>Karau</surname> <given-names>H</given-names></string-name>, <string-name><surname>Konwinski</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wendell</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zaharia</surname> <given-names>M</given-names></string-name> (<year>2015</year>). <source>Learning Spark: Lightning-Fast Big Data Analytics</source>. <publisher-name>O’Reilly Media, Inc.</publisher-name>, <edition>1st</edition> edition.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_035">
<mixed-citation publication-type="book"> <string-name><surname>Karau</surname> <given-names>H</given-names></string-name>, <string-name><surname>Warren</surname> <given-names>R</given-names></string-name> (<year>2017</year>). <source>High Performance Spark: Best Practices for Scaling and Optimizing Apache Spark</source>. <publisher-name>O’Reilly Media, Inc.</publisher-name>, <edition>1st</edition> edition.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_036">
<mixed-citation publication-type="journal"> <string-name><surname>Kleinberg</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ludwig</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mullainathan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Obermeyer</surname> <given-names>Z</given-names></string-name> (<year>2015</year>). <article-title>Prediction policy problems</article-title>. <source>The American Economic Review</source>, <volume>105</volume>(<issue>5</issue>): <fpage>491</fpage>–<lpage>495</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_037">
<mixed-citation publication-type="journal"> <string-name><surname>Leamer</surname> <given-names>EE</given-names></string-name> (<year>1985</year>). <article-title>Sensitivity analyses would help</article-title>. <source>The American Economic Review</source>, <volume>75</volume>(<issue>3</issue>): <fpage>308</fpage>–<lpage>313</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_038">
<mixed-citation publication-type="journal"> <string-name><surname>Millo</surname> <given-names>G</given-names></string-name> (<year>2017</year>). <article-title>Robust standard error estimators for panel models: a unifying approach</article-title>. <source>Journal of Statistical Software</source>, <volume>82</volume>(<issue>3</issue>): <fpage>1</fpage>–<lpage>27</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_039">
<mixed-citation publication-type="journal"> <string-name><surname>Mullainathan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Spiess</surname> <given-names>J</given-names></string-name> (<year>2017</year>). <article-title>Machine learning: an applied econometric approach</article-title>. <source>The Journal of Economic Perspectives</source>, <volume>31</volume>(<issue>2</issue>): <fpage>87</fpage>–<lpage>106</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_040">
<mixed-citation publication-type="journal"> <string-name><surname>Munnell</surname> <given-names>AH</given-names></string-name>, <string-name><surname>Tootell</surname> <given-names>GM</given-names></string-name>, <string-name><surname>Browne</surname> <given-names>LE</given-names></string-name>, <string-name><surname>McEneaney</surname> <given-names>J</given-names></string-name> (<year>1996</year>). <article-title>Mortgage lending in Boston: interpreting hmda data</article-title>. <source>The American Economic Review</source>, <volume>86</volume>(<issue>1</issue>): <fpage>25</fpage>–<lpage>53</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_041">
<mixed-citation publication-type="other"> <string-name><surname>Ng</surname> <given-names>S</given-names></string-name> (2017). <italic>Opportunities and Challenges: Lessons from Analyzing Terabytes of Scanner Data</italic>. Technical Report. National Bureau of Economic Research.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_042">
<mixed-citation publication-type="book"> <collab>R Core Team</collab> (<year>2019</year>). <source>R: A Language and Environment for Statistical Computing</source>. <publisher-name>R Foundation for Statistical Computing</publisher-name>, <publisher-loc>Vienna, Austria</publisher-loc>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_043">
<mixed-citation publication-type="journal"> <string-name><surname>Sala-I-Martin</surname> <given-names>XX</given-names></string-name> (<year>1997</year>). <article-title>I just ran two million regressions</article-title>. <source>The American Economic Review</source>, <volume>87</volume>(<issue>2</issue>): <fpage>178</fpage>–<lpage>183</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_044">
<mixed-citation publication-type="journal"> <string-name><surname>Samadi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zbakh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tadonki</surname> <given-names>C</given-names></string-name> (<year>2018</year>). <article-title>Performance comparison between hadoop and spark frameworks using hibench benchmarks</article-title>. <source>Concurrency and Computation: Practice and Experience</source>, <volume>30</volume>(<issue>12</issue>): <fpage>e4367</fpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_045">
<mixed-citation publication-type="other"> <string-name><surname>Sheppard</surname> <given-names>K</given-names></string-name> (2019). linearmodels: Models for Panel Data, 4.25 edition.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_046">
<mixed-citation publication-type="chapter"> <string-name><surname>Timmermann</surname> <given-names>A</given-names></string-name> (<year>2006</year>). <chapter-title>Forecast combinations</chapter-title>. In: <source>Handbook of Economic Forecasting</source> (<string-name><given-names>G</given-names> <surname>Elliott</surname></string-name>, <string-name><given-names>C</given-names> <surname>Granger</surname></string-name>, <string-name><given-names>A</given-names> <surname>Timmermann</surname></string-name>, eds.), volume <volume>1</volume> of <series><italic>Handbook of Economic Forecasting</italic></series>, <comment>Chapter 4</comment>. <fpage>135</fpage>–<lpage>196</lpage>. <publisher-name>Elsevier</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_047">
<mixed-citation publication-type="journal"> <string-name><surname>Varian</surname> <given-names>HR</given-names></string-name> (<year>2014</year>). <article-title>Big data: new tricks for econometrics</article-title>. <source>The Journal of Economic Perspectives</source>, <volume>28</volume>(<issue>2</issue>): <fpage>3</fpage>–<lpage>28</lpage>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_048">
<mixed-citation publication-type="book"> <string-name><surname>Wooldridge</surname> <given-names>JM</given-names></string-name> (<year>2010</year>). <source>Econometric Analysis of Cross Section and Panel Data</source>. <publisher-name>MIT press</publisher-name>.</mixed-citation>
</ref>
<ref id="j_jds1035_ref_049">
<mixed-citation publication-type="chapter"> <string-name><surname>Zaharia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chowdhury</surname> <given-names>M</given-names></string-name>, <string-name><surname>Franklin</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Shenker</surname> <given-names>S</given-names></string-name>, <string-name><surname>Stoica</surname> <given-names>I</given-names></string-name> (<year>2010</year>). <chapter-title>Spark: cluster computing with working sets</chapter-title>. In: <source>Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing</source>. <publisher-name>USENIX Association</publisher-name>, <publisher-loc>Boston, MA</publisher-loc>.</mixed-citation>
</ref>
</ref-list>
</back>
</article>
