Automating Data Analysis Methods in Epidemiology

Choueiry, George; Salameh, Pascale

doi:10.6339/JDS.201901_17(1).0003

Journal of Data Science

Automating Data Analysis Methods in Epidemiology

Volume 17, Issue 1 (2019), pp. 55–80

George Choueiry Pascale Salameh

https://doi.org/10.6339/JDS.201901_17(1).0003

Pub. online: 4 August 2022 Type: Research Article

Open Access

Published
4 August 2022

Abstract

Technological advances in software development effectively handled technical details that made life easier for data analysts, but also allowed for nonexperts in statistics and computer science to analyze data. As a result, medical research suffers from statistical errors that could be otherwise prevented such as errors in choosing a hypothesis test and assumption checking of models. Our objective is to create an automated data analysis software package that can help practitioners run non-subjective, fast, accurate and easily interpretable analyses. We used machine learning to predict the normality of a distribution as an alternative to normality tests and graphical methods to avoid their downsides. We implemented methods for detecting outliers, imputing missing values, and choosing a threshold for cutting numerical variables to correct for non-linearity before running a linear regression. We showed that data analysis can be automated. Our normality prediction algorithm outperformed the Shapiro-Wilk test in small samples with Matthews correlation coefficient of 0.5 vs. 0.16. The biggest drawback was that we did not find alternatives for statistical tests to test linear regression assumptions which are problematic in large datasets. We also applied our work to a dataset about smoking in teenagers. Because of the opensource nature of our work, these algorithms can be used in future research and projects.

No copyright data available.

Keywords

automation computer software machine learning normal distribution

Metrics

since February 2021

1551

Article info
views

812

PDF
downloads

RSS

Authors

Abstract

Export citation

Copy and paste formatted citation

Download citation in file