Predicting “Yes”: Machine Learning and Diverse Data to Boost Respondent Cooperation
Pub. online: 17 August 2026
Type: Data Science In Action
Open Access
Received
16 April 2026
16 April 2026
Accepted
14 July 2026
14 July 2026
Published
17 August 2026
17 August 2026
Abstract
In the evolving field of survey research, leveraging machine learning to predict response behavior has transformative potential for the efficiency of survey operations. Integrating multiple data sources may improve response prediction by providing more nuanced insights for household outreach. This study presents a model-driven approach to enhancing respondent cooperation in the Medical Expenditure Panel Survey (MEPS) by combining features from disparate data sources. MEPS is a longitudinal household survey with 5 rounds of interviewing over 2.5 years. Its sample is derived prior National Health Interview Survey (NHIS) participants. MEPS Round 1 response rates are critical for sustaining representativeness throughout each panel. In this study, we constructed a multimodal machine learning model to predict (1) the likelihood of a positive response for an upcoming contact attempt and (2) the likelihood that new panel households complete a Round 1 interview. The model integrates tract-level data from the American Community Survey (ACS), outcomes from the Advance Call Records (ACR) made prior to MEPS Round 1, and paradata from the early contact period. We also explored the relative contributions of these sources to model performance. Our model aims to help manage field labor by identifying complex cases needing specialized support. It can also assist in determining the optimal mode for the next contact to increase the chance of a completed interview. Beyond improving MEPS operations, this study offers a roadmap for incorporating additional data sources to support fieldwork.
Supplementary material
Supplementary MaterialR and Python code for sequence feature construction, clustering analysis, and multinomial logistic regression modeling.
References
Chun AY, Schouten B, Wagner J (2017). JOS special issue on responsive and adaptive survey design: Looking back to see forward. Journal of Official Statistics, 33(3): 571–577. https://doi.org/10.1515/jos-2017-0027
Durrant GB, Maslovskaya O, Smith PWF (2019). Investigating call record data using sequence analysis to inform adaptive survey designs. International Journal of Social Research Methodology, 22(1): 37–54. https://doi.org/10.1080/13645579.2018.1490981
Eckman S, Koch A (2019). Interviewer involvement in sample selection shapes the relationship between response rates and data quality. Public Opinion Quarterly, 83(2): 313–337. https://doi.org/10.1093/poq/nfz012
Groves RM, Peytcheva E (2008). The impact of nonresponse rates on nonresponse bias: A meta-analysis. Public Opinion Quarterly, 72(2): 167–189. https://doi.org/10.1093/poq/nfn011
Rybak A (2023). Survey mode and nonresponse bias: A meta-analysis based on the data from the International Social Survey Programme waves 1996–2018 and the European Social Survey rounds 1 to 9. PLOS ONE, 18(3): e0283092. https://doi.org/10.1371/journal.pone.0283092
Wuyts C, Durrant GB, Kreuter F (2023). Predicting fieldwork effort in face-to-face surveys using call record data. Journal of Survey Statistics and Methodology, 11(2): 367–391. https://doi.org/10.1093/jssam/smab036