Journal of Data Science logo


Login Register

  1. Home
  2. To appear
  3. Predicting “Yes”: Machine Learning and D ...

Journal of Data Science

Submit your article Information
  • Article info
  • Related articles
  • More
    Article info Related articles

Predicting “Yes”: Machine Learning and Diverse Data to Boost Respondent Cooperation
Rashi Saluja ORCID icon link to view author Rashi Saluja details   Hanyu Sun   Gizem Korkmaz     All authors (7)

Authors

 
Placeholder
https://doi.org/10.6339/26-JDS1242
Pub. online: 17 August 2026      Type: Data Science In Action      Open accessOpen Access

Received
16 April 2026
Accepted
14 July 2026
Published
17 August 2026

Abstract

In the evolving field of survey research, leveraging machine learning to predict response behavior has transformative potential for the efficiency of survey operations. Integrating multiple data sources may improve response prediction by providing more nuanced insights for household outreach. This study presents a model-driven approach to enhancing respondent cooperation in the Medical Expenditure Panel Survey (MEPS) by combining features from disparate data sources. MEPS is a longitudinal household survey with 5 rounds of interviewing over 2.5 years. Its sample is derived prior National Health Interview Survey (NHIS) participants. MEPS Round 1 response rates are critical for sustaining representativeness throughout each panel. In this study, we constructed a multimodal machine learning model to predict (1) the likelihood of a positive response for an upcoming contact attempt and (2) the likelihood that new panel households complete a Round 1 interview. The model integrates tract-level data from the American Community Survey (ACS), outcomes from the Advance Call Records (ACR) made prior to MEPS Round 1, and paradata from the early contact period. We also explored the relative contributions of these sources to model performance. Our model aims to help manage field labor by identifying complex cases needing specialized support. It can also assist in determining the optimal mode for the next contact to increase the chance of a completed interview. Beyond improving MEPS operations, this study offers a roadmap for incorporating additional data sources to support fieldwork.

Supplementary material

 Supplementary Material
R and Python code for sequence feature construction, clustering analysis, and multinomial logistic regression modeling.

References

 
Bianchi A, Biffignandi S, Lynn P (2015). Web–face-to-face mixed-mode design in a longitudinal survey: Effects on participation, composition, and costs. Survey Methodology, 41(2): 301–318.
 
Chun AY, Schouten B, Wagner J (2017). JOS special issue on responsive and adaptive survey design: Looking back to see forward. Journal of Official Statistics, 33(3): 571–577. https://doi.org/10.1515/jos-2017-0027
 
Dillman DA, Smyth JD, Christian LM (2014). Internet, Phone, Mail, and Mixed-Mode Surveys: The Tailored Design Method. Wiley, Hoboken, NJ, 4th edition.
 
Durrant GB, Maslovskaya O, Smith PWF (2019). Investigating call record data using sequence analysis to inform adaptive survey designs. International Journal of Social Research Methodology, 22(1): 37–54. https://doi.org/10.1080/13645579.2018.1490981
 
Eckman S, Koch A (2019). Interviewer involvement in sample selection shapes the relationship between response rates and data quality. Public Opinion Quarterly, 83(2): 313–337. https://doi.org/10.1093/poq/nfz012
 
Groves RM, Peytcheva E (2008). The impact of nonresponse rates on nonresponse bias: A meta-analysis. Public Opinion Quarterly, 72(2): 167–189. https://doi.org/10.1093/poq/nfn011
 
Kalton G, Flores-Cervantes I (2003). Weighting methods. Journal of Official Statistics, 19(2): 81–97.
 
Kreuter F, Kohler U (2009). Analyzing contact sequences in call record data: Potential and limitations of sequence indicators for nonresponse adjustments in the European Social Survey. Journal of Official Statistics, 25(2): 203–226.
 
Kreuter F, Olson K (2013). Paradata for nonresponse adjustment. The Annals of the American Academy of Political and Social Science, 645: 175–190.
 
Rybak A (2023). Survey mode and nonresponse bias: A meta-analysis based on the data from the International Social Survey Programme waves 1996–2018 and the European Social Survey rounds 1 to 9. PLOS ONE, 18(3): e0283092. https://doi.org/10.1371/journal.pone.0283092
 
Schouten B, Peytchev A, Wagner J (2017). Adaptive Survey Design. Chapman & Hall / CRC, 1st edition.
 
Schouten B, van den Brakel JA, Buelens B, Giesen D (2021). Adaptive mixed-mode survey designs. In: Mixed-Mode Official Surveys, 251–272. Taylor & Francis / CRC Press.
 
Wuyts C, Durrant GB, Kreuter F (2023). Predicting fieldwork effort in face-to-face surveys using call record data. Journal of Survey Statistics and Methodology, 11(2): 367–391. https://doi.org/10.1093/jssam/smab036

Related articles PDF XML
Related articles PDF XML

Copyright
2026 The Author(s). Published by the School of Statistics and the Center for Applied Statistics, Renmin University of China.
by logo by logo
Open access article under the CC BY license.

Keywords
case prioritization machine learning paradata predictive modeling

Metrics
since February 2021
11

Article info
views

2

PDF
downloads

Export citation

Copy and paste formatted citation
Placeholder

Download citation in file


Share


RSS

Journal of data science

  • Online ISSN: 1683-8602
  • Print ISSN: 1680-743X

About

  • About journal
  • Renmin University of China homepage
  • Academic Journal Management
    and Development Center homepage

For contributors

  • Submit
  • OA Policy
  • Become a Peer-reviewer

Contact us

  • JDS@ruc.edu.cn
  • Contact person: Jing Zhou
  • Phone: +86-10-62511318
  • No. 59 Zhongguancun Street, Haidian District Beijing, 100872, P.R. China
Powered by PubliMill  •  Privacy policy