Statistical Analysis and Machine Learning Integration for Disability Classification in Mexico: Comparative Methods for Public Health Analysis Based on the 2020 INEGI Population and Housing Census
Keywords:
Disability; data mining; machine learning; ensemble methods; INEGI; 2020 Population and Housing Census; territorial analysis; population aging; clustering; decision tree; predictive modeling; public policies.Abstract
Disability constitutes a complex sociodemographic phenomenon associated with territorial inequalities, population aging, differential access to services, and social vulnerability. The present study aimed to conduct a comprehensive statistical, territorial, and machine learning analysis of the population with disabilities in Mexico using official information from the 2020 Population and Housing Census conducted by the National Institute of Statistics and Geography (INEGI). A total of 660 valid records corresponding to the population with disabilities, activity limitations, and mental condition problems were analyzed, disaggregated by federal entity, five-year age group, and sex.
Two complementary analytical approaches were employed: (1) Traditional statistical methods including descriptive statistics, comparative analysis, inferential tests, correlations, outlier detection, and classical clustering through K-Means, hierarchical clustering, and DBSCAN; (2) Advanced machine learning techniques including supervised classification (Random Forests, Gradient Boosting, Support Vector Machines) and ensemble methods to enhance predictive performance and validate statistical findings.
The results showed that, within the analyzed universe, the population with disabilities represented 29.66%, the population with limitations represented 66.87%, and the population with mental conditions represented 7.63%. Disability was concentrated mainly among people aged 60 years and older, who represented 50.06% of the national total of people with disabilities. Territorial analysis showed that the State of Mexico, Mexico City, Veracruz, Jalisco, and Puebla concentrated the highest absolute numbers; however, Tabasco, Chiapas, Oaxaca, Guerrero, and Zacatecas presented the highest relative proportions.
The decision tree model (traditional statistical approach) achieved accuracy of 0.906 and F1-score of 0.909. Machine learning ensemble methods, specifically Stacking, achieved superior performance with accuracy of 0.950 and F1-score of 0.949, demonstrating that complex non-linear relationships enhance territorial vulnerability classification. Both methodological approaches identified the same three territorial clusters and confirmed the proportion of limitations as the most discriminative variable (r=−0.942; permutation importance=0.523).
It was concluded that the integration of traditional statistical methods with advanced machine learning approaches provides a more comprehensive and robust characterization of disability in Mexico as a territorially unequal and demographically conditioned phenomenon, providing useful evidence for the design of public policies focused on inclusion, accessibility, and support for vulnerable groups.





