Translate this page into:
Modeling stunting risk with multicollinear categorical predictors using a CATPCA-based binary logistic spline framework
* Corresponding author: E-mail address: annaislamiyati701@gmail.com (A Islamiyati)
-
Received: ,
Accepted: ,
Abstract
This study integrates categorical principal component analysis (CATPCA) with nonparametric binary logistic spline regression to simultaneously address multicollinearity and predictor heterogeneity. The proposed CATPCA-based binary logistic spline regression approach is designed to handle heterogeneous, categorical, and correlated predictors. Optimal knot selection in the spline component is determined using the minimum generalized cross validation (GCV) criterion. Model performance is evaluated using deviance, accuracy, and the akaike information criterion (AIC). Simulation studies are conducted to compare models with and without CATPCA, followed by an application to real stunting-risk family data consisting of categorical predictors. Model parameters are estimated using maximum likelihood estimation based on CATPCA-derived component scores. Simulation results show that the CATPCA-based binary logistic spline model yields lower deviance and AIC values than the model without CATPCA, indicating improved model fit. Application to family stunting-risk data from South Sulawesi Province, Indonesia, identifies five principal components that capture both linear and nonlinear patterns of predictor influence on the response. The integrated CATPCA-based binary logistic spline regression model effectively explains the influence of heterogeneous and interdependent predictors on the response. The combination of spline-based nonparametric modeling and CATPCA-driven dimensionality reduction enhances the model’s ability to capture complex and detailed relationships between family-level determinants and stunting risk.
Keywords
CATPCA
Categorical
Logistic spline
Stunting
1. Introduction
Multicollinearity, defined as a high correlation among predictor variables, can occur in both numerical and categorical data and may lead to unstable coefficient estimates, reduced interpretability, and lower predictive accuracy (Midi et al., 2010). This issue arises across various domains, including public health, where it may obscure key risk factors (Islamiyati et al., 2025), and materials science, where high-dimensional interactions can create complex interdependencies that pose challenges similar to those encountered in statistical models (Mishra et al., 2004). In logistic regression, various approaches have been proposed to address multicollinearity, including logistic principal component analysis (PCA) (Aguilera et al., 2006), sparse PCA (Lee et al., 2010; Erichson et al., 2020), Liu-type estimators (Sudjai and Duangsaphon, 2020; Ertan, E., and Akay, K.U., 2020), and penalized methods such as the least absolute shrinkage and selection operator (LASSO) and ridge regression (Bayman and Dexter, 2021; Šinkovec et al., 2021; Lukman et al., 2023, 2025.). However, these approaches largely remain within parametric modeling framework for quantitative predictors and have not fully explored dimensionality reduction with nonparametric regression approaches. Moreover, most existing studies represent categorical predictors using dummy coding, which may further exacerbate multicollinearity in high-dimensional settings.
To address these limitations, more flexible modeling approaches have been explored. With increasing data complexity, nonparametric regression has gained attention due to its ability to model nonlinear relationships without imposing strict functional forms. Several studies have incorporated penalty functions or weighting schemes to improve estimator stability under multicollinearity (Lestari et al., 2010; Chamidah et al., 2012; Mardianto et al., 2019; Islamiyati et al., 2020a; Islamiyati et al., 2020b; Hidayat et al., 2021; Islamiyati, 2022). For categorical responses, spline-based approaches such as generalized additive models have also been developed within a nonparametric framework (Wood et al. [2016]). Among these approaches, binary logistic spline regression is particularly useful because it captures nonlinear relationships while preserving the interpretability of logistic regression. Similar complexity in reactive and functional polymer systems underscores the need to integrate dimensionality reduction with nonparametric modeling.
Recent developments further integrate regularization with spline-based models to improve robustness under high correlation, including ridge-regularized splines (Zhang et al., 2019), adaptive penalties (Yu and Ruppert, 2021), LASSO-based splines (Mullah et al., 2021), structural penalties (Li, R., Liang, H., 2008), and double-penalized approaches (Guan and Fu, 2022). However, these methods generally assume homogeneous predictors and do not explicitly address the additional complexity introduced by categorical variables. In practice, categorical variables are often represented using dummy coding, which increases dimensionality and may exacerbate multicollinearity. To overcome this limitation, categorical principal component analysis (CATPCA) transforms categorical variables into orthogonal continuous components through optimal scaling while preserving their underlying structure (Linting et al., 2007; Chavent et al., 2017; Meulman et al., 2019; Abou-Senna et al., 2022). By reducing dimensionality and redundancy, CATPCA provides more stable inputs for regression modeling and has been widely applied in finance, health, and social sciences (Kemalbay and Korkmazoğlu, 2014; Lê et al., 2018; (Greenacre, 2018); Atkinson, 2024). Nevertheless, its integration with spline-based regression for binary outcomes remains limited.
This study analyzes family stunting-risk data from South Sulawesi Province, Indonesia, so the findings are region-specific. Previous spline-based studies on stunting (Islamiyati et al., 2022; Arifin et al., 2023; Islamiyati et al., 2024; Fadil et al., 2025; Rachmi et al., 2016) have not adequately addressed multicollinearity among categorical predictors. To fill this gap, we propose a CATPCA-based binary logistic spline framework that transforms multicollinear categorical predictors into orthogonal components and parametric and nonparametric effects. The novelty lies in the simultaneous treatment of multicollinearity, categorical heterogeneity, and parametric-nonparametric effects within a unified framework. Simulation and empirical results show improved performance over conventional and penalized logistic models. The remainder of the paper is organized as follows: Section 2 presents the methodology, Section 3 the results, and Section 4 the conclusions.
2. Materials and Methods
2.1 Data source
The study employs both simulation data and real-world data. The simulation dataset was generated for several sample sizes, where the binary response variable follows a binomial distribution, and five categorical predictor variables were constructed to exhibit strong intercorrelations.
For the real data analysis, we used family stunting-risk data from South Sulawesi Province, Indonesia, collected in 2023. The dataset consists of 23,150 families obtained from the National population and family agency of south Sulawesi Province. The study variables are presented in Table 1.
| Variables study | Variable name | Category |
|---|---|---|
| Families at risk of stunting | 1 = Yes | |
| 0 = No | ||
| Families with toddlers | 1 = Yes | |
| 2 = No | ||
| Main source of drinking water | 1 = Eligible | |
| 2 = Not eligible | ||
| Source of sanitation facility | 1 = Eligible | |
| 2 = Not eligible | ||
| Families with underage reproductive-age couples | 1 = Yes | |
| 2 = No | ||
| 3 = Not applicable | ||
| Families with older reproductive-age couples | 1 = Yes | |
| 2 = No | ||
| 3 = Not applicable | ||
| Families with closely spaced reproductive-age couples | 1 = Yes | |
| 2 = No | ||
| 3 = Not applicable | ||
| Families with multiple reproductive-age couples | 1 = Yes | |
| 2 = No | ||
| 3 = Not applicable | ||
| Family welfare ranking | 0 = Family not yet identified level his welfare | |
| 1 = Rank welfare 1 | ||
| 2 = Rank welfare 2 | ||
| 3 = Rank welfare 3 | ||
| 4 = Rank welfare 4 | ||
| 5 = Rank well-being > 4 | ||
| Families receiving social assistance mentoring | 1 = Yes | |
| 2 = No |
2.2 Analysis method
2.2.1 CATPCA
CATPCA is applied to transform a set of mutually correlated categorical predictors into optimally quantified numerical variables. In the CATPCA procedure, categorical variables are classified according to their measurement level. Nominal variables are transformed using unrestricted optimal scaling, whereas ordinal variables are quantified under monotonic constraints to preserve their inherent ordering. Let variable () denote a categorical variable with categories. The quantified version of is denoted by , and the relationship between the indicator matrix and its quantification is given by:
Let denote the matrix of quantified predictors. The optimal quantification is obtained by minimizing the loss function (Linting et al., 2007), defined at iteration as in Eq. (1).
The optimal value of is obtained by minimizing . The iteration continues until convergence, defined as:
Upon convergence, the final quantifications are denoted by , yielding the quantified data matrix . Principal components are then extracted from for further modeling.
2.2 CATPCA binary logistic spline regression
Let the binary response variable be , for . Let denote the -th principal component score for observation , where . Each component is defined as in Eq. (2).
where is the eigenvector associated with the -th component. To model the probability of success, define . A nonparametric logistic regression model with truncated spline basis functions is employed:
with the linear predictor given by Eq. (3)
where denotes the -th knot point , and
is the truncated linear spline basis function. The coefficients represent parametric effects, while capture nonparametric spline effects (Islamiyati et al., 2023).
The number and location of knots are selected using the generalized cross validation (GCV) criterion to capture local variations between predictors and response. This reflects a structured transformation where categorical variables are converted into continuous representations for flexible modeling. Similar principles are observed in macromolecular and colloidal systems, where discrete components assemble into functional structures (Kim et al., 2003). CATPCA thus integrates heterogeneous categorical predictors into a unified framework prior to spline-based modeling.
3. Results and discussion
3.1 Model parameter estimation
Heterogeneous or categorical predictor variables are first transformed into optimally quantified numerical variables through the CATPCA procedure, as described in Eq. (1), resulting in the quantified matrix . From this matrix, a set of principal components is extracted and used as predictors in the regression model. The binary response variable is denoted by , for . The relationship between the predictors and the probability of success is modeled using a binary logistic regression with truncated linear spline basis functions. The model can be written as Eq. (4) below.
where is the intercept, denotes the regression coefficient for the -th principal component, represents the coefficient associated with the truncated spline basis, and r denotes the number of knot points.
The following estimation procedure is derived from the proposed model using the maximum likelihood framework Eq. (5).
The corresponding log-likelihood function is in Eq. (6).
Let be the response vector and let denote the vector of predicted probabilities. Furthermore, define as the design matrix that includes a column of ones (intercept), the principal components , and the truncated spline basis terms .
The first derivative (gradient or score function) of the log-likelihood can then be expressed in matrix form as Eq. (7)
To obtain the parameter estimates, the Newton–Raphson iterative method is applied. The update formula is given by Eq. (8)
where is a diagonal weight matrix with elements:
The iteration continues until convergence is achieved, yielding the optimal parameter estimates for the CATPCA-based binary logistic spline regression model.
3.2 Simulation study
In this study, a simulation was conducted using observations, a binary response variable, and five categorical predictors, three of which were generated to exhibit strong intercorrelation . The initial model exhibited variance inflation factor (VIF) values exceeding 15 and high coefficient variance. Although the deviance explained reached 68%, the model had a relatively large akaike information criterion (AIC) of 210.4.
For benchmark comparison, penalized logistic regression models, namely LASSO and ridge regression, were evaluated using cross-validated tuning parameters. Ridge regression yielded an AIC of 185.7 with 72.5% deviance explained, while LASSO achieved an AIC of 178.9 with 74.2% deviance explained. Although both methods improved model stability, they remain within a parametric framework and are less capable of capturing complex relationships compared to the proposed CATPCA-based spline model.
After applying CATPCA, the proposed spline-based logistic model showed substantial improvement in all evaluation metrics. The coefficient variances decreased substantially, and all VIF values fell below 4, indicating effective mitigation of multicollinearity. Model performance improved significantly, with the AIC decreasing to 142.6 and the deviance explained increased to 80.3%. To explicitly compare the performance of models with and without CATPCA, as well as that of penalized logistic alternatives, the results are summarized in Table 2. The proposed CATPCA-based spline model achieved the lowest AIC and the highest deviance explained among all competing models.
| Model | AIC | Deviance explained |
|---|---|---|
| Logistic (without CATPCA) | 210.4 | 68.0% |
| Logistic (Ridge) | 185.7 | 72.5% |
| Logistic (LASSO) | 178.9 | 74.2% |
| CATPCA-based logistic spline | 142.6 | 80.3% |
To assess robustness beyond in-sample criteria, 5-fold cross-validation was performed across all models. The CATPCA-based spline model achieved the highest accuracy (0.78) and area under the curve (0.82) with the lowest prediction error, indicating strong out-of-sample generalization performance. The improvement achieved through CATPCA can be interpreted as an optimization of the predictor space, where correlated variables are transformed into more stable representations. Similar optimization principles are observed in materials chemistry and solid-state systems, where structural refinement enhances functional properties and system stability.
3.3 Analysis of stunting data using the nonparametric logistic spline CATPCA regression model
The data analyzed consisted of 23,150 families at risk of stunting in South Sulawesi Province, Indonesia, of which 76.45% were not at risk, and 23.55% were at risk. Several predictor variables exhibited strong correlations, particularly among and (e.g., ), indicating substantial multicollinearity. To address this, CATPCA was applied to transform categorical predictors into optimally scaled numerical variables. The dimensionality reduction resulted in five principal components , explaining 81.45% of the total variance.
The relationship between quantified variables and principal components is represented by loading values , indicating the strength of association. Loadings range from -1 to 1, with considered significant, as shown in Table 3.
| 0.13 | -0.43 | -0.47 | 0.67* | 0.18 | |
| 0.03 | -0.28 | 0.65* | 0.44 | -0.54* | |
| -0.02 | -0.39 | 0.60* | -0.06 | 0.67* | |
| 0.97* | -0.04 | -0.01 | -0.04 | -0.01 | |
| 0.94* | -0.04 | -0.01 | -0.04 | -0.01 | |
| 0.96* | -0.04 | -0.01 | -0.04 | -0.01 | |
| 0.95* | -0.03 | 0.01 | -0.04 | -0.01 | |
| -0.21 | -0.73* | -0.21 | -0.06 | 0.01 | |
| 0.10 | 0.64* | 0.11 | 0.53* | 0.30 |
Overall, the integration of CATPCA and binary logistic spline regression provides a flexible and informative framework. Environmental sanitation factors exhibit both parametric and nonparametric effects on stunting risk, whereas socioeconomic factors are primarily nonparametric, and demographic variables show predominantly parametric effects. The final model was constructed using the five CATPCA components, with the spline specification selected based on the minimum GCV criterion, as shown in Table 4.
| Knot Points | GCV | ||||||
|---|---|---|---|---|---|---|---|
| 1-knot point | 1.65 | 0.61 | 4.19 | -0.50 | -0.01 | 0.11 | |
| 2-knot points | -0.42 | -0.15 | 0.66 | 0.53 | -1.26 | 0.15 | |
| -0.33 | 0.59 | 1.27 | 0.60 | 0.24 | |||
| 3-knot points | -0.45 | -0.62 | -0.62 | 0.03 | -0.37 | 0.17 | |
| -0.42 | -0.09 | -0.25 | 0.39 | 0.20 | |||
| 2.44 | 1.15 | -0.02 | 0.45 | 0.22 |
Table 4 shows that the lowest GCV value (0.11) was obtained with a one-knot point. Accordingly, the optimal model selected was the CATPCA-based binary logistic spline regression model with one knot. The overall significance of the model was assessed using the Likelihood Ratio Test. The test statistic obtained was (p < 0.05). Thus, the model parameters collectively significant explanation for the variation in family stunting risk.
Furthermore, the significance of individual parameters was assessed using the Wald test. All coefficients were statistically significant at the 5% level indicating that each predictor contributes meaningfully to the probability of a family being at risk of stunting. The resulting optimal one-knot model is given in Eq. (9).
Eq. (9) shows that the five principal components meaningfully explain the risk of family stunting (AIC = 165; accuracy = 86.41%). The first component is associated with representing household demographic structure, particularly factors related to women of reproductive age ( and ). The third and fifth components primarily reflect environmental sanitation through access to drinking water and sanitation facilities ( and ), while the fourth component captures variation related to children under two years of age and social support ( and ). The second component represents socioeconomic conditions, including family welfare and assistance programs ( and ).
Overall, the loading structure enhances model interpretability by identifying key dimensions of stunting risk. The results highlight complex relationships among demographic, environmental, and socioeconomic factors. Methodologically, the proposed framework integrates CATPCA and binary logistic spline regression to address multicollinearity, categorical heterogeneity, and nonlinear predictor effects simultaneously.
4. Conclusions
This study proposes a CATPCA-based logistic spline framework for modeling stunting risk with multicollinear categorical predictors using data from South Sulawesi Province, Indonesia. The main contribution lies in integrating integration dimensionality reduction and nonparametric regression into a unified framework, enabling the simultaneous handling of multicollinearity, categorical heterogeneity, and parametric–nonparametric effects. The proposed model demonstrates improved stability and predictive performance compared with conventional and penalized logistic approaches, as supported by simulation and cross-validation results.
The findings indicate that environmental, socioeconomic, and demographic factors play distinct roles in stunting risk. However, the analysis is limited to a single regional dataset and specific model assumptions, which may affect generalizability. Future work should extend the framework to broader datasets and explore alternative spline specifications and machine learning-based dimensionality reduction methods to enhance robustness and generalization. In practice, the proposed model can support public health decision-making by improving the early identification of high-risk households and providing a structured tool for stunting risk assessment.
CRediT authorship contribution statement
Anna Islamiyati: Conceptualization, methodology, data curation, software, Formal analysis, writing - original draft, writing - review & editing; Anisa Kalondeng: Conceptualization, investigation, writing - review & editing; Muhammad Nur: Conceptualization, investigation, writing - review & editing; Muhammad Afdal: Conceptualization, software; Ummi Sari: Resources, validation; Syafrina Abdul Halim: Conceptualization, resources, validation.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Declaration of generative AI and AI-assisted technologies in the writing process
The authors confirm that there was no use of artificial intelligence (AI)-assisted technology for assisting in the writing or editing of the manuscript and no images were manipulated using AI.
Acknowledgment
Many thanks to the Directorate General of Research and Development, Ministry of Higher Education, Science, and Technology, Indonesia, with research contract No. 02209/UN4.22/PT.01.03/2025 dated June 2, 2025.
Funding
Research Grant Fund of the Ministry of Higher Education, Science, and Technology, Regular Fundamental Research Scheme
References
- Categorical principal component analysis (CATPCA) of pedestrian crashes in Central Florida. J Transportation Saf & Secur. 2022;14:1890-1912. https://doi.org/10.1080/19439962.2021.1988788
- [Google Scholar]
- Using principal components for estimating logistic regression with high-dimensional multicollinear data. Comput Stat Data Anal. 2006;50:1905-1924. https://doi.org/10.1016/j.csda.2005.03.011
- [Google Scholar]
- Ability of ordinal spline logistic regression model in the classification of nutritional status data. Commun Math Biol Neurosci. 2023;2023:1-11. https://doi.org/10.28919/cmbn/8072
- [Google Scholar]
- Charting fields and spaces quantitatively: from multiple correspondence analysis to categorical principal components analysis. Qual and Quan. 2024;58:829-848. https://doi.org/10.1007/s11135-023-01669-w.
- [Google Scholar]
- Multicollinearity in Logistic Regression Models. Anesth Analg. 2021;133:362-365. https://doi.org/10.1213/ANE.0000000000005593
- [Google Scholar]
- Designing of child growth chart based on multi-response local polynomial modeling. J Mathematics and Statistics. 2012;8:342-347. https://doi.org/10.3844/jmssp.2012.342.347
- [Google Scholar]
- Chavent, M., Kuentz, V., Labenne, A., Liquet, B., Saracco, J., 2017. PCAmixdata: Multivariate analysis of mixed data. R package version 3.1. https://doi.org/10.32614/CRAN.package.PCAmixdata
- Sparse principal component analysis via variable projection. SIAM J Appl Math. 2020;80:977-1002. https://doi.org/10.1137/18m1211350
- [Google Scholar]
- A new Liu-type estimator in binary logistic regression models. Commun Statistics - Theory Methods. 2022;51:4370-4394. https://doi.org/10.1080/03610926.2020.1813777
- [Google Scholar]
- Classification of nutritional status in toddlers using the support vector machine method. Commun Math Biol Neurosci. 2025;2025:1-12. https://doi.org/10.28919/cmbn/9126
- [Google Scholar]
- Compositional Data analysis in practice (1st Edition). CRC Press; 2018. https://doi.org/10.1201/9780429455537
- A double-penalized estimator to combat separation and multicollinearity in logistic regression. Mathematics. 2022;10:1-19. https://doi. org/10.3390/math10203824
- [Google Scholar]
- The regression curve estimation by using mixed smoothing spline and kernel (MsS-k) model. Commun Statistics - Theory Methods. 2021;50:3942-3953. https://doi.org/10.1080/03610926.2019.1710201
- [Google Scholar]
- Spline longitudinal multi-response model for the detection of lifestyle-based changes in blood glucose of diabetic patients. Curr Diabetes Rev. 2022;18:98-104. https://doi.org/10.2174/1573399818666211117113856
- [Google Scholar]
- Penalized spline estimator with multismoothing parameters in biresponse multipredictor regression model for longitudinal data. Songklanakarin J Sci Technol. 2020a;42:897-909. https://doi.org/10.14456/sjst-psu.2020.115.
- [Google Scholar]
- Biresponse nonparametric regression model in principal component analysis with truncated spline estimator. J King Saud Univ Sci. 2022;34:1-9. https://doi.org/10.1016/j.jksus.2022.101892
- [Google Scholar]
- Detecting age prone to growth retardation in children through a bi-response nonparametric regression model with a penalized spline estimator. Iran J Nurs Midwifery Res. 2024;29:549-554. https://doi.org/10.4103/ijnmr.ijnmr_342_22
- [Google Scholar]
- Risk factor analysis for stunting incidence using sparse categorical principal component logistic regression. MethodsX. 2025;14:1-9. https://doi.org/10.1016/j.mex.2025.103186
- [Google Scholar]
- Use of two smoothing parameters in penalized spline estimator for bi-variate predictor non-parametric regression model. J Sci Islam Repub Iran. 2020b;31:175-183. https://doi.org/10.22059/JSCIENCES.2020.286949.1007435.
- [Google Scholar]
- The use of the binary spline logistic regression model on the nutritional status data of children. Commun Math Biol Neurosci. 2023;2023:1-11. https://doi.org/10.28919/cmbn/7935
- [Google Scholar]
- Categorical principal component logistic regression: A case study for housing loan approval. Procedia - Soc Behav Sci. 2014;109:730-736. https://doi.org/10.1016/j.sbspro.2013.12.537
- [Google Scholar]
- Electrorheological characteristics of phosphate cellulose-based suspensions. Polymer. 2001;42:5005-5012. https://doi.org/10.1016/S0032-3861(00)00887-9
- [Google Scholar]
- FactoMineR: An R package for multivariate analysis. J Stat Softw. 2018;25:1-18. https://doi.org/10.18637/jss.v025.i01
- [Google Scholar]
- Sparse logistic principal components analysis for binary data. Ann Appl Stat. 2010;4:1579-1601. https://doi.org/10.1214/10-AOAS327
- [Google Scholar]
- Spline smoothing for multi-response nonparametric regression model in case of heteroscedasticity of variance. J Mathematics and Statistics. 2012;8:377-384. https://doi.org/10.3844/jmssp.2012.377.384
- [Google Scholar]
- Variable selection in semiparametric regression modeling. Ann Statist. 2008;36:261-286. https://doi.org/10.1214/009053607000000604
- [Google Scholar]
- Nonlinear principal components analysis: Introduction and application. Psychol Methods. 2007;12:336-358. https://doi.org/10.1037/1082-989X.12.3.336
- [Google Scholar]
- K-L estimator: Dealing with multicollinearity in the logistic regression model. Mathematics. 2023;11:1-14. https://doi.org/10.3390/math11020340
- [Google Scholar]
- Handling multicollinearity and outliers in logistic regression using the robust Kibria–Lukman estimator. Axioms. 2025;14:1-29. https://doi.org/10.3390/axioms14010019
- [Google Scholar]
- Semiparametric regression based on three forms of trigonometric function in fourier series estimator. J Phys: Conf Ser. 2019;1277:012052, 1-10. https://doi.org/10.1088/1742-6596/1277/1/012052
- [Google Scholar]
- ROS regression: Integrating regularization with optimal scaling regression. Statist Sci. 2019;34:361-390. https://doi.org/10.1214/19-STS697
- [Google Scholar]
- Collinearity diagnostics of binary logistic regression model. J Interdisciplinary Mathematics. 2010;13:253-267. https://doi.org/10.1080/09720502.2010.10700699
- [Google Scholar]
- Fenugreek mucilage for solid removal from tannery effluent. React Funct Polym. 2004;59:99-104. https://doi.org/10.1016/j.reactfunctpolym.2003.08.008
- [Google Scholar]
- LASSO type penalized spline regression for binary data. BMC Med Res Methodol. 2021;21:1-14. https://doi.org/10.1186/s12874- 021-01234-9
- [Google Scholar]
- Stunting, underweight and overweight in children aged 2.0–4.9 years in Indonesia: Prevalence trends and associated risk factors. PLoS One. 2016;11:e0154756, 1-17. https://doi.org/10.1371/journal.pone.0154756
- [Google Scholar]
- To tune or not to tune, a case study of ridge logistic regression in small or sparse datasets. BMC Med Res Methodol. 2021;21:1-19. https://doi.org/10.1186/s12874-021-01374-y
- [Google Scholar]
- Liu-type logistic regression coefficient estimation with multicollinearity using the bootstrapping method. Science, Eng Health Stud. 2020;14:203-214. https://doi.org/10.14456/sehs.2020.19.
- [Google Scholar]
- Smoothing parameter and model selection for general smooth models. J Am Stat Assoc. 2016;111:1548-1563. https://doi.org/10.1080/01621459.20 16.1180986
- [Google Scholar]
- Adaptive penalized splines for logistic regression. Stat Sin. 2021;31:1419-1442. https://doi.org/10.5705/ss.202018.0123.
- [Google Scholar]
- Regularization methods for high-dimensional spline regression. J Comput Graph Stat. 2019;28:851-862. https://doi.org/10.1080/10618600.2019.1598875.
- [Google Scholar]
