| Time | Event |
|---|---|
| 7:00AM-8:00AM | Breakfast |
| 8:00AM-8:15AM | Opening Remarks |
| 8:15AM-9:15AM | Keynote: Prof. Barry Arnold On bivariate power-function distributions and related models |
| 9:15AM-9:30AM | Refreshment break |
| 9:30AM-10:50AM | Parallel Session 1 |
| 10:50AM-11:00AM | Refreshment break |
| 11:00AM-12:20PM | Parallel Session 2 |
| 12:20PM-1:20PM | Lunch |
| 1:20PM-1:55PM | Plenary talk: Prof. Ruth Pfeiffer Predicting Second Cancer Risk in Cancer Survivors: From Model Development to Clinical Validation |
| 1:55PM-2:30PM | Plenary talk: Prof. HK Tony Ng When Is a More Flexible Model Actually Better? Lessons on Model Complexity and Model Selection from Degradation Data Analysis |
| 2:30PM-2:40PM | Refreshment break |
| 2:40PM-4:00PM | Parallel Session 3 |
| 4:00PM-4:10PM | Refreshment break |
| 4:10PM-5:30PM | Parallel Session 4 |
| 5:30PM-6:30PM | Dinner |
| 6:30PM-7:30PM | Keynote: Prof. Nilanjan Chatterjee Integrating Disparate Studies: Transfer and Federated Learning Across Heterogeneous Model Spaces |
| Time | Event |
|---|---|
| 7:00AM-8:00AM | Breakfast |
| 8:00AM-9:00AM | Keynote: Prof. Thomas Mathew A Joint Confidence Set for a Ranking Based on Ordered Means |
| 9:00AM-9:10AM | Refreshment break |
| 9:10AM-10:30AM | Parallel Session 5 |
| 10:30AM-10:40AM | Refreshment Break |
| 10:40AM-12:00PM | Parallel Session 6 |
| 12:00PM-1:00PM | Lunch |
| 1:00PM-2:00PM | Keynote: Prof. Susmita Datta Deciphering Disease Mechanisms: Statistical Modeling of Cell–Cell Communication Using Spatial Transcriptomics |
| 2:00PM-2:15PM | Refreshment Break |
| 2:15PM-3:00PM | Poster Session | 3:00PM-4:00PM | Keynote: Prof. Narayanaswamy Balakrishnan Discrete-time Signatures |
| 4:00PM-4:15PM | Refreshment break |
| 4:15PM-5:00PM | Closing, gift draw, and group picture |
| Name | Advisor | Title | Abstract |
|---|---|---|---|
| Gyamfi Charles (University of Nevada, Reno; Ph.D. student) | Alexander Boateng | Analysis of COVID-19 cases and comorbidities using machine learning algorithms: A case study of the Limpopo Province, South Africa | This study examined the biological, social, and clinical risk factors for mortality among hospitalized COVID-19 patients in the five districts of Limpopo Province, South Africa. Four supervised machine learning algorithms, logistic regression, random forest, support vector machine, and decision tree were implemented and compared using 20,592 records with twenty-one attributes obtained from the Limpopo Department of Health. Due to class imbalance, the Random Over-Sampling Examples (ROSE) technique was applied. The dataset was divided into 70% training and 30% testing sets, while StepAIC reduced insignificant variables in logistic regression. Among the algorithms, random forest achieved the highest recall rate of approximately 79% in predicting mortality. Important predictors included age, ventilation, oxygenation, intensive ward admission, Waterberg district, and private facility type. The findings demonstrate the usefulness of machine learning algorithms, particularly random forest, in identifying mortality risk factors among hospitalized COVID-19 patients. |
| Pius Addi (University of Nevada, Reno; Ph.D. student) | Tomasz J. Kozubowski | A Truncated Multivariate Lomax Distribution for Modeling Dependent Bounded Data | We introduce a truncated multivariate Lomax (TML) distribution for modeling dependent bounded risks arising in applications such as reliability theory, actuarial science, finance, and other areas where correlated yet bounded data must be analyzed. The proposed model is constructed as a mixture of truncated exponential components with a tilted gamma mixing variable, and reduces to multivariate Lomax distribution (see, e.g., Nayak, 1987) as the truncation parameter approaches infinity. We derive several fundamental properties of this new stochastic model and develop computational procedures for parameter estimation based on the expectation-maximization (EM) algorithm. A simulation study is conducted to assess the finite-sample performance of the proposed estimators and to illustrate their consistency as the sample size increases. Overall, the proposed model offers a flexible framework for modeling dependent heavy-tailed phenomena under truncation and provides practical tools for statistical inference in applications involving bounded risks. |
| Yuna Han (Ball State University ; Masters student) | Drew Lazar | NeuralProphet for Ordinal Time Series Forecasting | Time series forecasting methods typically produce continuous-valued predictions, yet many real-world applications involve ordinal outcomes with natural ordering among categories. We propose a practical approach for ordinal time series forecasting using NeuralProphet, a neural network-based forecasting framework that decomposes time series into interpretable components including trend, seasonality, and autoregressive effects. Our method applies NeuralProphet to ordinal data treated as continuous, then converts predictions to ordinal categories via threshold validation through Nelder-Mead optimization. We evaluate this approach on real and simulated datasets. Results demonstrate that this approach achieves strong predictive performance and identification of trend, seasonality and auto regressive components that determine the response. We investigate the use of lagged and future covariates and the handling of initial observations that lack sufficient autoregressive history. Our findings suggest that NeuralProphet provides an effective and interpretable framework for ordinal time series forecasting without requiring specialized ordinal classification methods. |
| Riku Hosonuma (Department of Applied Mathematics, Tokyo University of Science; PhD student) | Tamae Kawasaki and Takashi Seo | Rao’s U-type Statistic Test for the Two Sample Problem with Two-step Monotone Missing Data | This study proposes a novel test statistic for two-sample tests of sub-mean vectors under a two-step monotone missing data structure. The proposed procedure is constructed based on Rao’s U-statistic framework and efficiently utilizes the available information in incomplete observations. We derive the asymptotic expansion of the null distribution of the statistic and obtain approximations to its upper percentiles. In addition, Bartlett corrections are developed to improve the chi-squared approximation in finite samples. The accuracy of the proposed approximations and correction methods is investigated through Monte Carlo simulations. Numerical examples are also presented to demonstrate the applicability and practical usefulness of the proposed methodology for statistical inference on sub-mean vectors with monotone missing data. |
| Tetsuya Sato (Department of Applied Mathematics, Tokyo University of Science; PhD student) | Ayaka Yagi and Takashi Seo | Large-sample asymptotic expansions for sphericity testing under monotone incomplete data | This study discusses the sphericity test for a one-sample problem. For complete data, Muirhead (1982) provided large-sample asymptotic expansions for the null distribution of the modified likelihood ratio test (LRT) statistic. In practical applications such as clinical trials, however, complete datasets are rarely available, and monotone incomplete data frequently arise due to subject dropouts. Despite its practical importance, research on the sphericity test under monotone incomplete data remains highly limited. To address this issue, assuming a multivariate normal distribution under monotone incomplete data, we derive large-sample asymptotic expansions for the null distributions of both the LRT and modified LRT statistics by utilizing the general framework proposed by Box (1949). Furthermore, we obtain the asymptotic expansions for the upper percentiles of these test statistics and propose several approximate upper percentiles. Finally, Monte Carlo simulations are conducted to numerically evaluate the empirical Type I error rates for the proposed approximations. |
| Tanner Miller (University of Nevada, Reno; Ph.D. student) | Tomasz Kozubowski and Anna Panorska | A Hierarchical Model for Precipitation Events in the Western United States | Precipitation events in the western United States exhibit substantial variability in duration, magnitude, and frequency, complicating statistical modeling and simulation. Existing approaches often emphasize aggregated precipitation totals, but event-level modeling is essential for understanding water resources and extreme events in a region affected by both drought and atmospheric rivers. This poster presents a hierarchical stochastic framework for precipitation events based on the Bivariate Exponential-Geometric (BEG) distribution, which jointly models event duration and magnitude. Using daily precipitation records from 286 weather stations across the western United States, we examine interannual variability in BEG parameters and develop a hierarchical model in which annual parameters follow a bivariate Normal distribution. The resulting framework enables realistic simulation of annual precipitation event sequences, reproducing observed heavy-tailed behavior and extremes. Extensions incorporating spatial dependence through Gaussian Random Fields and environmental covariates are also briefly discussed. |
| Opoku Esther (University of Nevada, Reno; Ph.D. student) | Anna Panorska | A Skewed Logistic Distribution via Normal Mean-Variance Mixtures | The logistic distribution plays a central role in statistics because it underlies logistic regression, one of the most widely used methods for modeling binary outcomes. It is especially convenient for probabilistic modeling and interpretation of odds and probabilities. The classical logistic model is symmetric, which limits its application. While there is an extension to a skewed logistic model, it does not yield itself to the much needed multivariate extensions. We address this gap by introducing a new approach for constructing a skewed logistic distribution using a normal mean-variance mixture representation, which naturally extends to the multivariate setting. We derive key distributional properties and develop methods for parameter estimation. Potential applications of the proposed model are also explored. In addition, we discuss computational challenges associated with the practical implementation of the model, arising from the lack of explicit form of several of its fundamental characteristics such as probability density and cumulative distribution functions which only have infinite series representations. |
| Hettige Fernando (University of Louisville; Ph.D. student) | Audrey Q. Fu | Comparison of Causal Network Inference Methods on the Relationship Between DNA Methylation and Transcription | DNA methylation, a universal epigenetic mechanism, is pivotal in regulating transcription and suppressing gene expression in various ways, and its interference is associated with numerous complex diseases. Graphical networks illustrate the statistical dependence among multiple variables and are widely used in biology, such as gene regulatory networks. In this study, we explored and compared two causal network inference applications to analyze the relationship between DNA methylation and transcription. To achieve this, we generalized the MRTrios package to handle different cancer types, aiming to gain insights into the underlying mechanisms of gene regulation. Our analysis involved studying the relationships between transcription and methylation in these cancer types using data from The Cancer Genome Atlas (TCGA) consortium and the Genomics Data Common portal (GDC). The formulated trios consist of the Copy Number Alteration (CNA) of a gene, the expression (E) of the gene, and the methylation (M) of a site located nearby or within the same gene. We then applied MRGN, a novel causal network inference method that considers many confounding variables under the principle of Mendelian randomization. Using the Bayesian Inference, we calculated the posterior probabilities of edges in the inferred models using genetic variants and the identified confounders in trios for one of the cancer types. Comparing the results of the causal networks obtained from the Machine Learning method and Bayesian Learning method, we observe that most of the causal models generated for each trio are similar for both methods with minor differences. Our comparative analysis highlights the strengths and limitations of each causal network inference method in underlying the complex mechanisms of DNA methylation and transcription. Our findings provide an advancing understand to researchers on selecting appropriate inference methodologies for analyzing regulatory networks. |
| Thao Le (James Madison University, BS) | Ali Tavasoli | Predicting Diabetes Risk Using Statistical and Machine Learning Methods | Diabetes affects millions of people around the world and catching it early can make a real difference in someone's health outcomes. This project explores the use of statistical and machine learning methods to classify diabetes status from routinely collected patient health measurements. Using the Pima Indians Diabetes Database (768 records, 8 clinical predictors), we carried out data cleaning, exploratory analysis, and a comparison of five classification models: Logistic Regression, Decision Trees, Random Forests, and Gradient Boosting, and XGBoost with SHAP applied to interpret the final model. Particular attention was given to implausible zero values in variables such as glucose, blood pressure, skin thickness, insulin, and body mass index (BMI), which were treated as missing values and imputed before modeling, two imputation strategies were compared, and the outcome-based strategy was shown to inflate the apparent predictive value of the most heavily imputed variables. Glucose concentration emerged as the strongest predictor of diabetes status, followed by BMI and age, a ranking reproduced consistently by correlation analysis, the decision tree, ensemble feature importance, and SHAP. Model performance was assessed using accuracy, precision, recall, F1-score, and ROC analysis, with care taken to look beyond overall accuracy given the moderate class imbalance in the data. The findings reported here are exploratory and specific to this dataset and population; they are intended as a learning exercise in the applied data science workflow rather than as a clinically validated tool. |