Research Projects
CD-CAT Item Selection Method
Educational assessment has increasingly focused on formative evaluation and diagnostic feedback. Cognitive Diagnostic Assessment (CDA) meets this need by identifying examinees' mastery of specific skills. To enhance efficiency, Cognitive Diagnostic Computerized Adaptive Testing (CD-CAT) integrates cognitive diagnosis with adaptive testing. While various item selection algorithms exist, their high computational costs limit practical use. This project aims to investigate methods to reduce the computational demands of CD-CAT algorithms while preserving estimation accuracy.
Project Members: Xiuxiu Tang, Ying Cheng
Robust Estimation of the Latent Trait in Item Response Theory Models
Aberrant responses (e.g., careless responses, miskeyed items, etc.) often contaminate psychological assessments and surveys. Previous robust estimators for dichotomous IRT models have produced more accurate latent trait estimates with data containing response disturbances. This project is a series of studies proposing a robust estimator for (1) the graded response model (GRM) for Likert-type polytomous data, (2) the multidimensional IRT (MIRT) model for multidimensional dichotomous data, and (3) the multidimensional graded response model (MGRM) for multidimensional Likert-type polytomous data. By using weighting mechanisms, such as the Huber and the bisquare weight functions, “suspicious” responses are downweighted to lessen their influence on latent trait estimates. The reduction in bias and stable standard errors suggest that the robust estimators are effective in counteracting the harmful effects of response disturbances and providing more accurate scores on psychological assessments.
Project Members: Audrey Filonczuk, Ying Cheng
Statistical Permutation Tests and Fairness in Machine Learning Classification Algorithms
Classification is a fundamental challenge in many disciplines. For example, workplaces may want efficient ways to classify who is qualified for a position, clinicians may want to predict if a patient will commit self-harm, and educators may want efficient ways of knowing which students need additional help and which ones are thriving on their own. However, when we use machine learning for these tasks, we run the risk of those models making biased decisions against certain groups of individuals.
This project involves using a statistical permutation test to determine if a model is “fair” for different groups of individuals. The permutation test works by taking all possible permutations of groups labels, calculating a metric of fairness, and comparing the resulting distribution with the metric derived from the true labels. This allows us to conduct an unbiased inference of fairness. Eventually, we will build methods to examine all permutations flagged by the distribution to test fairness without knowing group affiliation.
Project Members: Walton Ferguson, Ying Cheng
R Package for Statistical Control of Data Quality
Functions in this R package will be able to handle response disturbances in data. Some capabilities include robust estimation of latent traits, and change-point analysis. Visualization tools are implemented to allow users to select a tuning parameter and compare robust estimates with maximum likelihood estimates.
Project Members: Audrey Filonczuk, Meredith Sanders, Cheng Liu, Ying Cheng
The Applications of Response Time in Educational and Psychological Measurement
In educational and psychological measurement, response time provides valuable insights beyond traditional accuracy scores. This Project explores three methodological approaches to improve the analysis of response time data in high stakes testing environments: (1) First, we develop a nonparametric regression approach to evaluate the fit of parametric response time models. By constructing smooth curves based directly on data, we can compare them with model-implied parametric curves to quantify departures from expected patterns. This method helps identify violations of parametric assumptions through visual inspection and statistical significance tests using parametric bootstrap; (2) Second, we propose a censored distribution for response time at the item level to address cases where responses are not observed due to time constraints. By incorporating tools from survival analysis, we develop a model that accounts for right-censored response times when test-takers reach time limits without responding. This approach improves the estimation of working speed parameters particularly for slower test-takers who are most affected by time limits; (3) Third, we employ L1 and L2 penalties to simultaneously test all test items for Differential Item Functioning. This approach addresses multiple testing issues and removes the assumption that non-tested items are DIF-free. Additionally, we implement permutation tests to determine variable importance when penalties do not shrink coefficients exactly to zero. Together, these three projects contribute to the methodological development of response time for high stake educational assessments.
Project Members: Qizhou Duan, Ying Cheng
Assessing ChatGPT's Proficiency in Generating Accurate Questions and Answers for Statistics Education
This project investigates the accuracy and potential of ChatGPT in generating questions and solutions for STEM subjects, with a particular focus on Advanced Placement (AP) Statistics. While AI tools like ChatGPT have shown promise in enhancing the educational experience, such as generating lesson plans or acting as virtual tutors, there is limited research on the reliability of AI-generated content, particularly in question creation and solution accuracy. Previous studies have produced mixed findings, with some indicating low accuracy in multiple-choice question generation, while others suggest high performance in specific contexts like medical exams. This study aims to evaluate ChatGPT 3.5’s ability to generate accurate and relevant questions and solutions in the context of AP Statistics, providing insights into its potential as a supportive tool for teachers preparing assessments.
Project members: Meredith Sanders , Sarah McDevitt, Nancy Le, Alison Cheng
AI-assisted Rubric Grading
This study examines the effectiveness of generative AI tools, such as ChatGPT, in assessing student responses to a college-level physics exam question. Specifically, it aims to evaluate ChatGPT’s ability to generate rubric-based scores and provide meaningful feedback. By comparing the AI-generated evaluations with those of content experts, the research will assess the tool’s grading accuracy, consistency, and potential biases. Additionally, the study will explore how ChatGPT can contribute to refining rubrics, enhancing human grading efficiency, and improving the quality of student feedback. The findings will provide insights into the role of AI in educational assessment and its implications for instructional support.
Project Members: Xiuxiu Tang, Ying Cheng, G. Alex Ambrose
AI Workforce in Higher Education
As AI continues to permeate various industries, there is a growing need to assess the impact of AI-related skill sets in the workplace. This project aims to evaluate the relevance of these skills in higher education job markets. To do so, job advertisements were collected from platforms like The Chronicle of Higher Education, HigherED Jobs, and Inside Higher ED Careers, filtered by specific keywords related to AI skills. Web scraping methods, including Python libraries such as BeautifulSoup and Selenium, will be employed to extract and analyze data from these job listings. Ultimately, this research seeks to understand how job requirements are evolving in response to the rise of artificial intelligence.
Project members: Alison Cheng, Nancy Le, Meredith Sanders , Cheng Liu
Performance Analytics for College STEM Education
This study explores whether non-thriving students—those at risk of earning C, D, F, or W grades—can be accurately identified early in the semester in a foundational STEM course. By analyzing pre-course preparation and early performance data from the first six weeks, the study will develop predictive models to support early interventions. The findings will help instructors provide timely, targeted support to at-risk students.
Project Members: Xiuxiu Tang, Ying Cheng, G. Alex Ambrose
Contextualising AI ethics in Higher Education: Comparing the Ethical Issues Raised by Large-Scale Models in Higher Education Across Countries and Subject Domains
The integration of artificial intelligence (AI) into higher education has sparked significant discussions about ethical use, governance, and best practices, leading many institutions to develop policies on the role of large-scale models (LSMs), such as ChatGPT, in the classroom. This project, a collaboration between the ND and UK teams, investigates the ethical issues these policies raise, such as academic misconduct and bias, and examines how they are interpreted across various disciplines, including Technology, Education, Physical Science, and Social Science. Structured into three phases, the project began with a literature review of institutional LSM policies (Phase 1), followed by interviews with faculty to assess the impact of LSMs on their practices (Phase 2). Phase 3 will extend this analysis through surveys of faculty. The goal is to provide insights that help higher education institutions integrate LSMs in ways that enhance teaching and learning.
Project members: Wayne Homes, Alison Cheng, Nancy Le, Meredith Sanders