Automated Essay Scoring In Educational AI: A Review Of Fairness, Bias, And Generalization
Keywords:
Automated Essay Scoring, Educational AI, Fairness, Bias, Generalization, Large Language Models, Transfer LearningAbstract
The automated essay scoring (AES) is a key use case of educational AI, and there has been an increasing adoption of this in the fields of formative support, large-scale assessment, and feedback-heavy writing instruction. While transformer-based models and large language models (LLMs) have shown potential for enhancing the technical properties of AES, they have also raised issues of fairness, bias, and generalization across prompts, populations, and educational settings. This review explores the shift from feature-based and hybrid models to neural and LLM-based systems and the importance of correctly deploying models that are trustworthy. The article is organized through the lens of literature review using PRISMA, which reviews studies on the basics of AES, benchmark corpora, fairness-centered analyses, transfer-learning studies, and newer studies of LLLM assessments. The review reveals the influence of benchmark datasets on methodological development, but also the limitation on the type of validity and equity claims that can be made. It also identifies a range of fairness outcomes that differ across model families, fairness measures, the composition of the training data and across various demographic subgroups, and transfer and domain-shift studies show that fairness limits are enduring across prompts and students. While the possibilities of scalability and richer contextual models are new, there are also concerns such as calibration issues, representation bias, and opacity and dependence on human-like yet unstable judgements in LLM-based AES. Significant areas of weakness in the following areas remain to be addressed: fairness portability, multimodal assessment, intersectional evaluation, and benchmark design for real-world educational settings. Future AES research should combine fairness-aware modeling, strong evaluation, and educational validity into a unified framework to enable systems that are accurate, fair, transparent, and trustworthy.




