Enhancing Automatic Answer Evaluation Using Bi-Directional Long Short-Term Memory-Based Deep Learning
Main Article Content
Abstract
Automated Essay Scoring improves the effectiveness, consistency and fairness of evaluation through the use of deep learning models such as Bi-Directional Long Short-Term Memory in order to overcome some of the shortcomings associated with manual evaluation in large-scale educational testing situations. Automated Essay Scoring systems apply artificial intelligence and natural language processing technology to assess written responses quickly and without biasIt addresses the challenges in the education sector concerning the efficient evaluation of diverse and complex student responses due to the increasing number of learners. Conventional automated scoring approaches, which believe on lexical and syntactic pattern matching, fail to capture the deeper semantic and contextual nuances of short and essay-type answers. To resolve these limitations, the proposed framework employs a deep learning architecture based on the Automated Student Assessment Prize Short Answer Scoring, dataset, encompassing stages such as data exploration, feature extraction, model training, and evaluation. The system performs text structure and lexical feature analysis through visualization, while feature extraction involves tokenization, elimination of stop words, application of Term Frequency–Inverse Document Frequency, and truncated Singular Value Decomposition for effective dimensionality reduction. The comparative analysis between unidirectional Long Short-Term Memory and Bi-Directional Long Short-Term Memory demonstrates that the latter processes sequences in both forward and backward directions, thereby capturing contextual information more comprehensively and improving accuracy. The constructed Bi-Directional Long Short-Term Memory model, consisting of nineteen layers with dropout mechanisms to avoid overfitting, exhibits superior performance by achieving lower mean absolute error and validation loss compared to the conventional Long Short-Term Memory model. The inclusion of callback functions during training allows dynamic performance monitoring and parameter optimization. Positioned within the broader context of automated essay scoring research, this study highlights the significance of deep learning and context-aware architectures in enhancing grading accuracy beyond simple lexical and syntactic evaluation. The findings reveal that the Bi-Directional Long Short-Term Memory based scoring system is an effective and reliable method that closely aligns with human assessment, offering a robust solution for modern educational evaluation.


