A Multimodal Machine Learning Framework for Bidirectional Indian Sign Language Conversion in Support of Inclusive Education
Main Article Content
Abstract
Indian Sign Language (ISL) is the native language of millions of Deaf and hard-of-hearing (DHH) people in India; however, the acute dearth of professional sign language interpreters and educational resources poses an obstacle to inclusion through accessible instruction. This research suggests a multimodal machine learning approach for bidirectional translation between Indian Sign Language — text/speech-to-ISL avatar synthesis and ISL-to-text/speech transcription — and solves four problems in literature — continuous sentence-level recognition, incorporation of non-manual facial signs, natural-looking co-articulation in avatars, and overall bidirectional architecture. Since there are no real-life video data and GPU computing facilities available, the architecture was built from scratch and tested using synthetic closed-vocabulary landmark dataset; all results must be interpreted as proof of concept evidence rather than actual performance. On average over five experiments, fusion of hand, face, and pose channels resolved recognition errors which hand-only model failed to fix (word error rate 0.0% ± 0.0% vs. 9.5% ± 1.6%; accuracy 100.0% ± 0.0% vs. 49.7% ± 1.4% in non-manual minimal pair), and learning co-articulation achieved 12.8% ± 2.2% improvement of avatar motion jerkiness. These findings motivate, but do not establish, extension to real corpora and evaluation with DHH learners before classroom use.


