A CNN-Based Hybrid Approach for Semantic Textual Similarity in Kannada–English Language Pair

Main Article Content

Megha V, Pooja M R

Abstract

Semantic similarity estimation across a pair sentence is one among the major concerns for NLP, and Information retrieval, machine translation, text summarization, and automatic evaluation are just a few of its many uses. While much work has been done in monolingual settings, (STS) cross-lingual semantic textual similarity remains an open issue, particularly when considering low-resource language pairs such as Kannada-English. This research proposes a semantic textual similarity assessment method using a hybrid technique of a brief English and Kannada text sentence pairs by considering lexical, syntactic, and semantic aspects. The approach incorporates word-based similarity measures along with lexical resources and vector-based similarity measures along with word embeddings. Cross-lingual word embedding alignment is done by utilizing VecMap and MUSE. A lexical decomposition approach is used to decompose similar and dissimilar semantic components. Similarity in word order is incorporated to record syntactic structure. Finally, a scoring-level fusion method is employed to generate similarity score. Moreover, CNN model is also applied for better classification results. Efficiency of the model is evaluated based on Samanantar Parallel Corpus, which contains sentence pairs in different languages generated through AI4Bharat.The hybrid method outperforms the standard word-level and embedding-level techniques, as shown in the experimental results. Furthermore, It has been noted that the embedding alignment approach performs better as the size of the bilingual dictionaries, and MUSE outperforms other techniques for large dictionaries. The proposed framework is effective for cross-lingual semantic similarity measurement in low-resource scenarios and can be extended to other multilingual tasks pertaining to NLP such as automatic grading, translation systems in machines, and cross-language IR systems.

Article Details

How to Cite
Megha V, Pooja M R. (2026). A CNN-Based Hybrid Approach for Semantic Textual Similarity in Kannada–English Language Pair. International Journal of Special Education, 41(4s), 689–708. Retrieved from https://internationalsped.com/index.php/ijse/article/view/2935
Section
General