A Hybrid Stroke-Aware CNN–Vision Transformer Framework with Contrastive Learning for Robust Handwritten Devanagari Character Recognition
Main Article Content
Abstract
The automated recognition of handwritten Devanagari script poses formidable challenges, driven primarily by intricate stroke typologies, pronounced inter-class morphological similarities, and idiosyncratic stylistic variations among writers. Although Convolutional Neural Networks (CNNs) are highly adept at extracting localized spatial representations, they demonstrate a conspicuous deficit in global contextual awareness. Conversely, Vision Transformers (ViTs) excel in establishing global receptive fields via self-attention mechanisms, yet remain hindered by an attenuated inductive bias. To reconcile these architectural dichotomies, we introduce a novel, hybrid computational framework that synergizes CNNs and ViTs, further augmented by stroke-aware attention mechanisms and contrastive learning paradigms.
Through rigorous empirical evaluation, we demonstrate a progressive trajectory of performance optimization, elevating recognition accuracy from a baseline of 83% to a robust 94% across varied architectural configurations. The proposed model exhibits superior generalizability, yielding optimal outcomes when conditioned on higher-resolution inputs and governed by adaptive optimization schemas. Crucially, our findings elucidate that asymptotic performance limits are fundamentally bounded by the fine-grained morphological proximity of character strokes, rather than by constraints in intrinsic model capacity. Ultimately, this study underscores the indispensable role of hierarchical feature representation and explicit structural modelling in advancing the recognition of complex orthographic systems.


