Phonology as Foundational Infrastructure: A Methodological Framework and Novel Architecture for Arabic Large Language Models
Main Article Content
Abstract
This paper argues that the orthography-centric approach dominating Arabic Large Language Model (LLM) development fails to capture the language's core morpho-phonological nature, where meaning is constructed through systematic sound patterns largely absent from unvowelized text. This oversight results in models that are semantically shallow, dialectally inflexible, and prosodically impoverished. We propose a paradigm shift: treating phonology as foundational infrastructure. The contribution is twofold. First, we introduce a comprehensive methodological framework for multi-modal data curation and phonological representation, encompassing segmental, articulatory-feature, prosodic, and dialectal tiers. Second, we present the novel Arabic Phonology-Aware Transformer Architecture (APATA), a dual-stream model that integrates orthographic and phonological encoders with cross-modal fusion. APATA is trained in a suite of auxiliary tasks—including masked phoneme modeling and root prediction—to internalize Arabic's phonological grammar. We detail expected empirical validation, projecting significant gains in diacritization, speech processing, dialect generalization, and machine translation. The paper concludes that phonological integration is essential for genuine Arabic language intelligence, outlining challenges and future directions like low-resource dialect adaptation and neuro-symbolic integration. This work reframes LLM development for Arabic, aiming to build systems that comprehend the intrinsic link between sound and meaning..


