An NLP Pipeline for Bangla Text Understanding and Linguistic Data Processing: Framework and Implementation
Asian Journal of Language, Literature and Culture Studies · pp. 17–38 · Published 19 Jan 2026
10.9734/ajl2c/2026/v9i1297Abstract
Raw text is the most prevalent form of human language in digital and electronic formats. This research proposes a comprehensive Bangla language processing framework that transitions raw data into structured data and value-added information through clearly defined annotation guidelines. Unlike existing fragmented approaches, this work offers a unified treatment of corpus development and annotation. It specifically detailed each processing phase and its input-output specifications. The pipeline focuses primarily on the text-understanding components and integrated essential tasks such as Parts of Speech (PoS) tagging, parsing and Named Entity Recognition (NER) etc. Moreover, to establish the semantic state of their linguistic inputs, the framework includes coreference resolution and word sense disambiguation. This end-to-end pipeline is designed for several different uses, including high-precision sentiment analysis, automated content moderation, and developing gold standard datasets that can be used in advanced Bangla NLP research.
Cited by 0
No indexed citations yet.
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
0
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.