•  
  •  
 

Document Type

Original Article

Abstract

Arabic Natural Language Processing (NLP) has made great strides, but the accurate extraction of mental state and physiological symptoms from patient-generated clinical consultations in dialectal Arabic, which is plagued with significant morphological complexity, colloquial noise and an extremely limited number of manually annotated clinical benchmarks, remains a significant challenge. In this work, we propose an end-to-end token-level symptom extraction system that is systematic and grounded in distant supervision paradigm and a strict gold standard evaluation with human verification to ensure real-world clinical generalization. A large weakly labelled corpus of 46,873 Arabic patient questions was created based on a clinical lexicon of 152 psychological and somatic symptoms, mapped to ICD-11 (Chapter 06) and DSM-5 criteria. To overcome the drawbacks of distant supervision and circular evaluation, a 500-consultation human-verified gold standard was developed, achieving an inter-annotator agreement of Cohen’s κ = 0.89 evaluated at the word token level across all 68,760 tokens. We have performed a thorough comparative study of four neural architectures: Bidirectional Gated Recurrent Unit (Bi-GRU), Pure Custom Transformer Encoder trained from scratch, Transformer with FastText augmentation, and AraBERT model fine-tuned using the data. We add a layer of 1D Convolutional Feature Representation (CFR) with k = 3 on all architectures to capture local subword/character-level morphological patterns and a Conditional Random Field (CRF) layer to enforce valid BIO syntax. AraBERT subword tokenization was aligned to word-level boundaries using first-subword pooling and loss masking (ignore_index = -100) to ensure a strictly standardized evaluation space across all models (712,287 original word tokens in the held-out test set, corresponding to 934,181 AraBERT subwords). Results & Contributions: On the human-verified gold standard, the fine-tuned AraBERT model demonstrates superior clinical semantic generalization, obtaining a Macro F1-score of 0.8735, outperforming Bi-GRU with 0.8214 by successfully resolving clinical negation and dialectal synonyms. Conversely, on the synthetic distant supervision test set, the Bi-GRU-CRF achieves a Macro F1-score of 0.9051 (compared to 0.8987 for AraBERT-CRF), providing empirical evidence of the effective rule replication and memorization of deterministic dictionary patterns by recurrent models. Crucially, complexity analysis reveals that Bi-GRU operates at a fraction of the computational footprint (6.31M parameters, 0.37 ms latency per sentence) compared to AraBERT (135.78M parameters, 11.77 ms latency), providing an ultra-efficient alternative for real-time edge clinical deployment. Conclusion: To the best of our knowledge, this work establishes the first comprehensive token-level Arabic clinical mental health NER benchmark. The integration of pre-trained language models with CRF sequential constraints, validated against a human-annotated benchmark, provides a robust methodology for Arabic clinical Named Entity Recognition (NER). Future directions include domain-adaptive clinical pretraining and active learning to further reduce manual annotation overhead.

Receive Date

26 June 2026

Revise Date

22 Aug 2026

Accept Date

22 Aug 2026

Publication Date

2026

Share

COinS