Authors :
Kinkete Mfumabi Hervé; Ngoy Déogracias; Ngoma Magema Arsene; Yamone Muisangie Clairette; Kapinga Nsapo Marisa; Mpoyi Kabulu Exauce; Likango Bokondo Christopher; Kapongo Musangu Israël
Volume/Issue :
Volume 11 - 2026, Issue 8 - August
Google Scholar :
https://tinyurl.com/yx3e8a5s
DOI :
https://doi.org/10.38124/ijisrt/26aug847
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
French and Lingala coexist in institutional, educational, and everyday communication in the Democratic
Republic of the Congo, where connectivity is uneven and entry-level Android devices remain widespread. This paper
specifies KOBONGOLA-Lite, a proposed bidirectional offline French–Lingala neural translation architecture, and defines
a prospective protocol for its implementation and confirmatory evaluation. The contribution is deliberately narrower than
claims of first bidirectionality or first compression, because recent French–Lingala systems and AfriNLLB already cover
these dimensions. KOBONGOLA-Lite instead coordinates four testable components: a governed and stratified Congolese
corpus; explicit treatment of natural French–Lingala code-switching; a context-gated lexical mechanism with neural
fallback; and physical-device validation of linguistic quality, memory, latency, energy, thermal behavior, reliability, and
network silence. Interface prototypes specify the intended user workflows, while the compact encoder–decoder model,
tokenizer, domain lexicon, span policies, hybrid reranker, and self-contained mobile bundle remain objects of prospective
implementation and validation. The planned study separates external baselines from controlled internal ablations that
isolate sequence-level distillation, lexical support, hybrid reranking, ONNX export, and dynamic INT8 quantization. The
primary automatic endpoint is chrF++; complementary evidence comprises SacreBLEU, locally validated AfriCOMET,
terminology and entity preservation, omission analysis, blinded human assessment, paired document-clustered bootstrap
intervals, effect sizes, and an explicitly defined ten-test Holm family. No implementation or confirmatory performance
result is claimed at this stage. The protocol establishes a falsifiable and reproducible basis for assessing offline neural
machine translation in a low-resource African and edge-computing context.
Keywords :
Code-Switching; Edge AI; French–Lingala Translation; INT8 Quantization; Knowledge Distillation; Low-Resource Languages; Neural Machine Translation; Offline Translation.
References :
- D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. ICLR, 2015.
- A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017.
- M. R. Costa-jussà et al., “No Language Left Behind: Scaling human-centered machine translation,” arXiv:2207.04672, 2022, doi: 10.48550/arXiv.2207.04672.
- S. Kabongo Kabenamualu, V. Marivate, and H. Kamper, “LiSTra automatic speech translation: English to Lingala case study,” in Proc. DCLRL, 2022, pp. 63–67, doi: 10.18653/v1/2022.dclrl-1.8.
- R. E. Mandiya et al., “A Transformer-based method for bidirectional French-Lingala machine translation in speech and text,” Applied Sciences, vol. 16, no. 7, art. 3399, 2026, doi: 10.3390/app16073399.
- Y. Moslem, A. K. Wassie, and A. G. Abebe, “AfriNLLB: Efficient translation models for African languages,” in Proc. AfricaNLP, 2026, doi: 10.18653/v1/2026.africanlp-main.30.
- R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proc. ACL, 2016, pp. 1715–1725, doi: 10.18653/v1/P16-1162.
- T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer,” in Proc. EMNLP System Demonstrations, 2018, pp. 66–71, doi: 10.18653/v1/D18-2012.
- D. I. Adelani et al., “A few thousand translations go a long way! Leveraging pre-trained models for African news translation,” in Proc. NAACL, 2022, doi: 10.18653/v1/2022.naacl-main.223.
- G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv:1503.02531, 2015.
- Y. Kim and A. M. Rush, “Sequence-level knowledge distillation,” in Proc. EMNLP, 2016, pp. 1317–1327, doi: 10.18653/v1/D16-1139.
- Y. Lin, X. Wang, Z. Zhang, M. Wang, T. Xiao, and J. Zhu, “MobileNMT: Enabling translation in 15MB and 30ms,” in Proc. ACL Industry Track, 2023, pp. 368–378, doi: 10.18653/v1/2023.acl-industry.36.
- M. Popović, “chrF: Character n-gram F-score for automatic MT evaluation,” in Proc. WMT, 2015, pp. 392–395, doi: 10.18653/v1/W15-3049.
- M. Popović, “chrF++: Words helping character n-grams,” in Proc. WMT, 2017, pp. 612–618, doi: 10.18653/v1/W17-4770.
- J. Wang et al., “AfriMTE and AfriCOMET: Enhancing COMET to embrace under-resourced African languages,” in Proc. NAACL, 2024, pp. 5997–6023, doi: 10.18653/v1/2024.naacl-long.334.
- E. M. Bender and B. Friedman, “Data statements for natural language processing,” Transactions of the ACL, vol. 6, pp. 587–604, 2018, doi: 10.1162/tacl_a_00041.
- T. Gebru et al., “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021, doi: 10.1145/3458723.
- M. Mitchell et al., “Model cards for model reporting,” in Proc. FAT*, 2019, pp. 220–229, doi: 10.1145/3287560.3287596.
- W. Nekoto et al., “Participatory research for low-resourced machine translation: A case study in African languages,” in Findings of EMNLP, 2020, pp. 2144–2160, doi: 10.18653/v1/2020.findings-emnlp.195.
- M. Post, “A call for clarity in reporting BLEU scores,” in Proc. WMT, 2018, pp. 186–191, doi: 10.18653/v1/W18-6319.
- R. Rei, C. Stewart, A. C. Farinha, and A. Lavie, “COMET: A neural framework for MT evaluation,” in Proc. EMNLP, 2020, pp. 2685–2702, doi: 10.18653/v1/2020.emnlp-main.213.
- P. Koehn, “Statistical significance tests for machine translation evaluation,” in Proc. EMNLP, 2004, pp. 388–395.
- M. Meeuwis, Lingala. Munich, Germany: LINCOM Europa, 1998.
- C. Myers-Scotton, Social Motivations for Codeswitching: Evidence from Africa. Oxford, U.K.: Oxford University Press, 1993.
- T. Nguyen, L. T. Nguyen, and D. Q. Nguyen, “How effective is machine translation on low-resource code-switching? A case study comparing human and automatic metrics,” in Findings of ACL, 2023, pp. 14186–14195, doi: 10.18653/v1/2023.findings-acl.893.
- C. Hokamp and Q. Liu, “Lexically constrained decoding for sequence generation using grid beam search,” in Proc. ACL, 2017, pp. 1535–1546, doi: 10.18653/v1/P17-1141.
- M. Post and D. Vilar, “Fast lexically constrained decoding with dynamic beam allocation for neural machine translation,” in Proc. NAACL-HLT, 2018, pp. 1314–1324, doi: 10.18653/v1/N18-1119.
- G. Dinu, P. Mathur, M. Federico, and Y. Al-Onaizan, “Training neural machine translation to apply terminology constraints,” in Proc. ACL, 2019, pp. 3063–3068, doi: 10.18653/v1/P19-1294.
French and Lingala coexist in institutional, educational, and everyday communication in the Democratic
Republic of the Congo, where connectivity is uneven and entry-level Android devices remain widespread. This paper
specifies KOBONGOLA-Lite, a proposed bidirectional offline French–Lingala neural translation architecture, and defines
a prospective protocol for its implementation and confirmatory evaluation. The contribution is deliberately narrower than
claims of first bidirectionality or first compression, because recent French–Lingala systems and AfriNLLB already cover
these dimensions. KOBONGOLA-Lite instead coordinates four testable components: a governed and stratified Congolese
corpus; explicit treatment of natural French–Lingala code-switching; a context-gated lexical mechanism with neural
fallback; and physical-device validation of linguistic quality, memory, latency, energy, thermal behavior, reliability, and
network silence. Interface prototypes specify the intended user workflows, while the compact encoder–decoder model,
tokenizer, domain lexicon, span policies, hybrid reranker, and self-contained mobile bundle remain objects of prospective
implementation and validation. The planned study separates external baselines from controlled internal ablations that
isolate sequence-level distillation, lexical support, hybrid reranking, ONNX export, and dynamic INT8 quantization. The
primary automatic endpoint is chrF++; complementary evidence comprises SacreBLEU, locally validated AfriCOMET,
terminology and entity preservation, omission analysis, blinded human assessment, paired document-clustered bootstrap
intervals, effect sizes, and an explicitly defined ten-test Holm family. No implementation or confirmatory performance
result is claimed at this stage. The protocol establishes a falsifiable and reproducible basis for assessing offline neural
machine translation in a low-resource African and edge-computing context.
Keywords :
Code-Switching; Edge AI; French–Lingala Translation; INT8 Quantization; Knowledge Distillation; Low-Resource Languages; Neural Machine Translation; Offline Translation.