Authors :
Apoorva S.; Usha B. S.; S. M. Darshan
Volume/Issue :
Volume 11 - 2026, Issue 8 - August
Google Scholar :
https://tinyurl.com/24dr6us5
Scribd :
https://tinyurl.com/5b7c56t6
DOI :
https://doi.org/10.38124/ijisrt/26aug453
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
The fast expansion of digital data has resulted in high demand for smart systems which can quickly obtain
compact and meaningful data from lengthy documents. Even though automation in text summarization has seen
significant developments in well-resourced languages, automated summarization of Kannada is still rare due to the
intricacies of the script of the language, poor computational power, and massive information available in printed and
scanned form. A web-based framework for automated summarization of Kannada documents is introduced in this paper,
utilizing AI by developing a processing platform that integrates OCR and NLP. This technique works not only with
originally typed text in Kannada but also with scanned documents. Scanned documents are converted to editable Unicode
by using Tesseract OCR engine before performing the language-specific NLP tasks that include normalization,
tokenization, sentence splitting, stopword removal, and extraction of summaries. Using technologies such as Python, Flask,
OpenCV, Tesseract OCR, and relational database management, the developed application can ensure secure
authentication of users, management of documents, and visualization of summaries using an interactive web interface. The
successful tests indicate that OCR and NLP technologies have been integrated into the process of performing various tasks
related to the examination of documents written in the Kannada language. The modular architecture of the project
enables applying transformer-based summaries, document processing in many languages, OCR of handwritten texts
written in the Kannada language, and various technologies for running applications in the cloud in the future. Thus, the
developed application is an example of the effective use of Artificial Intelligence in processing documents in regional
languages and lays the groundwork for creating automated systems for document management.
Keywords :
Artificial Intelligence, Optical Character Recognition, Natural Language Processing, Kannada Document Summarization, Extractive Summarization, Tesseract OCR, Flask, Intelligent Document Processing.
References :
- A. Vaswani et al., "Attention Is All You Need," Advances in Neural Information Processing Systems, vol. 30, pp. 5998–6008, 2017.
- J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," Proc. NAACL-HLT, pp. 4171–4186, 2019.
- C. Raffel et al., "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer," Journal of Machine Learning Research, vol. 21, no. 140, pp. 1–67, 2020.
- M. Lewis et al., "BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation," ACL, pp. 7871–7880, 2020.
- J. Zhang et al., "PEGASUS: Pre-training with Extracted Gap-Sentences for Abstractive Summarization," ICML, pp. 11328–11339, 2020.
- K. Heafield, "KenLM: Faster and Smaller Language Model Queries," Proc. WMT, pp. 187–197, 2011.
- R. Smith, "An Overview of the Tesseract OCR Engine," Proc. ICDAR, pp. 629–633, 2007.
- R. Smith, "History of the Tesseract OCR Engine: What Worked and What Didn't," Document Recognition and Retrieval, vol. 6815, 2009.
- Google, "Tesseract OCR Documentation," 2024.
- G. Bradski, "The OpenCV Library," Dr. Dobb's Journal of Software Tools, 2000.
- A. Rosebrock, Practical Python and OpenCV, PyImageSearch, 2022.
- A. K. Jain, Y. Zhong, and M. Dubuisson-Jolly, "Deformable Template Models: A Review," Signal Processing, vol. 71, no. 2, pp. 109–129, 1998.
- S. Bird, E. Klein, and E. Loper, Natural Language Processing with Python. O'Reilly Media, 2009.
- S. Bird, Natural Language Toolkit (NLTK) Documentation, 2024.
- M. Honnibal and I. Montani, "spaCy 3: Industrial-Strength Natural Language Processing," 2024.
- F. Pedregosa et al., "Scikit-learn: Machine Learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
- D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd ed., Pearson, 2024.
- J. Eisenstein, Introduction to Natural Language Processing. MIT Press, 2019.
- R. Mihalcea and P. Tarau, "TextRank: Bringing Order into Text," EMNLP, pp. 404–411, 2004.
- H. P. Luhn, "The Automatic Creation of Literature Abstracts," IBM Journal of Research and Development, vol. 2, no. 2, pp. 159–165, 1958.
- M. Allahyari et al., "Text Summarization Techniques: A Brief Survey," International Journal of Advanced Computer Science and Applications, vol. 8, no. 10, pp. 397–405, 2017.
- Y. Liu and M. Lapata, "Text Summarization with Pretrained Encoders," EMNLP-IJCNLP, pp. 3730–3740, 2019.
- A. K. Singh and B. B. Chaudhuri, "OCR for Indian Scripts: A Survey," Artificial Intelligence Review, vol. 55, pp. 1–31, 2022.
- S. K. Saha et al., "Recent Advances in Optical Character Recognition Using Deep Learning: A Survey," IEEE Access, vol. 10, pp. 110264–110298, 2022.
- P. Choudhary and R. K. Gupta, "Deep Learning-Based OCR for Indian Regional Languages: A Review," Multimedia Tools and Applications, vol. 82, pp. 14231–14262, 2023.
- A. Khanuja et al., "MuRIL: Multilingual Representations for Indian Languages," Findings of ACL, pp. 2658–2669, 2021.
- S. Doddapaneni et al., "IndicBERT: A Multilingual Language Model for Indian Languages," Findings of EMNLP, pp. 5153–5168, 2021.
- R. Kakwani et al., "AI4Bharat IndicNLP Corpus: Monolingual Corpora and Language Models for Indian Languages," ACL Findings, pp. 406–426, 2022.
- K. Krishna, M. S. Reddy, and P. Kumar, "Automatic Text Summarization for Low-Resource Indian Languages: A Review," Journal of King Saud University – Computer and Information Sciences, vol. 35, no. 5, pp. 101675, 2023.
- N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings Using Siamese BERT Networks," EMNLP-IJCNLP, pp. 3982–3992, 2019.
- M. Abadi et al., "TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems," Google Research, 2016.
- M. Grinberg, Flask Web Development, 2nd ed. O'Reilly Media, 2018.
- M. Mohri, A. Rostamizadeh, and A. Talwalkar, Foundations of Machine Learning, 3rd ed. MIT Press, 2024.
- S. Saaramsha, "Leveraging NLP for Efficient Kannada Text Summarization," International Journal of Computer Applications, vol. 186, no. 31, 2024.
The fast expansion of digital data has resulted in high demand for smart systems which can quickly obtain
compact and meaningful data from lengthy documents. Even though automation in text summarization has seen
significant developments in well-resourced languages, automated summarization of Kannada is still rare due to the
intricacies of the script of the language, poor computational power, and massive information available in printed and
scanned form. A web-based framework for automated summarization of Kannada documents is introduced in this paper,
utilizing AI by developing a processing platform that integrates OCR and NLP. This technique works not only with
originally typed text in Kannada but also with scanned documents. Scanned documents are converted to editable Unicode
by using Tesseract OCR engine before performing the language-specific NLP tasks that include normalization,
tokenization, sentence splitting, stopword removal, and extraction of summaries. Using technologies such as Python, Flask,
OpenCV, Tesseract OCR, and relational database management, the developed application can ensure secure
authentication of users, management of documents, and visualization of summaries using an interactive web interface. The
successful tests indicate that OCR and NLP technologies have been integrated into the process of performing various tasks
related to the examination of documents written in the Kannada language. The modular architecture of the project
enables applying transformer-based summaries, document processing in many languages, OCR of handwritten texts
written in the Kannada language, and various technologies for running applications in the cloud in the future. Thus, the
developed application is an example of the effective use of Artificial Intelligence in processing documents in regional
languages and lays the groundwork for creating automated systems for document management.
Keywords :
Artificial Intelligence, Optical Character Recognition, Natural Language Processing, Kannada Document Summarization, Extractive Summarization, Tesseract OCR, Flask, Intelligent Document Processing.