International Journal For Multidisciplinary Research

E-ISSN: 2582-2160     Impact Factor: 9.24

A Widely Indexed Open Access Peer Reviewed Multidisciplinary Bi-monthly Scholarly International Journal

Call for Paper Volume 8, Issue 4 (July-August 2026) Submit your research before last 3 days of August to publish your research paper in the issue of July-August.

Corpus Creation and Annotation For Under-Represented Languages

Author(s) Ms. Rakshitha s, Mr. Harish T A, Dr. Shantala C P
Country India
Abstract This project is developed to support the creation and annotation of corpora for regional and low-resource languages. Many local languages do not have sufficient digital datasets, making language research and NLP development difficult. The proposed tool provides a simple desktop application for managing language data. Users can upload text documents, perform tokenization, and annotate words with linguistic tags. The system also generates corpus statistics and visualizations to help users understand the dataset. It is developed using Python and Tkinter, along with supporting libraries for text processing and analysis. The annotated corpus can be exported in formats such as JSON, CSV, and CoNLL-U. This makes the data easy to reuse in other applications and research projects. The tool is useful for students, researchers, and language enthusiasts. It also supports the preservation and digital development of regional languages.
Keywords Natural Language Processing (NLP), Underrepresented Languages, Annotation Tool, Rule-Based Annotation
Field Engineering
Published In Volume 8, Issue 3, May-June 2026
Published On 2026-06-04

Share this