[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123439-en":3,"doc-seo-123439-105":29,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":11},123439,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Semi-automatic linguistic annotation for lexicography with machine learning - Proposal and evaluation","Machine-learning classifiers are used to semi-automatically annotate South African terminology lists for lexicographic purposes, focusing on part-of-speech and noun class information that is often missing when only translations are provided. Manual annotation is costly, time-intensive, and error-prone, so an expert-reviewed workflow is proposed to streamline grammatical enrichment. The approach targets isiXhosa resources and integrates results into the IsiXhosa.click online dictionary, and it is extended to additional domain glossaries. A bi-LSTM noun-class tagger and candidate part-of-speech classifiers are evaluated by accuracy and human review time to validate the process efficiency and quality.","Semi-automatic linguistic annotation for lexicography with machine learning\nBenjamin GUBB (\u0013 HYPERLINK \"mailto:gbbben@myuct.ac.za\" \\h \u0014gbbben\u0015\u0013 HYPERLINK \"mailto:gbbben@myuct.ac.za\" \\h \u0014001\u0015\u0013 HYPERLINK \"mailto:gbbben@myuct.ac.za\" \\h \u0014@myuct.ac.za\u0015)\nAffiliation: School of Languages and Literatures: African Languages, University of Cape Town, Cape Town, South Africa\nCael MARQUARD (\u0013 HYPERLINK \"mailto:cael.marquard@gmail.com\" \\h \u0014cael.marquard@gmail.com\u0015)\nAffiliation: Department of Computer Science, University of Cape Town, Cape Town, South Africa\nMany terminology lists for South African languages provide only the translations of each term without grammatical information such as part-of-speech and noun class. This makes it difficult to know how to use these words correctly in context. Annotating these terminology lists manually requires linguistic expertise and can be costly, time-intensive, and error prone. Instead, we propose annotating these terminology lists using a machine-learning classifier. An expert will then review the generated output to ensure accuracy. We will apply this approach to isiXhosa and integrate the results into the IsiXhosa.click online dictionary (Marquard 2024). This progresses the annotation of lexicographic works by making it easier to input linguistic information in dictionaries.\nWe build on IsiXhosa.click’s crowd-sourced lexicography and focus on streamlining linguistic annotation. Previously, terms were manually annotated, which is a labour-intensive, time-consuming process. With an automated machine-learning classifier, we hypothesise this process will be faster and less tedious. We propose applying machine-learning classifiers to annotate UCT’s mechanical engineering (Mechanical Engineering IsiXhosa Glossary 2023) and medical school glossaries.\nA variety of data-driven Human Language Technologies (HLTs) have been developed for isiXhosa (Agbeyangi & Jere 2024), including automatic classifiers for linguistic information such as part-of-speech and noun class. However, their performance is not yet good enough to rely on them solely. Additionally, most are developed for words in a sentence-level context, and not standalone as appearing in a dictionary.\nWe adopt a process similar to the one described by Gaustad & Puttkammer (2022) in developing a linguistically annotated corpus. The pipeline begins with words being automatically annotated by machine-learning classifiers with their part-of-speech and noun class, as well as the classifiers’ confidence. It is then reviewed by a human annotator and checked for quality control.\nKey tasks include developing the noun-class tagger, applying the taggers to the terminology lists, hiring an expert, native isiXhosa speaker to review the classifier’s output, and comparing the classifiers to each other and the manual corrections. The corrected terms will be added to the IsiXhosa.click database.\nTwo types of machine-learning classifiers are applied: one for part-of-speech annotation and one for noun-class annotation. Two options are considered for the part-of-speech classifier: the classifier developed in the NLAPOST21 Shared Task (Pannach et. al 2021), and the classifier developed by du Toit & Puttkammer (2021). Output from both will be analysed for accuracy at the end of the project. The classifier for noun-class annotation will be developed from scratch, we hypothesise that a purpose-built noun-class classifier will outperform general taggers like morphological analysers.\nThe noun-class classifier is a simple bi-LSTM model trained on linguistically annotated data (Gaustad & Puttkammer 2022). A bi-LSTM model combines two Long Short-Term Memory Models (LSTMs) (Hochreiter & Schmidhuber 1997), one reading the input forward and one backward. An LSTM is a type of Recurrent Neural Network, a class of models storing past events and are thus well-suited for sequence processing tasks. Bi-LSTMs have successfully been applied to similar tasks for Nguni languages, such as morph","cbCaimi43PCiwIOk","https://ap.wps.com/l/cbCaimi43PCiwIOk","docx",21503,1,3,"English","en",105,"# Background and motivation\n## Limitations of current terminology lists\n## Need for faster, accurate annotation\n# Proposed workflow\n## Machine-learning annotation with expert review\n## Integration into IsiXhosa.click\n# Scope and related work\n## Building on crowd-sourced lexicography\n## Prior isiXhosa HLTs and their limitations\n# Model and pipeline design\n## Part-of-speech classification options\n## Noun-class tagger development\n## Bi-LSTM architecture overview\n# Evaluation strategy\n## Metrics: accuracy and human review time\n# Expected outcomes\n## Efficiency, accuracy, and grammatical enrichment","[{\"question\":\"Why is manual annotation difficult for terminology lists?\",\"answer\":\"Many lists include only translations and lack grammatical information, so adding part-of-speech and noun class manually requires linguistic expertise. The work is costly, time-intensive, and prone to errors.\"},{\"question\":\"How does the proposed semi-automatic approach ensure annotation quality?\",\"answer\":\"Machine-learning classifiers generate annotations, including confidence where available, and then an expert/native speaker reviews the output. Quality control and manual corrections are fed back into the dictionary database.\"},{\"question\":\"What models are used for part-of-speech and noun-class annotation?\",\"answer\":\"Two classifier types are applied: one for part-of-speech using existing candidate classifiers, and a purpose-built noun-class classifier. The noun-class classifier is implemented as a bi-LSTM trained on linguistically annotated data.\"},{\"question\":\"How is the approach evaluated to test the main hypothesis?\",\"answer\":\"Evaluation uses two metrics: classifier accuracy and the time required for humans to review outputs. Together, these indicate whether the workflow is faster while maintaining acceptable quality.\"}]","Semi-automatic linguistic annotation for lexicography with machine learning - Proposal and evaluation | DOCX",1785816503,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":28},"semi-automatic-linguistic-annotation-for-lexicography-with-machine-learning-proposal-and-evaluation","",{"@graph":35,"@context":87},[36,52,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":21},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/semi-automatic-linguistic-annotation-for-lexicography-with-machine-learning-proposal-and-evaluation/123439/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/vnd.openxmlformats-officedocument.wordprocessingml.document","2026-08-04",true,{"@type":63,"interactionType":64,"userInteractionCount":4},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79,83],{"name":70,"@type":71,"acceptedAnswer":72},"Why is manual annotation difficult for terminology lists?","Question",{"text":73,"@type":74},"Many lists include only translations and lack grammatical information, so adding part-of-speech and noun class manually requires linguistic expertise. The work is costly, time-intensive, and prone to errors.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"How does the proposed semi-automatic approach ensure annotation quality?",{"text":78,"@type":74},"Machine-learning classifiers generate annotations, including confidence where available, and then an expert/native speaker reviews the output. Quality control and manual corrections are fed back into the dictionary database.",{"name":80,"@type":71,"acceptedAnswer":81},"What models are used for part-of-speech and noun-class annotation?",{"text":82,"@type":74},"Two classifier types are applied: one for part-of-speech using existing candidate classifiers, and a purpose-built noun-class classifier. The noun-class classifier is implemented as a bi-LSTM trained on linguistically annotated data.",{"name":84,"@type":71,"acceptedAnswer":85},"How is the approach evaluated to test the main hypothesis?",{"text":86,"@type":74},"Evaluation uses two metrics: classifier accuracy and the time required for humans to review outputs. Together, these indicate whether the workflow is faster while maintaining acceptable quality.","https://schema.org",{"og:url":50,"og:type":89,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":91,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":94},[95,99,103,107,112,117,122,125,130,133,137],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":113,"doc_module":4,"doc_module_name":45,"category_name":114,"show_sort_weight":115,"slug":116},6,"Technology",50,"technology",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":120,"slug":121},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":123,"slug":124},30,"research-report",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":128,"slug":129},9,"Religion & Spirituality",20,"religion-spirituality",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":128,"slug":132},"World Cup","world-cup",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":134,"slug":136},10,"Lifestyle","lifestyle",{"id":138,"doc_module":4,"doc_module_name":45,"category_name":139,"show_sort_weight":108,"slug":140},19,"General","general"]