[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128370-en":3,"doc-seo-128370-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128370,962085571259,"Theodora","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","THE DEVELOPMENT OF A TEXT GENERATION MODEL FOR SEPEDI LANGUAGE USING TRANSFORMER-BASED MACHINE LEARNING TECHNIQUES","Transformer-based machine learning techniques are used to model sequential input through an encoder–decoder architecture with parallel processing and attention over each token. Such models are reported to outperform recurrent neural networks like LSTM for natural language processing tasks, while avoiding RNN training issues such as vanishing and exploding gradients. A GPT-Sepedi transformer model is developed for Sepedi text generation using an NCHLT Sepedi corpus. Generated output is validated against Sepedi vocabulary and evaluated with ROUGE, achieving 61% vocabulary coverage and 83% precision, though comprehensibility remains limited by recall and F1-score results.","THE DEVELOPMENT OF A TEXT GENERATION MODEL FOR SEPEDI LANGUAGE USING TRANSFORMER-BASED MACHINE-LEARNING  \nTECHNIQUES  \nby  \nMAHLODI MERCY MOILA  \nDISSERTATION  \nSubmitted in fulfilment of the requirements for the degree of  \nMASTER OF SCIENCE  \nIn  \nDEPARTMENT OF COMPUTER SCIENCE  \nIn the  \nFACULTY OF SCIENCE AND AGRICULTURE  \n(School of Mathematical and Computer Sciences)  \nat the  \nUNIVERSITY OF LIMPOPO  \nSUPERVISOR: MR MJD MANAMELA  \nCO-SUPERVISOR: DR TI MODIPA  \n2025  \nDEDICATION  \nThis one is for me. I would like to thank myself for not giving up when it got tough.  \nDECLARATION  \nI declare that THE DEVELOPMENT OF A TEXT GENERATION MODEL FOR SEPEDI LANGUAGE USING TRANSFORMER-BASED MACHINE LEARNING TECHNIQUES is my own work and that all the sources that I have used or quoted have been indicated and acknowledged by means of complete references and that this work has not been submitted before for any other degree at any other institution.  \nSurname, Initials (Ms) Date  \nACKNOWLEDGEMENTS  \nI would like to acknowledge the following people for their assistance and contribution in ensuring that I complete this study:  \n• Davis Sialumba, thank you for pushing me to further my studies to this level. If it were not for you, I would have found comfort in only having an undergraduate degree.  \n• I would like to thank Mrs Rene Kotze for ensuring that I did not quit when things got tough due to the lack of finances. Because of you, I was able to get NRF funding through the National Institute for Theoretic and Computational Sciences (NITheCS) . If it was not for NITheCS, I would have dropped out.  \n• My Dad , who always encouraged me to give it my all and never give up when things got tough, thanks Pa.  \n• I also want to pass my sincere gratitude to my supervisor, Mr. MJDManamela, and Co-Supervisor Dr. TI Modipa. It has been a bumpy road towards completing this research project, but you never gave up on me.  \n• To Rethabile, thank you for being the pillar of my strength during this journey.  \n• To my daughter Luna, in my life, you gave me so many reasons not to give up when things got tough. Your presence became my encouragement to succeed.  \n• To everyone who has supported me, I thank you for your time and encouragement.  \nABSTRACT  \nThe transformer-based machine learning technique is a deep learning model that processes the sequential input data using an encoder-decoder process. Transformers process the input data simultaneously using a parallelism approach while paying attention to each word at the time by applying an attention mechanism to each unit text being processed. The transformer-based model has been known to provide more state-of-the-art performance in natural language processing (NLP) tasks than a recurrent neural network (RNN) such as Long Short-Term Memory (LSTM) . RNNs have the drawback of suffering from the problem of vanishing gradients and exploding gradients in implementation.  \nThe GPT-Sepedi transformer-based model has shown great success in dealing with the process of text generation for the Sepedi language. This has led to a limited text generation system developed using a transformer-based model for the underresourced African language , namely, the Sepedi language. This research project aimed to develop a text generation model for the Sepedi language using transformerbased machine learning techniques. The LSTM-Sepedi Attention-based model and the GPT-Sepedi Transformer-based model were developed and trained using a National Centre for Human Language Technology (NCHLT) Sepedi text corpus. The models were compared based on the results that they generated. A GPT-Sepedi Transformer-based model was used to generate the text.  \nThe generated text was then compared with the Sepedi language vocabulary to determine the validity of the text. It was found that 61% of the text within the generated texts is found in the Sepedi language vocabulary. The Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score was used t","cbCaigTuOsMEyGbb","https://ap.wps.com/l/cbCaigTuOsMEyGbb","pdf",1650929,1,96,"English","en",105,"# 1. CHAPTER 1: INTRODUCTION\n## 1.1 Problem Statement\n## 1.2 Motivation\n## 1.3 Aim\n## 1.4 Objectives\n## 1.5 Scientific Contribution\n## 1.6 Ethical Clearance\n## 1.7 Dissertation Structure","[{\"question\":\"What transformer principle enables effective text generation in the described model?\",\"answer\":\"Transformers process input in parallel and use an attention mechanism to focus on each token, enabling encoder–decoder-based sequence modeling for text generation.\"},{\"question\":\"How were the models trained and what language data was used?\",\"answer\":\"The LSTM-Sepedi Attention-based model and the GPT-Sepedi Transformer-based model were trained using an NCHLT Sepedi text corpus.\"},{\"question\":\"How was the generated text evaluated for quality and validity?\",\"answer\":\"The generated text was checked against Sepedi vocabulary validity and compared to human-written text using ROUGE, reporting 61% vocabulary presence and ROUGE-based precision metrics.\"}]","THE DEVELOPMENT OF A TEXT GENERATION MODEL FOR SEPEDI LANGUAGE USING TRANSFORMER-BASED MACHINE LEARNING TECHNIQUES | PDF",1785947132,242,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"the-development-of-a-text-generation-model-for-sepedi-language-using-transformer-based-machine-learning-techniques","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-development-of-a-text-generation-model-for-sepedi-language-using-transformer-based-machine-learning-techniques/128370/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What transformer principle enables effective text generation in the described model?","Question",{"text":76,"@type":77},"Transformers process input in parallel and use an attention mechanism to focus on each token, enabling encoder–decoder-based sequence modeling for text generation.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How were the models trained and what language data was used?",{"text":81,"@type":77},"The LSTM-Sepedi Attention-based model and the GPT-Sepedi Transformer-based model were trained using an NCHLT Sepedi text corpus.",{"name":83,"@type":74,"acceptedAnswer":84},"How was the generated text evaluated for quality and validity?",{"text":85,"@type":77},"The generated text was checked against Sepedi vocabulary validity and compared to human-written text using ROUGE, reporting 61% vocabulary presence and ROUGE-based precision metrics.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]