[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122199-en":3,"doc-seo-122199-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122199,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Prediction of Case Types from Nonsearchable PDF Documents in Arabic - Comparison of Machine Learning and Deep Learning with Image Processing","The study predicts different types of judicial cases presented to Moroccan administrative courts using court decisions stored as non-searchable Arabic PDF documents. The approach combines image processing and text cleaning with machine learning and deep learning models to extract usable text from scanned, unstructured inputs. Experiments run in two phases using 697 decisions and 14,207 decisions from the Administrative Court of Appeal in Marrakech. Despite challenges in Arabic OCR, machine learning reaches 91% and 97% accuracy, while deep learning reaches 100% and 96%.","Prediction of case types from nonsearchable pdf documents in arabic :  \ncomparison of machine learning and deep learning with image processing  \nMouad El Arrasse ¹ , Youness Khourdifi ² , Soufyane Mounir ¹ and Alae El Alami ³  \n¹ National School of Applied Sciences of Khouribga (Laboratory of Engineering Science and  \nTechnology), Morocco.  \n² University Sultan Moulay Slimane, Polydisciplinary Faculty of Khouribga (Laboratory of Materials Science, Mathematics and Environment), Morocco.  \n³ Higher School of Technology Meknès (Laboratory of Computer Engineering and Intelligent  \nElectrical Systems), Morocco  \nAbstract. The study conducted focuses on predicting the different types of judicial cases presented to Moroccan administrative courts by using court decisions in the form of non-searchable PDF documents in the Arabic language. To achieve this, we utilized image processing, text cleaning techniques, and machine learning algorithms.We carried out a comparative study using both machine learning and deep learning techniques. The experiment was conducted in two phases: first on 697 court decisions, and then on  \n14,207 decisions from the Administrative Court of Appeal in Marrakech. Despite the challenges associated with the Arabic language, our methods were able to efficiently extract text, leading to accurate predictions. For the experiment on 697 decisions, machine learning achieved an accuracy rate of 91%, while deep learning reached 100% . For the experiment on 14,207 decisions, machine learning obtained an accuracy of 97%, and deep learning achieved 96%.As a result, this study contributes to the existing literature on the digitization and processing of unstructured documents in the Arabic language, as well as on the prediction of judicial case types through the use of  \nmachine learning and deep learning algorithms.  \nKeywords :  \nMachine learning – Deep learning – Judicial case prediction – Nonsearchable PDFs – Image processing – Text extraction.  \n1.Introduction :  \nThe digitization of documents is a common and essential practice in various fields of research and application, enabling the transformation of unstructured data into manipulable structured data [1] . This transformation greatly facilitates the access, search, and manipulation of the information contained within these documents. However, a major challenge arises when dealing with non-searchable PDF documents,  \n© The Authors, published by EDP Sciences. This is an open access article distributed under the terms of the Creative Commons Attribution License 4.0 ([https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/)).  \nparticularly those generated through the digitization of paper documents [2] . These documents often contain images embedding text, making their processing and analysis difficult and labor-intensive. Furthermore, the reduced quality of scanned PDF files results in additional loss of information, which further complicates their use in research and data analysis. Thus, the need to develop efficient methods for extracting, processing, and analyzing the content of these documents has become a major concern in contemporary research [3] .  \nAdditionally, the integration of Arabic-language documents adds an extra layer of complexity due to the distinctive graphical characteristics of the language, which makes character and word recognition more challenging [4] . For instance, Arabic letters are often connected to form words, and a single letter can take on different forms depending on its position within the word, creating unique challenges for Optical Character Recognition (OCR) and textual analysis. These linguistic nuances require a sophisticated methodological approach in the development of document processing tools, highlighting the need to adopt specialized techniques to ensure precise text extraction from Arabic-language documents.  \nIn this context, it is also crucial to provide concrete examples of processed P","cbCaipDberJw0h1U","https://ap.wps.com/l/cbCaipDberJw0h1U","pdf",773693,1,20,"English","en",105,"# Introduction\n## Digitization and challenges of non-searchable PDFs\n## Arabic-specific difficulties for OCR and text analysis\n## Study objective and paper structure\n# State of the Art","[{\"question\":\"What problem does the study address?\",\"answer\":\"It addresses predicting judicial case types from Arabic court decisions provided as non-searchable PDF documents where the text is not directly retrievable.\"},{\"question\":\"Which methods are compared in the study?\",\"answer\":\"The study compares machine learning and deep learning approaches, supported by image processing and text extraction/cleaning techniques.\"},{\"question\":\"How is performance evaluated and what accuracy results are reported?\",\"answer\":\"Two experiments are conducted on 697 decisions and 14,207 decisions; machine learning achieves 91% and 97% accuracy, while deep learning achieves 100% and 96%.\"}]","Prediction of Case Types from Nonsearchable PDF Documents in Arabic - Comparison of Machine Learning and Deep Learning with Image Processing | PDF",1785809311,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"prediction-of-case-types-from-nonsearchable-pdf-documents-in-arabic-comparison-of-machine-learning-and-deep-learning-with-image-processing","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/prediction-of-case-types-from-nonsearchable-pdf-documents-in-arabic-comparison-of-machine-learning-and-deep-learning-with-image-processing/122199/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the study address?","Question",{"text":76,"@type":77},"It addresses predicting judicial case types from Arabic court decisions provided as non-searchable PDF documents where the text is not directly retrievable.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which methods are compared in the study?",{"text":81,"@type":77},"The study compares machine learning and deep learning approaches, supported by image processing and text extraction/cleaning techniques.",{"name":83,"@type":74,"acceptedAnswer":84},"How is performance evaluated and what accuracy results are reported?",{"text":85,"@type":77},"Two experiments are conducted on 697 decisions and 14,207 decisions; machine learning achieves 91% and 97% accuracy, while deep learning achieves 100% and 96%.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":29,"slug":114},6,"Technology","technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":21,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":21,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]