[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-1-en-105":3,"doc-detail-337623-en":53,"doc-seo-337623-105":75},{"code":4,"msg":5,"data":6},0,"success",[7,14,19,24,29,34,39,44,49],{"id":8,"doc_module":9,"doc_module_name":10,"category_name":11,"show_sort_weight":12,"slug":13},11,1,"Template","Presentations",90,"presentations",{"id":15,"doc_module":9,"doc_module_name":10,"category_name":16,"show_sort_weight":17,"slug":18},12,"Resumes",80,"resumes",{"id":20,"doc_module":9,"doc_module_name":10,"category_name":21,"show_sort_weight":22,"slug":23},14,"Invoices",70,"invoices",{"id":25,"doc_module":9,"doc_module_name":10,"category_name":26,"show_sort_weight":27,"slug":28},15,"Posters",60,"posters",{"id":30,"doc_module":9,"doc_module_name":10,"category_name":31,"show_sort_weight":32,"slug":33},16,"Social Media",50,"social-media",{"id":35,"doc_module":9,"doc_module_name":10,"category_name":36,"show_sort_weight":37,"slug":38},17,"Forms",40,"forms",{"id":40,"doc_module":9,"doc_module_name":10,"category_name":41,"show_sort_weight":42,"slug":43},18,"Letters",30,"letters",{"id":45,"doc_module":9,"doc_module_name":10,"category_name":46,"show_sort_weight":47,"slug":48},21,"Paper Templates",5,"papers-templates",{"id":50,"doc_module":9,"doc_module_name":10,"category_name":51,"show_sort_weight":4,"slug":52},158,"General","general-158",{"code":4,"msg":5,"data":54},{"doc_id":55,"user_id":56,"nickname":57,"user_avatar":58,"doc_module":9,"category_id":8,"category_name":11,"doc_title":59,"doc_description":60,"doc_content":61,"file_id":62,"file_url":63,"file_type":64,"file_size":65,"view_count":66,"is_deleted":4,"is_public":9,"is_downloadable":9,"audit_status":9,"page_count":30,"language":67,"language_code":68,"site_id":69,"html_lang":68,"table_of_contents":70,"faqs":71,"seo_title":72,"seo_description":60,"update_tm":73,"read_time":74},337623,2336475104736,"วิน","https://ap-avatar.wpscdn.com/avatar/22000c4c5e0e5b17e70?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786591360781797222","Extraction of Information from Invoices - Challenges in the Extraction Pipeline","Data from invoices is critical for enabling business processes, but value depends on capturing it in a digital, structured form. Current approaches using digital tools and AI/ML are state-of-the-art for invoice information extraction, yet they are often restricted to specific languages and layouts. They also tend to focus on isolated metrics, without demonstrating an end-to-end pipeline from raw data to processable information. This paper analyzes invoice information types and the resulting pipeline challenges, proposing a morphological framework to support pipeline design as part of a design science study.","Extraction of Information from Invoices – Challenges in the Extraction Pipeline  \nLukas-Walter Thiée , Felix Krieger and Burkhardt Funk 1  \nAbstract: Data from invoices are key information for business processes. In order to use the data and create business value, the information must be captured in a digital and structured form. Leveraging digital tools and AI/ML is state-of-the-art in the extraction of information from invoices. However, the existing approaches are trained on specific languages and layouts, and while focusing on the performance of individual metrics, they neglect the demonstration of the pipeline from raw data to processable information. In this paper, we investigate the types of information on invoices and address the challenges in the extraction pipeline. We contribute by providing a morphological framework for the problematization and design of a pipeline as part ofa design science study.  \nKeywords: Invoice recognition, Information extraction, Data pipeline.  \n1 Introduction, Motivation and Method  \nExtraction of information from business documents is an evolving area of research and practice, as structured, digital information support numerous business processes. While we focus on invoices in this paper, the research approaches can be applied to other document types, such as receipts or checks. Digitizing incoming invoices, i.e. , capturing structured information, can not only save a considerable amount of time, but can also add value. E.g., supply chain management can utilize these data by automatically integrating delivery dates and quantities into ERP systems [KAD04] . In addition, structured invoice data enable business analytics, e.g., for purchasing patterns [Fa04], [Ra21] . Furthermore, auditing firms can leverage the data to simplify and enhance financial audits from sample testing to substantive test of details [KAM95] . Despite the possibilities of the electronic creation, transfer and standardized integration of documents [KAD04], it is still common practice today to send invoices on a paper or pdf basis, so that the information must be extracted from the document or file. This refers to both B2B and B2C invoicing processes. In contrast to digital invoice recognition human invoice reading (as well as annotation) is error-prone and costly with average “processing costs of about 9 Euro” [KAD04] . Nevertheless, humans are good at the cognitive task of information extraction, i.e., infer abbreviations, link tabular data, and form composite information.  \nAs in many fields of digitalization and research, methods from artificial intelligence (AI) and machine learning (ML) are increasingly being examined and used. The ultimate goal  \n1 Leuphana University, Institute of Information Systems, Universitätsallee 1, Lüneburg, 21335, [lukas-walter.thiee@stud.leuphana.de](lukas-walter.thiee@stud.leuphana.de),  https://orcid.org/0000-0002-1998-376X  \nof this field of research is to digitally capture all (relevant) information from arbitrary business documents and make it available in a structured, processable format. In particular, the applications shall translate the raw data on a document into a machine-readable data type, so that the correct learning of the relationships in the data can be used to infer the original information. Existing industry solutions leverage these methods (see Chapter 2) . However, the solutions are still far from comprehensive recognition of all relevant information, because invoices are designed in a plethora of layouts and languages. In addition, data protection concerns arise in the external processing of business documents, as sensitive information must not end up on external/foreign storages. Although the use of external services can provide access to (pretrained) models, cloud- and development environments, it creates a dependency that reduces both the ability to influence and the understanding of the model output. For these reasons – extraction quality, data privacy, and mo","cbCaigTYS45NQ9TE","https://ap.wps.com/l/cbCaigTYS45NQ9TE","pdf",842212,4,"English","en",105,"# Introduction, Motivation and Method\n## Motivation for digitizing invoice information\n## AI/ML goals and limitations\n## Pipeline concept and challenges","[{\"question\":\"Why is invoice information extraction important for business processes?\",\"answer\":\"Structured invoice data supports numerous business processes by enabling the captured information to be used in digital, processable formats. It reduces manual effort and time while improving downstream analytics and system integration.\"},{\"question\":\"What limitations do existing invoice extraction approaches have?\",\"answer\":\"Many approaches are trained on specific languages and layouts and often optimize individual performance metrics. They typically do not provide an end-to-end demonstration from raw data to processable information, and generalization to unseen layouts remains limited.\"},{\"question\":\"What does the paper contribute to solving pipeline challenges?\",\"answer\":\"The paper investigates the types of information present in invoices and addresses extraction pipeline challenges. It contributes a morphological framework for problematization and for designing a pipeline as part of a design science study.\"}]","Extraction of Information from Invoices - Challenges in the Extraction Pipeline | PDF",1790019843,6,{"code":4,"msg":76,"data":77},"ok",{"site_id":69,"language":68,"slug":78,"title":59,"keywords":79,"description":60,"schema_data":80,"social_meta":135,"head_meta":137,"extra_data":139,"updated_unix":140},"extraction-of-information-from-invoices-challenges-in-the-extraction-pipeline","",{"@graph":81,"@context":134},[82,97,117],{"@type":83,"itemListElement":84},"BreadcrumbList",[85,89,92,95],{"item":86,"name":87,"@type":88,"position":9},"https://docshare.wps.com","Home","ListItem",{"item":90,"name":10,"@type":88,"position":91},"https://docshare.wps.com/template/",2,{"item":93,"name":11,"@type":88,"position":94},"https://docshare.wps.com/template/presentations/",3,{"item":96,"name":59,"@type":88,"position":66},"https://docshare.wps.com/template/extraction-of-information-from-invoices-challenges-in-the-extraction-pipeline/337623/",{"url":96,"name":59,"@type":98,"image":99,"author":104,"headline":59,"publisher":106,"fileFormat":109,"inLanguage":68,"description":60,"dateModified":110,"datePublished":111,"encodingFormat":109,"isAccessibleForFree":112,"interactionStatistic":113},"DigitalDocument",{"url":100,"@type":101,"width":102,"height":103},"https://docshare.wps.com/thumbnails/extraction-of-information-from-invoices-challenges-in-the-extraction-pipeline/337623.png","ImageObject",442,249,{"name":57,"@type":105},"Person",{"url":86,"name":107,"@type":108},"DocShare","Organization","application/pdf","2026-09-26","2026-09-21",true,{"@type":114,"interactionType":115,"userInteractionCount":66},"InteractionCounter",{"@type":116},"ViewAction",{"@type":118,"mainEntity":119},"FAQPage",[120,126,130],{"name":121,"@type":122,"acceptedAnswer":123},"Why is invoice information extraction important for business processes?","Question",{"text":124,"@type":125},"Structured invoice data supports numerous business processes by enabling the captured information to be used in digital, processable formats. It reduces manual effort and time while improving downstream analytics and system integration.","Answer",{"name":127,"@type":122,"acceptedAnswer":128},"What limitations do existing invoice extraction approaches have?",{"text":129,"@type":125},"Many approaches are trained on specific languages and layouts and often optimize individual performance metrics. They typically do not provide an end-to-end demonstration from raw data to processable information, and generalization to unseen layouts remains limited.",{"name":131,"@type":122,"acceptedAnswer":132},"What does the paper contribute to solving pipeline challenges?",{"text":133,"@type":125},"The paper investigates the types of information present in invoices and addresses extraction pipeline challenges. It contributes a morphological framework for problematization and for designing a pipeline as part of a design science study.","https://schema.org",{"og:url":96,"og:type":136,"og:title":59,"og:site_name":107,"og:description":60},"article",{"robots":138,"canonical":96},"index,follow",{"doc_id":55,"site_id":69},1790452663]