[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-1-en-105":3,"doc-seo-195221-105":53,"doc-detail-195221-en":126},{"code":4,"msg":5,"data":6},0,"success",[7,14,19,24,29,34,39,44,49],{"id":8,"doc_module":9,"doc_module_name":10,"category_name":11,"show_sort_weight":12,"slug":13},11,1,"Template","Presentations",90,"presentations",{"id":15,"doc_module":9,"doc_module_name":10,"category_name":16,"show_sort_weight":17,"slug":18},12,"Resumes",80,"resumes",{"id":20,"doc_module":9,"doc_module_name":10,"category_name":21,"show_sort_weight":22,"slug":23},14,"Invoices",70,"invoices",{"id":25,"doc_module":9,"doc_module_name":10,"category_name":26,"show_sort_weight":27,"slug":28},15,"Posters",60,"posters",{"id":30,"doc_module":9,"doc_module_name":10,"category_name":31,"show_sort_weight":32,"slug":33},16,"Social Media",50,"social-media",{"id":35,"doc_module":9,"doc_module_name":10,"category_name":36,"show_sort_weight":37,"slug":38},17,"Forms",40,"forms",{"id":40,"doc_module":9,"doc_module_name":10,"category_name":41,"show_sort_weight":42,"slug":43},18,"Letters",30,"letters",{"id":45,"doc_module":9,"doc_module_name":10,"category_name":46,"show_sort_weight":47,"slug":48},21,"Paper Templates",5,"papers-templates",{"id":50,"doc_module":9,"doc_module_name":10,"category_name":51,"show_sort_weight":4,"slug":52},158,"General","general-158",{"code":4,"msg":54,"data":55},"ok",{"site_id":56,"language":57,"slug":58,"title":59,"keywords":60,"description":61,"schema_data":62,"social_meta":119,"head_meta":121,"extra_data":123,"updated_unix":125},105,"en","document-metadata-extraction-multimodal-195221","Document Metadata Extraction (Multimodal)","","This document details the process and results of sentiment analysis on various text sources, including Twitter tweets, IMDB reviews, and Yelp reviews. It compares the effectiveness of different natural language processing techniques, namely Logistic Regression with TF-IDF and Recurrent Neural Networks (RNN) with Long Short-Term Memory (LSTM) and trainable or pre-trained embeddings. The tables present accuracy scores and classification results for negative and positive sentiments across these methods and data sources. Specifically, it highlights that RNN with LSTM and pre-trained embeddings generally achieve high accuracy, especially for Yelp reviews, while Logistic Regression with TF-IDF serves as a baseline. The document also includes a visualization of the text preprocessing pipeline, demonstrating steps like punctuation removal, unicode normalization, lemmatization, and stop-word removal, which are crucial for preparing text data for machine learning models. The latter part of the document outlines a neural network architecture for sequence processing, emphasizing the role of embedding layers, LSTM units, batch normalization, and a final sigmoid activation for prediction, illustrating a typical deep learning approach to text analysis.",{"@graph":63,"@context":118},[64,80,101],{"@type":65,"itemListElement":66},"BreadcrumbList",[67,71,74,77],{"item":68,"name":69,"@type":70,"position":9},"https://docshare.wps.com","Home","ListItem",{"item":72,"name":10,"@type":70,"position":73},"https://docshare.wps.com/template/",2,{"item":75,"name":51,"@type":70,"position":76},"https://docshare.wps.com/template/general/",3,{"item":78,"name":59,"@type":70,"position":79},"https://docshare.wps.com/template/document-metadata-extraction-multimodal-195221/195221/",4,{"url":78,"name":59,"@type":81,"image":82,"author":87,"headline":59,"publisher":90,"fileFormat":93,"inLanguage":57,"description":61,"dateModified":94,"datePublished":95,"encodingFormat":93,"isAccessibleForFree":96,"interactionStatistic":97},"DigitalDocument",{"url":83,"@type":84,"width":85,"height":86},"https://docshare.wps.com/thumbnails/document-metadata-extraction-multimodal-195221/195221.png","ImageObject",442,249,{"name":88,"@type":89},"Arica Lee","Person",{"url":68,"name":91,"@type":92},"DocShare","Organization","application/pdf","2026-09-21","2026-09-03",true,{"@type":98,"interactionType":99,"userInteractionCount":76},"InteractionCounter",{"@type":100},"ViewAction",{"@type":102,"mainEntity":103},"FAQPage",[104,110,114],{"name":105,"@type":106,"acceptedAnswer":107},"What are the main text preprocessing steps shown in the document?","Question",{"text":108,"@type":109},"The main text preprocessing steps include punctuation removal, unicode normalization, lemmatization, and stop-word removal.","Answer",{"name":111,"@type":106,"acceptedAnswer":112},"Which natural language processing techniques are compared in the document?",{"text":113,"@type":109},"The document compares Logistic Regression with TF-IDF and RNN with LSTM (using trainable and pre-trained embeddings).",{"name":115,"@type":106,"acceptedAnswer":116},"What is the purpose of the neural network architecture described?",{"text":117,"@type":109},"The neural network architecture, featuring embeddings and LSTMs, is designed for sequence processing, likely for tasks such as sentiment analysis, leading to a prediction via a sigmoid function.","https://schema.org",{"og:url":78,"og:type":120,"og:title":59,"og:site_name":91,"og:description":61},"article",{"robots":122,"canonical":78},"index,follow",{"doc_id":124,"site_id":56},195221,1788446462,{"code":4,"msg":5,"data":127},{"doc_id":124,"user_id":128,"nickname":88,"user_avatar":129,"doc_module":9,"category_id":50,"category_name":51,"doc_title":59,"doc_description":61,"doc_content":130,"file_id":131,"file_url":132,"file_type":133,"file_size":134,"view_count":76,"is_deleted":4,"is_public":9,"is_downloadable":9,"audit_status":9,"page_count":135,"language":136,"language_code":57,"site_id":56,"html_lang":57,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":61,"update_tm":125,"read_time":73},8796096645457,"https://ap-avatar.wpscdn.com/avatar/800003749518d68ffe3?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345340919836971","|  | Twitter tweet | IMDB review | Yelp review |\n| --- | --- | --- | --- |\n| Logistic regression with TF-IDF | 87.71% | 89.62% | 93.63% |\n| RNN with LSTM and trainable embed | 88.52% | 89.40% | 94.21% |\n| RNN with LSTM and pre-trained embed | 88.96% | 89.06% | 93.04% |\n| RNN with LSTM, trainable embed and avg pool | 87.44% | 88.92% | 94.09% |\n\n\n| RNN+LSTM+Pre-trained Embed\u003Cbr>for twitter tweets | Neg\u003Cbr>Pos | 85.88%\u003Cbr>91.62% | 89.87%\u003Cbr>88.23% | 87.83%\u003Cbr>89.89% |\n| --- | --- | --- | --- | --- |\n| Logistic regression+TF-IDF | Neg | 89.06% | 90.03% | 89.54% |\n| for IMDB reviews | Pos | 90.18% | 89.22% | 89.70% |\n| RNN+LSTM+Trainable Embed | Neg | 94.18% | 94.31% | 94.25% |\n| for Yelp reviews | Pos | 94.25% | 94.11% | 94.18% |\n\n\n| Neg Pos |  |  |  |\n| --- | --- | --- | --- |\n| RNN+LSTM+Pre-trained Embed\u003Cbr>for twitter tweets | Neg\u003Cbr>Pos | 736\u003Cbr>121 | 83\u003Cbr>907 |\n| Logistic regression+TF-IDF for IMDB reviews | Neg\u003Cbr>Pos | 2222\u003Cbr>273 | 246\u003Cbr>2259 |\n| RNN+LSTM+Trainable Embed for Yelp reviews | Neg\u003Cbr>Pos | 2654\u003Cbr>164 | 160\u003Cbr>2622 |","cbCaif84lZQ16CoZ","https://ap.wps.com/l/cbCaif84lZQ16CoZ","pdf",495441,6,"English","# Text Preprocessing Pipeline\n## Punctuation Removal\n## Unicode Normalization\n## Lemmatization\n## Stop-Word Removal\n# Neural Network Architecture","[{\"question\":\"What are the main text preprocessing steps shown in the document?\",\"answer\":\"The main text preprocessing steps include punctuation removal, unicode normalization, lemmatization, and stop-word removal.\"},{\"question\":\"Which natural language processing techniques are compared in the document?\",\"answer\":\"The document compares Logistic Regression with TF-IDF and RNN with LSTM (using trainable and pre-trained embeddings).\"},{\"question\":\"What is the purpose of the neural network architecture described?\",\"answer\":\"The neural network architecture, featuring embeddings and LSTMs, is designed for sequence processing, likely for tasks such as sentiment analysis, leading to a prediction via a sigmoid function.\"}]","Document Metadata Extraction (Multimodal) | PDF"]