[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126029-en":3,"doc-seo-126029-105":31,"detail-sidebar-cat-0-en-105":97},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126029,2336474466412,"Ezra","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","PRESERVING PRIVACY IN THE ERA OF BIG DATA - A MACHINE LEARNING-BASED ANONYMIZATION FRAMEWORK FOR SPATIOTEMPORAL TRAJECTORY DATASETS","Publishing open datasets supports research and government transparency, yet disclosure can expose users’ private information. Spatiotemporal trajectory datasets are especially sensitive because removing unique identifiers alone cannot prevent re-identification. Adversaries may infer trajectory fragments or link the released dataset with external sources to identify individuals. To address this risk, the paper presents a machine learning based anonymization framework (MLA) that clusters trajectories using k-means and a privacy-oriented variant, improving alignment via multiple sequence alignment, evaluated on T-Drive, Geolife, and Gowalla.","PRESERVING PRIVACY IN THE ERA OF BIG DATA: A MACHINE LEARNING-BASED ANONYMIZATION FRAMEWORK FOR SPATIOTEMPORAL TRAJECTORY DATASETS  \nC. Gazala Akhtar1, G. Lahari2, P. Shruthi2, V. Akshitha2  \n1Assistant Professor,2UG Students, Department of Cyber Security Engineering.  \n1,2Malla Reddy Engineering College for Women, Maisammaguda, Dhulapally, Kompally,  \nSecunderabad-500100, Telangana, India.  \nTo Cite this Article  \nC. Gazala Akhtar, G. Lahari, P. Shruthi, V. Akshitha ,“PRESERVING PRIVACY IN THE ERA OF BIG DATA: A MACHINE LEARNING-BASED ANONYMIZATION FRAMEWORK SPATIOTEMPORAL  \nTRAJECTORY DATASETS” Journal of Science and Technology, Vol. 08, Issue 12-Dec 2023, pp219-229 Article Info  \nReceived: 14-11-2023 Revised: 24-11-2023 Accepted: 04-12-2023 Published: 14-12-2023  \nABSTRACT  \nPublishing datasets plays an essential role in open data research and promoting transparency of government agencies. However, such data publication might reveal users’ private information. One of the most sensitive sources of data is spatiotemporal trajectory datasets. Unfortunately, merely removing unique identifiers cannot preserve the privacy of users. Adversaries may know parts of the trajectories or be able to link the published dataset to other sources for the purpose of user identification. Therefore , it is crucial to apply privacy preserving techniques before the publication of spatiotemporal trajectory datasets. In this paper, we propose a robust framework for the anonymization of spatiotemporal trajectory datasets termed as machine learning based anonymization (MLA) . By introducing a new formulation of the problem, we are able to apply machine learning algorithms for clustering the trajectories and propose to use k-means algorithm for this purpose. A variation ofk-means algorithm is also proposed to preserve the privacy in overly sensitive datasets. Moreover, we improve the alignment process by considering multiple sequence alignment as part of the MLA. The framework and all the proposed algorithms are applied to T-Drive, Geolife, and Gowalla location datasets. The experimental results indicate a significantly higher utility of datasets by anonymization based on MLA framework.  \nKeywords: Preserving Privacy, Big Bata, Machine Learning Algorithms, Spatiotemporal, Trajectory Datasets.  \n1. INTRODUCTION  \nIn the contemporary era of big data, the practice of publishing datasets plays a crucial role in advancing open data research and fostering transparency within government agencies. However, this seemingly beneficial practice raises concerns about the potential compromise of users' private information inherent in the published data. Among the most sensitive data sources are spatiotemporal trajectory datasets, which, when left unprotected, can expose individuals to privacy breaches. The conventional method of merely removing unique identifiers from such datasets proves insufficient to safeguard user privacy. Sophisticated adversaries could discern fragments of trajectories or establish connections between the published dataset and external sources, thereby facilitating user identification. Recognizing this vulnerability, it becomes imperative to employ privacy-preserving techniques prior to the dissemination of spatiotemporal trajectory datasets. In response to this challenge, this paper introduces a robust anonymization framework specifically tailored for spatiotemporal trajectory datasets, termed as the  \nMachine Learning-Based Anonymization (MLA) framework. The novelty of this approach lies in its innovative formulation of the problem, allowing for the application of machine learning algorithms to cluster trajectories effectively. The proposed framework leverages the widely-used k-means algorithm for trajectory clustering, presenting a refined variation to ensure privacy preservation in datasets with heightened sensitivity. Additionally, the alignment process is enhanced through the incorporation of multiple sequence alignment as an integral ","cbCaip2iOhT8yBDi","https://ap.wps.com/l/cbCaip2iOhT8yBDi","pdf",810112,6,1,11,"English","en",105,"# Abstract\n# Introduction\n## Problem of privacy leakage in published spatiotemporal trajectories\n## Proposed MLA framework for anonymization\n## Clustering with k-means and a privacy-preserving variant\n## Alignment improvement using multiple sequence alignment\n## Experimental evaluation on T-Drive, Geolife, and Gowalla","[{\"question\":\"Why is anonymization necessary for spatiotemporal trajectory datasets when publishing open data?\",\"answer\":\"Publishing such datasets can reveal private information. Identifier removal alone is insufficient because attackers can exploit trajectory fragments or perform linkage attacks with external data to re-identify users.\"},{\"question\":\"What is the core idea of the proposed MLA framework?\",\"answer\":\"The MLA framework formulates trajectory anonymization as a machine learning task. It clusters trajectories using the k-means algorithm and uses an adapted variation for highly sensitive datasets.\"},{\"question\":\"How does the framework improve the alignment process?\",\"answer\":\"It incorporates multiple sequence alignment as part of the MLA pipeline, enhancing how trajectories are aligned before/within anonymization.\"},{\"question\":\"On which datasets is the framework evaluated and what is the outcome?\",\"answer\":\"Experiments are conducted on T-Drive, Geolife, and Gowalla location datasets. Results show a significantly higher utility of the anonymized datasets using the MLA framework.\"}]","PRESERVING PRIVACY IN THE ERA OF BIG DATA - A MACHINE LEARNING-BASED ANONYMIZATION FRAMEWORK FOR SPATIOTEMPORAL TRAJECTORY DATASETS | PDF",1785902626,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":92,"head_meta":94,"extra_data":96,"updated_unix":29},"preserving-privacy-in-the-era-of-big-data-a-machine-learning-based-anonymization-framework-for-spatiotemporal-trajectory-datasets","",{"@graph":37,"@context":91},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/preserving-privacy-in-the-era-of-big-data-a-machine-learning-based-anonymization-framework-for-spatiotemporal-trajectory-datasets/126029/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83,87],{"name":74,"@type":75,"acceptedAnswer":76},"Why is anonymization necessary for spatiotemporal trajectory datasets when publishing open data?","Question",{"text":77,"@type":78},"Publishing such datasets can reveal private information. Identifier removal alone is insufficient because attackers can exploit trajectory fragments or perform linkage attacks with external data to re-identify users.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What is the core idea of the proposed MLA framework?",{"text":82,"@type":78},"The MLA framework formulates trajectory anonymization as a machine learning task. It clusters trajectories using the k-means algorithm and uses an adapted variation for highly sensitive datasets.",{"name":84,"@type":75,"acceptedAnswer":85},"How does the framework improve the alignment process?",{"text":86,"@type":78},"It incorporates multiple sequence alignment as part of the MLA pipeline, enhancing how trajectories are aligned before/within anonymization.",{"name":88,"@type":75,"acceptedAnswer":89},"On which datasets is the framework evaluated and what is the outcome?",{"text":90,"@type":78},"Experiments are conducted on T-Drive, Geolife, and Gowalla location datasets. Results show a significantly higher utility of the anonymized datasets using the MLA framework.","https://schema.org",{"og:url":53,"og:type":93,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":95,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":98},[99,103,107,111,116,120,125,128,133,136,140],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Exam",70,"exam",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},"Technology",50,"technology",{"id":121,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":123,"slug":124},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":126,"slug":127},30,"research-report",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":131,"slug":132},9,"Religion & Spirituality",20,"religion-spirituality",{"id":131,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":131,"slug":135},"World Cup","world-cup",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":137,"slug":139},10,"Lifestyle","lifestyle",{"id":141,"doc_module":4,"doc_module_name":47,"category_name":142,"show_sort_weight":112,"slug":143},19,"General","general"]