[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-169041-en":3,"doc-seo-169041-105":30,"detail-sidebar-cat-1-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":11,"category_id":12,"category_name":13,"doc_title":14,"doc_description":15,"doc_content":16,"file_id":17,"file_url":18,"file_type":19,"file_size":20,"view_count":4,"is_deleted":4,"is_public":11,"is_downloadable":11,"audit_status":11,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":15,"update_tm":28,"read_time":29},169041,962084926284,"Aria Callaghan","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",1,11,"Presentations","Doing Hoops With Hadoop - and friends","The presentation introduces Big Data, analytics, and data science, explaining how statistical methods such as clustering and machine learning drive insights. It positions cloud computing as a practical context for scalable experimentation, then walks through Hadoop’s role in distributed storage and parallel processing. The talk explains core Hadoop components, including HDFS and MapReduce, and clarifies major use cases and limitations such as handling small files and achieving low-latency access. Emphasis is placed on throughput, fault tolerance, and the Hadoop ecosystem.","Doing Hoops With Hadoop(and friends)\nGreg Rogers\nSystems Performance & Capacity Planning\nFinancial Services Industry\n(formerly DEC, Compaq, HP, MACP Consulting)\nHewlett-Packard Enterprise Systems Division\n2 August 2012\nWhat’s all the excitement about?\nIn a phrase:\nBig Data\nAnalytics\nData science:  Develop insights from Big Data\nOften statistical, often centered around Clustering and Machine Learning\nBig Brother Is Watching You (BBIWY)\nGovernments?  Corporations?  Criminals?\nSomebody else’s presentation topic!\nBruce Schneier, et al\nCloud Computing\nPublic: Amazon, et al, low-cost prototyping for startup companies\nElastic Compute Cloud (EC3)\nSimple Storage Service (S3)\nEBS\nHadoop services\nPrivate\nGrid or Cluster computing\nSay hello to Hadoop\nOne more time\nSay hello to Hadoop’s friendsa.k.a. The Hadoop Ecosystem ca. 2011-2012\nOverloading Terminology\nHadoop has become synonymous with Big Data management and processing\nThe name Hadoop is also now a proxy for both Hadoop and the large, growing ecosystem around it\nBasically, a very large “system” using Hadoop Distributed File System (HDFS) for storage, and direct or indirect use of the MapReduce programming model and software framework for processing\nHadoop: The High Level\nApache top-level project\n“…develops open-source software for reliable, scalable, distributed computing.”\nSoftware library\n“…framework that allows for distributed processing of large data sets across clusters of computers using a simple programming model…designed to scale up from single servers to thousands of machines, each offering local computation and storage…designed to detect and handle failures at the application layer…delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures.”\nHadoop Distributed File System (HDFS)\n“…primary storage system used by Hadoop applications. HDFS creates multiple replicas of data blocks and distributes them on compute nodes throughout a cluster to enable reliable, extremely rapid computations.”\nMapReduce\n“…a programming model and software framework for writing applications that rapidly process vast amounts of data in parallel on large clusters of compute nodes.”\nWhat’s Hadoop Used For?Major Use Cases per Cloudera (2012)\nData processing\nBuilding search indexes\nLog processing\n“Click Sessionization”\nData processing pipelines\nVideo & Image analysis\nAnalytics\nRecommendation systems (Machine Learning)\nBatch reporting\nReal time applications (“home of the brave”)\nData Warehousing\nHDFS\nGoogle File System (GFS) cited as the original concept\nCertain common attributes between GFS & HDFS\nPattern emphasis: Write Once, Read Many\nSound familiar?\nRemember WORM drives? Optical jukebox library?  Before CDs (CD-R)\nBack when dirt was new… almost…\nWORM optical storage was still about greater price:performance at larger storage volumes vs. the prevailing [disk] technology – though still better p:p than The Other White Meat of the day, Tape!\nVery Large Files: One file could span the entire HDFS\nCommodity Hardware\n2 socket; 64-bit; local disks; no RAID\nCo-locate data and compute resources\nHDFS:  What It’s Not Good For\nMany, many small files\nScalability issue for the namenode (more in a moment)\nLow Latency Access\nIt’s all about Throughput (N=XR)\nNot about minimizing service time (or latency to first data read)\nMultiple writers; updates at offsets within a file\nOne writer\nCreate/append/rename/move/delete - that’s it!  No updates!\nNot a substitute for a relational database\nData stored in files, not indexed\nTo find something, must read all the data (ultimately, by a MapReduce job/tasks)\nSelling SAN or NAS – NetApp & EMC need not apply\nProgramming in COBOL\nSelling mainframes & FICON\nHDFS Concepts & Architecture\nCore architectural goal: Fault-tolerance in the face of massive parallelism (many devices, high probability of HW failure)\nMonitoring, detection, fast automated recovery\nFocus is batch throughput\nHence support for very large files; large block size; sing","cbCaiuaxl27wjGep","https://ap.wps.com/l/cbCaiuaxl27wjGep","pptx",1749642,50,"English","en",105,"# What’s all the excitement about?\n## Big Data and data science\n## Cloud computing and examples\n## Hadoop services and ecosystem\n# Say hello to Hadoop’s friends\n## Overloading terminology\n# Hadoop: The high level\n## Apache project scope\n## HDFS and MapReduce definitions\n# What’s Hadoop used for?\n## Major use cases per Cloudera\n# Hadoop in context of file systems\n## GFS vs HDFS\n## WORM storage analogy\n# HDFS concepts & architecture\n## Fault tolerance and batch throughput\n## Master/slave components\n## Block size and metadata impact\n## Replication and reliability\n## NameNode availability concerns","[{\"question\":\"Why is Hadoop associated with Big Data and analytics?\",\"answer\":\"Hadoop has become synonymous with Big Data management and processing because it supports large-scale storage and parallel computation. It enables analytics and data science workflows that develop insights from massive datasets.\"},{\"question\":\"What are the main roles of HDFS and MapReduce?\",\"answer\":\"HDFS provides a distributed storage system that replicates data blocks across nodes for reliability and rapid computation. MapReduce provides a programming model and framework to process vast datasets in parallel across compute clusters.\"},{\"question\":\"What is HDFS not well-suited for?\",\"answer\":\"HDFS struggles with many small files and is not designed for low-latency access. It also is not a substitute for relational databases since data is stored in files rather than indexed for direct lookup.\"}]","Doing Hoops With Hadoop - and friends | PPTX",1788243685,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":14,"keywords":34,"description":15,"schema_data":35,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":28},"doing-hoops-with-hadoop-and-friends","",{"@graph":36,"@context":87},[37,54,70],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":11},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/template/","Template",2,{"item":49,"name":13,"@type":43,"position":50},"https://docshare.wps.com/template/presentations/",3,{"item":52,"name":14,"@type":43,"position":53},"https://docshare.wps.com/template/doing-hoops-with-hadoop-and-friends/169041/",4,{"url":52,"name":14,"@type":55,"author":56,"headline":14,"publisher":58,"fileFormat":61,"inLanguage":23,"description":15,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/vnd.openxmlformats-officedocument.presentationml.presentation","2026-09-05","2026-09-01",true,{"@type":66,"interactionType":67,"userInteractionCount":69},"InteractionCounter",{"@type":68},"ViewAction",5,{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why is Hadoop associated with Big Data and analytics?","Question",{"text":77,"@type":78},"Hadoop has become synonymous with Big Data management and processing because it supports large-scale storage and parallel computation. It enables analytics and data science workflows that develop insights from massive datasets.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What are the main roles of HDFS and MapReduce?",{"text":82,"@type":78},"HDFS provides a distributed storage system that replicates data blocks across nodes for reliability and rapid computation. MapReduce provides a programming model and framework to process vast datasets in parallel across compute clusters.",{"name":84,"@type":75,"acceptedAnswer":85},"What is HDFS not well-suited for?",{"text":86,"@type":78},"HDFS struggles with many small files and is not designed for low-latency access. It also is not a substitute for relational databases since data is stored in files rather than indexed for direct lookup.","https://schema.org",{"og:url":52,"og:type":89,"og:title":14,"og:site_name":59,"og:description":15},"article",{"robots":91,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":94},[95,98,103,108,113,117,122,126,130],{"id":12,"doc_module":11,"doc_module_name":46,"category_name":13,"show_sort_weight":96,"slug":97},90,"presentations",{"id":99,"doc_module":11,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},12,"Resumes",80,"resumes",{"id":104,"doc_module":11,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},14,"Invoices",70,"invoices",{"id":109,"doc_module":11,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},15,"Posters",60,"posters",{"id":114,"doc_module":11,"doc_module_name":46,"category_name":115,"show_sort_weight":21,"slug":116},16,"Social Media","social-media",{"id":118,"doc_module":11,"doc_module_name":46,"category_name":119,"show_sort_weight":120,"slug":121},17,"Forms",40,"forms",{"id":29,"doc_module":11,"doc_module_name":46,"category_name":123,"show_sort_weight":124,"slug":125},"Letters",30,"letters",{"id":127,"doc_module":11,"doc_module_name":46,"category_name":128,"show_sort_weight":69,"slug":129},21,"Paper Templates","papers-templates",{"id":131,"doc_module":11,"doc_module_name":46,"category_name":132,"show_sort_weight":4,"slug":133},158,"General","general-158"]