Book Cover
E-book
Author Miner, Donald

Title MapReduce design patterns / Donald Miner, Adam Shook
Published Beijing ; Sebastopol : O'Reilly, 2012

Copies

Description 1 online resource (xvi, 232 pages) : illustrations
Contents Design patterns and MapReduce -- Summarization patterns -- Filtering patterns -- Data organization patterns -- Join patterns -- Metapatterns -- Input and output patterns -- Final thoughts and the future of design patterns
Summary Until now, design patterns for the MapReduce framework have been scattered among various research papers, blogs, and books. This handy guide brings together a unique collection of valuable MapReduce patterns that will save you time and effort regardless of the domain, language, or development framework you're using. Each pattern is explained in context, with pitfalls and caveats clearly identified to help you avoid common design mistakes when modeling your big data architecture. This book also provides a complete overview of MapReduce that explains its origins and implementations, and why design patterns are so important. All code examples are written for Hadoop. Summarization patterns: get a top-level view by summarizing and grouping data Filtering patterns: view data subsets such as records generated from one user Data organization patterns: reorganize data to work with other systems, or to make MapReduce analysis easier Join patterns: analyze different datasets together to discover interesting relationships Metapatterns: piece together several patterns to solve multi-stage problems, or to perform several analytics in the same job Input and output patterns: customize the way you use Hadoop to load or store data "A clear exposition of MapReduce programs for common data processing patterns--this book is indespensible for anyone using Hadoop."--Tom White, author of Hadoop: The Definitive Guide
Notes Print version record
SUBJECT Apache Hadoop. http://id.loc.gov/authorities/names/n2013024279
MapReduce (Computer file) http://id.loc.gov/authorities/names/no2013077469
Apache Hadoop (Computer file) blmlsh
MapReduce (Computer program) blmlsh
Apache Hadoop fast
MapReduce (Computer file) fast
Subject Electronic data processing -- Distributed processing.
Cluster analysis -- Data processing
Software patterns.
Computer algorithms.
Algorithms
algorithms.
Apache Hadoop.
Cluster analysis -- Data processing
Computer algorithms
Electronic data processing -- Distributed processing
Software patterns
Form Electronic book
Author Shook, Adam
ISBN 9781449341954
1449341950
9781449341985
1449341985