Limit search to available items
Book Cover
E-book
Author Trujillo, George, author

Title Virtualizing Hadoop : how to install, deploy, and optimize Hadoop in a virtualized architecture / George Trujillo, Jr. [and four others]
Published New York : VMware Press, [2016]
©2016

Copies

Description 1 online resource (1 volume) : illustrations
Contents Understanding the big deal data world -- Hadoop fundamental concepts -- YARN and HDFS -- The modern data platform -- Data ingestion -- Hadoop SQL engines -- Multitenancy in Hadoop -- Virtualization fundamentals -- Best practices for virtualizing Hadoop -- Virtualizing Hadoop -- Virgualizing Hadoop master servers -- Virtualizing the Hadoop worker nodes -- Deploying Hadoop as a service in the private cloud -- Understanding the installation of Hadoop -- Configuring Linux for Hadoop -- Hadoop cluster creation : a prerequisite checklist -- Big data/Hadoop on VMware vSphere reference materials
Summary Plan and Implement Hadoop Virtualization for Maximum Performance, Scalability, and Business Agility Enterprises running Hadoop must absorb rapid changes in big data ecosystems, frameworks, products, and workloads. Virtualized approaches can offer important advantages in speed, flexibility, and elasticity. Now, a world-class team of enterprise virtualization and big data experts guide you through the choices, considerations, and tradeoffs surrounding Hadoop virtualization. The authors help you decide whether to virtualize Hadoop, deploy Hadoop in the cloud, or integrate conventional and virtualized approaches in a blended solution. First, Virtualizing Hadoop reviews big data and Hadoop from the standpoint of the virtualization specialist. The authors demystify MapReduce, YARN, and HDFS and guide you through each stage of Hadoop data management. Next, they turn the tables, introducing big data experts to modern virtualization concepts and best practices. Finally, they bring Hadoop and virtualization together, guiding you through the decisions you'll face in planning, deploying, provisioning, and managing virtualized Hadoop. From security to multitenancy to day-to-day management, you'll find reliable answers for choosing your best Hadoop strategy and executing it. Coverage includes the following: • Reviewing the frameworks, products, distributions, use cases, and roles associated with Hadoop • Understanding YARN resource management, HDFS storage, and I/O • Designing data ingestion, movement, and organization for modern enterprise data platforms • Defining SQL engine strategies to meet strict SLAs • Considering security, data isolation, and scheduling for multitenant environments • Deploying Hadoop as a service in the cloud • Reviewing the essential concepts, capabilities, and terminology of virtualization • Applying current best practices, guidelines, and key metrics for Hadoop virtualization • Managing multiple Hadoop frameworks and products as one unified system • Virtualizing master and worker nodes to maximize availability and performance • Installing and configuring Linux for a Hadoop environment
Bibliography Includes bibliographical references and index
Notes Print version record
SUBJECT Apache Hadoop. http://id.loc.gov/authorities/names/n2013024279
Apache Hadoop fast
Subject Virtual computer systems.
Electronic data processing -- Distributed processing.
File organization (Computer science)
File processing (Computer science)
Electronic data processing -- Distributed processing
File organization (Computer science)
File processing (Computer science)
Virtual computer systems
Form Electronic book
ISBN 9780133812350
0133812359
0133811026
9780133811025
Other Titles How to install, deploy, and optimize Hadoop in a virtualized architecture