Produktbild: Professional Hadoop Solutions

Professional Hadoop Solutions

59,99 €

inkl. gesetzl. MwSt., Versandkostenfrei


Beschreibung

Produktdetails

Einband

Taschenbuch

Erscheinungsdatum

23.09.2013

Verlag

John Wiley & Sons

Seitenzahl

504

Maße (L/B/H)

23,4/18,9/2,8 cm

Gewicht

857 g

Sprache

Englisch

ISBN

978-1-118-61193-7

Beschreibung

Produktdetails

Einband

Taschenbuch

Erscheinungsdatum

23.09.2013

Verlag

John Wiley & Sons

Seitenzahl

504

Maße (L/B/H)

23,4/18,9/2,8 cm

Gewicht

857 g

Sprache

Englisch

ISBN

978-1-118-61193-7

Herstelleradresse

Libri GmbH
Europaallee 1
36244 Bad Hersfeld
DE

Email: gpsr@libri.de

Noch keine Bewertungen vorhanden

Verfassen Sie die erste Bewertung zu diesem Artikel

Helfen Sie anderen Kundinnen und Kunden durch Ihre Meinung.

Kundinnen und Kunden meinen

Bewertungen (0)

Die Leseprobe wird geladen.
  • Produktbild: Professional Hadoop Solutions
  • Introduction xviiChapter 1: Big Data and the Hadoop Ecosystem 1Big Data Meets Hadoop 2Hadoop: Meeting the Big Data Challenge 3Data Science in the Business World 5The Hadoop Ecosystem 7Hadoop Core Components 7Hadoop Distributions 10Developing Enterprise Applications with Hadoop 12Summary 16Chapter 2: Storing Data in Hadoop 19HDFS 19HDFS Architecture 20Using HDFS Files 24Hadoop-Specific File Types 26HDFS Federation and High Availability 32HBase 34HBase Architecture 34HBase Schema Design 40Programming for HBase 42New HBase Features 50Combining HDFS and HBase for Effective Data Storage 53Using Apache Avro 53Managing Metadata with HCatalog 58Choosing an Appropriate Hadoop Data Organization for Your Applications 60Summary 62Chapter 3: Processing Your Data with MapReduce 63Getting to Know MapReduce 63MapReduce Execution Pipeline 65Runtime Coordination and Task Management in MapReduce 68Your First MapReduce Application 70Building and Executing MapReduce Programs 74Designing MapReduce Implementations 78Using MapReduce as a Framework for Parallel Processing 79Simple Data Processing with MapReduce 81Building Joins with MapReduce 82Building Iterative MapReduce Applications 88To MapReduce or Not to MapReduce? 94Common MapReduce Design Gotchas 95Summary 96Chapter 4: Customizing MapReduce Execution 97Controlling MapReduce Execution with InputFormat 98Implementing InputFormat for Compute-Intensive Applications 100Implementing InputFormat to Control the Number of Maps 106Implementing InputFormat for Multiple HBase Tables 112Reading Data Your Way with Custom RecordReaders 116Implementing a Queue-Based RecordReader 116Implementing RecordReader for XML Data 119Organizing Output Data with Custom Output Formats 123Implementing OutputFormat for Splitting MapReduceJob's Output into Multiple Directories 124Writing Data Your Way with Custom RecordWriters 133Implementing a RecordWriter to Produce Outputtar Files 133Optimizing Your MapReduce Execution with a Combiner 135Controlling Reducer Execution with Partitioners 139Implementing a Custom Partitioner for One-to-Many Joins 140Using Non-Java Code with Hadoop 143Pipes 143Hadoop Streaming 143Using JNI 144Summary 146Chapter 5: Building Reliable MapReduce Apps 147Unit Testing MapReduce Applications 147Testing Mappers 150Testing Reducers 151Integration Testing 152Local Application Testing with Eclipse 154Using Logging for Hadoop Testing 156Processing Applications Logs 160Reporting Metrics with Job Counters 162Defensive Programming in MapReduce 165Summary 166Chapter 6: Automating Data Processing with Oozie 167Getting to Know Oozie 168Oozie Workflow 170Executing Asynchronous Activities in Oozie Workflow 173Oozie Recovery Capabilities 179Oozie Workflow Job Life Cycle 180Oozie Coordinator 181Oozie Bundle 187Oozie Parameterization with Expression Language 191Workflow Functions 192Coordinator Functions 192Bundle Functions 193Other EL Functions 193Oozie Job Execution Model 193Accessing Oozie 197Oozie SLA 199Summary 203Chapter 7: Using Oozie 205Validating Information about Places Using Probes 206Designing Place Validation Based on Probes 207Designing Oozie Workflows 208Implementing Oozie Workflow Applications 211Implementing the Data Preparation Workflow 212Implementing Attendance Index and Cluster StrandsWorkflows 220Implementing Workflow Activities 222Populating the Execution Context from a java Action 223Using MapReduce Jobs in Oozie Workflows 223Implementing Oozie Coordinator Applications 226Implementing Oozie Bundle Applications 231Deploying, Testing, and Executing Oozie Applications 232Deploying Oozie Applications 232Using the Oozie CLI for Execution of an Oozie Application 234Passing Arguments to Oozie Jobs 237Using the Oozie Console to Get Information about OozieApplications 240Getting to Know the Oozie Console Screens 240Getting Information about a Coordinator Job 245Summary 247Chapter 8: Advanced Oozie FEATURES 249Building Custom Oozie Workflow Actions 250Implementing a Custom Oozie Workflow Action 251Deploying Oozie Custom Workflow Actions 255Adding Dynamic Execution to Oozie Workflows 257Overall Implementation Approach 257A Machine Learning Model, Parameters, and Algorithm 261Defining a Workflow for an Iterative Process 262Dynamic Workflow Generation 265Using the Oozie Java API 268Using Uber Jars with Oozie Applications 272Data Ingestion Conveyer 276Summary 283Chapter 9: Real-Time Hadoop 285Real-Time Applications in the Real World 286Using HBase for Implementing Real-Time Applications 287Using HBase as a Picture Management System 289Using HBase as a Lucene Back End 296Using Specialized Real-Time Hadoop Query Systems 317Apache Drill 319Impala 320Comparing Real-Time Queries to MapReduce 323Using Hadoop-Based Event-Processing Systems 323HFlame 324Storm 326Comparing Event Processing to MapReduce 329Summary 330Chapter 10: Hadoop Security 331A Brief History: Understanding Hadoop Security Challenges 333Authentication 334Kerberos Authentication 334Delegated Security Credentials 344Authorization 350HDFS File Permissions 350Service-Level Authorization 354Job Authorization 356Oozie Authentication and Authorization 356Network Encryption 358Security Enhancements with Project Rhino 360HDFS Disk-Level Encryption 361Token-Based Authentication and Unified Authorization Framework 361HBase Cell-Level Security 362Putting it All Together -- Best Practices for Securing Hadoop 362Authentication 363Authorization 364Network Encryption 364Stay Tuned for Hadoop Enhancements 365Summary 365Chapter 11: Running Hadoop Applications on AWS 367Getting to Know AWS 368Options for Running Hadoop on AWS 369Custom Installation using EC2 Instances 369Elastic MapReduce 370Additional Considerations before Making Your Choice 370Understanding the EMR-Hadoop Relationship 370EMR Architecture 372Using S3 Storage 373Maximizing Your Use of EMR 374Utilizing CloudWatch and Other AWS Components 376Accessing and Using EMR 377Using AWS S3 383Understanding the Use of Buckets 383Content Browsing with the Console 386Programmatically Accessing Files in S3 387Using MapReduce to Upload Multiple Files to S3 397Automating EMR Job Flow Creation and Job Execution 399Orchestrating Job Execution in EMR 404Using Oozie on an EMR Cluster 404AWS Simple Workflow 407AWS Data Pipeline 408Summary 409Chapter 12: Building Enterprise Security Solutions for Hadoop Implementations 411Security Concerns for Enterprise Applications 412Authentication 414Authorization 414Confidentiality 415Integrity 415Auditing 416What Hadoop Security Doesn't Natively Provide for Enterprise Applications 416Data-Oriented Access Control 416Differential Privacy 417Encrypted Data at Rest 419Enterprise Security Integration 419Approaches for Securing Enterprise Applications Using Hadoop 419Access Control Protection with Accumulo 420Encryption at Rest 430Network Isolation and Separation Approaches 430Summary 434Chapter 13: Hadoop's Future 435Simplifying MapReduce Programming with DSLs 436What Are DSLs? 436DSLs for Hadoop 437Faster, More Scalable Processing 449Apache YARN 449Tez 452Security Enhancements 452Emerging Trends 453Summary 454APPENDIX : Useful Reading 455Index 463