When will I receive my Course Certificate?

If you complete the course successfully, your electronic Course Certificate will be added to your Accomplishments page - from there, you can print your Course Certificate or add it to your LinkedIn profile.

Why can’t I audit this course?

This course is currently available only to learners who have paid or received financial aid, when available.

Is financial aid available?

Yes. In select learning programs, you can apply for financial aid or a scholarship if you can’t afford the enrollment fee. If fin aid or scholarship is available for your learning program selection, you’ll find a link to apply on the description page.

The Ultimate Hands-On Hadoop

The Ultimate Hands-On Hadoop

本课程是 Big Data Foundations with Hadoop and Spark 专项课程的一部分

位教师：Packt - Course Instructors

包含在中

了解更多

12个模块

深入了解一个主题并学习基础知识。

中级等级

推荐体验

2 周完成

在 10 小时一周

灵活的计划

自行安排学习进度

12个模块

深入了解一个主题并学习基础知识。

中级等级

推荐体验

2 周完成

在 10 小时一周

灵活的计划

自行安排学习进度

您将学到什么

Remember Hadoop setup and configuration steps.
Understand the Hadoop ecosystem, including HDFS, MapReduce, and YARN.
Apply queries using Pig, Hive, and Spark.
Evaluate Hadoop cluster performance and optimize it.

您将获得的技能

您将学习的工具

要了解的详细信息

可分享的证书

添加到您的领英档案

作业

5 项作业

授课语言：英语（English）

91% of learners achieved a positive career outcome

了解顶级公司的员工如何掌握热门技能

了解关于 Coursera for Business 的更多信息

Petrobras, TATA, Danone, Capgemini, P&G 和 L'Oreal 的徽标

积累特定领域的专业知识

本课程是 Big Data Foundations with Hadoop and Spark 专项课程专项课程的一部分

在注册此课程时，您还会同时注册此专项课程。

向行业专家学习新概念
获得对主题或工具的基础理解
通过实践项目培养工作相关技能
获得可共享的职业证书

该课程共有12个模块

Updated in May 2025.

This course now features Coursera Coach — your interactive learning companion that helps you test your knowledge, challenge assumptions, and deepen your understanding as you progress. Build a strong, hands-on foundation in Hadoop and big data processing with this comprehensive course designed for data engineers, developers, and IT professionals. From installation to advanced analytics, you’ll learn how to work confidently with Hadoop’s ecosystem and design scalable solutions for real-world data challenges. You’ll begin by installing the Hortonworks Data Platform (HDP) Sandbox on your local machine, giving you an isolated environment to explore Hadoop’s core components. Through guided exercises, you’ll work with the Hadoop Distributed File System (HDFS) and build your understanding of MapReduce, learning how large-scale distributed processing works behind the scenes. As you progress, you’ll move into advanced Hadoop programming with Pig, Hive, and Spark. You’ll write complex queries, analyze large datasets, and work with real-world data to build scalable data workflows. You’ll also explore machine learning with Spark MLLib, giving you a practical introduction to distributed ML techniques. In the final modules, you’ll learn how to manage and optimize Hadoop clusters using YARN, ZooKeeper, Oozie, and Kafka. You’ll practice feeding data into your cluster, orchestrating workflows, managing resources, and analyzing streaming data in real time — essential skills for production-grade environments. By the end of this course, you will have: - Installed and configured the Hortonworks Sandbox for Hadoop development. - Worked with HDFS, MapReduce, and Hadoop’s core data processing concepts. - Written queries and pipelines using Pig, Hive, and Spark. - Performed distributed machine learning with Spark MLLib. - Integrated relational and non-relational data sources with Hadoop. - Managed clusters and streaming workflows with YARN, ZooKeeper, Oozie, and Kafka. - Gained the confidence to design and implement Hadoop-based data solutions. This course is ideal for data engineers, developers, and IT professionals with basic programming or data management experience. Familiarity with Java, SQL, or the Linux command line is helpful but not required.

In this module, we will dive into the world of Hadoop, starting with its installation and setup using the Hortonworks Data Platform Sandbox. You'll explore the key buzzwords and technologies that make up the Hadoop ecosystem, learn about the historical context and impact of the Hortonworks and Cloudera merger, and begin working with real data to get a feel for Hadoop's capabilities.

涵盖的内容

4个视频2篇阅读材料

In this module, we will explore the core components of Hadoop: the Hadoop Distributed File System (HDFS) and MapReduce. You'll learn how HDFS reliably stores massive data sets across a cluster and how MapReduce enables distributed data processing. Through hands-on activities, you'll import datasets, set up a MapReduce environment, and write scripts to analyze data, including breaking down movie ratings and ranking movies by popularity.

涵盖的内容

10个视频

10个视频总计94分钟

Hadoop Distributed File System (HDFS): What it is and How it Works14分钟
Installing the MovieLens Dataset6分钟
Activity - Installing the MovieLens Dataset into Hadoop's Distributed File System (HDFS) using the Command Line8分钟
MapReduce: What it is and How it Works11分钟
How MapReduce Distributes Processing13分钟
MapReduce Example: Breaking Down the Movie Ratings by Rating Score12分钟
Activity - Installing Python, MRJob, and Nano8分钟
Activity - Coding Up and Running the Ratings Histogram MapReduce Job8分钟
Exercise - Ranking Movies by Their Popularity7分钟
Activity - Checking Results8分钟

In this module, we will delve into Pig, a high-level scripting language that simplifies Hadoop programming. You'll start by exploring the Ambari web-based UI, which makes working with Pig more accessible. The module includes practical examples and activities, such as finding the oldest five-star movies and identifying the most-rated one-star movies using Pig scripts. You'll also learn about the capabilities of Pig Latin and test your skills through challenges and result comparisons.

涵盖的内容

7个视频1个作业

7个视频总计56分钟

Introducing Ambari10分钟
Introducing the Pig6分钟
Example - Finding the Oldest Movie with Five-Star Rating Using the Pig15分钟
Activity - Finding the Old Five-Star Movies with Pig10分钟
More Pig Latin8分钟
Exercise - Finding the Most-Rated One-Star Movie2分钟
Pig Challenge - Comparing Results6分钟

1个作业总计15分钟

Assessment 115分钟

In this module, we will explore the power of Apache Spark, a key technology in the Hadoop ecosystem known for its speed and versatility. You’ll start by understanding why Spark is a game-changer in big data. The module will cover Resilient Distributed Datasets (RDDs) and Datasets, showing you how to use them to analyze movie ratings data. You'll also delve into Spark's machine learning library (MLLib) to create a movie recommendation system. Through hands-on activities, you'll practice writing Spark scripts and refining your data analysis skills.

涵盖的内容

8个视频

8个视频总计74分钟

Why Spark?10分钟
The Resilient Distributed Datasets (RDD)10分钟
Activity - Finding the Movie with the Lowest Average Rating with the Resilient Distributed Datasets (RDD)16分钟
Datasets and Spark 2.06分钟
Activity - Finding the movie with the Lowest Average Rating with DataFrames10分钟
Activity - Recommending a Movie with Spark's Machine Learning Library (MLLib)12分钟
Exercise - Filtering the Lowest-Rated Movies by Number of Ratings3分钟
Activity - Checking Results7分钟

In this module, we will explore the integration of relational datastores with Hadoop, focusing on Apache Hive and MySQL. You'll start by learning how Hive enables SQL queries on data within HDFS, followed by hands-on activities to find popular and highly-rated movies using Hive. The module also covers the installation and integration of MySQL with Hadoop, using Sqoop to seamlessly transfer data between MySQL and Hadoop's HDFS/Hive. Through practical exercises, you'll gain proficiency in managing and querying relational data within the Hadoop ecosystem.

涵盖的内容

9个视频

9个视频总计63分钟

What is Hive?7分钟
Activity - Using Hive to Find the Most Popular Movie11分钟
How Hive Works?9分钟
Exercise - Using Hive to Find the Movie with the Highest Average Rating2分钟
Comparing Solutions4分钟
Integrating MySQL with Hadoop8分钟
Activity - Installing MySQL and Importing Movie Data8分钟
Activity - Using Sqoop to Import Data from MySQL to HFDS/Hive8分钟
Activity - Using Sqoop to Export Data from Hadoop to MySQL7分钟

In this module, we will explore the use of non-relational (NoSQL) data stores within the Hadoop ecosystem. You'll learn why NoSQL databases are crucial for scalability and efficiency, and dive into specific technologies like HBase, Cassandra, and MongoDB. Through a series of activities, you'll practice importing data into HBase, integrating it with Pig, and using Cassandra and MongoDB alongside Spark. The module concludes with exercises to help you choose the most suitable NoSQL database for different scenarios, empowering you to make informed decisions in big data management.

涵盖的内容

12个视频1个作业

12个视频总计148分钟

Why NoSQL?14分钟
What is HBase?13分钟
Activity - Importing Movie Ratings into HBase13分钟
Activity - Using HBase with Pig to Import Data at Scale11分钟
Cassandra - Overview15分钟
Activity - Installing Cassandra12分钟
Activity - Writing Spark Output into Cassandra11分钟
MongoDB - Overview17分钟
Activity - Installing MongoDB and Integrating Spark with MongoDB13分钟
Activity - Using the MongoDB Shell8分钟
Choosing Database Technology16分钟
Exercise - Choosing a Database for a Given Problem5分钟

1个作业总计15分钟

Assessment 215分钟

In this module, we will focus on interactive querying tools that allow you to quickly access and analyze big data across multiple sources. You'll explore technologies like Drill, Phoenix, and Presto, learning how each one solves specific challenges in querying large datasets. The module includes hands-on activities where you'll set up these tools, execute queries that span across databases such as MongoDB, Hive, HBase, and Cassandra, and integrate these tools with other Hadoop ecosystem components. By the end of this module, you'll be equipped to perform efficient, real-time data analysis across varied data stores.

涵盖的内容

9个视频

9个视频总计82分钟

Overview of Drill8分钟
Activity - Setting Up Drill11分钟
Activity - Querying Across Multiple Databases with Drill7分钟
Overview of Phoenix9分钟
Activity - Installing Phoenix and Querying HBase7分钟
Activity - Integrating Phoenix with the Pig12分钟
Overview of Presto7分钟
Activity - Installing Presto and Querying Hive12分钟
Activity - Querying Both Cassandra and Hive Using Presto9分钟

In this module, we will explore the critical components involved in managing a Hadoop cluster. You'll learn about YARN's resource management capabilities, how Tez optimizes task execution using Directed Acyclic Graphs, and the differences between Mesos and YARN. We'll dive into ZooKeeper for maintaining reliable operations and Oozie for orchestrating complex workflows. Hands-on activities will guide you through setting up and using Zeppelin for interactive data analysis and using Hue for a more user-friendly interface. The module also touches on other noteworthy technologies like Chukwa and Ganglia, providing a comprehensive understanding of cluster management in Hadoop.

涵盖的内容

13个视频

13个视频总计119分钟

Yet Another Resource Negotiator (YARN)10分钟
Tez5分钟
Activity - Using Hive on Tez and Measuring the Performance Benefit9分钟
Mesos7分钟
ZooKeeper13分钟
Activity - Simulating a Failing Master with ZooKeeper7分钟
Oozie12分钟
Activity - Setting Up a Simple Oozie Workflow17分钟
Zeppelin - Overview5分钟
Hands-On with Zeppelin for Spark and MovieLens Analysis12分钟
SQL and Data Visualization in Zeppelin: MovieLens Analysis with Spark10分钟
Hue - Overview8分钟
Other Technologies Worth Mentioning5分钟

In this module, we will explore the essential tools for feeding data into your Hadoop cluster, focusing on Kafka and Flume. You'll learn how Kafka supports scalable and reliable data collection across a cluster and how to set it up to publish and consume data. Additionally, you'll discover how Flume's architecture differs from Kafka and how to use it for real-time data ingestion. Through hands-on activities, you'll configure Kafka to monitor Apache logs and Flume to watch directories, publishing incoming data into HDFS. These skills will help you manage and process streaming data effectively in your Hadoop environment.

涵盖的内容

6个视频1个作业

In this module, we will focus on analyzing streams of data using real-time processing frameworks such as Spark Streaming, Apache Storm, and Flink. You’ll start by learning how Spark Streaming processes micro-batches of data in real-time and participate in activities that include analyzing web logs streamed by Flume. The module then introduces Apache Storm and Flink, providing hands-on exercises to implement word count applications with these tools. By the end of this module, you will be able to build continuous applications that efficiently process and analyze streaming data.

涵盖的内容

8个视频

8个视频总计76分钟

Spark Streaming: Introduction14分钟
Activity - Analyzing Web Logs Published with Flume using Spark Streaming14分钟
Exercise - Monitor Flume-Published Logs for Errors in Real Time2分钟
Exercise Solution: Aggregating the Hypertext Transfer Protocol (HTTP) Access Codes with Spark Streaming4分钟
Apache Storm: Introduction9分钟
Activity - Counting Words with Storm15分钟
Flink: Overview7分钟
Activity - Counting Words with Flink10分钟

In this module, we will focus on designing and implementing real-world systems using a combination of Hadoop ecosystem tools. You'll start by exploring additional technologies like Impala, NiFi, and AWS Kinesis, learning how they fit into broader Hadoop-based solutions. The module then guides you through the process of understanding system requirements and designing applications that consume and analyze large-scale data, such as web server logs or movie recommendations. By the end of this module, you’ll be equipped to design and build complex, efficient, and scalable data systems tailored to specific business needs.

涵盖的内容

7个视频1个作业

7个视频总计53分钟

The Best of the Rest9分钟
Review: How the Pieces Fit Together?6分钟
Understanding Your Requirements8分钟
Sample Application: Consuming Web Server Logs and Keeping Track of Top-Sellers10分钟
Sample Application: Serving Movie Recommendations to a Website11分钟
Exercise - Designing a System to Report Web Sessions Per Day3分钟
Exercise Solution: Designing a System to Count Daily Sessions4分钟

1个作业总计15分钟

Assessment 415分钟

In this final module, we will provide you with a selection of books, online resources, and tools recommended by the author to further your knowledge of Hadoop and related technologies. This module serves as a guide for continued learning, offering you the means to stay updated with the latest developments in the Hadoop ecosystem and expand your skills beyond this course.

涵盖的内容

1个视频1篇阅读材料1个作业

获得职业证书

将此证书添加到您的 LinkedIn 个人资料、简历或履历中。在社交媒体和绩效考核中分享。

位教师

Packt - Course Instructors

Packt

1,895 门课程531,687 名学生

提供方

Packt

从 Data Management 浏览更多内容

Packt
Apache Spark with Scala – Hands-On with Big Data!
课程
Packt
Streaming Big Data with Spark Streaming, Scala, and Spark 3!
课程
Packt
Apache Kafka Series - Learn Apache Kafka for Beginners v3
课程

人们为什么选择 Coursera 来帮助自己实现职业发展

Felipe M.

自 2018开始学习的学生

''能够按照自己的速度和节奏学习课程是一次很棒的经历。只要符合自己的时间表和心情，我就可以学习。'

Jennifer J.

自 2020开始学习的学生

''我直接将从课程中学到的概念和技能应用到一个令人兴奋的新工作项目中。'

Larry W.

自 2021开始学习的学生

''如果我的大学不提供我需要的主题课程，Coursera 便是最好的去处之一。'

Chaitanya A.

''学习不仅仅是在工作中做的更好：它远不止于此。Coursera 让我无限制地学习。'

常见问题

Yes, you can preview the first video and view the syllabus before you enroll. You must purchase the course to access content not included in the preview.

If you decide to enroll in the course before the session start date, you will have access to all of the lecture videos and readings for the course. You’ll be able to submit assignments once the session starts.

Once you enroll and your session begins, you will have access to all videos and other resources, including reading items and the course discussion forum. You’ll be able to view and submit practice assessments, and complete required graded assignments to earn a grade and a Course Certificate.

The Ultimate Hands-On Hadoop

The Ultimate Hands-On Hadoop

您将学到什么

您将获得的技能

您将学习的工具

要了解的详细信息

了解顶级公司的员工如何掌握热门技能

积累特定领域的专业知识

该课程共有12个模块

Learning All the Buzzwords and Installing the Hortonworks Data Platform Sandbox

涵盖的内容

Using the Hadoop's Core: Hadoop Distributed File System (HDFS) and MapReduce

涵盖的内容

Programming Hadoop with Pig

涵盖的内容

Programming Hadoop with Spark

涵盖的内容

Using Relational Datastores with Hadoop

涵盖的内容

Using Non-Relational Data Stores with Hadoop

涵盖的内容

Querying Data Interactively

涵盖的内容

Managing Your Cluster

涵盖的内容

Feeding Data to Your Cluster

涵盖的内容

Analyzing Streams of Data

涵盖的内容

Designing Real-World Systems

涵盖的内容

Learning More

涵盖的内容

获得职业证书

位教师

提供方

从 Data Management 浏览更多内容

Apache Spark with Scala – Hands-On with Big Data!

Streaming Big Data with Spark Streaming, Scala, and Spark 3!

Apache Kafka Series - Learn Apache Kafka for Beginners v3

人们为什么选择 Coursera 来帮助自己实现职业发展

Felipe M.

Jennifer J.

Larry W.

Chaitanya A.

Coursera Plus 3 个月课程 4 折优惠，让您省钱无忧

推动业务发展，增强团队能力

常见问题

Can I preview a course before enrolling?

When will I have access to the lectures and assignments?

What will I get when I enroll?

更多问题