Big Data Integration and Processing

Big Data Integration and Processing

Archived
Course
en
English
20 h
This content is rated 0 out of 5

You can't access an archived course

Source
  • From www.coursera.org
Conditions
  • Self-paced
  • Free Access
  • Fee-based Certificate
More info
  • 6 Sequences
  • Introductive Level

You can't access an archived course

Their employees are learning daily with Edflex

  • Safran
  • Air France
  • TotalEnergies
  • Generali
Learn more

Course details

Syllabus

  • Week 1 - Welcome to Big Data Integration and Processing
    Welcome to the third course in the Big Data Specialization. This week you will be introduced to basic concepts in big data integration and processing. You will be guided through installing the Cloudera VM, downloading the data sets to be used for this course, ...
  • Week 1 - Retrieving Big Data (Part 1)
    This module covers the various aspects of data retrieval and relational querying. You will also be introduced to the Postgres database.
  • Week 2 - Retrieving Big Data (Part 2)
    This module covers the various aspects of data retrieval for NoSQL data, as well as data aggregation and working with data frames. You will be introduced to MongoDB and Aerospike, and you will learn how to use Pandas to retrieve data from them.
  • Week 3 - Big Data Integration
    In this module you will be introduced to data integration tools including Splunk and Datameer, and you will gain some practical insight into how information integration processes are carried out.
  • Week 4 - Processing Big Data
    This module introduces Learners to big data pipelines and workflows as well as processing and analysis of big data using Apache Spark.
  • Week 5 - Big Data Analytics using Spark
    In this module, you will go deeper into big data processing by learning the inner workings of the Spark Core. You will be introduced to two key tools in the Spark toolkit: Spark MLlib and GraphX.
  • Week 6 - Learn By Doing: Putting MongoDB and Spark to Work
    In this module you will get some practical hands-on experience applying what you learned about Spark and MongoDB to analyze Twitter data.

Prerequisite

None.

Instructors

Ilkay Altintas
Chief Data Science Officer
San Diego Supercomputer Center

Amarnath Gupta
Director, Advanced Query Processing Lab
San Diego Supercomputer Center (SDSC)

Editor

The University of California, San Diego is a public land-grant research university in San Diego, California. Established in 1960 near the pre-existing Scripps Institution of Oceanography, UC San Diego is the southernmost of the University of California's ten campuses and offers more than 200 undergraduate and graduate degree programs, enrolling 33,096 undergraduate students and 9,872 graduate students. 

UC San Diego is considered one of the best universities in the world. Several publications have ranked UC San Diego's Departments of Biological Sciences and Computer Science among the top 10 in the world.

Platform

Coursera is a digital company offering massive open online course founded by computer teachers Andrew Ng and Daphne Koller Stanford University, located in Mountain View, California. 

Coursera works with top universities and organizations to make some of their courses available online, and offers courses in many subjects, including: physics, engineering, humanities, medicine, biology, social sciences, mathematics, business, computer science, digital marketing, data science, and other subjects.

This content is rated 4.5 out of 5
(no review)
This content is rated 4.5 out of 5
(no review)
Complete this resource to write a review