Features

Features
Certification:
  • TAK
Dedicated training:
Number of training hours:
  • 16
Producer:
Training language:
  • polski
Training level:
  • Średniozaawansowany
Type of training:
  • stacjonarnie; online; warsztat

Description

Company Description

InfoShare Academy is a leading IT academy offering comprehensive educational programs in new technologies for companies. Since 2015, we have been supporting organizations in developing technology teams through dedicated courses in Machine Learning, DevOps, Data Engineering, Python, UX/UI Design, AWS, and Kubernetes. Our training is based on practical skills and real business cases. We collaborate with over 300 industry practitioners, ensuring that our programs are tailored to current market needs. We specialize in reskilling and upskilling employees. With us, you will build effective teams implementing new technologies that will accelerate innovation and strengthen your company's competitiveness in the market. Check out our training offerings designed for companies, created to enhance your employees' competencies in the IT field.

Training Description
  • The Hadoop training is an intensive two-day course focused on the practical application of this popular framework for processing and analyzing large datasets. The training program is designed so that participants gain solid theoretical foundations (20%) and develop their practical skills (80%) through numerous workshops and projects. The course is ideal for those looking to understand and utilize Hadoop in their projects.
  • Required technical skills:
  • Basic programming knowledge in Java or Python
  • Basic knowledge of data processing
  • Ability to work in a Unix/Linux environment
Who the training is for
  • Programmers and data engineers who want to expand their skills with Hadoop
  • Data scientists and data analysts wishing to process large datasets efficiently
  • IT specialists and big data professionals looking to utilize Hadoop in their projects
Goals

 

Benefits
  • You will learn:
  • How to effectively manage data in HDFS and create MapReduce tasks
  • How to process and analyze data using Hive and Pig
  • How to optimize MapReduce tasks and manage Hadoop resources
  • How to deploy and monitor Hadoop applications in a production environment
Training Program

Day 1: Basics of Hadoop and Data Processing

  1. Hadoop Architecture

  • Overview of main Hadoop components: HDFS, MapReduce, YARN

  • Interaction between components

  1. Basics of HDFS and MapReduce

  • File management in HDFS

  • Creating and running basic MapReduce tasks

  1. Basics of Apache Hive and Apache Pig

  • Introduction to Hive: table structure, SQL queries

  • Analyzing file structure under Hive

  • Introduction to Pig: Pig Latin scripts

  1. Workshop: Data Processing using MapReduce

  • Implementing a simple MapReduce task

  • Analyzing results and optimizing the task

Day 2: Advanced Techniques and Practical Applications

  1. Advanced Data Processing

  • Writing advanced Hive queries

  • Creating complex Pig scripts

  1. Optimization and Performance Tuning

  • Techniques for optimizing MapReduce tasks

  • Managing resources in a Hadoop cluster

  1. Workshop: Data Analysis using Hive and Pig

  • Implementing Hive queries on real datasets

  • Creating Pig scripts for data processing

  1. Deploying and Managing Hadoop Clusters

  • Preparing and deploying Hadoop applications

  • Monitoring and managing Hadoop clusters in a production environment

  1. Cost Optimization – how to control and optimize costs arising from Redshift warehouses

Duration

16 h/2 days

Price includes
  • Certificate of completion
  • Monthly access to the training recording (in case of online format)
  • Customization of the training program to client needs

Zamów szkolenie