Post 39 | HDPCD | Load data into a Hive table from an HDFS directory

Hello, everyone. Thanks for returning for the next tutorial in the HDPCD certification series. In the last tutorial, we saw how to load data into a Hive table from a local directory. In this tutorial, we are going to see how to load the data from the local Directory into the Hive table. Let us begin then.Continue reading “Post 39 | HDPCD | Load data into a Hive table from an HDFS directory”

Post 37 | HDPCD | Specifying delimiter of a Hive table

Hello, everyone. Thanks for coming back for one more tutorial in this HDPCD certification series. In the last tutorial, we saw how to specify the storage format of a Hive table. In this tutorial, we are going to see how to specify the delimiter of a Hive table. We are going to follow the processContinue reading “Post 37 | HDPCD | Specifying delimiter of a Hive table”

Post 34 | HDPCD | Defining Hive Table using an ORC File Format

Hi, everyone. Thanks for joining me today for this tutorial. In the last tutorial, we saw how to create a hive table using the SELECT query. In this tutorial, we are going to see how to create a hive table which stores the data in the ORC File Format. The process of creating this tableContinue reading “Post 34 | HDPCD | Defining Hive Table using an ORC File Format”

Post 29 | HDPCD | Define a Hive-managed Table

Hello, everyone. Welcome to the second post in the Data Analysis section of the HDPCD certification series. In the last tutorial, we saw the three ways in which we run the hive commands. In this tutorial, we are going to create the hive-managed table i.e. hive internal table. For creating a hive-managed or internal table,Continue reading “Post 29 | HDPCD | Define a Hive-managed Table”

Post 26 | HDPCD | Define an ALIAS for a User Defined Function

Hi, everyone. Thank you for returning again to this certification series. In the last tutorial, we saw the process of registering the jar file in the Apache PIG session. This tutorial is an extension to the previous one and in this, we are going to see how to define an alias for the UDF presentContinue reading “Post 26 | HDPCD | Define an ALIAS for a User Defined Function”

Post 25 | HDPCD | Register a Jar file of UDF in Apache Pig

Hello, everyone. Thanks for coming back again to continue with this certification series. In the last tutorial, we saw how to run any pig script with TEZ as the execution mode. In this tutorial, we are going to see how to register a JAR file to use the User Defined Function written and packages inside it.Continue reading “Post 25 | HDPCD | Register a Jar file of UDF in Apache Pig”

Post 2 | Machine Learning | Installations – R and Python

Hello, everyone, we are going to start off learning the concepts of Machine Learning. If you are following my blog posts on Hadoop and Big Data Analytics, then you will come to know I do give more importance on performing the hands-on exercises. Same is going to be the case for these tutorials. Here, weContinue reading “Post 2 | Machine Learning | Installations – R and Python”

Post 1 | Machine Learning | Introduction

Hello, people. In this new tutorial series, we are going to talk about the different aspects of the Machine Learning. As an aspiring Data Scientist, I always wanted to get my hands dirty with the concepts of Machine Learning and the Summar Break gave me exactly what I wanted – “TIME TO LEARN MACHINE LEARNINGContinue reading “Post 1 | Machine Learning | Introduction”