Chaining multiple MapReduce jobs in Hadoop

Question

2 Answers

Shivangi · Answer 1 · 2019-05-24T06:47:46+0000

There are multiple methods for the same, I’ll explain one of them here, you will be able to easily chain jobs together in this manner by writing multiple driver methods, one for each of them. First Call the first driver method, that uses JobClient.runJob() to run the job and wait for its completion. When the job has completed, then call next driver method, which will create a new JobConf object referring to different instances.

Create the JobConf object "one" for the first job.

Execute this job:

JobClient.run(one).

Then, create another JobConf object "two" for the next job.

Execute this job:

JobClient.run(two).

Amit Rawat · Answer 2 · 2019-09-18T10:12:32+0000

You apply the JobClient.runJob(). The output path of the data from the original job becomes the input path to your second job. These need to be passed in as parameters to your jobs with appropriate code to parse them and set up the parameters for the job.

I think that the above method might, however, be the way the now older mapred API did it, but it should still work. There will be a similar method in the new MapReduce API but I'm not sure what it is.

As far as eliminating intermediate data after a job has finished you can do this in your code. The way I've done it before is using something like:

FileSystem.delete(Path f, boolean recursive);

Where the path is the location on HDFS of the data. You need to make sure that you only delete this data once no other job requires it.

Refer the following video regarding Hadoop:

Chaining multiple MapReduce jobs in Hadoop

2 Answers

Related questions

Browse Categories

Browse By Domains

Popular Courses

Popular Tutorials

Popular Resources