Home | About | Sematext search-lucene.com search-hadoop.com
 Search Hadoop and all its subprojects:

Switch to Plain View
MapReduce, mail # user - Re: Retrieve and compute input splits


Copy link to this message
-
Re: Retrieve and compute input splits
Sai Sai 2013-09-27, 05:25
Hi
I have attached the anatomy of MR from definitive guide.

In step 6 it says JT/Scheduler  retrieve  input splits computed by the client from hdfs.

In the above line it refers to as the client computes input splits.
1. Why does the JT/Scheduler retrieve the input splits and what does it do.
If it is retrieving the input split does this mean it goes to the block and reads each record 
and gets the record back to JT. If so this is a lot of data movement for large files.
which is not data locality. so i m getting confused.

2. How does the client know how to calculate the input splits.

Any help please.
Thanks
Sai
+
Sonal Goyal 2013-09-27, 08:41
+
Peyman Mohajerian 2013-09-27, 23:02
+
Jay Vyas 2013-09-28, 00:05
+
Sai Sai 2013-09-30, 12:00