Home | About | Sematext search-lucene.com search-hadoop.com
 Search Hadoop and all its subprojects:

Switch to Plain View
Pig, mail # user - group by clickstream


Copy link to this message
-
group by clickstream
Steve Bernstein 2012-08-29, 23:27
Hi all,
I have a bag, clickstreams: {clickStream: {pageName: chararray}}, for which each row represents a sequence of pages and events in a single session on a website.  The interior bag, clickstream, represents this as a sequence of one or more single element tuples, e.g.,

{(homepage),(pg1),(pg2),...,(pgN)}

I'd like to group by the sequences so I can get counts and ultimately sort to find the most common clickstreams.  A bag can't be a key for grouping, I've discovered, but it seems like it ought to be easy to flatten the clickstream bag into some other form such that the sequences can be used as keys for grouping.  But I can't figure it out.

Any ideas?

Thanks!
Steve

+
Steve Bernstein 2012-08-30, 16:06
+
Steve Bernstein 2012-08-30, 16:22
+
=?KOI8-U?B?96bUwcymyiD0yc... 2012-08-31, 20:48
+
Steve Bernstein 2012-08-31, 23:14