Wednesday, June 13, 2012

Apache Flume 1.2.x and HBase

The newest (and first) HBase sink was committed into trunk one week ago and was my point at the HBase workshop @Berlin Buzzwords. The slides are available in my slideshare channel.

Let me explain how it works and how you get an Apache Flume - HBase flow running. First, you've got to checkout trunk and build the project (you need git and maven installed on your system):

git clone git://git.apache.org/flume.git && cd flume && git checkout trunk && mvn package -DskipTests && cd flume-ng-dist/target

Within trunk, the HBase sink is available in the sinks - directory (ls -la flume-ng-sinks/flume-ng-hbase-sink/src/main/java/org/apache/flume/sink/hbase/)

Please note a few specialities:
The sink controls atm only HBase flush (), transaction and rollback. Apache Flume reads out the $CLASSPATH variable and uses the first available hbase-site.xml. If you use different versions of HBase on your system please keep that in mind. The HBase table, columns and column family have to be created. Thats all.

The using of an HBase sink is pretty simple, an valid configuration could look like:

host1.sources = src1
host1.sinks = sink1 
host1.channels = ch1 
host1.sources.src1.type = seq 
host1.sources.src1.port = 25001
host1.sources.src1.bind = localhost
host1.sources.src1.channels = ch1
host1.sinks.sink1.type = org.apache.flume.sink.hbase.HBaseSink 
host1.sinks.sink1.channel = ch1
host1.sinks.sink1.table = test3
host1.sinks.sink1.columnFamily = testing
host1.sinks.sink1.column = foo
host1.sinks.sink1.serializer = org.apache.flume.sink.hbase.SimpleHbaseEventSerializer
host1.sinks.sink1.serializer.payloadColumn = pcol
host1.sinks.sink1.serializer.incrementColumn = icol 
host1.channels.ch1.type=memory

In this example we start a Seq interface on localhost with a listening port, point the sink to the HBase sink jar and define the event serializer. Why? HBase needs the data in a HBase format, to achieve that we need to transform the input into a HBase compilant format. Apache Flume's HBase sink uses synchronous / blocking client, asynchronous support will follow (FLUME-1252). 

Links:

10 comments:

  1. Hi !
    Thanks for informations. I tried that and i have the following result http://pastebin.com/0YZBw8YL . Flume and ZK don't seem to communicate together :(
    Have you an idea ?

    ReplyDelete
  2. Hi,

    didn't work on a SASL Zookeeper I guess. Hmm, the certificate are readable? If yes, file pls a jira about.
    http://hbase.apache.org/configuration.html#zk.sasl.auth

    Thanks,
    Alex

    ReplyDelete
  3. Sorry but I don't use SASL ZK. :/

    ReplyDelete
    Replies
    1. The pastebin wasn't clear about, so I fired the gun. The nio exception means _mostly_ a heavy used zookeeper cluster. Are HBase and flume installed on the same host? I would point into a problem there. If the error persists please write a mail to the mailinglist.

      Delete
    2. I use HBase on 8 nodes, but it's just for benchmark so my cluster is never used... In addition Zookeeper is deployed on 3 nodes and i have the same problem on all nodes.
      Which mailinglist do you want i notify ?
      Thanks for your work and your help :)

      Delete
    3. That's work now ! Thanks for your helpfull informations.
      Your work is fantasic :)

      Delete
  4. thank you so much alo for the wonderful work..using this writeup I was able to use hbase-sink..but being a newbie I was left some questions..could you please tell something about following 2 line -

    host1.sinks.sink1.serializer.payloadColumn = pcol
    host1.sinks.sink1.serializer.incrementColumn = icol

    Many thanks

    ReplyDelete
  5. I was very pleased to find this site. I wanted to thank you for this unique read. I definitely savoured all bits and pieces of it including all the comments and I have added you to my bookmark list to check out new articles you post.
    http://celabright.biz

    ReplyDelete
  6. I just started looking at Flume and it seems great!
    However I still doesn't find a complete example to start play with..
    after building Flume, what do I need to do to start feeding the HBase Sink?
    Where do I have to put the configuration file..?
    Is there an Avro Source to HBase complete example..?

    Best,
    Flavio

    ReplyDelete
  7. Flavio, you can easily add a Avro source and deliver the catched data into HBase, if you want. I guess we have some examples in our user guide, http://flume.apache.org/FlumeUserGuide.html

    ReplyDelete