Work | Intel Shanghai Software Engineer, Smart Storage ManagementApril 2017 - Present -
SSM is a storage management system built on top of Apache HDFS. It collects file related metrics and
automatically optimizes HDFS storage policies based on these metrics data along with user's pre-defined rules.
As a core contributor, I participated in multiple rounds of project architecture design, implemented
the data-collecting, data-maintenance, distributed execution service and other core features, built
the dashboard prototype, integrated Travis and Codecov to the project.
- github.com/Intel-bigdata/SSM
Software Engineer, MPICH on YarnFeb 2017 - Mar 2017 -
MPICH is a high performance and widely portable implementation of the Message Passing Interface (MPI) standard.
Some deep learning frameworks, such as Intel-Caffe, use MPICH as their underlying communication framework.
To bring these frameworks into Hadoop world and leverage Yarn's resource management, I built a prototype that
can run MPICH applications on Yarn by myself. Now the two examples of MPICH, CPI and Parent & Child, can
run on Yarn successfully.
- github.com/huafengw/mpich-on-yarn
Software Engineer, HiBenchJun 2016 - Sep 2016 Software Engineer, Apache GearpumpJuly 2014 - Present -
Gearpump is an Intel open sourced lightweight real-time big data streaming engine built on top of Akka.
I contributed high availability, dynamic DAG, message loss detection and recovery, CGroup support and other
core features to Gearpump. I also performed several rounds of performance tuning of Gearpump, optimized
critical hotspot functions and stabilized the JVM’s GC behavior which brought about 35% performance improvement
against the initial result. Currently, Gearpump is an incubating project under Apache Software Foundation.
- github.com/apache/incubator-gearpump
Software Engineer, Open SourceJuly 2014 - Present -
Work in open source community. Contributed patches to Apache Hadoop, Apache Beam, Apache Flink.
Software Engineer Intern, Intel Hadoop DistributionJuly 2013 - April 2014 -
NativeTask is a high-performance C++ API & runtime for Hadoop MapReduce, I contributed its support for
Apache Mahout and Hadoop 2.0 version, and evaluated its performance improvement using HiBench.
The final result shows 30% boost up compared with the original MapReduce implementation. NativeTask
will be released in Hadoop 3.0.
|