Wang Huafeng

18621585140
unicor_n@163.com

Work

Intel

Shanghai

Software Engineer, Smart Storage Management

April 2017 - Present
  • SSM is a storage management system built on top of Apache HDFS. It collects file related metrics and automatically optimizes HDFS storage policies based on these metrics data along with user's pre-defined rules. As a core contributor, I participated in multiple rounds of project architecture design, implemented the data-collecting, data-maintenance, distributed execution service and other core features, built the dashboard prototype, integrated Travis and Codecov to the project.
  • github.com/Intel-bigdata/SSM

Software Engineer, MPICH on Yarn

Feb 2017 - Mar 2017
  • MPICH is a high performance and widely portable implementation of the Message Passing Interface (MPI) standard. Some deep learning frameworks, such as Intel-Caffe, use MPICH as their underlying communication framework. To bring these frameworks into Hadoop world and leverage Yarn's resource management, I built a prototype that can run MPICH applications on Yarn by myself. Now the two examples of MPICH, CPI and Parent & Child, can run on Yarn successfully.
  • github.com/huafengw/mpich-on-yarn

Software Engineer, HiBench

Jun 2016 - Sep 2016

Software Engineer, Apache Gearpump

July 2014 - Present
  • Gearpump is an Intel open sourced lightweight real-time big data streaming engine built on top of Akka. I contributed high availability, dynamic DAG, message loss detection and recovery, CGroup support and other core features to Gearpump. I also performed several rounds of performance tuning of Gearpump, optimized critical hotspot functions and stabilized the JVM’s GC behavior which brought about 35% performance improvement against the initial result. Currently, Gearpump is an incubating project under Apache Software Foundation.
  • github.com/apache/incubator-gearpump

Software Engineer, Open Source

July 2014 - Present
  • Work in open source community. Contributed patches to Apache Hadoop, Apache Beam, Apache Flink.

Software Engineer Intern, Intel Hadoop Distribution

July 2013 - April 2014
  • NativeTask is a high-performance C++ API & runtime for Hadoop MapReduce, I contributed its support for Apache Mahout and Hadoop 2.0 version, and evaluated its performance improvement using HiBench. The final result shows 30% boost up compared with the original MapReduce implementation. NativeTask will be released in Hadoop 3.0.

Skills

  • Java - Scala - C/C++ - Shell - Javascript - SQL
  • Apache Hadoop - Akka - Apache Kafka - Apache Storm

Education

Nanjing University

Nanjing Jiangsu

Undergraduate Software Engineering, GPA 4.35/5

Sep 2010 - Jun 2014

    Reference

    Speech

    Apache Gearpump: Next-Gen Streaming Engine

    Apache Big Data Europe 2016

    Streaming Report: Functional Comparison and Performance Evaluation

    Apache Big Data Europe 2016