QUALIFICATIONS
- Minimum 7-8 years System/Database Administration experience with at least 3+ Years of Big Data Systems Administration experience required,
- Bachelor’s degree in Computer Science, Engineering or related field,
- Strong knowledge of Hadoop Architecture (HDFS), Hadoop Cluster installation, configuration, monitoring, cluster security, cluster resources management, maintenance and performance tuning,
- Strong knowledge of Hadoop components such as HDFS, Sentry, Kafka, Impala, Hue, Hive, ZooKeeper, HBase,
- Strong knowledge of key scripting and programming languages such as MapReduce, Spark, Python, Bash, Scala,
- Good knowledge of relational databases, industry practices, techniques and standards, (hands-on experience at least one of MySQL, MSSQL, PostgreSQL, Oracle)
- Good knowledge of NoSQL databases, industry practices, techniques and standards, (hands-on experience at least one of MongoDB, Cassandra, Elastic Search, Neo4j, Redis)
- Good knowledge of Linux Administration,
- Work individually, with minimal supervision, as well as in team environments,
- Detail oriented, well organized work skills,
- Multitasking, time and stress management,
- Strong ability to identify, prioritize and solve problems,
- Passionate about learning big data, new technologies , open source technologies,
- Fluent in English is an asset.
JOB DESCRIPTION
- Design, installation, configuration and administration of Hadoop platform and maintaining the Hadoop infrastructure and operations,
- Forecasting, planning, capacity arrangement and scaling of Hadoop clusters,
- Design and implement for 7/24 uptime requirements of Hadoop cluster connectivity, monitoring and security, conducting performance tuning of Hadoop clusters,
- Setting and implementing Backup and Disaster Recovery strategy,
- Automation and integration of monitoring and server job processes,
- Assist the team in designing and developing a scalable real-time infrastructure and data pipelines,
- Design and implement distributed data processing pipelines using Hadoop/Spark and Hive,
- Focus on performance, throughput, and latency to build and maintain our architecture,
- Analyze and optimize performance of systems, eliminate bottlenecks, plan for capacity upgrades.

