Experience Required: 10+ Years
Required Technical/Functional Skills
- 10+ years of experience working with the Hadoop ecosystem and large-scale distributed data platforms.
- Strong hands-on experience with:
- HDFS
- Apache Hive
- Apache Spark
- Apache Ozone
- Strong experience with Hive Metastore management, including table and partition maintenance.
- Experience with Hadoop data lifecycle management, storage management, and data retention.
- Strong knowledge of data ingestion, ETL processes, data validation, and troubleshooting.
- Proficiency in SQL, Shell scripting, and Python automation.
- Experience with Cloudera Data Platform (CDP) or similar Hadoop distributions.
- Strong data analysis, troubleshooting, and root cause analysis skills.
Roles & Responsibilities
- Monitor, troubleshoot, and resolve HDFS and Apache Ozone issues.
- Investigate Hadoop cluster performance, storage utilization, capacity, and availability concerns.
- Manage storage allocation, quota requests, capacity planning, and storage forecasting.
- Perform Hive Metastore administration, including table and partition maintenance.
- Support Hadoop platform maintenance, upgrades, integrations, onboarding, and decommissioning activities.
- Support data ingestion, ETL workflows, and data validation processes.
- Develop and maintain Shell and Python automation scripts for operational activities.
- Collaborate with application, data engineering, and platform teams during incident, problem, and change management activities.
- Perform root cause analysis and implement corrective and preventive actions.
- Create and maintain runbooks, operational documentation, knowledge articles, and support procedures.
- Ensure effective Hadoop data lifecycle, storage, and quota governance.
Preferred / Desired Skills
- Experience supporting large-scale, multi-petabyte Hadoop environments.
- Strong understanding of distributed storage architecture, replication, and fault tolerance.
- Experience with capacity management, storage forecasting, quota governance, and utilization optimization.
- Experience with Hadoop platform onboarding, integration, migration, and decommissioning.
- Familiarity with AWS, Azure, or GCP cloud platforms.
- Experience with monitoring and observability tools such as Cloudera Manager, Grafana, Splunk, and Prometheus.
- Strong analytical, troubleshooting, and problem-solving skills.
- Experience performing Root Cause Analysis (RCA) for complex production issues.
- Strong communication and collaboration skills.