The common responsibilities for this position include designing, developing, and maintaining scalable ETL/ELT data pipelines for both structured and unstructured data, utilizing tools such as Apache Airflow, Python, and PySpark. Collaborating with business users and cross-functional teams to gather data requirements and provide actionable insights is essential. Implementing data quality, governance, and security measures throughout the data lifecycle, as well as optimizing performance and reliability of data solutions, is critical. Monitoring and troubleshooting data pipeline issues, conducting validation and compliance checks, and documenting data assets and workflows to promote transparency are key duties. Additionally, managing cloud infrastructure for data processing workloads and providing technical guidance and mentorship to junior engineers are important responsibilities. Supporting AI and machine learning initiatives by preparing high-quality datasets and integrating model outputs into operational workflows is also a significant part of the role.
The percentages next to each skill reflect the sector’s demands in these respective skills. E.g., 30% means this skill has been listed in 30% of all the job postings in this sector.
The skills distribution tells you what specific skill sets are in demand. E.g., Skills with a distribution of “More than 50%” means that these skills are wanted in more than 50% of the job postings.
Job classifications that have advertised a position
Academic degree required as indicated by all job postings
Job subclassifications that have advertised a position