Design, develop, and maintain scalable ETL pipelines and data infrastructure for large-scale structured and unstructured datasets. Integrate and govern data across sources, optimize pipeline performance, and deploy cloud-based solutions. Collaborate with data scientists, analysts, and engineers to deliver accessible datasets. Implement automation, testing, security, compliance, monitoring, and continuous improvements across data engineering workflows.
Responsibilities:
- Design, develop, and maintain scalable ETL pipelines to process and transform large-scale datasets.
- Integrate structured and unstructured data from multiple sources, ensuring quality, security, and consistency.
- Collaborate with data scientists, analysts, and software engineers to deliver well-structured and accessible datasets.
- Build and optimize data infrastructure using big data technologies such as Apache Spark, Hadoop, and Kafka.
- Deploy and manage cloud-based data solutions on AWS, GCP, or Azure.
- Monitor and troubleshoot data pipeline performance, ensuring reliability and efficiency.
- Implement data governance, security, and compliance best practices.
- Drive automation, testing strategies, and continuous improvements in data engineering workflows.
Qualifications:
- 5 years of related experience with a Bachelor’s degree or equivalent work experience.
- Advanced proficiency in SQL and experience with relational and NoSQL databases (PostgreSQL, MySQL, MongoDB, etc.).
- Strong programming skills in Python, Java, or Scala for data processing and automation.
- Deep expertise in ETL processes, data modeling, and data warehousing.
- Hands-on experience with big data frameworks such as Apache Spark, Hadoop, or Kafka.
- Proficiency in cloud platforms (AWS Redshift, Google BigQuery, Azure Synapse) and data infrastructure automation.
- Experience optimizing data pipeline performance and scalability.
- Strong problem-solving skills with the ability to work on complex, large-scale datasets.
- Knowledge of data governance, security, and compliance best practices.
- Excellent leadership, collaboration, and communication skills to work effectively across teams.
Similar Jobs
Digital Media • Information Technology • News + Entertainment
Designs, develops, and manages advanced data architectures, ingestion frameworks, and pipelines. Ensures data quality, lineage, integrity, accessibility, privacy, and regulatory compliance across AWS, Databricks, Kubernetes, and Teradata environments. Builds APIs, database views, and data extracts; selects appropriate storage platforms; applies transformation rules; monitors data quality; and collaborates with cross-functional teams to optimize data processing and resolve issues. Requires independent judgment and availability for variable schedules, including nights and weekends.
Top Skills:
Apache AirflowAPIsAWSDatabricksKubernetesPythonTeradata
Digital Media • Information Technology • News + Entertainment
Design, implement, and administer SAP Analytics Cloud (SAC) planning and reporting solutions. Lead SAC modelling, story creation, SAC admin, and migrations from BPC to SAC. Collaborate with data, product, and functional teams to define data acquisitions, virtual models in HANA, pipelines, and security. Provide support, training, and stabilization efforts while managing multiple projects and mentoring team members.
Top Skills:
AbapAWSAzureEpmGCPSap Analytics Cloud (Sac)Sap AoeSap BpcSap BtpSap BwSap DatasphereSap HanaSap S/4 Hana
Information Technology
Design, develop, and maintain data solutions, pipelines, and ETL processes. Ensure data quality, support data migration across systems, optimize workflow performance, troubleshoot data issues, and promote data governance. Serve as a subject matter expert, make team decisions, mentor colleagues, and coordinate cross-functional data strategy efforts.
Top Skills:
Data GovernanceData IntegrationData MigrationData PipelinesETLInformatica Data Quality
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


