Design, develop, and maintain scalable ETL pipelines and data infrastructure for large-scale structured and unstructured datasets. Integrate and govern data across sources, optimize pipeline performance, and deploy cloud-based solutions. Collaborate with data scientists, analysts, and engineers to deliver accessible datasets. Implement automation, testing, security, compliance, monitoring, and continuous improvements across data engineering workflows.
Responsibilities:
- Design, develop, and maintain scalable ETL pipelines to process and transform large-scale datasets.
- Integrate structured and unstructured data from multiple sources, ensuring quality, security, and consistency.
- Collaborate with data scientists, analysts, and software engineers to deliver well-structured and accessible datasets.
- Build and optimize data infrastructure using big data technologies such as Apache Spark, Hadoop, and Kafka.
- Deploy and manage cloud-based data solutions on AWS, GCP, or Azure.
- Monitor and troubleshoot data pipeline performance, ensuring reliability and efficiency.
- Implement data governance, security, and compliance best practices.
- Drive automation, testing strategies, and continuous improvements in data engineering workflows.
Qualifications:
- 5 years of related experience with a Bachelor’s degree or equivalent work experience.
- Advanced proficiency in SQL and experience with relational and NoSQL databases (PostgreSQL, MySQL, MongoDB, etc.).
- Strong programming skills in Python, Java, or Scala for data processing and automation.
- Deep expertise in ETL processes, data modeling, and data warehousing.
- Hands-on experience with big data frameworks such as Apache Spark, Hadoop, or Kafka.
- Proficiency in cloud platforms (AWS Redshift, Google BigQuery, Azure Synapse) and data infrastructure automation.
- Experience optimizing data pipeline performance and scalability.
- Strong problem-solving skills with the ability to work on complex, large-scale datasets.
- Knowledge of data governance, security, and compliance best practices.
- Excellent leadership, collaboration, and communication skills to work effectively across teams.
Similar Jobs
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Design and build data services and pipelines supporting machine learning products. Develop automation tools for model deployment, maintain production data systems, improve infrastructure, participate in code reviews, and collaborate across engineering, data science, and product teams. The role requires expertise in distributed systems, large-scale data processing, CI/CD, container orchestration, and AI-enabled workflow improvements.
Top Skills:
AirflowAWSAws BatchCi/CdDockerEmrGlueGoKafkaKubernetesKv StoresLinuxPythonRelational DatabasesSagemakerSparkSpinnaker
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Provide L1-L3 production support for enterprise MongoDB Atlas environments, including monitoring, scaling, backups, upgrades, performance tuning, security, networking, disaster recovery, and incident resolution. Build Terraform-based infrastructure automation, support multi-cloud deployments across Azure, AWS, and GCP, partner with application and engineering teams, coordinate vendor escalations, and participate in 24/7 on-call coverage.
Top Skills:
AWSAzureGCPGitGithub ActionsGoJavaKubernetesMongodb AtlasPowershellPythonShellSQL ServerTerraform
Automotive
Designs and maintains scalable ETL pipelines, data models, databases, and cloud-based analytics workflows. Manages PostgreSQL, MySQL, MongoDB, and Google Cloud SQL environments, performs database optimization, develops AutoML solutions, and deploys data engineering pipelines to production. Works within Agile teams using Scrum or Kanban, participating in sprint planning, backlog grooming, and daily stand-ups.
Top Skills:
AutomlData ModelingETLGoogle BigqueryGoogle Cloud DataflowGoogle Cloud DataprocGoogle Cloud PlatformGoogle Cloud SqlGoogle Cloud StorageKanbanMongoDBMySQLPostgresScrum
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.



