Designs, builds, maintains, and optimizes data pipelines, ETL processes, data warehouses, and data lakes. Collaborates with data scientists, analysts, software engineers, and stakeholders to develop scalable data models and infrastructure. Responsibilities include ensuring data quality, security, governance, reliability, and performance; testing and troubleshooting pipelines; documenting systems; and evaluating emerging data engineering technologies.
Job Description:
As a Data Engineer, you will play a critical role in the development, implementation, and maintenance of data infrastructure and systems. Your primary responsibility will be to design, build, and optimize data pipelines and data warehouses, ensuring the efficient and reliable collection, storage, and processing of large volumes of data. You will collaborate with cross-functional teams, including data scientists, analysts, and software engineers, to understand data requirements and translate them into scalable solutions. Your work will enable the organization to extract valuable insights, drive data-based decision-making, and support various business initiatives.
Responsibilities:
- Design, develop, and maintain data pipelines and ETL processes to efficiently ingest, transform, and load data from various sources into data warehouses and data lakes.
- Collaborate with data scientists, analysts, and business stakeholders to understand data requirements and design data models that facilitate efficient data retrieval and analysis.
- Optimize data pipeline performance, ensuring scalability, reliability, and data integrity.
- Implement data governance and security measures to ensure compliance with data privacy regulations and protect sensitive information.
- Identify and implement appropriate tools and technologies to enhance data engineering capabilities and automate processes.
- Conduct thorough testing and validation of data pipelines to ensure data accuracy and quality.
- Monitor and troubleshoot data pipelines to identify and resolve issues, ensuring minimal downtime.
- Develop and maintain documentation, including data flow diagrams, technical specifications, and user guides.
- Collaborate with software engineers and infrastructure teams to optimize data infrastructure, including storage, processing, and retrieval systems.
- Stay up-to-date with emerging trends and technologies in the field of data engineering, and recommend innovative solutions to improve efficiency and performance.
Requirements:
- Bachelor's degree in Computer Science, Engineering, or a related field. A master's degree is a plus.
- Proven experience as a Data Engineer or in a similar role, with a strong understanding of data engineering concepts, practices, and tools.
- Proficiency in programming languages such as Python, Java, or Scala, and experience with data manipulation and transformation frameworks/libraries (e.g., Apache Spark, Pandas, SQL).
- Solid understanding of relational databases, data modeling, and SQL queries.
- Experience with distributed computing frameworks, such as Apache Hadoop, Apache Kafka, or Apache Flink.
- Knowledge of cloud platforms (e.g., AWS, Azure, GCP) and experience with cloud-based data engineering services (e.g., Amazon Redshift, Google BigQuery, Azure Data Factory).
- Familiarity with data warehousing concepts and technologies (e.g., dimensional modeling, columnar databases).
- Strong problem-solving skills and the ability to analyze complex data-related issues.
- Excellent communication and collaboration skills, with the ability to work effectively in cross-functional teams.
- Attention to detail and a commitment to delivering high-quality work within specified timelines.
Similar Jobs
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Analyze business and user requirements across financial services value streams, define technical requirements, develop cloud solution options, and manage business and IT stakeholders. The role requires understanding existing framework capabilities, complex data models, cross-functional solution design, clear documentation, team coordination across regions, and familiarity with the SDLC from design through testing and implementation.
Top Skills:
Google Cloud Platform
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Provide L1-L3 production support for enterprise MongoDB Atlas environments, including monitoring, scaling, backups, upgrades, performance tuning, security, networking, disaster recovery, and incident resolution. Build Terraform-based infrastructure automation, support multi-cloud deployments across Azure, AWS, and GCP, partner with application and engineering teams, coordinate vendor escalations, and participate in 24/7 on-call coverage.
Top Skills:
AWSAzureGCPGitGithub ActionsGoJavaKubernetesMongodb AtlasPowershellPythonShellSQL ServerTerraform
Digital Media • Information Technology • News + Entertainment
Design, implement, and administer SAP Analytics Cloud (SAC) planning and reporting solutions. Lead SAC modelling, story creation, SAC admin, and migrations from BPC to SAC. Collaborate with data, product, and functional teams to define data acquisitions, virtual models in HANA, pipelines, and security. Provide support, training, and stabilization efforts while managing multiple projects and mentoring team members.
Top Skills:
AbapAWSAzureEpmGCPSap Analytics Cloud (Sac)Sap AoeSap BpcSap BtpSap BwSap DatasphereSap HanaSap S/4 Hana
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.



