Design, build, and optimize end-to-end data pipelines for large structured and unstructured data. Implement near-real-time ETL, data validation, monitoring, and performance optimization. Collaborate with stakeholders, document designs and workflows, and provide technical guidance to the team.
Location: Pune
Responsibilities include:- Design, implement, and optimize end-to-end data pipelines for ingesting, processing, and transforming large volumes of structured and unstructured data.
- Develop data pipelines to extract and transform data in near real-time using cloud-native technologies.
- Implement data validation and quality checks to ensure accuracy and consistency.
- Monitor system performance, troubleshoot issues, and implement optimizations to enhance reliability and efficiency.
- Collaborate with business users, analysts, and other stakeholders to understand data requirements and deliver tailored solutions.
- Document technical designs, workflows, and best practices to facilitate knowledge sharing and maintain system documentation.
- Provide technical guidance and support to team members and stakeholders as needed.
- 8+ years of work experience.
- Proficiency in writing complex SQL queries on MPP systems (Snowflake/Redshift).
- Experience in Databricks and Delta tables.
- Data engineering experience with Spark/Scala/Python.
- Experience in Microsoft Azure stack (Azure Storage Accounts, Data Factory, and Databricks).
- Experience in Azure DevOps and CI/CD pipelines.
- Working knowledge of Python.
- Comfortable participating in 2-week sprint development cycles.
Photon Chennai, Tamil Nadu, IND Office
DLF IT Park 1/124 Mount Poonamallee Road Sivaji Gardens Manapakkam , Chennai, India, 600089
Similar Jobs
Agency • Information Technology
Design, build, optimize, and maintain high-performance Spark-based data pipelines using Scala/Java and Hive on Hadoop/CDP. Own full project lifecycle, enforce coding best practices, troubleshoot Spark/Hive/YARN performance, and collaborate with stakeholders to deliver scalable data solutions.
Top Skills:
SparkCloudera Data Platform (Cdp)HadoopHiveJavaScalaYarn
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads scalable UI and API test automation using Python, PyTest, Playwright, Robot Framework, and Selenium. Develops comprehensive test suites, validates APIs and databases, tests microservices and healthcare applications, and integrates quality checks into CI/CD pipelines. Oversees defect triage, risk-based testing, quality metrics, release readiness, and Agile collaboration. Applies AI tools, LLMs, machine-learning-based test optimization, prompt engineering, and intelligent failure analysis. Supports cloud, containerization, infrastructure-as-code, and healthcare interoperability testing.
Top Skills:
AlmAWSAzureAzure AiAzure BoardsAzure PipelinesClaudeCloudFormationDicomDockerDynamoDBEmbeddingsFhirGCPGitGithub ActionsGithub CopilotGitlab CiHipaaHl7Hugging FaceInsomniaJenkinsJIRAKafkaM365 CopilotMongoDBOpenaiPlaywrightPostmanPytestPythonRagRallyRequestsRest AssuredRobot FrameworkSauce LabsSeleniumSQLTerraformVector Databases
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Develops and maintains ETL/ELT pipelines, legacy SSIS and SSAS solutions, reporting platforms, and relational and dimensional data models. Tunes SQL, data pipelines, semantic layers, and reporting workloads for performance and reliability. Supports Snowflake and cloud-native data platforms, documents data lineage and procedures, responds to ServiceNow incidents, and follows HIPAA, ITIL, governance, and change-control standards. Collaborates across technical and analytics teams while supporting healthcare data systems and modernization initiatives.
Top Skills:
Apache AirflowAzure Data FactoryAzure Machine LearningAzure SynapseCozyroc Ssis+DbtKronosPower BIPythonServicenowSnowflakeSQL ServerSsasSsisSsrsT-SqlTableauUkg
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

