AstraZeneca Logo

AstraZeneca

Associate Data Engineer

Posted 10 Days Ago
Be an Early Applicant
In-Office
Chennai, Tamil Nadu, IND
Senior level
In-Office
Chennai, Tamil Nadu, IND
Senior level
Operate, monitor, and improve cloud ETL pipelines to ensure timely, high-quality data across global markets. Maintain Python/PySpark transformations, manage AWS S3/Redshift/EMR storage and processing, trigger and validate API jobs, troubleshoot incidents, perform root-cause analysis, collaborate with data providers/DevOps, keep documentation and schedules current, and drive automation and quality controls to increase reliability and throughput.
The summary above was generated by AI
Job Title: Associate Data Engineer
 
GCL : C3
 

Introduction to role:
 

Are you ready to keep mission-critical data flowing at global scale and turn incidents into improvements that boost reliability? As an Associate Data Engineer, you will run and enhance the cloud pipelines that power decisions across 85+ markets, ensuring timely, high-quality data reaches the people who need it most.

You will join a high-performing, digitally savvy team that partners across the enterprise to drive speed and precision. Your focus on automation, monitoring, and rapid incident response will translate into trusted data and smoother releases—accelerating how we deliver life-changing medicines. Can you picture yourself orchestrating robust pipelines that help colleagues move faster with confidence?
 

Accountabilities:
 

Pipeline Operations: Implement and monitor end-to-end data pipelines and ETL jobs across multiple stages to ensure on-time, high-quality delivery at scale.

Data Transformation: Maintain and modify Python (Pandas, PySpark) scripts in line with evolving business needs to improve data quality and performance.

Cloud Data Management: Manage data storage and protected data exchanges across AWS S3, Redshift, and EMR, keeping data flows accurate and compliant.

API Orchestration: Trigger and validate jobs using Postman and other API interfaces to keep schedules on track and detect issues early.

Data Flow Governance: Track inbound and outbound files, log exceptions, and maintain observability to prevent and detect data breaks.

Incident Response and Root Cause Analysis: Investigate and remediate pipeline failures or delays, implement durable fixes, and drive automation that reduces repeat incidents.

Teamwork and Collaborator Management: Work closely with data providers, data custodians, and DevOps teams to assure pipeline health and data accuracy across global collaborators.

Documentation and Versioning: Keep pipeline documentation, job schedules, and technical configurations up to date; support code enhancements and environment updates.

Quality Control: Participate in data quality procedures with Data Stewards to validate releases and safeguard trust in data products.

Continuous Improvement: Identify and implement opportunities to standardize, simplify, and automate operations, increasing reliability and throughput over time.
 

Essential Skills/Experience:
 

Python (PyCharm, Pandas, PySpark) for maintaining ETL scripts and automation routines
 

Postman for testing and triggering API-based job executions
 

SQL proficiency using tools such as DBeaver to query and validate relational data

AWS services proficiency across Redshift, S3, and EMR for processing and storage
 

WinSCP or equivalent tools for secure file transfers
 

Proactive, structured approach to monitoring and troubleshooting

Strong programming and analytical problem-solving abilities

Excellent documentation and organizational skills

Ability to work independently and coordinate across functional teams

Desirable Skills/Experience:
 

Familiarity with Git and version control systems
 

5–8 years of experience in data engineering, production support, or data operations
 

Background handling large-scale data workflows in cloud environments
 

Experience working in pharmaceutical or healthcare data ecosystems

Consistent track record resolving performance bottlenecks and job failures

Familiarity with DevOps principles and agile ways of working
 

Why AstraZeneca:
 

Here, data engineering fuels real-world impact. You’ll work with modern cloud platforms and digital tools, side by side with unexpected combinations of experts—engineers, data stewards, and market teams in the same room—turning bold ideas into operational reality. We move with urgency and clarity, blending imagination with rigor to strengthen how the business runs today while preparing for tomorrow. Your contribution will help colleagues across the globe focus on what matters most, translating into faster, smarter decisions that ultimately benefit patients. We value patience alongside ambition, and we back curiosity with the support and autonomy needed to deliver significant results.

Call to Action:

If you’re ready to build resilient workflows that drive faster decisions and tangible patient impact, step forward and build what reliable data can make possible!

Date Posted

12-Aug-2026

Closing Date

27-Aug-2026

AstraZeneca embraces diversity and equality of opportunity.  We are committed to building an inclusive and diverse team representing all backgrounds, with as wide a range of perspectives as possible, and harnessing industry-leading skills.  We believe that the more inclusive we are, the better our work will be.  We welcome and consider applications to join our team from all qualified candidates, regardless of their characteristics.  We comply with all applicable laws and regulations on non-discrimination in employment (and recruitment), as well as work authorization and employment eligibility verification requirements.

AstraZeneca Chennai, Tamil Nadu, IND Office

Ramanujan IT City, 10th & 11th Floor Neville Tower 2nd Floor Hardy Towers, Tharamani, Chennai, Tamil Nadu, India, 600113

Similar Jobs

3 Days Ago
In-Office
Chennai, Tamil Nadu, IND
Mid level
Mid level
Artificial Intelligence • Information Technology • Machine Learning • Software • Virtual Reality • Analytics
Design, develop, deploy, and maintain scalable real-time data pipelines using Java, Apache Kafka, and Apache Flink. Build low-latency, high-throughput streaming solutions and event-driven applications for analytics and operational reporting. Collaborate with architects and stakeholders, monitor and troubleshoot production systems, optimize performance and fault tolerance, and implement data quality, security, governance, and operational best practices.
Top Skills: Apache FlinkApache KafkaCloud PlatformsContainerizationJavaMicroservicesRest Apis
Yesterday
In-Office
Chennai, Tamil Nadu, IND
Senior level
Senior level
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Define, package, price and drive go-to-market for telco cloud and HCP-based managed services. Lead service lifecycle, enable sales and delivery, engage customers and partners, prioritize tools/resources, and optimize profitability and operational scalability.
Top Skills: CnfHybrid CloudHyperscaler Cloud PlatformsMulti-CloudOrchestration And Automation LayersPlatform-As-A-ServiceTelco CloudVnfZero-Touch Operations
Yesterday
Hybrid
Chennai, Tamil Nadu, IND
Junior
Junior
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Processes and fulfills transactional, batch, and self-service data requests within service-level agreements. Validates inputs, system setups, data extraction, cleansing, outputs, and reports using Ab Initio on UNIX. Coordinates with customer engagement and sales teams, resolves data concerns and complaints, ensures quality and turnaround targets, follows documented procedures, and identifies process improvements. The role requires SQL and Python expertise, attention to detail, problem-solving ability, and willingness to work US shifts.
Top Skills: Ab InitioAutosysMS OfficePythonSQLUnix

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account