Build and lead reliable healthcare data pipelines using AWS, Databricks, Spark, PySpark, and SQL. Develop ingestion frameworks for clinical and semi-structured data, optimize Delta Lake and Iceberg tables, and create FHIR, HL7, and OMOP-based data models. Implement data quality, metadata, lineage, PHI privacy, CI/CD, infrastructure as code, observability, testing, and incident recovery. Govern data consumption across analytical engines while applying healthcare and regulated-data controls.
Requisition Number: 2384843
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
Required Qualifications:
Preferred Qualification:
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
- Build reliable batch, micro-batch, and event-driven pipelines on AWS and Databricks
- Develop reusable ingestion frameworks for REST APIs, FHIR Bulk Export, HL7 interfaces, databases, SFTP/file exchange, JSON/NDJSON, CSV, XML, PDFs, and clinical text
- Implement scalable Spark/PySpark and SQL transformations, including schema inference/evolution, checkpointing, idempotency, retries, backfills, and replay
- Design and maintain Delta Lake and Apache Iceberg tables, including physical design, partitioning/clustering, compaction, file sizing, incremental reads/writes, performance tuning, and cost optimization
- Build curated healthcare data models and transformations using FHIR, HL7, OMOP CDM, and clinical terminology mappings
- Implement data-quality frameworks: schema validation, referential-integrity checks, business-rule testing, anomaly detection, reconciliation, completeness checks, and data-quality observability
- Build metadata and lineage capture across source systems, pipeline runs, code versions, transformation rules, mappings, and published data products
- Implement privacy-aware data processing for PHI, including access controls, masking, tokenization/pseudonymization, de-identification, and auditable handling patterns
- Deliver CI/CD pipelines, automated unit/integration/data tests, Terraform or CloudFormation, containerized services, monitoring, alerting, runbooks, and incident-recovery procedures
- Integrate and optimize governed data consumption through Databricks SQL, Athena, Snowflake, Trino, or equivalent engines
- Comply with all applicable Company policies, procedures, and business directives, changes including those relating to work location, team assignments, work schedules, and flexible work arrangements
Required Qualifications:
- Undergraduate degree or equivalent experience
- 4+ years of production data-engineering experience
- Production Databricks experience with Auto Loader, Delta Lake, Workflows, Unity Catalog, notebooks/jobs, SQL Warehouses, and cluster/job optimization
- Hands-on AWS experience with S3, IAM, VPC, ECS/EKS or Lambda, Step Functions, EventBridge, CloudWatch, Secrets Manager, and KMS
- Experience with Git, pull requests, CI/CD, Docker, Terraform/CloudFormation, automated testing, observability, and incident response
- Experience processing semi-structured data and building resilient ingestion pipelines with quality controls, error handling, and replay capability
- Practical Apache Iceberg knowledge, including tables, catalogs, snapshots, schema/partition evolution, compaction, and interoperability with query engines
- Healthcare data knowledge: FHIR R4 and/or HL7 v2, OMOP CDM, clinical terminologies, PHI, HIPAA-aligned engineering controls, and de-identification concepts
- Advanced Python, SQL, and Apache Spark/PySpark; solid understanding of distributed processing and performance tuning
- Demonstrated responsible use of AI coding tools and the ability to critically review, test, and productionize generated code.
Preferred Qualification:
- Healthcare or regulated-data experience
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Optum Chennai, Tamil Nadu, IND Office
Chennai, India, India
Similar Jobs at Optum
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Facilitates behavioral, professional, and capability development programs; coordinates learning calendars, stakeholders, vendors, and program metrics; manages learning projects from planning through completion; owns onboarding readiness and delivery for the Chennai site; maintains learning documentation and reports; and builds trusted partnerships with business and HR leaders to improve employee development experiences.
Top Skills:
ExcelMicrosoft 365Microsoft TeamsOutlookPowerPointSharepoint
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads enterprise software design, development, architecture, deployment, and production support using Trizetto Facets, SQL, Python, or Java. Drives REST APIs, microservices, cloud adoption, CI/CD, DevOps, security, reliability, and AI/ML initiatives. Provides technical leadership, stakeholder collaboration, mentoring, code reviews, architecture ownership, root-cause analysis, and continuous improvement across large-scale healthcare applications.
Top Skills:
Azure Ai StudioAzure DevopsCloud ComputingDistributed SystemsDockerGitGithub ActionsInfrastructure As CodeJavaJenkinsKubernetesLangchainMicroservicesPythonPyTorchRestful ApisScikit-LearnSQLTensorFlowTrizetto Facets
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Develop full-stack applications using React or Flutter and Python microservices, integrate AI technologies, manage databases, implement security, and build AWS CI/CD infrastructure. Ensure adherence to the AWS Well-Architected Framework across scalability, reliability, performance, and cost optimization. Collaborate with cross-functional teams, own the application lifecycle, and mentor junior engineers.
Top Skills:
APIsAWSCi/CdFlutterMicroservicesNext.JsPythonReact
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

