AT&T Logo

AT&T

Senior App/Prod Support (Tier 3 Site Reliability Engineer (SRE) / Platform Engineer)

Posted 3 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Hyderabad, Telangana
Senior level
In-Office or Remote
Hiring Remotely in Hyderabad, Telangana
Senior level
Lead platform reliability and DevOps automation: implement CI/CD with GitHub Actions, automate JFrog/Helm and image migrations, enable microservices deployments, and operate observability and logging stacks. Provide Tier 3 troubleshooting and incident leadership, manage cloud infrastructure governance, capacity and DR planning, cost/license governance, and maintain SOPs and reliability best practices.
The summary above was generated by AI

Key Responsibilities

- Own platform reliability practices for availability, resilience, latency, and operational efficiency.

- Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.

- Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.

- Lead JFROG Helm chart automation and JFROG images/ACR migration work.

- Support microservices deployment enablement and platform/tooling upgrades.

- Own and optimize monitoring, alerting, observability, and logging stack components:

  Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.

- Support health-check frameworks including Airflow health-check requirements.

- Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.

- Collaborate with architecture and delivery teams on reliability and scalability patterns.

- Lead cloud infrastructure creation, maintenance, governance, and access controls.

- Drive capacity planning, DR planning/exercises, and platform best-practice documentation.

- Support cost management, role enforcement, and license management governance.

- Maintain SOP documentation for established alerts and incident patterns.

Required Qualifications / Must-Have Skills

- 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.

- Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations.

- Advanced experience with CI/CD engineering and GitHub Actions.

- Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.

- Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.

- End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).

- Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).

- Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.

- Experience in governance controls: access management, role enforcement, and separation of duties.

- Proven high-severity incident leadership and post-incident reliability improvement execution.

Good-to-Have / Nice-to-Have

- Postgres performance and reliability operations.

- Telecom-scale high-availability systems experience.

Experience Level

Senior to Lead IC (typically 10 to 17 years)

Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)

What We Offer

- Opportunity to define and scale platform reliability standards.

- High technical ownership and strong cross-functional influence.

- Enterprise-scale impact across observability, automation, and resilience engineering.

Weekly Hours:

40

Time Type:

Regular

Location:

IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City, IND:KA:Bangalore / Intl Tech Park, Navigator Bldg, Whitefield Road: Whitefield Road:Intl Tech Park, Navigator Bldg

It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

Similar Jobs

37 Minutes Ago
Remote or Hybrid
India
Mid level
Mid level
Cloud • Information Technology • Security • Software • Cybersecurity
Own technical relationships across the customer lifecycle: lead discovery, design bespoke architectures, run demos and PoCs across security, networking and developer platforms, drive adoption and expansion with quota responsibility, run technical QBRs, and leverage AI-augmented workflows to automate tasks and focus on strategic advisory.
Top Skills: Ai AgentsApi ConnectorsCloud InfrastructureCloudflare Ai GatewayCloudflare WorkersDdos MitigationDnsJavaScriptLlm OrchestrationPrompt EngineeringPythonRetrieval-Augmented Generation (Rag)Routing (Bgp)Telemetry Monitoring
Yesterday
Remote or Hybrid
India
Expert/Leader
Expert/Leader
Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Develop and maintain statistical and ML forecasting models for SKU-level demand using Python/PySpark and Databricks; analyze model performance (MAPE/bias), post-process outputs, collaborate with demand planners for explainability, produce dashboards and KPIs, and drive continuous improvement in forecasting and data models.
Top Skills: DatabricksGoogle AdwordsGoogle AnalyticsGoogle Cloud PlatformGoogle Tag ManagerJuliaExcelPower BIPysparkPythonRSASSQLTableau
Yesterday
Easy Apply
Remote
India
Easy Apply
Senior level
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Lead the Help Center product roadmap to increase automation rate and CSAT by building AI-powered search, personalized article experiences, and self-service workflows. Partner with ML, engineering, CX ops, content, and data science to maintain global knowledge, define funnel metrics, and integrate chat/agent tooling and third-party vendors to deflect contacts and proactively resolve issues.
Top Skills: Ai/MlCmsContentfulFreshdeskGenerative AiLlmRagSalesforceZendesk

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account