Develop and administer observability platforms supporting e-commerce, supply-chain, and enterprise services. Build monitoring integrations, manage telemetry agents, configure secure access controls, automate operations, and improve platform scalability, reliability, performance, and cost efficiency. Deploy distributed monitoring solutions, onboard applications, integrate enterprise systems, conduct health checks and problem isolation, and collaborate with engineering, security, product, and vendor teams in a SAFe environment.
As a Senior Software Engineer specializing in Observability, you will play a key role in Drive innovation in infrastructure and application monitoring solutions to support Staples’ B2C/B2B platforms, back-end supply chain, and related services. Enhance and administer the observability platforms like Zabbix, NewRelic, FullStory, Splunk, Grafana, PagerDuty by applying best practices, recommending improvements, and solving complex business challenges to boost service quality, availability, and performance. Identify and implement opportunities for automation, efficiency, and cost savings. Integrate external tools, participate in strategic planning, and collaborate with eCommerce, engineering, security, and product teams to deliver critical business objectives.
Requirements
Basic Qualifications
- NewRelic/Dynatrace/Zabbix/FullStory/Splunk - Certified Admin / Architect
- Proficient in scripting languages (Python, Bash, etc.) and automation tools (Puppet/Ansible/Terraform/Jenkins)
- Experience working within a SAFe environment, including participation in PI (Program Increment) Planning, Agile Release Trains (ARTs), and cross-functional collaboration across teams.
- Proficient in managing telemetry agents like Zabbix Agent, Splunk Forwarders, NewRelic Agents, etc.
- Proficient in setting up users, roles, and authentication protocols to ensure secure access control
- Experience in building custom integrations for extending monitoring capabilities for business applications.
- Skilled in monitoring, problem isolation, and system health checks to maintain performance
- Deep understanding of cloud platforms like Azure and GCP
- Demonstrated expertise in sizing, planning, and deploying distributed monitoring solution.
- Demonstrated expertise in onboarding diverse applications and optimizing platform performance and scalability.
- Experience integrating observability platforms with other enterprise systems (CMDB, ticket tools, etc.)
- Implementing a robust monitoring strategy for the monitoring platform to ensure platform stability and license controls.
- Experience working with external vendors.
Similar Jobs
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Build, test, deploy, operate, and improve software services, APIs, applications, integrations, and automation within an AI-enabled delivery pod. Use specification-driven development and approved AI coding tools to produce secure, maintainable software while validating AI-generated output. Develop comprehensive automated tests, integrate quality and security checks into CI/CD, establish observability and resilience, collaborate with cross-functional stakeholders, and continuously improve engineering quality, delivery speed, reliability, and production support.
Top Skills:
APIsC#Ci/CdClaude CodeCloud PlatformsContainersDistributed SystemsEvent-Driven ArchitectureGenerative AiGitGitGithub CopilotGithub Spec-KitGoInfrastructure As CodeJavaJavaScriptKubernetesLarge Language ModelsMcpMicroservicesNosql DatabasesObservabilityOpenai CodexPrompt EngineeringPythonRelational DatabasesRetrieval-Augmented GenerationTypescript
Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Build and deploy AI-powered solutions for enterprise and campus customers. Scope environments, prototype agentic AI applications, design secure multi-tenant and hybrid architectures, and harden successful prototypes into production. Own identity, encryption, networking, CI/CD, observability, and production support within customer environments. Collaborate with product, AI, data, and program teams to standardize repeatable solutions. The role is highly customer-facing and requires regular domestic and occasional international travel.
Top Skills:
Api GatewaysAWSAws PrivatelinkCi/CdDockerDuckdbEncryptionFastmcpIamJavaKafkaKey ManagementKubernetesLangchainLanggraphMcpMtlsObservabilityPgvectorPostgresPythonRagTypescriptVpc Peering
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Design, develop, test, deploy, monitor, and maintain secure, scalable software components and services. Own complex features end to end, shape architecture and APIs, implement CI/CD and quality automation, diagnose production issues, and improve observability and reliability. Lead risk-based testing across multiple levels, apply governed AI-assisted development, integrate LLM and agentic capabilities where valuable, mentor engineers, conduct technical reviews, and promote engineering standards and continuous delivery improvements.
Top Skills:
Agentic WorkflowsAPIsC#Ci/CdClaude CodeCloud PlatformsCodexContainersDistributed SystemsGitGithub CopilotInfrastructure As CodeJavaJavaScriptLlm IntegrationNosql DatabasesObservabilityPythonRelational DatabasesRetrieval-Augmented GenerationTypescript
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


