MX Build Technologies India Private Limited Logo

MX Build Technologies India Private Limited

Senior Observability Engineer

Posted 22 Days Ago
Be an Early Applicant
In-Office
Chennai, Tamil Nadu, IND
Senior level
In-Office
Chennai, Tamil Nadu, IND
Senior level
Build and operate an observability control plane: automate Datadog/Terraform monitors and dashboards, run audits and maturity assessments, own monthly observability health reports, reduce alert fatigue, enable self-serve onboarding, rotate on shared on-call as Incident Commander, and close detection gaps after incidents.
The summary above was generated by AI

We are fueled by a moral imperative to advance mankind, and it all begins with our people, our product, and our purpose. Passion isn’t something we turn on and off; it’s woven into everything we do. If you thrive in high-challenge environments, are inspired by exceptional teammates, and are driven to grow beyond what you thought possible, MX is where you belong.

Come build the future with us. Join an award-winning company that isn’t just shaping the financial industry, but transforming it in ways that create meaningful, lasting impact for millions of people.


At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.

We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.

We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.

This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.

Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.

Job Duties
  • Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
  • After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
  • Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
  • Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
  • Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
  • Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
  • Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
  • Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
  • After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
  • Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.
Basic Requirements
  • BS in Computer Science or equivalent experience
  • 5+ years running production observability, SRE, or DevOps. Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast.
  • Automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency. You encode monitoring standards as code rather than clicking the UI.
  • AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs.
  • Alerting and SLO strategy: burn-rate and error-budget thinking, with a track record of cutting alert fatigue on evidence.
  • Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis.
  • Shared on-call, Incident Commander-capable. You've run or supported incident bridges and written postmortems, and you'll take shifts on the shared IR & Observability rotation.
Preferred Requirements
  • Fintech experience with MX-like architectures
  • Google SRE practices: toil elimination, incident management, automation for self-healing
  • Cross-functional influence without authority. You've improved teams that don't report to you.
  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends).
  • OpenTelemetry instrumentation
  • Datadog cost optimization at scale (cardinality, log indexing, sampling)
  • Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
  • Golang and Ruby on Rails (the MX stack)
What Success Looks Like

By six months, you're a full participant on the shared on-call rotation and a capable Incident Commander on live SEVs, teams you've engaged have alerts and dashboards that answer "what's broken and where do I look?", and the monthly health report runs largely on its own. By twelve months, incidents get caught earlier because of instrumentation the loop added, new services get baseline observability from a self-serve template on day one, and no team depends on a shepherd for day-one coverage.

Compensation

The expected earnings for this role could be comprised of a base salary and other forms of cash compensation, such as bonus or commissions as applicable.

This pay range is just one component of MX’s total rewards package. MX takes a number of factors into account when determining individual starting pay, including job and level they are hired into, location, skillset, peer compensation.


**Please note applicants applying for this position must have the legal right to work in India without the need for sponsorship. We are unable to provide work sponsorship for this role, and candidates should be able to verify their eligibility to work in the country independently. Proof of eligibility to work in India will be required as part of the hiring process.


Work Environment

In this role, a significant aspect of the job involves working in the office for a standard 40-hour workweek. We believe that the collaborative nature of our work and the face-to-face interactions among team members are essential for fostering a dynamic and productive work environment. Being present in the office enables seamless communication, facilitates quick decision-making, and encourages spontaneous collaboration that contributes to the overall success of our projects. We value the synergy that comes from having our team members physically together, allowing for immediate problem-solving, idea exchange, and team building.

Similar Jobs

8 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Information Technology • Software • Infrastructure as a Service (IaaS)
Build ingestion pipelines for logs and metrics, scalable alerting engines, and observability APIs. Interface with product teams and develop microservices using Golang and Rust.
Top Skills: AnsibleGoGraphQLGrpcRustTerraformTypescript
8 Days Ago
In-Office or Remote
India
Senior level
Senior level
Software
The Senior Infra Engineer will build and maintain ingestion pipelines, scalable alerting engines, and observability APIs, while ensuring resilience and scalability in infrastructure. They will work with tools like Golang, Rust, Terraform, and Ansible, documenting requirements and interfacing with product teams.
Top Skills: AnsibleGoGraphQLGrpcRustTerraformTypescript
23 Minutes Ago
Hybrid
Chennai, Tamil Nadu, IND
Entry level
Entry level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Perform clinical data monitoring and management activities: data review, query management, investigate logic checks, and ensure database design quality through documentation, testing, validation, and implementation of data collection tools. Act as first point of contact for CTMS issues, follow SOPs, collaborate to resolve TMF/document discrepancies, and contribute to process improvements and project work.
Top Skills: ChatgptCtmsElectronic Documentation Management SystemsMicrosoft CopilotMicrosoft Office SuiteWeb-Based Data Management Systems

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account