Oscilar Logo

Oscilar

Sr./Staff - Infrastructure/Site Reliability Engineer (SRE)

Posted Yesterday
Remote
2 Locations
Senior level
Remote
2 Locations
Senior level
Seeking a seasoned SRE to lead reliability for a cloud-native platform, overseeing infrastructure, CI/CD pipelines, observability, and mentoring engineers.
The summary above was generated by AI

Shape the future of trust in the age of AI
At Oscilar, we're building the most advanced AI Risk Decisioning™ Platform. Banks, fintechs, and digitally native organizations rely on us to manage their fraud, credit, and compliance risk with the power of AI. If you're passionate about solving complex problems and making the internet safer for everyone, this is your place.

Why join us:
  • Mission-driven teams: Work alongside industry veterans from Meta, Uber, Citi, and Confluent, all united by a shared goal to make the digital world safer.

  • Ownership and impact: We believe in extreme ownership. You'll be empowered to take responsibility, move fast, and make decisions that drive our mission forward.

  • Innovate at the cutting edge: Your work will shape how modern finance detects fraud and manages risk.

About the Role

Oscilar is growing fast, and so is the complexity of our systems. We’re looking for a experienced SRE to take ownership of reliability across our multi-region, cloud-native platform. You’ll have the mandate and autonomy to design, implement, and evolve systems that stay performant and resilient—through traffic spikes, dependency failures, and global deployments. You’ll be shaping how we scale, how we build observability, and how we run infrastructure that supports billions of events and large-scale data pipelines.

What You’ll Own
  • Architect and operate resilient cloud infrastructure (AWS, Pulumi, Kubernetes).

  • Lead initiatives to improve availability, latency, and performance at scale.

  • Design and evolve our CI/CD pipelines to optimize for speed, safety, and repeatability.

  • Define the metrics, alerts, and runbooks that form our observability backbone.

  • Run chaos experiments and failure simulations to harden the platform.

  • Mentor engineers and set best practices for SRE across the company.

What You Bring
  • Proven track record as a senior SRE, DevOps, or infrastructure engineer in high-scale environments.

  • Expert-level skills in AWS and Infrastructure as Code (Pulumi, Terraform).

  • Strong programming ability in Go and Java.

  • Deep understanding of distributed systems (Kafka, ClickHouse) and microservices architecture.

  • Mastery of container orchestration (Kubernetes) and production debugging.

  • Strong sense of ownership, and the judgment to balance velocity with reliability.

Benefits
  • Compensation: Competitive salary and equity packages, including a 401k plan

  • Flexibility: Remote-first culture — work from anywhere

  • Health: 100% Employer covered comprehensive health, dental, and vision insurance with a top tier plan for you and your dependents (US and Canada)

  • Balance: Unlimited PTO policy

  • Technical: AI First company; both Co-Founders are engineers at heart; and over 50% of the company is Engineering and Product

  • Culture: Family-Friendly environment; Regular team events and offsites

  • Development: Unparalleled learning and professional development opportunities

  • Gear: Home office setup assistance

  • Impact: Making the internet safer by protecting online transactions

Top Skills

AWS
Clickhouse
Go
Java
Kafka
Kubernetes
Pulumi
Terraform

Similar Jobs

2 Days Ago
Remote or Hybrid
Calgary, AB, CAN
Senior level
Senior level
eCommerce • Payments • Software
The Senior Site Reliability Engineer will design, build, and maintain systems supporting our SaaS infrastructure, ensuring platform reliability and performance.
Top Skills: AnsibleAWSAzureDockerGCPGitopsGoGrafanaKubernetesLinuxOpentelemetryPrometheusPythonRubyTerraformUnix
2 Days Ago
Remote or Hybrid
Winnipeg, MB, CAN
Senior level
Senior level
eCommerce • Payments • Software
As a Senior Site Reliability Engineer, you'll design and maintain reliable SaaS infrastructure, optimize monitoring processes, and enhance software reliability through collaboration.
Top Skills: AnsibleAWSAzureDockerGCPGitopsGoGrafanaKubernetesLinuxOpentelemetryPrometheusPythonRubyShell ScriptingTerraformUnix
4 Days Ago
Remote
30 Locations
Senior level
Senior level
Information Technology
As a Senior Site Reliability Engineer, you'll build and maintain infrastructure, tackle operational challenges, and automate processes to enhance reliability.
Top Skills: DockerDocker ComposeGoLinuxPerlPython

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account