Site Reliability Engineer

permanent
Fully Remote

Only accepting applications from: United States

  • Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
  • Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
  • Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
  • Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
  • Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning in the CI/CD process
  • Implement and maintain policy-as-code to block unsafe infrastructure and Kubernetes changes at admission time
  • Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings
  • Partner with Security Engineer and broader Security & Reliability disciplines
  • Participate in and improve the on-call rotation; build the runbooks and automation
  • Coach team members and engineers across the org on reliability patterns and operational best practices

Experience

  • 5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
  • Strong programming skills for automation and tooling (Go, Python, Typescript or similar)
  • Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
  • Proven track record leading incident response and building SLO-driven reliability practices.
  • Working fluency with observability tooling (Datadog is a plus)
  • Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
  • Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management)
  • The judgment and communication skills to raise a security or reliability finding with a senior engineer
  • Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant)
  • Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
  • Chaos engineering or game-day experience is a plus
  • Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus

Salary and Perks

Pay range: $120K - $165K

  • Healthcare, parental planning, mental health benefits
  • Annual performance bonus, a 401(k) plan and match
  • Responsible time off, monthly wellness and technology allowances
  • Opportunities for face-to-face connections and annual gatherings
  • Flexible time-off policy with 'Responsible Time Off' benefit
  • Volunteer days off
  • Mentorship program
  • Paid maternity and paternity leave
  • Monthly Wellness Allowance
  • Reward and recognition platform
  • MyFitnessPal Premium access
  • Virtual learning and development library
  • DEI Committee initiatives
  • Medical, dental, and vision benefits
  • Retirement savings program with employer match

About MyFitnessPal

MyFitnessPal provides tools that make it easy for everyone to live a healthier life by tracking meals and physical activity.

MyFitnessPal provides tools that make it easy for everyone to live a healthier life by tracking meals and physical activity.

View all devops and sysadmin jobs

Workster

Remote Jobs for US Residents

We've built a new platform specifically for US residents to find remote work.

Discover Workster

Power Search

Find the jobs that don't get advertised

We've built a tool to help you discover all of the remote jobs that never get advertised.

Discover Power Search