Opportunity overview
What you should know
We checked this opportunity from Supabase. It is listed as Worldwide.
The work location is Remote — worldwide. Employer-stated compensation: Salary not included in the job posting.
Job description
About this role
Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform. You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable — not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves.
Responsibilities
- Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions
- Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation
- Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes
- Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design
- Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it
- Help teams design sustainable on-call practices: alert quality, escalation paths, runbook coverage, and noise reduction
- Track and report on org-wide operational maturity, surfacing systemic gaps and driving remediation
Qualifications and requirements
- Have 7+ years of experience in SRE, production engineering, or reliability-focused roles, including experience shaping SRE practices and driving adoption across engineering teams
- Have a software engineering mindset — you write code and build tools, not just configure them
- Have hands-on experience defining and operationalizing SLOs/SLIs at scale, including error budget policies that actually influenced engineering decisions
- Have deep experience with incident response, postmortem facilitation, and turning incident learnings into systemic improvements
- Have worked with large-scale multi-tenant systems (bonus: managed database platforms or Postgres)
- Are proficient with cloud infrastructure (AWS preferred) and infrastructure-as-code (Pulumi preferred, Terraform/CDK also acceptable)
- Communicate clearly and persuasively — this role requires influencing without authority across a distributed org
- Have experience in async or globally distributed teams
- Are energized by making other teams more effective rather than being the one who fixes everything
- Experience with Kubernetes-based platform operations
- Familiarity with OpenTelemetry, VictoriaMetrics, Grafana, or similar observability tooling
- Experience building developer-facing reliability tooling (SLO dashboards, ORR frameworks, toil tracking, DORA metrics)
The description, responsibilities and qualifications used for this page and its JobPosting data appear above. Use the application action in the Job snapshot to confirm any later changes before applying.
Why you may be eligible
The source says applicants worldwide may apply. You should still confirm time-zone, work-authorization and experience requirements.
Career Studio · Early access
Want help preparing your application?
Use this job to tailor your CV and prepare a focused cover letter based on the employer’s requirements.
- Job-specific CV guidance
- Focused cover-letter draft
- You review, edit and submit it yourself
How this listing was checked
- Source status
- Applicant-tracking source confirmed
- Employer or source
- Supabase
- Source checked
- September 7, 2026
- Application destination
- Employer or official applicant-tracking website
- Work setup and location
- Remote · Remote — worldwide
- Who can apply
- Worldwide · Confirmed by source
- Pay
- Not stated by employer
- Deadline
- Not stated by employer
We check the source and visible requirements, but the employer controls changes and the final hiring decision. Confirm the latest requirements before applying.