Stratus

Senior Site Reliability Engineer

Stratus, deriving from the Latin term meaning 'layer', offers an advanced set of MEP specific solutions that seamlessly layer across a contractor's entire workflow from design to fabrication to installation. Our team of seasoned industry experts, skilled technology leaders, innovators, and entrepreneurs understands that fabrication does not occur in isolation, and increasingly, it may not happen within your own fabrication shop. Through close relationships with our customers—who include some of the most innovative and largest MEP contractors—we have developed a suite of Stratus tools to digitize, automate, and optimize piping, plumbing, sheet metal, and electrical contracting. Stratus provides the software layer an MEP Contractor needs to optimize profits with true "Data Driven Contracting."

GENERAL DESCRIPTION:

The Senior Site Reliability Engineer is accountable for how Stratus behaves in production. Reporting to the Senior Director, Platform Engineering, this role brings genuine SRE discipline to a platform that MEP contractors run their fabrication shops on — where downtime does not mean a degraded experience, it means work stops on a job site. Stratus is a ~50-person, primarily remote, Series B software company growing quickly.

Unrelenting Reliability is one of our company values, and this is the role that operationalizes it. You will define what reliable means in numbers, instrument the system so we can see it, and close the loop from production signal back into engineering priority. This is an engineering role with a production mandate: you will write code, tune queries, build alerting, run load tests, and lead incidents — and you will be measured on customer-visible availability and latency, not on tickets closed.

The load-bearing problems on your plate are: (1) establishing service level objectives and an error budget the whole engineering organization operates against, with the measurement infrastructure to back them; (2) building the detection and response capability — high-signal alerting, clear on-call and escalation paths, per-system incident ownership, and well-exercised recovery paths — so that production problems are caught and resolved fast; and (3) building the performance and capacity engineering practice for our data and messaging layers, with headroom measured rather than assumed.

You will work across every engineering pod, with our platform and security functions, and with the customer-facing teams who see reliability problems first. The right candidate is comfortable being the person who says a number out loud and defends it, and is drawn to a place where the reliability practice is being built rather than maintained.

KEY RESPONSIBILITIES:

  • Define and own service level indicators, objectives, and error budgets for customer-facing services; build the measurement pipeline that makes them trustworthy and publish availability and latency against target on a regular cadence.
  • Build and own production observability: instrumentation standards, dashboards, and actionable alerting on our stack — AKS on Azure, Prometheus, Loki, Tempo, and Grafana, with Istio as the service mesh and Flux for GitOps — extending coverage across every environment.
  • Establish and run incident response practice — on-call rotation and paging paths, severity definitions, per-system incident ownership, escalation into and out of customer support, and blameless post-mortems in incident.io.
  • Set the recovery standard: define what rollback and recovery must demonstrate, and verify through regular exercise that every service meets it.
  • Lead performance and capacity engineering: query and index tuning, connection-pool sizing, capacity modeling against real load, and designing for graceful degradation under partial failure.
  • Work with the existing team that runs our k6 load, stress, spike, and soak testing against a production-like environment; help set per-endpoint latency thresholds tied to our SLOs and make the results a gate on delivery.
  • Drive production readiness: review new services and significant changes against readiness criteria covering instrumentation, alerting, failure modes, resource limits, and rollback before they ship.
  • Close the loop on remediation: track post-mortem action items through to verified production change.
  • Contribute to business continuity and disaster recovery planning, including backup and restore validation, failover design, and recovery objectives.
  • Partner with engineering pods to push reliability ownership outward — coach teams on instrumenting and operating their own services.
  • Write code: tooling, automation, instrumentation libraries, and fixes in the production codebase.

QUALIFICATIONS:

  • 6+ years of professional engineering experience, with at least 3+ years in a dedicated SRE or production engineering role at a B2B SaaS company.
  • Demonstrated ownership of SLOs and error budgets in production — you have defined them, measured them, argued about them with product leadership, and changed engineering behavior with them.
  • Deep, practical observability skills with Prometheus, Loki, Tempo, and Grafana or comparable: you have instrumented real systems and built alerting that pages on customer impact rather than on CPU.
  • Hands-on incident command experience at meaningful severity, including building on-call and escalation practice from the ground up.
  • Strong database performance skills — query profiling, index design, connection pooling, and diagnosing saturation under load. MongoDB experience strongly preferred; comparable document or relational depth acceptable.
  • Production Kubernetes experience (AKS preferred) sufficient to debug a live problem — pod scheduling, resource limits, networking, and service mesh behavior (Istio preferred).
  • Solid coding ability in at least one general-purpose language (Go, Python, C#, or TypeScript); willingness to work in a C#/.NET codebase.
  • Experience with load and performance testing tooling (k6, JMeter, Gatling, or comparable) and with turning results into engineering priority.
  • Production experience on Azure or AWS, with real understanding of the failure modes of managed services.
  • Fluency with AI-assisted engineering tooling and a track record of designing AI-leveraged workflows for your team — this is a graded expectation at every level at Stratus.
  • Excellent written communication: you write post-mortems, runbooks, and reliability reports that executives and engineers both read and act on.
  • Judgment and steadiness under pressure, and the credibility to tell engineering and product leadership something they do not want to hear.

NICE TO HAVE:

  • Experience establishing or maturing an SRE practice — defining the discipline, not inheriting it.
  • Experience with Sentry or comparable application error-monitoring platforms.
  • Experience operating event-driven and real-time systems — message brokers (Azure Service Bus, Kafka) and websocket or push layers (SignalR or comparable).
  • Experience operating MongoDB Atlas at production scale, including replica set topology and Atlas performance tooling.
  • Experience with durable workflow orchestration (Temporal or comparable).
  • Background in multi-region or multi-zone architecture and DR design against stated recovery objectives.
  • Experience with incident.io, PagerDuty, or comparable incident management platforms.
  • Familiarity with DORA metrics and with reliability work inside SOC 2 or NIST 800-171 scope.
  • Experience with legacy monolith reliability — improving the operational behavior of an ASP.NET or comparable application you cannot rewrite.
  • Domain interest in MEP, BIM, AEC, or construction technology.
  • Prior experience in a Series B / growth-stage company navigating the transition from product-market fit to scale.


BENEFITS:

  • Comprehensive and competitive health benefits plan
  • Matching 401k contributions
  • 20 days annual PTO
  • Primarily remote work with occasional annual team onsites.


E-VERIFY STATEMENT 
Stratus participates in E-Verify. After you join the team, we'll verify your eligibility to work in the U.S. by submitting information from your Form I-9 to the Social Security Administration and, if needed, the Department of Homeland Security. This process happens post-hire only - we never use E-Verify to pre-screen applicants. 
E-Verify Notice 
Right to Work Notice 

Product & Development

Remote (United States)

Udostępnij w:

Warunki korzystania z usługPrywatnośćPliki cookieUsługa działa z technologią Rippling