
About Nexcess
Nexcess provides specialty cloud solutions for organizations where performance and compliance have to coexist. We serve businesses worldwide, from agencies scaling client sites to enterprises running mission-critical operations. We've built our reputation on deep technical expertise and genuine partnership with every client we work with. Behind every environment we manage is a team of people who take the craft seriously and keep showing up when it matters.
The Platform SRE is a member of the Platform SRE team, the group responsible for the operational health, resiliency, and modernization of our hosting and cloud infrastructure. This role sits at the intersection of systems engineering, site reliability engineering, and automation : combining hands-on technical debt remediation with large-scale automation and reliability practice.
The Platform SRE team owns three core mandates:
1. Technical debt reduction : systematically identifying and remediating aging firmware, kernels, operating systems, and software across the managed hosting fleet and managed application environments.
2. Automation of critical operational workflows : provisioning, patching, remediation, and release processes across both managed apps and managed hosting fleets, with automated release planning and execution under human supervision (not “automation for automation’s sake,” but automation with a human checkpoint before production impact).
3. Reliability and incident support : defining, instrumenting, and tracking SLIs/SLOs for platform engineering and operations, visualized through dashboards and reporting tools, and providing fast, expert frontline response and remediation during service-impacting events (incident command and process ownership sit with the Incident Management team; this team is the technical responder, not the incident owner).
This position serves as a senior technical point of contact for platform engineering, driving initiatives around scalability, fault tolerance, automation, and operational excellence across production infrastructure.
Key Responsibilities
Technical Debt & Platform Modernization
● Own the lifecycle of firmware, kernel, OS, and software patching across the managed hosting fleet and managed application environments
● Build a standing inventory and risk model of technical debt (end-of-life OS versions, unpatched firmware, deprecated software) and drive prioritized remediation plans
● Evaluate and implement infrastructure modernization initiatives, replacing manual or legacy processes with supportable, automated alternatives
Automation & Release Engineering
● Design and build automation for provisioning, deployment, patching, remediation, and configuration management across managed apps and managed hosting fleets
● Own the design of automated release pipelines : planning, staging, and executing releases with defined human-in-the-loop approval gates
● Develop self-healing and auto-remediation capability for common failure modes to reduce manual operational load
● Support and extend CI/CD workflows and infrastructure-as-code practices across the platform
Reliability Engineering, SLIs/SLOs & Observability
● Define SLIs and SLOs for platform engineering and operations in partnership with engineering and product stakeholders
● Instrument systems to measure SLIs accurately and build SLO tracking into standard reporting
● Build and maintain dashboards (e.g., Grafana, Datadog, or equivalent visualization tooling) to make SLI/SLO performance, error budgets, and platform health visible to engineering and leadership
● Continuously improve platform observability : monitoring, alerting, logging, and tracing : across distributed and containerized environments
Incident Response & Remediation (Support Role)
● Serve as the frontline technical responder: acknowledge pages quickly, diagnose, and remediate platform-level issues
● Partner with the Incident Management team, who own incident command, severity classification, and customer communication : this role provides the technical hands and expertise, not incident ownership
● Contribute technical findings to blameless root cause analysis (RCA) and own follow-through on corrective actions for platform systems
● Maintain runbooks and on-call readiness for platform and infrastructure systems
● Track incident trends on platform systems and feed them back into the technical debt and automation roadmap
Collaboration & Technical Leadership
● Partner with software engineering teams on platform architecture, operational readiness reviews, and scalability initiatives
● Support platform security, compliance, and operational governance requirements
● Mentor engineers and contribute to technical leadership and knowledge-sharing across the team
● Maintain clear operational documentation and contribute to team standards and process improvement
● Other duties as assigned
Requirements
● 3–5+ years of experience in platform engineering, systems engineering, SRE, or infrastructure operations (level based on experience and scope)
● Advanced Linux systems administration and troubleshooting expertise, including kernel and firmware-level familiarity
● Strong experience with Kubernetes, Docker, and container orchestration/distributed systems
● Hands-on automation and infrastructure-as-code experience (e.g., Terraform, Ansible, Puppet/Chef, or equivalent)
● Experience building or maintaining CI/CD and automated release/deployment pipelines
● Experience defining and tracking SLIs/SLOs and working with observability/visualization tools (e.g., Grafana, Datadog, Prometheus, or equivalent)
● Experience supporting enterprise-scale, high-concurrency, or customer-impacting production environments
● Demonstrated experience as a technical responder in production incidents, including root cause analysis and corrective action follow-through
● Strong scripting ability (e.g., Python, Bash, Go) for automation and tooling
● Strong troubleshooting skills across compute, network, storage, and application layers
● Experience supporting cloud-hosted, managed hosting, or hybrid infrastructure environments
● Ability to lead technical initiatives, mentor others, and communicate clearly across teams
Preferred Qualifications
● Experience owning fleet-wide firmware/OS patch management programs at scale
● Experience designing human-in-the-loop release automation or progressive delivery systems (canary, blue/green)
● Familiarity with error budgets and SLO-driven prioritization frameworks
● Experience with configuration/patch management at scale across heterogeneous hardware fleets
Development
India
Share on: