Careers at Cirrascale

Planned Maintenance Lead (Night Shift)

About Cirrascale

Cirrascale Cloud Services provides high-performance cloud infrastructure purpose-built for deep learning, generative AI, and large-scale AI inference workloads. We specialize in dedicated GPU cloud solutions tailored to the unique needs of startups, research labs, and enterprise AI teams. Our mission is to accelerate AI innovation by combining powerful hardware with white-glove service and flexible, custom-built environments.

Position Overview

The After-Hours Planned Maintenance Lead (Data Center Operations) serves as the primary technical gatekeeper, change coordinator, and governance lead for nighttime data center operations. This advanced role blends the technical expertise of a senior NOC contributor with the strict, process-driven oversight required to manage complex, high-stakes infrastructure maintenance.

The primary focus of this position is the comprehensive, end-to-end lifecycle management of all Planned Maintenances (PMs). Working non-traditional after-hours shifts, you will ensure all scheduled internal, carrier, and vendor activities are rigorously vetted, tracked, risk-mitigated, and seamlessly executed without unplanned impact to clients, while simultaneously guiding junior shift technicians through standard operating procedures.

Hours 11:30pm – 8:30am


Key Responsibilities

PM Queue Intake & Vetting: Serve as the final arbiter for incoming vendor, carrier, and internal engineering maintenance notices. Audit complex technical notifications to dissect and verify the exact scope, chronological windows, geographic locations, and equipment impact levels.

Change Governance & Jira Schema Management: Own the translation of raw technical maintenance advisories into highly structured, detailed Jira tickets and change requests. Ensure absolute tracking precision of execution dates, specific data center suites, circuit IDs, physical/virtual server dependencies, fallback windows, and vendor escalation paths.

Risk Mitigation & Impact Isolation: Utilize internal data center monitoring systems, network topology maps, and asset management platforms to identify the exact customer cross-connects, VLANs, and hardware instances affected by upcoming maintenance windows. Assess risk profiles and verify that backup links or redundant pathways are operational prior to window commencement.

Pre- & Post-Window Stakeholder Communication: Draft and dispatch explicit, audience-appropriate technical notices to affected enterprise clients and internal executive stakeholders. Translate intricate multi-layered networking updates into clear customer-facing updates regarding potential service degradation or downtime. Deliver clear "All Clear" or rollback notices post-maintenance.

Maintenance Runbooks & Process Optimization: Author, review, and continuously optimize data center operational maintenance runbooks, pre/post-maintenance health checks, and Knowledge Base articles to streamline the execution of repeated or routine maintenance schedules.

Automated Pre-/Post-Checks: Utilize a strong command-line working knowledge of Linux environments to execute specialized scripts that gather real-time system and environmental health telemetry immediately before and after vendor windows.

Maintenance Script Maintenance: Write, troubleshoot, and optimize Bash or Python scripts aimed at reducing repetitive manual verification tasks during maintenance setup and post-window handoffs.

Alert Suppression & Telemetry Correlation: Leverage enterprise network and data center monitoring platforms to actively isolate infrastructure alerts from expected maintenance noises, preventing false-positive escalations while maintaining total visibility on un-related anomalies.

First-Line Defense During Windows: Act as the premier after-hours technical point of contact for unpredicted data center infrastructure degradation or failure occurring inside or outside a maintenance window (power, environmental dependencies, routing loops, or fiber cuts).

Rollback Management: Perform advanced incident isolation and diagnostics when a maintenance window goes out of scope. Swiftly execute and coordinate fallback/rollback procedures with vendors to protect client SLAs.

Technical Vendor Escalations: Own the end-to-end incident management pipeline during high-pressure window overruns. Package thorough diagnostic timelines and technical context to cleanly escalate to Tier IV engineering, physical facilities teams, or external carriers to accelerate Service Level Agreement (SLA) restorations.

Maintenance Standards Enforcement: Provide ongoing technical mentorship and peer reviews to junior NOC Technicians. Validate the accuracy of their PM ticket documentation, customer communications, and script execution to elevate overall shift competency and adherence to change management policies.

Strategic Cross-Department Liaison: Serve as the strategic midnight liaison between Infrastructure Engineering, Customer Support, DevOps, Facilities management, and external vendor teams, ensuring PM action items are pursued proactively.

PM Shift Hand-Off Integrity: Deliver meticulously structured, comprehensive shift reports summarizing the outcomes, open windows, rollbacks, and unresolved issues of all planned maintenance and incident pipelines to incoming daytime teams.

 

Required Qualifications

Professional Experience: 4+ years of hands-on experience in a Network Operations Center (NOC), Change Management Group, Data Center Operations, or high-tier enterprise infrastructure environment, with a demonstrable track record of overseeing complex maintenance windows.

Ticketing & Change Management Systems: High-level mastery of enterprise Jira instances or change management tools (e.g., ServiceNow), including detailed field auditing, custom queue filtering, and task linking.

OS & Scripting Proficiency: Deep familiarity with Linux/Unix command-line interfaces, shell navigation, and basic scripting skills (Bash, Python, or similar configurations) for manual task reduction during maintenance testing.

Infrastructure Knowledge: Practical understanding of basic routing/switching models, structuring of network cables, system virtualization, SAN storage frameworks, power grids, and cooling metrics.

Crisis & Change Communication: Exceptional English verbal and written capabilities, featuring a proven track record of synthesizing stressful technical infrastructure disruptions into calm, structured customer updates.

Schedule & Availability: Ultimate flexibility to work dedicated, non-traditional after-hours schedules, consisting of consistent night shifts, rotating weekends, and holidays to align with standard maintenance windows.

Preferred Qualifications

Preferred Certifications: Possession of or actively pursuing ITIL Foundation (or higher) with a focus on Change Management, alongside intermediate industry certifications (e.g., CCNA, CompTIA Network+/Linux+, LPI, or specialized infrastructure certificates).

 

Cirrascale Cloud Services is an Equal Opportunity Employer committed to fostering an inclusive workplace and providing equal employment opportunities to all qualified applicants.

 

Salary Range

The base salary range for the PM NOC Lead is $80,000 to $90,000 USD. This pay range reflects the broad, minimum to maximum, pay range for this job for the location for which it has been posted. Compensation decisions are dependent on several factors including, but not limited to, an individual’s qualifications, location where the role is to be performed, internal equity, and alignment with market data.

 

Benefits

Comprehensive benefits package, including health, dental, and vision insurance, retirement plans, paid time off, and opportunities for professional development.

 

Why Join Cirrascale?

Join a growing team that's pushing the boundaries of AI infrastructure. At Cirrascale, you’ll contribute to projects powering next-generation AI applications while working with top-tier hardware in a collaborative and innovative environment. From custom deployments to hands-on customer support, every role here plays a part in enabling breakthroughs in AI.

 

Compensation Package details

Bonus

Stock Options

 

              

Operations

Austin, TX

Teilen auf:

NutzungsbedingungenDatenschutzCookiesPowered by Rippling