About Us:
At apexanalytix, we’re lifelong innovators! Since the date of our founding nearly four decades ago we’ve been consistently growing, profitable, and delivering the best procure-to-pay solutions to the world. We’re the perfect balance of established company and start-up. You will find a unique home here.
And you’ll recognize the names of our clients. Most of them are on The Global 2000. They trust us to give them the latest in controls, audit and analytics software every day. Industry analysts consistently rank us as a top supplier management solution, and you’ll be helping build that reputation.
Read more about apexanalytix - https://www.apexanalytix.com/about/
Job Details
The Role
Quick Take -
apexanalytix is seeking an experienced Infrastructure Operations Leader to serve as Global IT Managed Services Leader, accountable for the day-to-day operation of apexanalytix's global technology estate — the Kubernetes fleet, the physical compute and network infrastructure beneath it across two datacenters, and the platform services deployed across it.
apexanalytix runs its own private cloud: more than 60 Kubernetes clusters across two datacenters, every workload declared in Git and reconciled through GitOps, networked with Cilium's eBPF datapath and BGP control plane, backed by Ceph software-defined storage, and hosted as KubeVirt tenant clusters managed through Cluster API. Beneath that sit Cisco UCS and Supermicro compute, a multi-vendor network fabric spanning 100/200/400G leaf switches and InfiniBand, redundant WAN edge appliances carrying multiple carrier circuits, a large VMware estate and a Citrix tier fronting client access.
This role owns the operational health, availability, cost and service quality of all of it — in-house teams and outsourced providers alike — and sets the strategy for what is delivered in-house versus by a partner. It works in close collaboration with Engineering, Architecture, Security & Compliance, Finance and Procurement, and reports to the SVP, Global Head of Engineering & Infrastructure.
This is a run-the-estate role, not a service-desk role. The successful candidate is equally comfortable with a BGP session that will not establish, a storage cluster in a degraded state, a GitOps reconciliation that will not converge, and a vendor missing an SLA.
The Work -
Trusted Advisor
- Leverages subject-matter expertise in Kubernetes fleet operations, datacenter and network operations, and managed-services governance to advise engineering leadership on operational strategy, run-cost and reliability investment.
- Acts as the voice of operational reality — escalating capacity, reliability, lifecycle and vendor performance issues with evidence, and driving them to resolution across Engineering, Security and Finance.
- Partners with Finance, Procurement and regional leaders to translate business priorities into the in-house versus outsourced service model and vendor commercial terms.
Roles and Requirements
- Kubernetes fleet operations. Owns the operational health of a fleet of more than 60 production, UAT and development clusters, including cluster lifecycle through Cluster API — tenant provisioning, hosted and self-managed control planes, node pool scaling, machine health, and version upgrades across the fleet — and the management clusters that host it.
- GitOps delivery. Operates GitOps as the delivery mechanism: reconciliation health, image update automation and drift detection. Git is the only path to production; the successful candidate runs operations through the git source and treats direct cluster mutation as a time-boxed mitigation with the durable fix named.
- Platform services. Owns the operational health of the platform components deployed across the fleet — CNI and ingress (Cilium, Multus, MetalLB, Traefik, NGINX, Kong, external DNS, IPAM), storage and data protection (Rook-Ceph, Longhorn, Portworx, the CSI driver set, Velero, Commvault), database operators (OceanBase, CockroachDB, MariaDB, PostgreSQL, Vitess, ClickHouse, Elasticsearch, Valkey, RabbitMQ, Kafka), identity and secrets (Keycloak, OpenBao, cert-manager, policy enforcement), and the Prometheus/Thanos/Grafana/Loki/Tempo observability stack.
- Datacenter and physical infrastructure. Owns datacenter operations across both sites plus regional and cloud footprints: rack, power, cabling, provisioning, firmware and hardware lifecycle across Cisco UCS, Supermicro, Dell and NVIDIA platforms; the VMware estate; and the network fabric spanning Cisco NX-OS and IOS-XE, Pica8, MikroTik RouterOS, NVIDIA Cumulus and MLNX-OS, and Aruba wireless.
- WAN, remote access and telephony. Owns the WAN edge — VRRP, IPsec, BGP and NAT across redundant appliance pairs and multiple carrier circuits — along with SSL VPN, the WireGuard overlay, the Citrix NetScaler tier fronting client access, and the enterprise PBX and dialer estate.
- Reliability and recoverability. Establishes provable backup and restore across every tenant and platform system, with dated evidence produced and witnessed independently of the accountable team. Establishes verified high availability across production — replica counts, host anti-affinity and disruption budgets, checked continuously rather than after an incident. Owns disaster recovery posture and exercises it on a defined cadence.
- Service management. Owns end-to-end IT support delivery: incident, problem and change management (ITIL), major incident command, end-user computing, asset management, access provisioning, and escalation paths into Infrastructure, Engineering and Security. Ensures a consistent, high-quality support experience across all locations.
- Vendor and commercial governance. Leads RFP, vendor selection and contract negotiation for managed service providers and OEMs. Manages ongoing vendor and OEM performance against SLAs. Builds a vendor governance framework — QBRs, scorecards, escalation paths — giving leadership clear visibility into service health and cost.
- Budget and reporting. Owns the global infrastructure and managed-services operating budget and drives run-cost optimisation without compromising reliability. Provides transparent reporting on availability, incident SLAs, capacity, patch and backup posture, and cost to executive leadership, tied to the company's broader value-creation plan.
The Must-Haves -
- 8+ years of infrastructure and IT operations leadership experience, including direct accountability for production availability across a multi-site estate.
- Hands-on production Kubernetes operations experience at fleet scale — tens of clusters, not one — including cluster lifecycle, upgrades, capacity management and incident response. Bare-metal or private-cloud experience strongly preferred over managed-cloud-only.
- Working fluency with GitOps operations (FluxCD or ArgoCD with Kustomize and Helm), and the discipline that Git is the only durable path to production.
- Datacenter and physical infrastructure operations experience — server hardware lifecycle, out-of-band management, firmware, power and capacity planning.
- Enterprise network operations at Layer 2 and Layer 3 — BGP, MLAG/vPC, LACP, VLAN and VRRP design and troubleshooting — across more than one switch operating system. Cisco NX-OS required; exposure to Linux-based network operating systems (Cumulus, Pica8, SONiC, RouterOS) a strong plus.
- Storage operations experience — software-defined storage (Ceph strongly preferred), SAN/NAS arrays, snapshots, replication, and backup and restore verification.
- Virtualization operations experience — VMware vSphere at scale, and ideally KubeVirt or another Kubernetes-native virtualization platform.
- Strong working knowledge of ITIL-based incident, problem and change management, ITSM platforms, and major incident command.
- Strong commercial acumen — comfortable negotiating MSP and OEM contracts and owning a budget with clear cost accountability.
- Demonstrated experience managing global, multi-region operations, ideally including support for an India-based Global Capability Center.
- An evidence-driven operating style: able to describe, from direct experience, a control they proved was broken by testing it rather than by reading a status report.
- Strong communicator, able to translate operational and technical issues into business terms for executive audiences.
Preferred
- Experience operating in a regulated or attested environment (SOC 2, FedRAMP/StateRAMP, ISO 27001) where recoverability and change control are audited.
- Automation ability in at least one of Python, Go, Bash or PowerShell — enough to build and review operational tooling, not only to consume it.
- Experience in a PE-backed or private equity value-creation context.
- Willingness to adopt AI-assisted operations tooling; apexanalytix operates a substantial automation and agentic runbook capability.
- Education Levels/Credentials (Degree types and Emphasis) - Undergraduate degree in information technology, computer science, engineering or a related field. ITIL Foundation, CKA or equivalent certification a strong plus but not required.
Physical/Remote/Travel Work Environment
apexanalytix Global Capability Center, Noida, India.
Travel required — occasional, including periodic travel to apexanalytix datacenters and other global offices.
Scope Boundaries:
This role is accountable for running the estate. Platform strategy, target architecture and the capital plan sit with the SVP, Global Head of Engineering & Infrastructure; technical design authority with the Architect; product software development with the VP, Head of Software Development; security policy and compliance attestation with the CISO; and AI and automation delivery with a separate AI & Automation organization.
Over the years, we’ve discovered that the most effective and successful associates at apexanalytix are people who have a specific combination of values, skills, and behaviors that we call “The apex Way”. Read more about The apex Way - https://www.apexanalytix.com/careers/
Benefits
At apexanalytix we know that our associates are the reason behind our successes. We truly value you as an associate and part of our professional family. Our goal is to offer the very best benefits possible to you and your loved ones. When it comes to benefits, whether for yourself or your family the most important aspect is choice. And we get that. apexanalytix offers competitive benefits for the countries that we serve, in addition to our BeWell@apex initiative that encourages employees’ growth in six key wellness areas: Emotional, Physical, Community, Financial, Social, and Intelligence.
With resources such as a strong Mentor Program, Internal Training Portal, plus Education, Tuition, and Certification Assistance, we provide tools for our associates to grow and develop.