CBTS India

Sr. Engineer – Site Reliability Engineering

CBTS serves enterprise and midmarket clients in all industries across the United States and Canada. CBTS combines deep technical expertise with a full suite of flexible technology solutions--including Application Modernization, Managed Hybrid Cloud, Cybersecurity, Unified Communications, and Infrastructure solutions. From developing and deploying modern applications and the secure, scalable platforms on which they run, to managing, monitoring, and optimizing their operations, CBTS delivers comprehensive technology solutions for its clients' transformative business initiatives. For more information, please visit www.cbts.com.


OnX is a leading technology solution provider that serves businesses, healthcare organizations, and government agencies across Canada. OnX combines deep technical expertise with a full suite of flexible technology solutions—including Generative AI, Application Modernization, Managed Hybrid Cloud, Cybersecurity, Unified Communications, and Infrastructure solutions. From developing and deploying modern applications and the secure, scalable platforms on which they run, to managing, monitoring, and optimizing their operations, OnX delivers comprehensive technology solutions for its clients’ transformative business initiatives. For more information, please visit www.onx.com.



1.      Role Purpose (1–3 lines): Ensures high availability, performance, scalability, and resilience of cloud and infrastructure platforms by applying SRE engineering principles, automation-first practices, observability, and continual reliability improvements across services and platforms.

2.      Key Responsibilities:

·       Implement SRE frameworks, SLIs/SLOs/SLAs, error budgets, performance engineering, and reliability guardrails across cloud platforms & services.

·       Drive automation for provisioning, deployment, configuration management, drift control, patching, recovery, and operations workflows.

·       Build observability stack, dashboards, anomaly detection, synthetic tests, runbooks, incident readiness and RCA automation.

·       Partner with DevOps, Platform Engineering, Cloud Engineering & application squads to define reliability patterns, capacity planning & scalable workload landing models.

·       Manage incident response, major incident coordination, postmortem improvement actions, resiliency testing, fault injection, chaos engineering initiatives.

·       Ensure infra security alignment, vulnerability remediation, compliance & secure configuration baselines in cloud infrastructure.

Staff Augmentation Resources

Chennai, India

Compartir en:

Términos de servicioPrivacidadCookiesPatrocinado por Rippling