Era4

Observability Platform Engineer

Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations


Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations


Role Summary:

As an Observability Platform Engineer, you will design, implement, and operate Era4’s enterprise observability platform across its AI infrastructure environment. You will build and maintain highly available telemetry, monitoring, alerting, and dashboarding capabilities using the Grafana ecosystem, Kubernetes, OpenTelemetry, and Infrastructure-as-Code.


Working across compute, GPU, storage, networking, and application platforms, you will establish scalable observability standards, automate platform operations, support incident detection and root-cause analysis, and ensure the platform is secure, resilient, and maintainable as Era4 scales.

 

Key Responsibilities:


Observability Platform Implementation:

  • Design and implement highly available observability services across multiple co-location and production sites.
  • Define platform standards for telemetry collection, labelling, metadata enrichment, retention policies, and data governance.
  • Implement multi-tenant observability controls and tenant isolation strategies
  • Configure telemetry ingestion pipelines for metrics, logs, and future distributed tracing workloads.
  • Configure and maintain object-storage-backed telemetry platforms for long-term retention and scalability.

 

Telemetry Collection & Integration:

  • Deploy and manage Grafana Alloy collectors across Kubernetes clusters, Linux hosts, network infrastructure, storage platforms, and hardware management systems.
  • Integrate telemetry from Kubernetes, GPU infrastructure, HPE hardware, storage platforms, network devices, and cloud-native services.
  • Develop and maintain observability integrations using OpenTelemetry standards and protocols.
  • Establish onboarding processes for new platforms, applications, and infrastructure services.
  • Collaborate with application teams to define observability requirements and future tracing adoption strategies.

 

 Alerting & Operational Insights:

  • Design and implement alerting frameworks using recording rules, AlertManager, and operational best practices.
  • Develop operational dashboards and service health views for infrastructure, platform, and application services.
  • Support integration of observability events with ITSM and incident-management platforms.
  • Define SLIs, SLOs, alert thresholds, and operational KPIs.
  • Continuously improve platform observability, incident detection, and root-cause analysis capabilities.

 

Reliability & Automation:

  • Implement Infrastructure-as-Code and GitOps practices for observability platform deployment and configuration management.
  • Develop automation for dashboard provisioning, alert deployment, tenant onboarding, and telemetry configuration.
  • Design and validate disaster recovery, resilience, and failover capabilities across observability services.
  • Contribute to platform security, compliance, and operational governance initiatives.
  • Work with operational teams to ensure observability services remain reliable, scalable, and maintainable.

 

Required Experience & Skills:

  • Significant experience implementing and operating enterprise observability or monitoring platforms.
  • In depth understanding of metrics, logs, traces, OpenTelemetry, and modern observability principles.
  • Relevant experience of:
    • Grafana ecosystem technologies - Grafana, Prometheus, Grafana Mimir, Grafana Loki, Grafana Tempo, Grafana Alloy 
    • Designing Kubernetes-native solutions and operating distributed platforms at scale.
    • Linux systems administration and cloud-native infrastructure.
    • Infrastructure-as-Code and GitOps approaches (preferably including Ansible).
    • Automation and operational tooling using Python and/or Go.
    • Technical architecture, operational documentation, and deployment designs.
    • Object storage technologies and distributed data platforms.

 

Why Join Era4:

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale. 

 

Diversity & Inclusion:  

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.  

Technology

United Kingdom: (Occasional office visit required)

Partilhar em:

Termos de serviço.PrivacidadeCookiesDesenvolvido pela Rippling