Team Leader At at Remote
Recruit Finds
Nairobi, Kenya
Job summary
Team Leader, SRE at Remote
About this role
Join Our Whatsapp Channel - CLICK HERE
Job Overview
Remote is seeking an SRE Team Leader to lead its Site Reliability Engineering team while remaining actively involved in technical work. The position combines approximately 60% individual technical contribution with 40% people leadership.
The SRE team supports engineering velocity and product reliability by managing Kubernetes, AWS, PostgreSQL, CI infrastructure, observability, and reliability practices. The Team Leader will guide team members' career development, set priorities in line with company objectives, represent the team across engineering, and remain technically engaged enough to provide direction and identify issues early.
The reliability function is continuing to mature, with SLOs already being introduced across teams and further work required to expand the framework and balance operational responsibilities with project delivery.
What You Bring
People Leadership
Experience leading an SRE, infrastructure, or platform engineering team and managing employee development, performance, and career progression.
Ability to coach technical skills as well as interpersonal and professional skills, with evidence of helping team members grow.
Ability to address performance concerns early, clearly, and with empathy.
Experience hiring engineers and assessing candidates effectively during technical recruitment.
Strong understanding of team dynamics and the ability to resolve workplace conflicts.
Ability to build commitment around company and team objectives.
Technical Experience
Hands-on experience in site reliability, DevOps, or cloud infrastructure engineering, with sufficient technical depth to review work, challenge designs, and contribute effectively during incidents.
Practical production experience with Kubernetes.
Strong AWS experience at meaningful scale.
Hands-on experience building, enabling, or scaling AI infrastructure.
Strong understanding of observability practices and principles.
Experience using Terraform for infrastructure as code.
Experience with CI/CD platforms such as GitLab CI, GitHub Actions, or Jenkins.
Experience with Docker and shell scripting.
Experience operating a reliability practice covering incident response, on-call support, SLOs, and error budgets, as well as implementing lasting improvements after incidents.
Experience working within regulated environments.
Ways of Working
Excellent prioritisation skills when operational demands compete with project commitments, while maintaining the team's operational responsibilities.
Strong written communication skills, particularly in a fully distributed and asynchronous working environment.
Ability to develop productive relationships across different teams and encourage early collaboration when problems arise.
Additional Experience
Knowledge of a backend programming language, preferably Elixir, or alternatives such as Java, Clojure, Node.js, Python, or similar.
Experience with modern observability technologies, including OpenTelemetry, distributed tracing, and platforms such as Honeycomb.
Database operations experience, particularly PostgreSQL or Aurora performance, connection pool management, and query optimisation.
Experience administering and configuring Linux systems outside cloud environments.
Security experience covering both defensive and offensive practices.
Knowledge of cloud cost management and FinOps.
Experience expanding a team from a small starting point, including establishing hiring standards as the team grows.
Key Responsibilities
People Management
Manage the complete career lifecycle of team members, including onboarding, feedback, performance reviews, progression, and recruitment.
Monitor team health and dynamics and encourage effective retrospective practices.
Represent the SRE team to engineering teams and senior leadership.
Delivery
Define SRE objectives, including priorities, commitments, and the reasoning behind them.
Manage the support rotation and on-call structure.
Platform & Reliability
Oversee Remote's core infrastructure, including Kubernetes, AWS, PostgreSQL, DNS, TLS, and CI infrastructure.
Lead reliability practices covering SLOs, error budgets, incident response, and observability.
Work with the Security team on threats, patching, infrastructure controls, audits, and compliance requirements.
Manage platform-related vendor relationships, including renewals and commercial discussions with support from the Director.
