Engineering Job At Qure.ai
Recruit Finds
Nairobi, Kenya
Job summary
Site Reliability Engineer - Kenya At Qure.ai
About this role
Join our Whatsapp Channel - CLICK HERE
About the Role
Our client is seeking an experienced Site Reliability Engineer (SRE) to support the deployment, reliability, and performance of AI-powered healthcare solutions across customer environments. The successful candidate will be responsible for maintaining highly available infrastructure, supporting production deployments, optimizing system performance, and ensuring the seamless operation of both cloud-based and on-premises platforms. This position requires frequent travel within Kenya and other African countries to support customer implementations and provide on-site technical assistance.
Key Responsibilities
Deploy, configure, and support AI and computer vision applications across customer environments running Linux and Windows operating systems in accordance with established deployment standards.
Install, administer, and maintain on-premises server infrastructure within hospitals and healthcare facilities, including server provisioning, operating system updates, security hardening, configuration management, and capacity planning for GPU-enabled and compute-intensive workloads.
Troubleshoot, diagnose, and resolve technical issues affecting production systems by performing root cause analysis across operating systems, networks, infrastructure, and application layers.
Implement and maintain infrastructure automation using Infrastructure-as-Code (IaC) tools such as Ansible, Terraform, or equivalent technologies to improve deployment consistency and operational efficiency.
Monitor system health, infrastructure performance, and application availability across cloud and on-premises environments, identifying performance bottlenecks and implementing optimization strategies.
Configure and maintain monitoring, logging, and alerting platforms such as Prometheus, Grafana, ELK Stack, or similar tools to enable proactive issue detection and rapid incident response.
Participate in incident management activities, including troubleshooting, service restoration, post-incident reviews, and implementation of corrective and preventive measures.
Collaborate with information security teams to strengthen infrastructure security, manage user access controls, and ensure compliance with healthcare data protection and security standards.
Communicate technical issues, recommendations, and project updates effectively to both technical teams and non-technical stakeholders, including healthcare professionals and client representatives.
Work closely with Engineering, Product Management, Customer Success, and other cross-functional teams to promote system reliability, operational excellence, and continuous service improvement.
Travel locally and internationally, as required, to support customer deployments, infrastructure maintenance, implementation projects, and technical engagements.
Qualifications and Experience
Applicants should possess the following qualifications and competencies:
Bachelor's degree in Computer Science, Information Technology, Software Engineering, Information Systems, or a related field.
Demonstrated experience in Site Reliability Engineering (SRE), Systems Administration, DevOps, Infrastructure Engineering, or a similar technical role.
Strong knowledge of Linux and Windows server administration.
Experience deploying, configuring, and supporting enterprise applications in both cloud and on-premises environments.
Hands-on experience with Infrastructure-as-Code (IaC) tools such as Terraform, Ansible, or equivalent automation technologies.
Working knowledge of monitoring, logging, and observability platforms including Prometheus, Grafana, ELK Stack, or comparable solutions.
Experience troubleshooting infrastructure, networking, operating systems, and enterprise application issues within production environments.
Familiarity with GPU-based computing environments, virtualization technologies, and cloud platforms will be an added advantage.
Understanding of cybersecurity best practices, access management, and infrastructure hardening techniques.
Excellent analytical, troubleshooting, and problem-solving skills with strong attention to detail.
Strong interpersonal, written, and verbal communication skills with the ability to engage effectively with both technical and non-technical stakeholders.
Willingness and ability to travel extensively within Kenya and across Africa to support customer implementations and technical operations.
Core Competencies
The successful candidate should demonstrate:
A collaborative mindset with the ability to work effectively across multidisciplinary teams.
High levels of ownership, accountability, and initiative in managing technical operations.
Strong critical thinking and structured problem-solving abilities.
Adaptability and resilience when working in dynamic, customer-facing environments.
Commitment to continuous learning, innovation, and operational excellence.
Professionalism, integrity, and a customer-focused approach to service delivery.