CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Senior Site Reliability Engineer

Infosys • Bangalore, India

Job Description

Reliability Engineering Design, build, and maintain highly available and fault-tolerant production systems. Define and monitor SLIs, SLOs, and SLAs for critical services. Drive reliability improvements through automation and proactive engineering. Conduct capacity planning and performance optimization activities. Production Support & Operations Manage production environments and ensure service uptime. Lead incident response, troubleshooting, and root cause analysis (RCA). Develop runbooks, operational playbooks, and disaster recovery procedures. Participate in on-call rotations and major incident management processes. Technical Requirements: Observability & Monitoring Build monitoring, logging, tracing, and alerting solutions. Implement observability frameworks using industry-standard tools. Monitor application health, performance metrics, and infrastructure utilization. Drive continuous improvements in platform visibility and diagnostics. Automation & DevOps Automate deployments, infrastructure management, and operational workflows. Improve CI/CD pipelines and release processes. Implement self-healing, auto-scaling, and operational automation solutions. Promote DevOps and SRE best practices across engineering teams. Responsibilities: Cloud & Infrastructure Deploy and manage cloud-native infrastructure across AWS, Azure, or GCP. Automate infrastructure provisioning using Infrastructure as Code (IaC). Implement scalable and secure infrastructure solutions. Support Kubernetes-based platforms and containerized workloads. Preferred Skills: Technology->DevOps->Site Reliability Engineering(SRE)

Job Reference ID: CF-91522 • Posted on CloudFrame Job Scanner