CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

SRE / Production Engineering

Infosys • Bangalore, India

Job Description

Join a high-impact Production Engineering/SRE team where reliability, speed, and customer trust are at the center of everything we do. In this role, you’ll lead the engineering effort to keep large-scale, business-critical platforms stable, performant, and continuously improving—while enabling product teams to ship confidently. You’ll work closely with developers, infrastructure, and security partners to design resilient systems, automate operational toil, and build pragmatic observability that turns signals into action. If you enjoy solving complex production challenges, driving incident excellence, and mentoring teams toward modern reliability practices, this is a place where your leadership will be visible and valued. Expect a collaborative culture, ownership-driven execution, and the opportunity to shape reliability standards across services and environments. Technical Requirements: SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana) Responsibilities: Key Responsibilities: Reliability & Production Ownership • Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms. • Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities. • Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements. Incident Management & Operational Excellence • Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed. • Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence. • Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes. Cloud Operations & Automation • Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions. • Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency. • Standardize operational tooling and workflows to improve service health, change success rate, and MTTR. Leadership & Collaboration • Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices. • Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives. Minimum Qualifications: • BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience). • 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems. • Strong hands-on experience in cloud operations, incident management, and production support for critical services. • Proven ability to drive automation initiatives that reduce manual effort and improve system reliability. • Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices. Preferred Skills: Technology->DevOps->Site Reliability Engineering(SRE),Domain->Manufacturing->Production Planning

Job Reference ID: CF-128383 • Posted on CloudFrame Job Scanner