CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Senior Splunk Engineer

Avepoint • Singapore

Job Description

<p><strong>Senior Splunk Engineer for Automation and Reliability Engineering Project</strong></p> <p><strong>Project Summary</strong></p> <ul> <li>Support Automation and Reliability Engineering project and operations.</li> <li>Responsibilities:</li> <li>Observability Engineering and Governance</li> <li>Architect and maintain enterprise SIEM solutions aligned with operational resilience mandates (e.g., MAS TRM, DORA, APRA CPS 230).</li> <li>Lead deployment, configuration, and optimization of Splunk for full-stack visibility across infrastructure, applications, networks, and user experience.</li> <li>Define and enforce telemetry data governance standards—metrics, logs, and traces—ensuring consistency, retention compliance, and security.</li> <li>Integrate Splunk with incident management, ITSM, and AIOps systems to enable predictive alerting and anomaly detection.</li> <li>Act as the SIEM/Splunk subject matter expert (SME) for architecture reviews, platform upgrades, and performance tuning.</li> <li>Reliability Engineering and Automation</li> <li>Implement and champion SRE frameworks and reliability practices for mission-critical systems.</li> <li>Design and automate runbooks, alerts, and self-healing workflows using Python, Ansible, and Terraform.</li> <li>Collaborate with Application, Infrastructure, and Cyber teams to embed reliability principles into the delivery lifecycle.</li> <li>Conduct resilience, chaos, and capacity testing aligned with business continuity and disaster recovery standards.</li> <li>Define and track error budgets, reliability scorecards, and service health indicators for production workloads.</li> <li>Cloud & Platform Integration</li> <li>Engineer SIEM for cloud-native workloads in AWS and Azure, ensuring visibility across compute, storage, and network layers.</li> <li>Integrate Splunk and cloud observability tools into CI/CD pipelines and landing zones to ensure continuous compliance.</li> <li>Implement infrastructure-as-code (IaC) models using Terraform and Ansible for consistent, auditable provisioning.</li> <li>Collaborate with Cloud, DevOps, and Security teams to ensure telemetry aligns with audit, compliance, and operational risk requirements.</li> <li>Operational Excellence and Collaboration</li> <li>Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.</li> <li>Partner with Service Delivery, Cyber, and Application teams to enable predictive incident prevention and root cause transparency.</li> <li>Develop and maintain executive dashboards and reports showcasing availability, reliability KPIs, and operational risk indicators.</li> <li>Provide technical leadership during major incidents, post-incident reviews, and audits, ensuring lessons learned are codified into automation and process improvements.</li> </ul> <p><strong>Skillset (Must have)</strong></p> <ul> <li>Possess a degree in Computer Science, Engineering, or related discipline.</li> <li>Minimum 8 years of experience in Infrastructure, Cloud, or Site Reliability Engineering related roles, with at least 5 years of experience specializing in SIEM/Splunk engineering or observability in financial or regulated environments.</li> <li>Proven hands-on expertise in the following technical areas:</li> </ul> <ul> <li>SIEM Platforms: Splunk (must), EL/Elastic</li> <li>Automation/IaC, Terraform, Ansible, Python, CI/CD tools</li> <li>Cloud and other platforms and integrations: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Ins

Job Reference ID: CF-158339 • Posted on CloudFrame Job Scanner