Senior Site Reliability Engineer
Oracle · Bangalore, India · 4d ago
At Oracle Cloud Infrastructure (OCI), we are building the future of cloud for enterprises with the agility of a startup and the scale of a global enterprise leader. Compute is one of OCI’s foundational organizations, responsible for delivering the core infrastructure powering Virtual Machines (VMs) and Bare Metal (BM) services.
As an OCI Site Reliability Engineer (SRE), you will work closely with development and product teams in a shared full-stack ownership model across multiple services and technology domains. You will develop deep expertise in service architecture, dependencies, configurations, and operational behavior across large-scale production environments.
You will be responsible for improving the reliability, scalability, performance, and operational efficiency of OCI Compute services. The role includes handling critical customer incidents, supporting deployments, performing validation and operational testing, troubleshooting complex infrastructure issues, conducting root cause analysis (RCA), and driving service reliability improvements.
You will act as a key escalation point for complex production issues, leveraging strong knowledge of distributed systems, service topology, and infrastructure dependencies to identify mitigations and restore service health while partnering with development teams to meet SLA commitments.
The role also involves leveraging AIOps and intelligent automation to enhance monitoring, anomaly detection, event correlation, predictive alerting, RCA, and remediation workflows. Using observability platforms, telemetry analytics, and automation frameworks, you will help reduce operational toil, improve incident response, and enhance overall service reliability.
This is an opportunity to combine deep technical expertise with operational excellence to solve complex cloud infrastructure challenges at massive scale within Oracle’s next-generation cloud platform.
Job Responsibilities
- Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
- Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
- Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
- Build automation and tooling to reduce operational toil and improve production safety.
- Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
- Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
- Support upgrades, migrations, patching, capacity planning, performance tuning, and production rollouts.
- Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
- Contribute to incident-management practices, operational readiness, and service ownership improvements.
- Share technical knowledge and support team members through documentation, reviews, and collaboration.
- Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.
Mandatory Skills
- 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
- Experience operating and improving highly available production systems.
- Strong programming or scripting skills in Python, Go or similar languages.
- Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
- Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
- Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
- Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
- Experience with deployment pipelines, release validation, automation, and change-management practices.
- Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
- Ability to work independently on technical problems and collaborate effectively with engineering teams.
- Strong written and verbal communication skills.
Preferred Skills
- Experience with OCI and cloud infrastructure services.
- Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
- Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
- Experience with service migrations, fleet maintenance, upgrades, patching, security vulnerability management or production rollouts.
- Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
- Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
- Familiarity with security, compliance, and access-control practices in production environments.