Site Reliability Engineer - Istanbul, Türkiye
MetLife · No · 2mo ago
When you join MetLife’s Global Technology team, you’ll be part of a forward-thinking group dedicated to shaping the future of digital solutions for customers worldwide. You’ll develop, maintain and support technology applications and delivery, leveraging AI, automation, and contemporary ways of working to enhance experiences and drive business outcomes. Your work will simplify complex processes, improve tech resiliency, and ensure high-performing, seamless solutions that power life’s most important moments. In this dynamic environment, you’ll collaborate with talented peers across teams and functions, expanding your skills in impactful ways. Ready to push boundaries and set new industry standards? Join us and help drive the future of technology forward.
- Work on enterprise-scale, mission-critical systems serving real business operations
- Build and enhance observability capabilities (logs, metrics, traces) to improve system reliability and transparency
- Utilize AI-powered analytics on observability data (logs, metrics, traces) to detect anomalies, accelerate root cause analysis, and improve operational intelligence
- Design and implement scalable and resilient platform solutions in cloud and hybrid environments
- Collaborate with cross-functional teams (development, infrastructure, security) to improve system reliability and performance
- Contribute to automation-first operations, reducing manual effort and increasing efficiency
- Gain hands-on experience with modern SRE practices including SLOs, incident management, and reliability engineering
- Participate in building a data-driven engineering culture using observability insights
Be part of a global organization with modern engineering standards, tools, and practices
- Opportunity to work on enterprise-scale, mission-critical systems
- Ownership of advanced observability and monitoring platforms
- A culture of engineering excellence, automation, and continuous improvement
- Collaboration with global teams and exposure to modern SRE practices
- Continuous learning and professional growth opportunities
- Design, implement, and manage end-to-end observability solutions (metrics, logs, traces)
- Develop dashboards, alerts, and visualization layers for proactive issue detection
- Build and maintain Elastic Stack (ELK / OpenSearch) based logging and monitoring platforms
- Define and continuously improve SLIs, SLOs, and alerting strategies
- Enable log, metric, and trace correlation to improve troubleshooting efficiency
- Ensure high availability, scalability, and performance of distributed systems
- Drive adoption of reliability practices such as incident retrospectives and proactive monitoring
- Participate in incident response, root cause analysis, and resilience improvement initiatives
- Implement automated remediation and self-healing mechanisms
- Integrate security monitoring and logging (SIEM-like use cases) into observability platforms
- Collaborate with security teams on threat detection, anomaly monitoring, and audit logging
- Contribute to DevSecOps practices, embedding security into CI/CD pipelines
- Support audit readiness and compliance reporting through structured logging and monitoring
- Automate operational workflows to reduce toil and increase efficiency
- Contribute to the improvement of CI/CD pipelines and release processes
- Support on-call operations and continuously improve alert quality and signal-to-noise ratio
- Develop Python scripts for synthetic monitoring and testing
- Bachelor’s degree in Computer Science, Engineering, or related field
- 3+ years of experience in SRE, DevOps, or production engineering roles
- Good command of English
- Strong hands-on experience with Elastic Stack (Elasticsearch, Logstash, Kibana)
- Proficiency in other monitoring tools (e.g., Prometheus, Grafana, Azure Monitor, App Insights, Splunk)
- Experience with observability frameworks (metrics, distributed tracing, logging)
- Experience working with cloud platforms (Azure preferred)
- Strong scripting/programming skills (Python, Bash, etc.)
- Understanding of distributed systems and microservices architecture
- Solid understanding of security logging, audit trails, and system hardening
- Experience in tools such as Visual Studio, Azure DevOps, GitHub Enterprise, GitLab, CI/CD
- Experience working in Financial Services / Insurance sector is an advantage
- Experience building advanced automation scripts or tooling is a plus
- Experience of working in an Agile environment and using Agile methodologies
Sunduğumuz Yan Haklar
Yan haklarımız; fiziksel ve ruhsal sağlığı, finansal iyilik hâlini ve ailelere yönelik destekleri kapsayan programlarla bütünsel iyilik hâlinizi desteklemek üzere tasarlanmıştır.
Size ve ailenize özel sağlık sigortası, hayat sigortası, işveren katkılı emeklilik planı, yemek ve ulaşım ödeneği ile evden çalışma ödeneği sunuyoruz. Ayrıca Kültürel Miras Günü izni ile “okula dönüş” ve “karne günü” izinleri ve çok daha fazlasını sağlıyoruz!