Lead Site Reliability Engineer
JPMorgan Chase · Jersey City, United States · 13h ago
Staff / Lead / PrincipalOn-siteDevOps, SRE & Infrastructure
J.P. Morgan Asset & Wealth Management delivers industry-leading investment management and private banking solutions. Asset Management provides individuals, advisors and institutions with strategies and expertise that span the full spectrum of asset classes through our global network of investment professionals. Wealth Management helps individuals, families and foundations take a more intentional approach to their wealth or finances to better define, focus and realize their goals.
As a Lead Site Reliability Engineer at JPMorgan Chase within the Asset and Wealth Management, Tech Production and Infrastructure Delivery team, you will be responsible for improving reliability, resilience, and operational performance across a hybrid technology environment spanning modern distributed platforms and mainframe systems.
Job Responsibilities
- Lead adoption and operationalization of SRE practices, including SLIs/SLOs, error budgets, reliability reviews, and blameless post-incident processes.
- Design, implement, and continuously improve monitoring and observability capabilities across metrics, logs, traces, and event telemetry to support faster detection and diagnosis.
- Establish actionable alerting standards, dashboards, and runbooks to improve operational readiness and reduce noise.
- Drive automation initiatives (self-service, self-healing, automated remediation, CI/CD operational controls, and standardized tooling) to reduce manual effort and improve consistency.
- Identify, measure, and reduce operational toil through process optimization, tooling enhancements, and platform improvements.
- Improve incident management practices, including incident response coordination, escalation paths, and continuous improvement based on root cause analysis.
- Support capacity planning, performance engineering, and resilience testing to strengthen availability and service stability.
- Partner with application, infrastructure, and operations teams across distributed and mainframe domains to standardize reliability patterns and operational controls.
- Contribute to governance and operational excellence, including documentation, control evidence where applicable, and operational health reporting.
Required qualifications, capabilities and skills
- Relevant experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or a similar reliability-focused role, including leadership of technical initiatives.
- Strong knowledge of operating and supporting distributed systems in production (e.g., Linux, networking, middleware, containers and/or cloud platforms).
- Hands-on experience with monitoring/observability platforms and practices (metrics, logs, traces), including dashboarding and alert engineering.
- Demonstrated ability to automate operational workflows using one or more scripting/programming languages (e.g., Python, Go, Shell) and standard automation approaches (CI/CD, infrastructure-as-code).
- Experience supporting or integrating mainframe systems into enterprise operations (monitoring, incident response, operational processes).
- Strong communication and stakeholder management skills, with ability to lead cross-team reliability improvements.
Preferred Qualifications
- Experience establishing SLIs/SLOs and reliability reporting at service or platform level.
- Familiarity with ITSM/incident tooling, on-call operations, and operational maturity improvements.
- Experience with resilience patterns (graceful degradation, failover, rate limiting) and reliability testing (chaos testing, load/performance testing).
- Exposure to regulated or high-control environments and operational risk management practices.