Lead Software Engineer - AWS/SRE Engineer
JPMorgan Chase · United States · 16h ago
We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.
As a Lead Software Engineer at JPMorgan Chase within the Deposits team of Consumer & Community Banking Division, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.
Job responsibilities
- Lead Live Site Management and cloud production management practices, including production readiness, change governance, incident response, problem management, and operational risk reduction.
- Define and improve SRE practices, including service-level indicators, service-level objectives, error budgets, capacity planning, and reliability engineering.
- Advance observability across applications and infrastructure through effective logging, metrics, distributed tracing, dashboards, alerting, and automated anomaly detection.
- Drive major incident response, root-cause analysis, corrective actions, and the systematic elimination of recurring production issues.
- Establish production health, availability, performance, and resiliency measures that provide actionable insight to engineering and business leadership.
- Promote team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes, while establishing consistent validation standards and promoting reuse of effective patterns across the team.
- Apply knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized through automation.
- Shape the end-to-end AWS cloud architecture strategy across all product components ; Strengthen cloud cost strategy and optimization across the product.
- Advance the resiliency strategy, including multi-region, active-active, failover/failback, and disaster recovery.
- Evolve the automation and infrastructure-as-code strategy ; Mature AI integration across engineering and product workflows.
- Raise the bar on shared architectural standards and decision-making across teams ; Sharpen how complex technical strategy is distilled into leadership-facing narratives.
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Deep, broad AWS expertise across compute, data, networking, security, and observability, with proven production architecture experience.
- Systems thinking, including the ability to reason about trade-offs, failure modes, and dependencies across an entire product.
- Strategic command of cost, scalability, automation, reliability, and AI integration as connected levers.
- Hands-on depth to validate and prototype ideas, not just advise.
- Executive communication and rigorous, independent judgment in a regulated, high-stakes environment.
- Proven experience managing cloud-hosted production services and leading Live Site Management practices in complex, business-critical environments.
Strong SRE and observability expertise, including SLI/SLO definition, error budgets, telemetry, distributed tracing, alert design, capacity management, and performance engineering. - Demonstrated ability to lead high-severity incident response, perform root-cause analysis, and drive durable remediation across multiple engineering teams.
- Experience establishing production readiness standards, operational controls, runbooks, dashboards, and measurable service health objectives.
- Demonstrated experience leading effective use of approved AI-assisted software development tools, with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity, secure handling of inputs and outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices.