SRE III
JPMorgan Chase · Buenos Aires, Argentina · 15h ago
As a Senior Production Engineer, you are the named, single point of accountability for the operational health of an assigned business service or payment flow. You defend its reliability budget, own its controls evidence, and represent its client impact to executive stakeholders — converting the question "is the ticket closed?" into "did the client experience degrade, and what did it cost us?"
Job Responsibilities-
Owns end-to-end operational health of a named business-service or payment-flow portfolio, accountable for availability, MTTD, MTTM, MTTR, and client-impact trend against published SLOs and error budgets
-
Defends the reliability (error) budget for owned services and makes change-governance decisions based on budget consumption, not ticket urgency alone
-
Leads incident, problem, and change management for full-stack payment systems, with authority to escalate cross-functionally and represent client impact directly to executive stakeholders
-
Owns controls and regulatory evidence for assigned services — patch compliance, certificate expiry, vulnerability aging, EOL exposure — and drives automation of manual evidence collection
-
Converts manual, repetitive work into automation against a named toil backlog, measuring hours removed and capacity returned to change-the-bank work
-
Partners with application development and central SRE/observability teams under a shared reliability services agreement, without absorbing their engineering backlog
-
Builds depth of coverage on owned services, eliminating single-person dependencies and naming a backup owner
-
5+ years operating payment or financial-services production systems with named accountability for service availability and incident outcomes
-
Demonstrated ownership of SLOs, error budgets, or an equivalent reliability engineering discipline in a large-scale, regulated technology environment
-
Working knowledge of payment flows — authorization, clearing, settlement, billing, payment/rebate — and their client and regulatory impact
-
Experience partnering with controls, risk, and audit functions on production evidence and regulatory findings
-
Experience with observability/monitoring tooling and automation or scripting for toil elimination
-
Experience standing up or operating under a formal SRE/reliability operating model
-
Experience with AI-assisted operations (RCA, ECC, runbook automation) under outcome-gated governance
-
Cross-business-unit or multi-region service ownership experience