Technology Support Lead, Athena Core
JPMorgan Chase · Singapore · 3h ago
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
Join our dynamic team to innovate and refine technology operations, impacting the core of our business services.
As a Technology Support Lead in Commercial and Investment Banking (CIB) Technology - Athena Core, you will play a leadership role in ensuring the operational stability, availability, resiliency, and performance of our production services and core infrastructure platforms. You will demonstrate strong knowledge across multiple technical domains, advise others on technical and business issues, and apply critical thinking while overseeing day-to-day maintenance of the firm’s systems to ensure a seamless user experience.
Job Responsibilities
- Lead end-to-end service delivery for Athena Core application platform and infrastructure operations, including acting as a technical lead for infrastructure operations and partnering with teams to deliver resilient outcomes.
- Partner with development and infrastructure teams through the product lifecycle to ensure compliant, resilient deployments, including leading and conducting resiliency design reviews.
- Balance operational support with engineering delivery by automating toil, improving platform reliability, and building/maintaining tools and services across multiple technology domains.
- Lead performance and resiliency testing, bottleneck remediation, and capacity planning using data-driven service insights and operational telemetry.
- Serve as escalation during major incidents, restoring service quickly, reducing business impact, and driving incident, problem, and change management across full stack technology systems.
- Monitor production environments for anomalies, evolve usage of standard observability tools (monitoring, SLO-based alerting, telemetry collection), and communicate status, impact, and remediation to business and technology stakeholders through service restoration.
- Provide technical leadership and mentorship, breaking down complex problems into actionable work for engineers, promoting SRE culture, and driving adoption of SRE best practices (reliability, scalability, performance, security, and enterprise architecture).
- Lead team adoption of enterprise-authorized AI capabilities to improve incident triage speed and consistency (e.g., synthesizing operational signals into prioritized actions), with human-in-the-loop validation and appropriate handling of sensitive data.
- Leverage enterprise-authorized AI to accelerate patching, migrations, and testing with secure data handling, strong validation habits, and auditability.
- Apply reuse-first, AI-assisted practices across incident/problem/change routines to identify recurring interruption patterns and validate remediation actions aligned to resiliency and security expectations.
- Follow a 4 weekdays + 1 weekend day schedule
Required Qualifications, Capabilities, and Skills
- Formal training or certification on troubleshooting, resolving, and maintaining information technology services concepts and 5+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services
- Advanced Linux expertise, including system provisioning and configuration management using tools such as Puppet, Ansible, and Terraform.
- Strong software engineering skills in one or more programming languages (Python), covering design, coding, testing, and delivery.
- Proven ability to build automated tools, systems, and services across multiple technology domains to reduce manual toil.
- Solid understanding of core infrastructure domains, including networking, cloud services, orchestration, container platforms, and compute/storage.
- Demonstrated capability in reliability, scalability, performance, security, enterprise architecture, and other SRE best practices.
- Hands-on observability experience (monitoring, SLO-based alerting, telemetry collection) and proficiency in monitoring tools and techniques.
- Strong networking troubleshooting experience across common technologies and failure modes.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support production operations workflows, with strong validation habits and awareness of data sensitivity.
- Ability to review and validate AI-assisted operational/incident recommendations before action, assess correctness and risk, define team guardrails, escalate when uncertain, and ensure outcomes align to operational, resiliency, security, and auditability expectations.
- Experience managing applications or infrastructure in a large-scale technology environment (on premises and public cloud) and executing on processes in scope of the ITIL framework.
Preferred Qualifications, Capabilities, and Skills
- Proficiency with CI/CD practices and associated tooling.
- Demonstrated ability to implement service-level changes and troubleshoot end-to-end systems and components.
- Strong understanding of software applications and technical processes, with developing depth in at least one discipline and a continuous-learning mindset to evaluate and recommend emerging technologies.
- Experience designing and enhancing performance monitoring and capacity management solutions.