Lead Principal Systems Engineer
Oracle · Cairo, Egypt · 3d ago
Key Responsibilities
System Installation & Configuration – Software Administration:
-Evaluates and sets standards for the performance and installation requirements of operating systems to ensure optimal installation and performance.
-Provides guidance on the administration of middleware products in environments.
-Facilitates and implements strategies for the deployment, maintenance, and operation of internal applications, ensuring the efficiency and performance of these systems.
-Influences and leads regular administration and conducts highly complex performance trend analyses and manages server capacity to ensure service performance meets and exceeds standards.
-Serves as a subject matter expert in utilizing application monitoring tools to optimize and ensure efficiency.
System Installation & Configuration – Installation and Configuration:
-Leads the team installing and configuring servers, cloud infrastructure, and all software and environments.
-Facilitates the optimization of system configurations and backups to ensure the optimal performance and stability of the server infrastructure.
-Leads hardware maintenance, auditing, installation, and provisioning, ensuring all tasks are performed efficiently and effectively.
-Partners with internal technical experts and third-party vendors to resolve integration challenges, providing expert guidance and industry insights for innovative solutions.
System Installation & Configuration – Identity & Access Management:
-Provides additional support and guidance for the administration of access privileges in the identity and access management system, ensuring accurate and secure access to IT resources.
-Interprets user activity data in the identity and access management system and leverages expertise to provide insights on system activity and recommend improvements.
-Designs and optimizes access management systems.
Service Lifecycle Management – Batch Processing:
-Drives the monitoring and assessment of the batch process to ensure updates are applied and proactively resolves any issues that arise.
-Leverages industry insights to drive improvements in batch management techniques using different work schedulers to configure jobs and job streams, define dependencies, and report job performance.
-Influences and collaborates with teams to ensure scheduling and budgets of batch monitoring services align with and support Service Level Agreements (SLA).
Service Lifecycle Management – Security Maintenance:
-Drives strategic improvements to procedures to ensure that compute and storage devices are secure.
-Influences and collaborates with teams to maintain privileged accounts/secrets integrity of systems and compute and file system security for the compute and storage environment.
-Analyzes and evaluates highly complex service and infrastructure dashboards, taking the lead in addressing identified anomalies.
-Coordinates and manages long-term implementation strategies for monthly, quarterly, or hotfix patches to address security vulnerabilities or bugs.
Service Lifecycle Management – System & Security Improvements:
-Analyzes system performance data and insights to drive enhancements to improve the performance, reliability, and security of systems and environments.
-Influences and collaborates with Service teams to proactively identify, address, and predict gaps in operational capabilities, enhancing scalability and resiliency.
Incident Management & Support – Incident Management:
-Leads end-to-end incident management lifecycle to ensure systems are stable, secure, and performing accurately.
-Evaluates results from incident-based data analyses for team metrics and key performance indicators (KPIs) to identify patterns, root causes, and solutions to prevent system and network incidents.
-Facilitates incident review meetings to provide strategic oversight for operational performance and long-term solution implementation.
-Influences third party vendors and cross-functional teams (e.g., Development, Cloud Engineering, Product Engineering, other IT teams) to develop and implement long-term solutions for high-severity incidents, risks, or migrations.
-Serves as a subject matter expert in the investigation of highly complex system issues and facilitates high-severity incident triage by designing Corrective and Preventative Action plans (CAPA) to drive incident resolution and prevention.
Incident Management & Support – Escalation Cases:
-Provides expertise for escalated support cases by collaborating with internal technical teams and third party vendors to drive issue resolution for a wide range of production environment problems (e.g., immense growth, scaling, leveraging the cloud, extremely high performance, high availability requirements).
Incident Management & Support – Technical Support:
-Drives and evaluates the production environment by analyzing system error logs and ticket queues, and coordinating with multiple teams involved in maintaining the environments.
-Adheres to team schedule to drive ongoing technical support and service objectives.
-Facilitates strategic resolution and long-term solutions for highly complex, critical customer system issues and develops and documents comprehensive technical solutions.
Incident Management & Support – Backups and Disaster Recovery:
-Drives the execution and effectiveness of backup, restore, and disaster recovery processes.
-Influences the strategic planning and coordination of disaster recovery drills.
-Designs disaster recovery solutions to ensure preparedness and regulatory compliance.
Communication & Documentation – Technical Communication:
-Communicates highly complex technical information to both technical and nontechnical personnel including management.
-Develops and implements training programs to ensure personnel are well-versed in domain-specific knowledge and practices.
-Drives technical strategies and solutions for cross-organization projects, programs, and activities by leveraging domain-specific expertise and interpreting highly complex technical information.
Communication & Documentation – Documentation & Reporting:
-Drives the creation of documentation on ticket updates, code contributions, infrastructure, configurations, processes, and procedures (e.g., Disaster Recovery plans, Standard Operating Procedures, Corrective and Preventative Action Plans).
-Analyzes and interprets weekly and monthly reports on system performance and incident progress to provide insights on operational and management outcomes and business impacts.
-Reviews and refines technical documentation standards and best practices for internal use.
Additional Responsibilities (as needed)
Cloud Infrastructure Support:
-Serves as a subject matter expert in collaborations with DevOps and Site Reliability Engineer (SRE) teams to manage large-scale infrastructure.
-Manages continuous integration and continuous deployment (CI/CD) pipelines.
-Outlines and plans for patching and version upgrades to support cloud infrastructure.
Automation:
-Approves recommendations and determines plans for improvements to reduce incidents and problems with automation and simplify server management.
-Owns and leads the implementation of reusable frameworks, standards, and automation to support Oracle Cloud Infrastructure.
-Establishes and oversees Workload Automation tools through design support, administration, and optimization efforts.
-Champions troubleshooting complex issues with automation tools, agents, and other connectivity to 3rd party applications.
-Manages cloud technologies.
Core Responsibilities
Planning & Execution:
-Manages and provides direction on timelines, deliverables, and budgets when applicable for critical high-impact projects or initiatives that impact the line of business, ensuring timely completion and adherence to requirements.
-Anticipates and plans for shifts in resources or timelines based on changing business priorities, ensuring optimal outcomes.
Collaboration & Partnership:
-Influences cross-functional leaders and external stakeholders to gain alignment on strategic objectives.
-Fosters partnerships with key business leaders, stakeholders, and/or customers, identifying opportunities for expanding partnerships and promoting long-term organizational success.
-Champions transparency and inclusivity by actively seeking, listening to, and incorporating diverse perspectives.
Problem Solving:
-Leads specialized, advanced problem-solving efforts, serving as an escalation point for complex issues.
-Guides others to leverage innovative data-driven techniques to address ambiguous or novel issues, identify root causes, and drives the implementation of solutions that prevent future issues.
Continuous Learning:
-Leverages deep industry knowledge and expertise to serve as a thought leader within the organization.
-Contributes to the advancement of the field or industry through thought leadership (e.g., conference presentations, white papers, research contributions).
-Maintains and evolves expertise in relevant areas by proactively monitoring emerging trends, technologies, and industry standards, ensuring the organization remains current with best practices.
-Champions continuous learning and knowledge sharing, promoting professional development across teams. Applies new knowledge to drive advancement and mentors others to do the same.
Continuous Improvement:
-Develops innovative solutions and drives the implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the organization.
-Evaluates effectiveness of updated approaches and methods for continued improvement to enhance efficiencies and ensure changes align with organizational goals.
-Designs and develops metrics to measure success of improvement initiatives.
Performance and Development:
-Serves as a subject matter expert regarding talent needs and organizational talent strategy.
-Imparts leadership and expert knowledge throughout the talent development pipeline including candidate interviews, candidate assessment, and hiring decisions, ensuring alignment with organizational talent strategy.