Lead Principal Software Engineer
Oracle · Nashville, United States · 2mo ago
Oracle Cloud Infrastructure (OCI) delivers mission-critical cloud services to enterprises worldwide. The Physical Networking Automation and Tooling team builds and operates the software platforms that enable Network Engineers to manage OCI’s global physical network through automation, observability, and actionable insights.
We are building intelligent network automation platforms that combine network telemetry, topology, device state, operational workflows, and AI/ML to automate work across the network lifecycle. You will set the technical direction for AI agents that analyze network context, interact with approved tools, and execute controlled workflows with appropriate human oversight. You will define how these capabilities integrate into reliable, scalable software platforms that support production network operations.
As a Lead Principal Software Engineer, you will own architecture and technical direction for these platforms, aligning engineering teams and driving complex initiatives from concept through production. You will remain hands-on in critical design and implementation, establish engineering standards for reliability and safe automation, and mentor engineers and technical leaders. Working closely with engineering and network operations leaders, you will shape how OCI operates its physical network as it grows.
Key Responsibilities
- Define the long-term architecture and technical roadmap for physical network automation, validation, observability, and AI-driven operations across OCI.
- Lead the design of distributed, event-driven systems that ingest and correlate network failures, performance, power, capacity, topology, and change data at global scale.
- Architect AI decision systems for anomaly detection, fault diagnosis, capacity forecasting, and remediation. Establish how models and agents use network context, express uncertainty, and escalate decisions to engineers when needed.
- Direct the development of agent workflows that discover and invoke approved capabilities, coordinate tools and services, maintain execution state, and recover safely from partial failures.
- Establish controls for actions affecting production networks, including identity and authorization, approval gates, validation, limits on the impact of failures, cancellation, rollback, and audit trails.
- Define evaluation and observability standards for AI-driven capabilities. Measure diagnostic accuracy, false positives, action quality, latency, cost, and operational outcomes before and after deployment.
- Set architecture and engineering standards across APIs, backend services, data pipelines, workflow execution, testing, deployment, and production operations.
Advise senior leaders on technical investments and risk. Mentor principal and senior engineers and strengthen engineering practices across teams.
Minimum Qualifications
- At least 15 years of experience in software engineering, network automation, cloud infrastructure, or a related field, including technical leadership of initiatives spanning multiple teams.
- Bachelor’s degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience.
- Strong programming skills in Java and Python, with experience building and operating production software.
- Deep experience designing distributed systems, APIs, event-driven data pipelines, workflow systems, or cloud-native services.
- Experience defining technical architecture and leading complex projects from concept through production operation.
- Experience with network operations, data center infrastructure, or other large-scale infrastructure systems.
- Experience building AI/ML or LLM-based capabilities to operational problems, including evaluating their output and defining when human review is required.
- Strong knowledge of Linux, software development practices, CI/CD, observability, and production incident response.
Demonstrated ability to influence senior stakeholders, resolve architectural tradeoffs, and mentor experienced engineers.
Preferred Qualifications
- Experience developing AI/ML automation workflows with network telemetry, device connectivity, configuration validation, or automated remediation in large-scale networks.
- Track record of architecting agentic systems, including planner and executor patterns, tool orchestration, agent handoffs, state management, checkpointing, retries, and durable workflows.
- Hands-on work integrating agents with enterprise tools through MCP, A2A, or similar protocols, including authentication, tool schemas, context handling, approval gates, and failure handling.
- Knowledge of AI evaluation, quality, observability, and governance practices for production systems.
- Proficiency with Kubernetes, containers, and operating cloud-native platforms in production.
- Practical use of AI-assisted development tools for coding, testing, review, and debugging.
- Contributions to industry standards groups or open source projects related to network automation or data center infrastructure, such as the IETF, IEEE 802, or the Open Compute Project.
This is an individual contributor role with broad technical influence. Success means delivering cross org impact that increases OCI’s operational capabilities, and reduce the time required to resolve and deploy network infrastructure safely at global scale.
Disclaimer:Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.