Senior Network Developer
Oracle · Bangalore, India · 22d ago
We are the AI Infrastructure - Network Operations team at OCI (Oracle Cloud Infrastructure). We support and operate the RDMA/RoCE/InfiniBand network fabrics for OCI's largest AI and HPC customers. These fabrics are the foundation underneath OCI's AI, GPU and HPC services, and support major tier-0 vendors in the generative AI industry. If you're running an AI workload at OCI, we're running the RDMA network underneath your workload.
A Network Operations Engineer on our team supports the design, deployment, and operations of a large-scale global Oracle cloud computing environment (Oracle Cloud Infrastructure - OCI). Primarily focused on operation and support of RDMA/RoCE/InfiniBand network fabrics and systems, through a combination of a deep network understanding and automation skills to operate a production environment. As OCI is a cloud-based network with a global footprint, this support will include hundreds of thousands of network devices supporting millions of servers, connected over a mix of dedicated backbone infrastructure and the Internet.
Key Responsibilities
Design, operate, validate, and scale advanced network fabrics for large-scale cloud, AI, and data center environments.
Develop automation, scripts, and tooling to streamline network testing, operations, deployment, and troubleshooting.
Build and enhance telemetry, dashboards, alerting, and monitoring to improve network health, reliability, and SLO performance.
Develop test strategies, lead pre-production validation, and drive root cause analysis (RCA) for network issues and changes.
Analyze network performance, capacity, latency, throughput, and packet loss to identify issues and support infrastructure growth.
Participate in incident response and operational support, resolving complex production and customer issues.
Partner with engineering teams, vendors, and stakeholders on network architecture, deployments, standards, and operational readiness.
Identify design and operational risks, drive mitigations, and continuously improve network processes and reliability.
Mentor engineers, provide technical guidance, and contribute to architecture, roadmap, and engineering best practices.
Independently manage priorities and deliverables while collaborating across teams to achieve shared objectives.
- Act as a Tier 2 and specialized escalation point for network incidents, driving root cause analysis, corrective actions, and long-term reliability improvements.
Preferred Skills & Experience
- 6+ years of experience.
- Bachelor’s degree (Master’s preferred) in Computer Science, Electrical Engineering or related field.
- Strong experience with large-scale network operations, design, and troubleshooting.
- Expertise in routing and switching technologies, including BGP, OSPF, EVPN-VXLAN, MPLS, and data center networking. Prion experience with RDMA/RoCE/InfiniBand would be a plus.
- Experience with network automation using Python, Ansible, APIs, or similar technologies, building AI agents to automate.
- Strong understanding of observability, monitoring, telemetry, and incident management.
- Experience working with cloud infrastructure, hyperscale environments, or large-scale distributed systems.
- Ability to lead technical projects and influence outcomes across multiple teams.
- Strong written and verbal communication skills with the ability to work effectively across engineering, operations, and leadership teams.