Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
At Nscale, our Engineering team plays a critical role in designing, deploying, and operating the infrastructure and software platforms that power our customers and enable AI workloads at scale.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
We are hiring a Principal Network Engineer to act as a senior technical authority for Nscale's AI-optimised network infrastructure.
Our Network Engineering team is responsible for the design, validation, and ongoing operation of the networking services underpinning both our internal management platform and customer-facing cloud infrastructure. This includes high-performance Ethernet fabrics, InfiniBand, RoCE, WAN connectivity, and large-scale data centre networking.
In this role, you'll set technical direction across the low-latency, high-bandwidth networks supporting large-scale AI training and inference workloads. You'll own critical technical domains end-to-end and help raise the bar for architecture, automation, operational rigour, and engineering standards across Nscale.
This is a deeply technical Principal-level role combining hands-on engineering with broad architectural influence. You'll define reference architectures, drive consistency across sites, lead complex technical decisions and escalations, and mentor engineers while partnering closely with Deployment, Data Centre Operations, Platform Engineering, Systems, Storage, and technology vendors.
Define, design, validate, and evolve large-scale InfiniBand, RoCE, and Ethernet fabric architectures at rack, row, and data centre scale.
Design networks that integrate closely with bare-metal provisioning and cluster management systems.
Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS.
Establish reference architectures and engineering standards that can be implemented consistently across Nscale's data centre estate.
Identify systemic risks and architectural gaps and drive durable solutions that improve scalability, reliability, and operational simplicity.
Lead Nscale's network automation strategy using a GitOps operating model.
Build and guide Python and Ansible tooling for provisioning, configuration validation, compliance, and operational workflows.
Drive version-controlled configuration and CI/CD-based network change across multi-vendor environments.
Apply Infrastructure-as-Code and Network-as-Code principles to reduce manual intervention and improve operational consistency.
Continuously identify opportunities to automate repetitive operational tasks and reduce reactive toil.
Design and engineer perimeter and network security infrastructure across WAN and data centre edge environments.
Own architecture across firewalls, NAT, VPN, security policies, and multi-tenant segmentation.
Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure.
Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues.
Establish measurable SLOs and operational standards for network services.
Set technical direction for network observability, telemetry, monitoring, and alerting.
Ensure clear visibility into fabric health, traffic patterns, performance, and capacity.
Develop runbooks, automation, and engineering improvements that systematically reduce operational toil.
Act as a senior 3rd/4th line escalation point for complex networking issues.
Ensure the accuracy and reliability of source-of-truth network inventory and configuration data.
Establish structured engineering and change-management practices for network configuration.
Ensure network changes are controlled, auditable, repeatable, and scalable across multiple sites.
Partner with Deployment, Data Centre Operations, Platform Engineering, Systems, Storage, and vendors on new site delivery and platform evolution.
Lead architecture and design reviews for significant network initiatives.
Mentor engineers and raise technical capability across the wider networking organisation.
Lead complex technical decisions and incidents spanning networking, systems, storage, and AI/HPC workloads.
Influence engineering strategy and standards across teams without relying on formal authority.
10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data centre environments.
Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE.
Experience with subnet managers and fabric orchestration technologies such as OpenSM or NVIDIA UFM.
Expert-level understanding of modern data centre routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures.
Production experience with network platforms such as Cumulus, Nokia, or Arista EOS.
Strong network automation expertise using Python and Ansible.
Experience with Git-based workflows and modern Infrastructure-as-Code and CI/CD tooling such as Terraform, GitLab CI, or GitHub Actions.
Deep experience designing and engineering firewall infrastructure using platforms such as Juniper SRX and/or Palo Alto.
Experience designing telemetry and observability solutions for high-throughput, performance-sensitive environments.
Proven ability to lead complex technical decisions and incidents across multiple engineering disciplines.
Demonstrated experience defining network architecture, engineering standards, and technical strategy beyond a single project or data centre.
Ability to balance performance, reliability, operability, scalability, and delivery velocity when making technical decisions.
Experience influencing engineering teams and technical direction without formal authority.
Strong communication skills with the ability to explain complex technical trade-offs to engineers, operators, and senior stakeholders.
Proven ability to mentor and develop other engineers through design reviews, technical guidance, incident leadership, and knowledge sharing.
Deeply technical and comfortable remaining hands-on at Principal level.
Highly analytical with a structured approach to solving complex infrastructure problems.
Strong sense of ownership and accountability.
Comfortable operating in a fast-paced environment where priorities and requirements evolve quickly.
Pragmatic and able to balance architectural excellence with business and delivery requirements.
Passionate about building next-generation infrastructure for AI and ML at scale.
At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
Highly competitive package (base + equity) with reviews every 12 months.
Join one of the fastest-growing AI infrastructure companies — your opportunity to design and build the high-performance networks powering some of the world's most demanding AI workloads. ✨
Expect a dynamic progression plan tailored to your ambitions. Shape network architecture, establish engineering standards, solve complex infrastructure challenges, and influence the technical direction of Nscale's global AI platform.
Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there's anything we can do to accommodate your specific situation, please let us know.
The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.