EngRadardirect-apply

Cluster Engineer

Stninc

Remote Remote Full-time Posted 8d ago

Position Summary
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization.
The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.

Responsibilities

  • Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.

  • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.

  • Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.

  • Build and support production AI infrastructure running hundreds to thousands of GPUs.

  • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.

  • Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.

  • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.

  • Configure and tune distributed AI software stacks including:

    • PyTorch

    • NCCL

    • CUDA

    • UCX

    • MPI

    • Slurm

    • Pyxis/Enroot

  • Optimize GPU scheduling and resource allocation for both training and inference environments.

  • Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.

  • Identify performance regressions and troubleshoot distributed training issues at scale.

  • Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.

  • Work closely with ML engineers to improve training scalability and inference efficiency.

  • Create automation to deploy, validate, benchmark, and monitor GPU clusters.

  • Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.



    Required Qualifications

  • 7+ years designing or operating large-scale Linux infrastructure.

  • 5+ years supporting production GPU clusters for AI or HPC workloads.

  • Demonstrated experience building multi-node GPU training environments from the ground up.

  • Deep expertise with distributed PyTorch training.

  • Extensive experience troubleshooting and optimizing NCCL communications.

  • Strong understanding of distributed AI communication patterns, including:

    • AllReduce

    • ReduceScatter

    • AllGather

    • Broadcast

    • Point-to-point communications

  • Experience benchmarking distributed training using tools such as:

    • nccl-tests

    • NVIDIA DCGM

    • Nsight Systems

    • MLPerf (preferred)

  • Strong understanding of GPU memory management, including:

    • KV Cache

    • Activation checkpointing

    • Tensor Parallelism

    • Pipeline Parallelism

    • Data Parallelism

  • Experience optimizing LLM inference throughput, including:

    • Tokens/sec optimization

    • Batch sizing

    • Continuous batching

    • KV cache tuning

    • Memory bandwidth optimization

  • Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.

  • Expert-level Linux systems administration skills.

  • Experience with Slurm workload manager.

  • Experience using Pyxis and Enroot for containerized GPU workloads.

  • Strong scripting skills using Python and Bash.



Technical Expertise
AI Frameworks

  • PyTorch

  • CUDA

  • NCCL

  • Triton (preferred)

  • TensorRT-LLM (preferred)

Cluster Scheduling

  • Slurm

  • Pyxis

  • Enroot

GPU Networking
Strong understanding of:

  • InfiniBand

  • RoCE v2

  • RDMA

  • GPUDirect RDMA

  • GPUDirect Storage

  • UCX

  • MPI

  • Network topology optimization

  • Congestion control

  • QoS

  • ECN/PFC

  • High-speed Ethernet (200/400/800 GbE)

Storage
Experience designing or tuning storage for AI workloads, including:

  • Parallel file systems

  • Distributed storage

  • Object storage

  • NVMe

  • Checkpoint optimization

  • Dataset staging

  • GPUDirect Storage

  • Storage bandwidth optimization

  • Metadata performance

Performance Engineering
Experience with:

  • NCCL benchmarking

  • Multi-node scaling analysis

  • GPU utilization optimization

  • Communication/computation overlap

  • NUMA optimization

  • CPU affinity

  • PCIe topology

  • GPU topology (NVLink/NVSwitch)

  • Memory bandwidth analysis

  • End-to-end performance profiling



Preferred Qualifications

  • Experience deploying AI workloads on Kubernetes.

  • Experience with NVIDIA GPU Operator.

  • Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).

  • Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.

  • Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.

  • Familiarity with MLPerf benchmarking.

  • Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.

  • Experience automating infrastructure using Ansible, Terraform, or similar tools.

  • Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.

Posted by Stninc on their own careers page — you apply directly, no recruiter in between. View original / apply →

More at Stninc