Description
Location: US Remote
Salary Range: 100-130K USD
Join a Team Where You'll Keep Learning
If you're an experienced Platform Engineer with a strong Linux and High Performance Computing background looking for your next technical challenge, this could be the opportunity you've been looking for.
At X-ISS, part of n² Group, we deliver managed services and project expertise for customers running complex, large-scale High-Performance Computing (HPC) environments across scientific research, artificial intelligence and enterprise workloads.
You'll join a highly skilled, collaborative team where knowledge is shared, ideas are welcomed and everyone is encouraged to keep developing their expertise. You'll work alongside experienced engineers, solve challenging technical problems and deepen your HPC specialism day-to-day.
This is a hands-on engineering role where no two days are quite the same. You'll work directly with customers, helping them maintain and improve business-critical HPC platforms while developing specialist skills in a highly specialized area of infrastructure.
If you're passionate about Linux, enjoy solving complex technical challenges and are excited by the opportunity to build specialist HPC expertise, we'd love to hear from you.
RequirementsWhat We’re Looking For Essential Experience and Skills
- A bachelor’s degree, or equivalent qualification, in computer science.
- A minimum of six years' professional experience implementing and supporting Linux server solutions.
- Experience using service desk or ticketing systems to a high standard.
- Exceptional troubleshooting skills across Linux servers and networking within large-scale environments, with the ability to work through interconnected hardware, network and software components across on-premise and cloud infrastructure to identify and resolve root causes.
- Expertise in HPC platform engineering, with practical experience deploying, automating, and maintaining compute infrastructure in both on-premise and cloud environments, including containerization (Docker, Kubernetes) and Linux/Rocky 8 administration.
- Hands-on experience with HPC job scheduling and workload management (e.g., Slurm, PBS Pro, or LSF), and automating bare-metal cluster provisioning (e.g., Ansible, xCAT, Warewulf, or Foreman/Cobbler) alongside equivalent automation in cloud environments.
- Experience supporting and utilizing continuous integration and delivery (CI/CD) pipelines, enabling seamless, repeatable deployments across hybrid on-premise and cloud infrastructure. Familiarity with GitLab and DevOps pipelines is a plus.
- Proficient programming skills in scripting and automation languages, such as bash or Python, and infrastructure-as-code tooling (e.g., Ansible or Terraform), for automating operational tasks, system configuration, and cluster provisioning at scale.
- Familiarity with core HPC infrastructure components, including parallel file systems (e.g., Lustre, GPFS/Spectrum Scale, or BeeGFS), high-speed interconnects (InfiniBand/RDMA), and environment/module management (e.g., Lmod, Spack, or EasyBuild).
- Strong interpersonal and communication skills to effectively engage and collaborate with both technical and non-technical team members, ensuring clarity and understanding across all collaborators.
Preferred Skills
- Familiarity with databases, including deployment, optimization, and automation. MySQL experience preferred.
- Linux and/or networking certifications.
- Hands-on experience with cloud and on-premise infrastructure monitoring and observability tools (e.g., Prometheus/Grafana alongside cloud-native monitoring stacks), to ensure system observability and performance tracking.
- Experience with GPU-accelerated computing environments, including NVIDIA drivers, CUDA, and GPU scheduling, is a plus.
Who You AreWe're looking for someone who enjoys learning and takes pride in delivering great technical solutions. You'll likely be:
- Curious, enthusiastic and motivated to learn new technologies.
- A creative problem-solver who enjoys tackling complex technical challenges.
- Adaptable and able to manage multiple priorities effectively.
- Comfortable learning new tools and technologies quickly.
- A collaborative team player who enjoys sharing knowledge and supporting colleagues.