Editorial summary

The company is seeking an AI Infrastructure Engineer to manage and maintain high-density multi-GPU compute clusters and container orchestration platforms optimized for AI and machine learning workloads. The role involves monitoring GPU health, optimizing high-performance networking and storage, and ensuring efficient cluster operation. Candidates should have experience with GPU architectures such as NVIDIA HGX/DGX and familiarity with orchestration tools like Kubernetes, Slurm, or Ray. The position offers a salary range of $5,000 to $7,000 and is based at the company's location in Singapore, with a standard Monday to Friday work schedule.

This summary is AI-generated and may contain inaccuracies. Please refer to the full job description below.

Job description

[This job id 27298 first appeared in Job-Q.com on 24 Aug 2026]

AI Infrastructure Engineer

5 days, Mon - Fri 8.30am to 5.30pm

Salary: $5,000 to $7,000

Location: 2 Kaki Bukit Ave 1, Singapore 417938

Job scopes:

Compute & Cluster Management

  • Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).
  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.

High-Performance Networking & Storage

  • Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
  • Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.

Automation & Infrastructure as Code (IaC)

  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
  • Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.

Operations, Observability & Performance

  • Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
  • Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
  • Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.

Requirements:

  • Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
  • Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
  • Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
  • High-Speed Networking: Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.
  • Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
  • Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
  • Bachelor’s Degree in Computer Science, Information Technology, Computer
    Engineering, or equivalent practical experience.
  • 3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
  • Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified
    Associate/Professional, AWS/Azure/GCP Solutions Architect).

If you are keen to apply, please send me your resume and job applied for on WhatsApp at 9789 3505 or email me at email address (˶ᵔ ᵕ ᵔ˶)

❄️Anabel Boon Xue Qi | Recruitment Consultant (R25159272) |📍The Supreme HR Advisory EA No: 14C7279

Scam prevention reminder: You should not make any pre-payment when applying for any job.

Illegal practices reminder: It is illegal for recruiter to collect payment (kickback) from the worker https://www.mom.gov.sg/-/media/mom/documents/publications/foreign-workers/what-are-kickbacks.pdf

Login is optional, you may send application via email

Login to Save Login to Apply

Get AI to assess your suitability to this job

Assess My Fit with AI Beta — Free during trial period

Login to upload your resume and get an instant match score, strengths, and gaps.


Or use your preferred AI chat tool manually:

Use AI chat of your choice: ChatGPT, Gemini, Claude — and:

  1. Paste this into the prompt:
    I am a jobseeker. Below is a job posting. Please: 1. Give a match score (0–100) based on my resume vs the job requirements 2. List my 3–5 key strengths that align with this role 3. List 2–3 areas to improve or gaps to address before applying 4. Give a one-sentence verdict: should I apply, apply with adjustments, or skip? Job posting URL: https://singapore.job-q.com/jobs/detail/ab03-ai-infrastructure-engineer-27298 After reading the job, ask me to upload or paste my resume.
  2. Upload your resume in the same chat.

Similar Jobs

AB03 - Head of Security Operations & Delivery

Head of Security Operations & Delivery5 days, Mon - Fri 8.30am to...

On site

Permanent

THE SUPREME HR ADVISORY PTE. LTD.

AB03 - AI Security & Automation Lead

AI Security & Automation Lead5 days, Mon - Fri 8.30am to 5.30pmSalary:...

On site

Permanent

THE SUPREME HR ADVISORY PTE. LTD.

AB03 - Quality Control Executive

Quality Control ExecutiveWorking days: 5-day work week Working hours: 9am to 6pmWork...

On site

Permanent

THE SUPREME HR ADVISORY PTE. LTD.

AB03 - Food Processing Worker (Kitchen)

Food Processing Worker (Kitchen)Working days: 6 days (Mon-Sat) [No need to work...

On site

Permanent

THE SUPREME HR ADVISORY PTE. LTD.

Job Summary

  • Published on: 24 Aug, 2026
  • Category: Others
  • Vacancy: 1
  • Job type: Permanent
  • Salary: 7000
  • Location: On site
  • Job Nature: Permanent

Company Details