Infrastructure/GPU Engineer

Remote Full-time
Cognizant is seeking a highly skilled hands-on Infrastructure Engineer with proven experience in the physical and technical deployment of AI-ready environments optimized for AI and machine learning workloads. This role focuses on NVIDIA DGX or similar systems, GPU-accelerated compute clusters, high-speed networking, and scalable storage solutions. The ideal candidate will have deep expertise in infrastructure design ,deployment, workload orchestration, and performance optimization in enterprise environments. This is a remote role in the US. Salary range for this role is between $99,000 and $116,000 depending on skills and qualifications of the candidate. Applications will be accepted till 10/21/2025. Key Responsibilities System Design & Deployment Help in rightsizing GPU investment Architect and deploy NVIDIA DGX systems and GPU-based compute clusters. Design and implement scalable parallel filesystems (e.g., Lustre, BeeGFS, GPFS). Integrate high-speed interconnects using InfiniBand, RoCE, and RDMA. Collaborate on rack planning and airflow optimization. Cluster & Infrastructure Management Configure and manage Slurm Workload Manager for job scheduling. Deploy and maintain cluster orchestration tools Automate provisioning using PXE boot, Terraform, Redfish, and Kubernetes. Perform firmware updates, BIOS/IPMI/BMC configuration, and OS provisioning Knowledge of Run.ai, ClearML or similar platform Networking & Performance Optimization Design and validate network topologies including IPMI, internal/external networks, and InfiniBand fabrics. Optimize RDMA and RoCE configurations for low-latency, high-throughput data transfers. Conduct performance benchmarking using GPU-Burn, NCCL, and NVSM. Monitoring & Troubleshooting Implement system health checks and diagnostics across compute, storage, and network layers. Troubleshoot hardware/software issues and ensure reliable infrastructure operation. Required Skills & Qualifications Technical Expertise Deep understanding of NVIDIA DGX architecture, CUDA, and GPU compute. Strong Linux system administration and shell scripting skills. Experience with Slurm, parallel filesystems, and high-speed networking (InfiniBand/RDMA/RoCE). Familiarity with containerization (Docker), orchestration (Kubernetes), and automation tools (Ansible, Redfish). Preferred Qualifications Experience with BBCM, and DGX BasePOD/SuperPOD configuration Certifications by Nvidia or equivalent OEM.
Apply Now

Similar Opportunities

Experienced Registered Behavior Technician for In-Home ABA Therapy - Atlanta, GA

Remote Full-time

Immediate Hiring: Experienced Registered Behavioral Technician (RBT) for Clinic-Based ABA Therapy Services

Remote Full-time

Experienced Registered Behavioral Technician (RBT) - ABA Therapy for Children with Autism Spectrum Disorder

Remote Full-time

Experienced Registered Nurse - Telehealth: Providing Remote Care Coordination and Patient Support

Remote Full-time

Experienced Substitute Teacher for Riverside County Schools - Join Scoot Education's Innovative Team

Remote Full-time

Experienced Substitute Teacher for San Bernardino County - Flexible Schedules & Competitive Pay

Remote Full-time

Experienced School Year Instructional Coach for High-Dosage Tutoring Programs in Edgewater Park, NJ

Remote Full-time

Experienced School Year Tutor for K-8 Students in Math and Literacy - Mickleton, NJ

Remote Full-time

Experienced Secondary Social Studies Teacher for Kansas - Flexible Hybrid Remote Arrangement

Remote Full-time

USPS Office Helper

Remote Full-time

Economic Researcher

Remote Full-time

Audio Recording Project – English (Romania)

Remote Full-time

**Experienced Live Chat Support Agent – Same Day Pay Remote Job Opportunity**

Remote Full-time

**Experienced Customer Support Specialist – Real Estate Industry**

Remote Full-time

GTM Strategic Planning, Senior Manager

Remote Full-time

Experienced Senior Data Product Manager I – Driving Business Value through Data-Driven Solutions in the Advertising Technology Space

Remote Full-time

Experienced Virtual Customer Care Representative - Work from Home Opportunity at blithequark

Remote Full-time

`Work From Home – Entry Level Benefits Representative | No Experience – USA Remote Jobs

Remote Full-time

Electronic Records Analyst for AM 100 Law Firm

Remote Full-time

Public Relations Manager – Adventures by Disney | National Geographic Expeditions

Remote Full-time
← Back to Home