NVIDIA CUDA Infrastructure Optimization
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
Our NVIDIA CUDA-based servers are online but still missing a production-ready software stack. I need the environment installed, tuned, and kept rock-steady so my research team can start running heavy models without delays.
Scope
• System setup & configuration – install the latest CUDA toolkit, drivers, cuDNN, NCCL and required OS dependencies, then provision Docker/Container runtime so future upgrades are painless.
• Performance optimization – profile current throughput, adjust BIOS, kernel, power and GPU settings, implement mixed-precision or other CUDA tweaks, and document the gains with repeatable benchmarks.
• Ongoing maintenance & troubleshooting – create health-checks, monitoring hooks (Prometheus/Grafana preferred), and a rapid-roll-back procedure so downtime is measured in minutes, not hours.
Acceptance criteria
1. “nvidia-smi” shows all GPUs at expected PCIe lanes, ECC status and correct clock levels.
2. Training workload provided (PyTorch script,
Scope
• System setup & configuration – install the latest CUDA toolkit, drivers, cuDNN, NCCL and required OS dependencies, then provision Docker/Container runtime so future upgrades are painless.
• Performance optimization – profile current throughput, adjust BIOS, kernel, power and GPU settings, implement mixed-precision or other CUDA tweaks, and document the gains with repeatable benchmarks.
• Ongoing maintenance & troubleshooting – create health-checks, monitoring hooks (Prometheus/Grafana preferred), and a rapid-roll-back procedure so downtime is measured in minutes, not hours.
Acceptance criteria
1. “nvidia-smi” shows all GPUs at expected PCIe lanes, ECC status and correct clock levels.
2. Training workload provided (PyTorch script,
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.