Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
TypeContract
LocationPoland
Posted3 hours ago
Job Overview
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms.
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
Strong understanding of server hardware, including:
BIOS/UEFI
RAID controllers
Firmware management
iLO/iDRAC/IPMI
NICs and SmartNICs
HBA cards
Hardware diagnostics and troubleshooting
Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
Understanding of NVIDIA GPU technologies including:
A100, H100, H200, B200, or equivalent GPU platforms
NVIDIA DGX and OEM GPU servers
GPU provisioning and lifecycle management
GPU monitoring and performance optimization
Knowledge of AI Factory architecture and infrastructure requirements
Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
Understanding of:
GPU resource allocation and scheduling
Multi-GPU systems
GPU networking requirements
High-bandwidth, low-latency infrastructure design
Familiarity with NVIDIA ecosystem technologies such as:
CUDA
NCCL
GPUDirect Storage
NVIDIA Fabric Manager
NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
Advanced Linux storage administration:
LVM
XFS, EXT4
NFS
iSCSI
Fibre Channel SAN
Multipath I/O
Strong hands-on experience with Ceph, including:
Cluster architecture
MON, OSD, MDS
RBD, CephFS, RGW
Capacity planning
Performance tuning
Failure recovery
Experience with high-performance AI storage platforms such as:
WEKA
VAST Data
Dell PowerScale
Pure Storage FlashBlade
NetApp
Understanding of:
NVMe-over-Fabrics (NVMe-oF)
RDMA
GPUDirect Storage
Parallel file systems
AI data pipelines
Networking & Infrastructure
Strong networking knowledge:
Bonding
VLANs
Routing
MTU optimization
DNS
DHCP
Experience with high-performance data center networking:
100G/200G/400G Ethernet
RoCE
RDMA
Spine-Leaf architectures
Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
Experience with high availability, clustering, and disaster recovery
Strong troubleshooting skills across:
Linux operating systems
Hardware platforms
GPU infrastructure
Networking
Enterprise storage
Experience supporting mission-critical production environments
Bash and Python scripting for automation and operational efficiency
Experience creating operational documentation, runbooks, and infrastructure standards
Nice to Have
Kubernetes infrastructure (especially AI/ML and GPU integration)
KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
Ansible automation
NVIDIA Base Command Manager
Slurm or HPC workload schedulers
Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
Data Center Infrastructure Management (DCIM) tools
IPAM solutions
AWS, Azure, or hybrid cloud exposure
We Are Not Looking For
Candidates whose experience is primarily CI/CD pipeline engineering
Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
Cloud-only administrators with limited bare metal, storage, or hardware experience
Professionals whose primary expertise is software development rather than infrastructure engineering
Ideal Candidate
Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.
Originally posted on Himalayas
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms.
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
Strong understanding of server hardware, including:
BIOS/UEFI
RAID controllers
Firmware management
iLO/iDRAC/IPMI
NICs and SmartNICs
HBA cards
Hardware diagnostics and troubleshooting
Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
Understanding of NVIDIA GPU technologies including:
A100, H100, H200, B200, or equivalent GPU platforms
NVIDIA DGX and OEM GPU servers
GPU provisioning and lifecycle management
GPU monitoring and performance optimization
Knowledge of AI Factory architecture and infrastructure requirements
Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
Understanding of:
GPU resource allocation and scheduling
Multi-GPU systems
GPU networking requirements
High-bandwidth, low-latency infrastructure design
Familiarity with NVIDIA ecosystem technologies such as:
CUDA
NCCL
GPUDirect Storage
NVIDIA Fabric Manager
NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
Advanced Linux storage administration:
LVM
XFS, EXT4
NFS
iSCSI
Fibre Channel SAN
Multipath I/O
Strong hands-on experience with Ceph, including:
Cluster architecture
MON, OSD, MDS
RBD, CephFS, RGW
Capacity planning
Performance tuning
Failure recovery
Experience with high-performance AI storage platforms such as:
WEKA
VAST Data
Dell PowerScale
Pure Storage FlashBlade
NetApp
Understanding of:
NVMe-over-Fabrics (NVMe-oF)
RDMA
GPUDirect Storage
Parallel file systems
AI data pipelines
Networking & Infrastructure
Strong networking knowledge:
Bonding
VLANs
Routing
MTU optimization
DNS
DHCP
Experience with high-performance data center networking:
100G/200G/400G Ethernet
RoCE
RDMA
Spine-Leaf architectures
Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
Experience with high availability, clustering, and disaster recovery
Strong troubleshooting skills across:
Linux operating systems
Hardware platforms
GPU infrastructure
Networking
Enterprise storage
Experience supporting mission-critical production environments
Bash and Python scripting for automation and operational efficiency
Experience creating operational documentation, runbooks, and infrastructure standards
Nice to Have
Kubernetes infrastructure (especially AI/ML and GPU integration)
KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
Ansible automation
NVIDIA Base Command Manager
Slurm or HPC workload schedulers
Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
Data Center Infrastructure Management (DCIM) tools
IPAM solutions
AWS, Azure, or hybrid cloud exposure
We Are Not Looking For
Candidates whose experience is primarily CI/CD pipeline engineering
Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
Cloud-only administrators with limited bare metal, storage, or hardware experience
Professionals whose primary expertise is software development rather than infrastructure engineering
Ideal Candidate
Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.
Originally posted on Himalayas
Apply on Himalayas →
Job sourced from Himalayas. Applications happen directly on the original platform — we never collect your data.