Jobgether

HPC Storage Engineer - West Coast

Jobgether

GoPythonRustCephMinIOLustreGPFS/Spectrum ScaleLinuxKubernetesPrometheus

About the Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a HPC Storage Engineer - West Coast based in the United States. This is a senior, hands-on infrastructure engineering role focused on building and scaling a multi-region storage platform for demanding AI workloads. You’ll own critical storage systems spanning network volumes, local NVMe, and S3-compatible object storage at petabyte scale. Your work will directly influence training, fine-tuning, and inference performance, including cold-start speed, data streaming, and workload reliability. You’ll operate at the intersection of storage, networking, hardware, and software, with substantial ownership from architecture through production operations. The role offers significant latitude to automate manual processes, establish SLOs, optimize performance, and shape long-term storage strategy. You’ll collaborate closely with SRE, network engineering, supply chain teams, and infrastructure partners in a fast-moving remote environment. This opportunity is ideal for an engineer who enjoys solving complex production problems and building infrastructure that serves millions of developers. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a HPC Storage Engineer - West Coast based in the United States. This is a senior, hands-on infrastructure engineering role focused on building and scaling a multi-region storage platform for demanding AI workloads. You’ll own critical storage systems spanning network volumes, local NVMe, and S3-compatible object storage at petabyte scale. Your work will directly influence training, fine-tuning, and inference performance, including cold-start speed, data streaming, and workload reliability. You’ll operate at the intersection of storage, networking, hardware, and software, with substantial ownership from architecture through production operations. The role offers significant latitude to automate manual processes, establish SLOs, optimize performance, and shape long-term storage strategy. You’ll collaborate closely with SRE, network engineering, supply chain teams, and infrastructure partners in a fast-moving remote environment. This opportunity is ideal for an engineer who enjoys solving complex production problems and building infrastructure that serves millions of developers. Accountabilities Own the capacity, durability, availability, and performance of network volumes, local NVMe, and S3-compatible object storage. Tune the complete I/O path, including device and filesystem configuration, caching, read-ahead strategies, replication, erasure coding, and client-side mount behavior. Diagnose complex storage and performance issues end to end, identifying root causes and implementing durable solutions. Lead capacity expansions, hardware refreshes, migrations, and data rebalancing while minimizing or eliminating customer-visible disruption. Design and optimize the networking infrastructure supporting storage workloads, including high-throughput east-west fabrics, MTU and jumbo-frame configuration, congestion and flow control, multipath, and NIC/offload settings. Optimize storage traffic across RDMA/RoCE and high-speed InfiniBand or Ethernet environments, collaborating with network engineering on topology, oversubscription, and cross-region data movement. Develop and ship production software in Go, Python, Rust, or similar languages for storage control-plane services, provisioning, data movement, and monitoring. Build and extend integrations with internal control-plane services, S3-compatible interfaces, CSI drivers, Kubernetes APIs, vendor platforms, and cloud-provider APIs. Replace manual operational procedures with reliable automation and infrastructure-as-code, while participating fully in code reviews, testing, and CI. Instrument storage infrastructure with meaningful metrics covering IOPS, throughput, latency, errors, retries, capacity utilization, and tenant consumption. Build dashboards, SLOs, and alerts that identify degradation proactively and support reliable production operations. Participate in an on-call rotation and lead blameless post-incident follow-through, ensuring lessons learned translate into measurable system improvements. Requirements 8+ years of experience in infrastructure, storage, or systems engineering, including substantial ownership of production storage environments at scale. Deep practical experience with at least one distributed storage platform such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or a comparable technology. Strong knowledge of Linux internals and the storage stack, including block devices, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI, and NVMe-oF. Hands-on experience building or operating S3-compatible object storage services. Strong networking fundamentals and demonstrated experience tuning networks specifically for storage workloads. Proven ability to write and ship production-quality software using Go, Python, Rust, or a similar programming language, beyond scripting alone. Experience with observability platforms such as Prometheus, Grafana, Datadog, or equivalent, including designing meaningful metrics and monitoring strategies. Demonstrated ability to analyze and resolve performance problems under real production pressure. Self-directed approach, with the ability to take broad infrastructure goals, develop an options analysis, diagnose problems, and execute solutions with minimal supervision. Strong continuous-improvement mindset, with a track record of eliminating operational toil and replacing recurring manual work with automation. High ownership and accountability, including the willingness to follow problems across team boundaries through to resolution. Collaborative, low-ego communication style combined with confidence in technical decision-making. Experience with AI/ML storage workloads, including checkpointing, dataset streaming, model-weight distribution, GPU-adjacent data locality, or GPUDirect Storage, is preferred. Familiarity with Kubernetes storage internals, including CSI drivers, PV/PVC lifecycles, StatefulSets, and local persistent volumes, is a plus. Bare-metal or colocation experience, including hardware selection, vendor management, firmware, and physical failure domains, is beneficial. Experience operating multi-tenant environments where isolation, fairness, and quality of service are critical is preferred. Background in a rapidly scaling cloud or infrastructure provider is advantageous. Benefits Base salary: $180,000–$260,000, with the final range determined based on career level, experience, qualifications, and location. Meaningful equity through stock options, giving employees an opportunity to share in the company’s growth. Generous medical, dental, and vision coverage. Flexible paid time off. Remote-first work environment with collaborative teams and Slack as a primary internal communication channel. $1,200 home office and equipment stipend to help create an effective remote workspace. Opportunity to work on cutting-edge AI infrastructure with a strong emphasis on ownership, learning, and technical impact. Inclusive workplace committed to equal opportunity and respect for people from diverse backgrounds. Candidates must be legally authorized to work in the United States; employment visa sponsorship is not currently available. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

You'll be redirected to Jobgether's application page

Job Details

Salary

$180K–$260K

Location

United States

Job type

Full-time

Category

Infrastructure Engineering

Experience

8+ years

Posted

Today

Job Highlights

  • $180K–$260K salary
  • 8+ years level role
  • 100% Remote — open to candidates in United States

About Jobgether

This job is hosted by Jobgether. Clicking Apply opens their site.

More jobs from Jobgether on RC9

Remote Work Style

Mixed

Mix of flexible and scheduled meetings

Your Match

See how well your skills line up with this role, and what you're missing.

AI Cover Letter

Generate a cover letter tailored to this job from your profile.