Jobgether

Senior Staff Software Engineer, Serving

Jobgether

GoTensorRTmachine learningdistributed systemsGPU computingReal-time data processingPerformance Optimizationsystem architecturedata structuresalgorithms

About the Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Staff Software Engineer, Serving based in United States. This role offers the opportunity to shape mission-critical infrastructure powering real-time advertising at massive global scale. As part of the Serving team, you will design and operate high-throughput, low-latency systems that process millions of requests per second under demanding performance constraints. You will work at the intersection of distributed systems, machine learning infrastructure, GPU computing, and real-time data processing. Your work will directly influence the scalability, reliability, and efficiency of ad-serving platforms used across a global mobile ecosystem. You will collaborate closely with machine learning engineers to bring sophisticated models into production and continuously optimize their serving performance. This is a highly technical role for an experienced engineer who thrives on complex systems, performance optimization, and large-scale infrastructure challenges. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Staff Software Engineer, Serving based in United States. This role offers the opportunity to shape mission-critical infrastructure powering real-time advertising at massive global scale. As part of the Serving team, you will design and operate high-throughput, low-latency systems that process millions of requests per second under demanding performance constraints. You will work at the intersection of distributed systems, machine learning infrastructure, GPU computing, and real-time data processing. Your work will directly influence the scalability, reliability, and efficiency of ad-serving platforms used across a global mobile ecosystem. You will collaborate closely with machine learning engineers to bring sophisticated models into production and continuously optimize their serving performance. This is a highly technical role for an experienced engineer who thrives on complex systems, performance optimization, and large-scale infrastructure challenges. Accountabilities: Design, build, and operate high-throughput, low-latency services that process bid requests, execute real-time bidding logic, and respond to ad exchanges. Improve the performance, scalability, and reliability of systems handling millions of requests per second while operating within strict latency requirements. Develop and optimize GPU-powered inference services using TensorRT to efficiently execute neural network models in production. Profile complete inference pipelines and identify opportunities to optimize GPU utilization, batching, memory access, serialization, networking, and request handling. Partner with machine learning engineers to productionize new model architectures and features, translating modeling requirements into efficient, reliable serving implementations. Design and operate large-scale feature store fleets providing low-latency access to real-time and precomputed features. Build benchmarking and performance-analysis tools that enable engineers to compare models, evaluate latency and throughput trade-offs, and detect performance regressions before deployment. Improve experimentation infrastructure that supports offline training runs, candidate model evaluation, and model leaderboards. Own engineering changes throughout their complete lifecycle, including system design, implementation, testing, deployment, observability, capacity planning, incident response, and ongoing optimization. Requirements: 10+ years of professional software engineering experience, with a strong track record working on complex, large-scale systems. Strong Computer Science fundamentals, including data structures, algorithms, distributed systems, and system architecture. M.S. or higher in Computer Science or a related field, or equivalent professional experience. Strong ability to design and operate highly scalable, reliable, and performance-sensitive production systems. Experience working with machine learning infrastructure, model serving, inference systems, or other high-performance computing environments is highly valuable. Strong analytical and performance-engineering skills, with the ability to identify bottlenecks and optimize systems across multiple layers of the stack. Ability to collaborate effectively with machine learning engineers and other technical teams to translate research and modeling requirements into production-ready systems. Experience with Go is a plus. Comfortable taking ownership of technically complex projects from architecture and implementation through production operations and continuous improvement. Benefits: Full-time remote work available across approved US states, including CA, CO, ID, IL, FL, GA, MA, MI, MN, MO, NJ, NV, NY, OR, PA, TX, UT, and WA. Competitive base salary based on location and experience. Base salary of $255,000–$310,000 for the San Francisco Bay Area, New York City, and Los Angeles/Orange County. Base salary of $234,000–$285,000 for Seattle/Olympia, Austin, San Diego, Santa Barbara, and Boston. Base salary of $219,000–$266,000 for other cities and towns within approved states. Equity as part of the overall compensation package. Health, vision, and dental benefits based on country of residence. Potential additional benefits such as wellness stipends and other location-specific perks. Remote-first work environment with opportunities for in-person project meetings, regional meetups, and company-wide gatherings. Expected participation in in-person team gatherings at least once per quarter to support collaboration, communication, and team building. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

You'll be redirected to Jobgether's application page

Job Details

Salary

$219K–$310K

Location

United States

Job type

Full-time

Category

Backend Development

Experience

10+ years

Posted

Today

Job Highlights

  • $219K–$310K salary
  • 10+ years level role
  • 100% Remote — open to candidates in United States

About Jobgether

This job is hosted by Jobgether. Clicking Apply opens their site.

More jobs from Jobgether on RC9

Remote Work Style

Mixed

Mix of flexible and scheduled meetings

Your Match

See how well your skills line up with this role, and what you're missing.

AI Cover Letter

Generate a cover letter tailored to this job from your profile.