Jobgether

Staff ML Engineer

Jobgether

PythonLinuxDockerKubernetesCI/CDAWSTerraformPyTorchML InfrastructureDistributed Computing

About the Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff ML Engineer based in Brazil. This role focuses on accelerating Machine Learning research by building the engineering foundations, workflows, and shared tools researchers need to iterate at scale. You will work closely with ML researchers and engineers to turn evolving research needs into reliable, reusable engineering capabilities. The position spans ML infrastructure, platform engineering, MLOps, automation, experimentation, and distributed computing. You will support complex workloads involving large datasets, GPU clusters, distributed computation, and interconnected ML components. A key focus will be improving research reproducibility, reliability, scalability, development speed, and cost efficiency. You will also help bridge research and platform teams, translating ambiguous technical challenges into practical and reusable solutions. The role offers the opportunity to contribute to cutting-edge AI initiatives with a strong focus on accessibility, inclusion, collaboration, and global impact. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff ML Engineer based in Brazil. This role focuses on accelerating Machine Learning research by building the engineering foundations, workflows, and shared tools researchers need to iterate at scale. You will work closely with ML researchers and engineers to turn evolving research needs into reliable, reusable engineering capabilities. The position spans ML infrastructure, platform engineering, MLOps, automation, experimentation, and distributed computing. You will support complex workloads involving large datasets, GPU clusters, distributed computation, and interconnected ML components. A key focus will be improving research reproducibility, reliability, scalability, development speed, and cost efficiency. You will also help bridge research and platform teams, translating ambiguous technical challenges into practical and reusable solutions. The role offers the opportunity to contribute to cutting-edge AI initiatives with a strong focus on accessibility, inclusion, collaboration, and global impact. Accountabilities: Partner with ML researchers working on generative AI teams to identify bottlenecks and improve the speed, scalability, and reliability of research iteration. Design, build, and maintain reusable research infrastructure, workflows, models, interfaces, and automation supporting experimentation, training, evaluation, data processing, and model packaging. Enable reproducible experiments through consistent environments, dependency management, artifact and model versioning, configuration management, observability, and CI/CD practices. Support scalable ML workloads involving large datasets, GPU clusters, distributed computing, and multiple interconnected models, services, and algorithmic components. Deliver pragmatic research-enablement capabilities for immediate needs while keeping solutions aligned with the broader ML platform architecture and roadmap. Act as a technical bridge between researchers and ML platform teams by translating research pain points into clear platform requirements, validating new capabilities, and supporting adoption of shared infrastructure. Improve the path from research to production by making research outcomes easier to reproduce, integrate, test, and operationalize. Contribute to shared ML engineering standards and architecture while promoting strong engineering practices through hands-on collaboration, technical guidance, and knowledge sharing. Evaluate and introduce technologies that can materially improve research velocity, reliability, scalability, and cost efficiency. Requirements: Proven experience building reusable infrastructure, developer tools, or platforms that enable multiple engineers or researchers rather than supporting only a single predefined model or pipeline. Strong proficiency in Python and Linux, including the ability to build sustainable software, automation, and services. Hands-on experience with Docker, Kubernetes, CI/CD pipelines, cloud environments such as AWS, and Infrastructure as Code tools such as Terraform. Practical understanding of the end-to-end ML lifecycle, including data preparation, experimentation, training, evaluation, model and artifact management, packaging, deployment, and monitoring. Experience supporting compute-intensive or distributed workloads and diagnosing reliability, performance, resource, and cost bottlenecks. Practical knowledge of modern ML frameworks such as PyTorch, with sufficient understanding of model behavior and constraints to collaborate effectively with ML researchers; this is an engineering and infrastructure role rather than a research scientist position. Ability to work with ambiguous and evolving requirements, uncover underlying needs, and translate them into simple, reusable engineering capabilities. Strong communication and collaboration skills across research, engineering, and platform teams. Advanced English proficiency is required across conversational communication, writing, and reading. Strong interest in diversity, inclusion, and accessibility, combined with curiosity, collaboration, an action-oriented approach, and comfort working in ambiguous, early-stage environments. Impact-oriented mindset, with a focus on meaningful outcomes rather than output alone. Experience supporting ML research in areas such as language modeling, machine translation, computer vision, multimodal or generative AI, robotics, or autonomous systems is desirable. Experience scaling GPU clusters, distributed training, and large-scale data processing is a plus. Experience building researcher-focused ML platforms, self-service experimentation environments, or tools supporting multiple evolving research workflows is advantageous. Experience helping research teams migrate to or adopt shared ML platforms is desirable. Knowledge of inference optimization techniques, such as custom GPU kernels, is a plus but secondary to strong research-enablement and ML infrastructure experience. Benefits: CLT employment contract, providing statutory employment rights and protections. Caju Benefits Card with R$1,160 monthly allowance, with flexibility across meal, food, home office, culture, and mobility expenses. Fully remote work from anywhere in Brazil. SulAmérica medical insurance. SulAmérica dental insurance. SulAmérica life insurance. Conexa Saúde online consultations and telemedicine. Wellhub membership. Extended year-end holiday break. Birthday day off. Extended parental leave. Continuous professional development through LinkedIn Learning and an annual allowance for courses and training. University partnerships offering discounts and additional learning opportunities. Libras training. English Pass language-learning program. Work equipment provided as part of the onboarding kit. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

You'll be redirected to Jobgether's application page

Job Details

Salary

Not disclosed

Location

Brazil

Job type

Full-time

Category

MLOps

Experience

Senior

Posted

Today

Job Highlights

  • Senior level role
  • 100% Remote — open to candidates in Brazil
  • Full-time position

About Jobgether

This job is hosted by Jobgether. Clicking Apply opens their site.

More jobs from Jobgether on RC9

Remote Work Style

Mixed

Mix of flexible and scheduled meetings

Your Match

See how well your skills line up with this role, and what you're missing.

AI Cover Letter

Generate a cover letter tailored to this job from your profile.