Senior ML Specialist
machine learningNLPGenerative AIPythonPyTorchScikit-learnHugging FaceFastAPIstatistical analysisModel Evaluation
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior ML Specialist based in Brazil. This is a machine learning-focused role at the heart of a production agentic AI platform, shaping how models are selected, evaluated, optimized, and deployed. You will bring deep expertise in machine learning, statistics, NLP, and generative AI to make model decisions measurable and defensible. The role combines research-driven thinking with hands-on engineering, requiring you to design, ship, operate, and continuously improve production AI systems. You will develop evaluation methodologies, fine-tuned models, classifiers, quality gates, and agent capabilities that solve real business problems. Working across data engineering, software engineering, and product teams, you will influence the technical direction of AI products and raise the team’s machine learning standards. You will also investigate model behavior from first principles, balancing accuracy, reliability, latency, and cost across different approaches and providers. This opportunity is ideal for an autonomous ML professional who enjoys turning rigorous experimentation and research into measurable production outcomes. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior ML Specialist based in Brazil. This is a machine learning-focused role at the heart of a production agentic AI platform, shaping how models are selected, evaluated, optimized, and deployed. You will bring deep expertise in machine learning, statistics, NLP, and generative AI to make model decisions measurable and defensible. The role combines research-driven thinking with hands-on engineering, requiring you to design, ship, operate, and continuously improve production AI systems. You will develop evaluation methodologies, fine-tuned models, classifiers, quality gates, and agent capabilities that solve real business problems. Working across data engineering, software engineering, and product teams, you will influence the technical direction of AI products and raise the team’s machine learning standards. You will also investigate model behavior from first principles, balancing accuracy, reliability, latency, and cost across different approaches and providers. This opportunity is ideal for an autonomous ML professional who enjoys turning rigorous experimentation and research into measurable production outcomes. Accountabilities: Model Strategy & Selection: Own model strategy across the platform by evaluating and selecting models from major and emerging providers for specific tasks, using rigorous benchmarks, LLM-as-judge frameworks, and appropriate statistical analysis to make decisions defensible. AI Evaluation: Design and maintain the evaluation methodology used across the team, including offline evaluation sets, judge calibration, reliability testing, regression benchmarks, quality metrics, and statistical significance testing. Machine Learning Development: Build classification and fine-tuned models for routing, categorization, detection, and other use cases where trained models can outperform prompting in accuracy, cost, or latency. Agentic AI Development: Design, develop, and ship agentic AI products, including agent identities, reusable skills, tool integrations, structured outputs, and platform-level enhancements when required. Prompt & Context Engineering: Develop sophisticated prompt and context strategies covering system prompt assembly, live-data context injection, thread and memory management, and structured outputs using Pydantic schemas. Output Quality & Guardrails: Own AI output quality through deterministic enforcement mechanisms, validation gates, LLM-as-judge evaluators, and ongoing measurement of evaluator reliability. Model Diagnostics: Investigate hallucinations, inconsistencies, model drift, prompt sensitivity, and other behaviors from first principles, translating findings into improvements across prompts, retrieval, guardrails, and model selection. Agent Tooling & Integration: Build and extend tools that connect agents with Databricks, MongoDB, external systems, and broader agent ecosystems through the Model Context Protocol and similar integration standards. Production Operations: Operate and continuously improve the systems you ship by instrumenting services with OpenTelemetry and monitoring quality, latency, cost, and reliability through observability tooling. Collaboration & Technical Leadership: Partner with data engineers, software engineers, and product teams to deliver end-to-end AI solutions while strengthening the team’s ML capabilities through technical reviews, documentation, and knowledge sharing. Research & Engineering Excellence: Critically evaluate new machine learning research, apply techniques that demonstrate value on real-world data, and maintain high standards for code quality, testing, documentation, and engineering best practices. Requirements Machine Learning Experience: 5+ years of applied machine learning experience, including substantial experience before the current generative AI wave, with hands-on responsibility for training, evaluating, and deploying models rather than primarily integrating LLM APIs. Statistics & Probability: Strong foundation in hypothesis testing, confidence intervals, sampling and sample-size reasoning, bias and variance, calibration, and determining whether observed model differences are statistically meaningful. LLM & Generative AI Expertise: Deep understanding of transformer architectures, attention, tokenization, embeddings, pretraining, supervised fine-tuning, RLHF/DPO, decoding strategies, scaling behavior, and common failure modes such as hallucination, sycophancy, long-context degradation, and prompt sensitivity. Classical ML & NLP: Hands-on experience with classification, clustering, feature engineering, text classification, embeddings, semantic similarity, and the ability to determine when a smaller trained model is preferable to an LLM. Model Fine-Tuning: Experience with LoRA/PEFT, supervised fine-tuning, or full fine-tuning, including dataset construction and the judgment to determine when prompting or retrieval is a more effective solution. AI Evaluation: Experience designing evaluation methodologies for AI systems, including offline evaluation sets, LLM-as-judge approaches and their biases, inter-rater agreement, regression benchmarks, and statistical significance testing. Production Python: Strong production-grade Python skills with PyTorch, scikit-learn, Hugging Face, and modern asynchronous Python, plus experience shipping code into FastAPI/Pydantic environments with tests and CI. LLM Provider Experience: Hands-on experience with at least one major LLM provider API, such as OpenAI, Anthropic, or Google, including structured outputs and tool/function calling, as well as an understanding of model and provider trade-offs. Education: Master’s or PhD in Machine Learning, Statistics, Computer Science, Mathematics, or another quantitative discipline, or equivalent demonstrated expertise through publications, competition results, released models, or research code. Autonomy & Ownership: Demonstrated senior-level ownership, with the ability to take a problem from initial framing through implementation, deployment, measurement, and measurable business or technical outcomes. Preferred Expertise: Publications, research, or open-source contributions in machine learning or NLP; strong Kaggle or equivalent competition results; retrieval systems; agent evaluation; model cost and latency optimization; Databricks and Azure; Docker; observability platforms such as OpenTelemetry, Grafana, and Loki; MCP; and experience taking products from 0 to 1 in rapidly changing environments. Benefits Contract Type: PJ contractor engagement based in Brazil. Work Arrangement: Remote position, allowing you to work from Brazil. Technical Environment: Opportunity to work on a production agentic AI platform involving modern LLMs, machine learning, retrieval, model evaluation, automation, and AI agents. Innovation: Exposure to technologies and approaches including OpenAI, Anthropic, Google models, Databricks, Azure, MLflow, MCP, LLM-as-judge evaluation, fine-tuning, and advanced AI observability. Ownership & Impact: High degree of autonomy to shape model strategy, evaluation methodology, AI product capabilities, and production engineering practices. Cross-Functional Collaboration: Opportunity to work closely with data engineering, software engineering, and product professionals on end-to-end AI solutions. Professional Growth: Opportunity to deepen expertise across machine learning research, generative AI, agentic systems, production engineering, and applied AI innovation. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
Not disclosed
Location
Brazil
Job type
Full-time
Category
Machine Learning / AI
Experience
5+ years
Posted
Today
Job Highlights
- 5+ years level role
- 100% Remote — open to candidates in Brazil
- Full-time position
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.