Jobgether

Senior AI Systems Quality Engineer

Jobgether

PythonTypeScriptDatabricksMLflowAWSCI/CDautomated testingtesting methodologiesobservability toolsagentic workflows

About the Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI Systems Quality Engineer based in United States. This role focuses on engineering quality, reliability, and trust into production-grade AI systems used in mission-critical healthcare environments. You will design automated testing frameworks, evaluation pipelines, and scalable quality practices for agentic and LLM-driven applications. The position goes beyond traditional QA, emphasizing quality-by-design, safe failure behavior, measurable evaluation signals, and continuous validation throughout the development lifecycle. You will help build an AI testing platform integrated with Databricks and MLflow to support traceability, lineage, and auditability at scale. Working closely with AI engineering, platform, security, and delivery teams, you will influence release readiness and operational confidence. This is an automation-first engineering opportunity for someone passionate about making advanced AI systems reliable, governed, and production-ready. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI Systems Quality Engineer based in United States. This role focuses on engineering quality, reliability, and trust into production-grade AI systems used in mission-critical healthcare environments. You will design automated testing frameworks, evaluation pipelines, and scalable quality practices for agentic and LLM-driven applications. The position goes beyond traditional QA, emphasizing quality-by-design, safe failure behavior, measurable evaluation signals, and continuous validation throughout the development lifecycle. You will help build an AI testing platform integrated with Databricks and MLflow to support traceability, lineage, and auditability at scale. Working closely with AI engineering, platform, security, and delivery teams, you will influence release readiness and operational confidence. This is an automation-first engineering opportunity for someone passionate about making advanced AI systems reliable, governed, and production-ready. Accountabilities: Build and deploy production-grade automated validation frameworks, test harnesses, and evaluation pipelines across the full AI development lifecycle. Design and evolve an AI testing platform integrated with Databricks and MLflow to enable repeatable testing, traceability, lineage, and auditability. Create large-scale, scenario-based test suites covering hundreds or thousands of cases, including edge cases, long-tail scenarios, and system failure modes. Validate agentic orchestration behaviors such as tool usage, memory, decision logic, and non-deterministic outputs before production deployment. Embed quality-by-design principles by defining system contracts, guardrails, safe-degradation patterns, and validation requirements at key system boundaries. Define measurable quality signals for LLM systems, including grounding, hallucination rates, relevance, latency, cost, accuracy, and explainability. Integrate automated quality gates into CI/CD pipelines and ensure validation runs continuously following model, prompt, or code changes. Build reusable testing libraries, frameworks, and components that enable engineering teams to adopt consistent AI quality practices. Establish measurable release-readiness criteria and support go/no-go decisions based on defined quality thresholds. Partner with AI, platform, security, and delivery teams to translate business and mission requirements into clear quality criteria, trade-offs, and confidence levels. Evaluate system behavior, reliability, security, privacy, and operational risk in regulated and mission-critical environments. Requirements: 7+ years of software engineering experience, primarily focused on backend or platform systems. Proven experience designing and implementing automated AI testing and validation solutions in production environments. Demonstrated ability to build custom testing, validation, or evaluation frameworks for complex and distributed systems. Strong proficiency in Python and/or TypeScript within modern AI engineering environments. Hands-on experience with AI-powered systems, including LLM-based or agentic workflows and non-deterministic behavior. Experience designing AI testing at scale, including regression frameworks, long-tail evaluations, and broad test coverage. Deep understanding of CI/CD practices and experience embedding automated tests and quality gates into deployment pipelines. Solid knowledge of AWS cloud-native architectures. Strong track record of engineering for quality, reliability, governance, safety, and operational resilience as core system principles. Working knowledge of security, privacy, and operational risk within regulated or mission-critical environments, including failure modes and recovery strategies. Experience with AI testing methodologies such as non-deterministic output evaluation, drift detection, bias and fairness testing, and robust regression strategies. Ability to establish measurable trust thresholds and operationalize metrics such as query accuracy, hallucination limits, explainability, and PHI-safe behavior as release criteria. Experience collaborating with domain experts to define correctness and real-world validation scenarios that reflect genuine production use cases. Experience with Databricks and Medallion architecture is preferred but not required. Familiarity with MLflow for model evaluation, lineage, and auditability is a plus. Exposure to observability tools such as Datadog, Prometheus, or Grafana is desirable. Familiarity with LLM evaluation techniques, guardrails, and policy enforcement frameworks is beneficial. Experience evaluating AI performance, latency, and cost regressions is a plus. Ability to clearly communicate system behavior and quality trade-offs to both technical and business audiences. Formal AI/ML training or certifications, such as ISTQB AI Testing, AWS ML Specialty, or Google ML Engineer, are welcome. Experience designing prompts, agent behaviors, and orchestration logic as versioned, testable artifacts is advantageous. Familiarity with using AI systems to generate and expand diverse, adversarial, and large-scale test scenarios is a plus. Benefits: Compensation based on experience, skills, and location, including base salary, performance bonus eligibility, and equity grants. Unlimited paid time off. Work-from-anywhere flexibility. Comprehensive health coverage with multiple plan options. Equity for every employee. Growth-focused environment with opportunities for professional development. One-time home office setup allowance. Monthly cell phone allowance. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

You'll be redirected to Jobgether's application page

Job Details

Salary

Not disclosed

Location

United States

Job type

Full-time

Category

QA / Testing

Experience

7+ years

Posted

Today

Job Highlights

  • 7+ years level role
  • 100% Remote — open to candidates in United States
  • Full-time position

About Jobgether

This job is hosted by Jobgether. Clicking Apply opens their site.

More jobs from Jobgether on RC9

Remote Work Style

Mixed

Mix of flexible and scheduled meetings

Your Match

See how well your skills line up with this role, and what you're missing.

AI Cover Letter

Generate a cover letter tailored to this job from your profile.