MLOps Engineer
PythonAWSCI/CDInfrastructure as Codecontainerizationmonitoringautoscalingself-service toolingobservabilityLLMs
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a MLOps Engineer based in India. This role offers the opportunity to own the infrastructure powering production AI systems in a rapidly scaling mortgage lending environment. You will transform early-stage AI deployments into reliable, scalable, and predictable production platforms capable of handling real-world workloads. The position spans AI serving infrastructure, cloud deployment, CI/CD, infrastructure as code, internal developer tooling, and platform reliability. You will design systems that enable engineers to ship AI changes quickly and safely without relying on infrastructure support for every release. Observability, capacity planning, cost optimization, security, access control, and environment isolation will be central to your work. You will operate at the intersection of software engineering, infrastructure, and AI, working with non-deterministic workloads and real production constraints. This is a high-ownership opportunity for an engineer who enjoys building platforms from the ground up and creating leverage for the wider engineering organization. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a MLOps Engineer based in India. This role offers the opportunity to own the infrastructure powering production AI systems in a rapidly scaling mortgage lending environment. You will transform early-stage AI deployments into reliable, scalable, and predictable production platforms capable of handling real-world workloads. The position spans AI serving infrastructure, cloud deployment, CI/CD, infrastructure as code, internal developer tooling, and platform reliability. You will design systems that enable engineers to ship AI changes quickly and safely without relying on infrastructure support for every release. Observability, capacity planning, cost optimization, security, access control, and environment isolation will be central to your work. You will operate at the intersection of software engineering, infrastructure, and AI, working with non-deterministic workloads and real production constraints. This is a high-ownership opportunity for an engineer who enjoys building platforms from the ground up and creating leverage for the wider engineering organization. Accountabilities: Design and operate scalable infrastructure for production AI services, moving workloads from single-host deployments to horizontally scalable, orchestrated environments. Architect infrastructure specifically for LLM-backed workloads, including long-running requests, streaming responses, bursty concurrency, expensive downstream calls, and upstream rate limits. Own capacity planning, autoscaling, resource utilization, and infrastructure unit economics to ensure systems scale efficiently. Build and maintain consistent promotion paths across development, pre-production, and production environments. Implement infrastructure as code and ensure infrastructure, configuration, and application logic remain version-controlled and reviewable. Build CI/CD pipelines supporting automated testing, environment promotion, safe deployments, staged rollouts, rollback, and auditable change history. Develop internal self-service tooling that enables engineers and technically minded teams to define, modify, and test AI workflow logic without depending on infrastructure specialists. Establish appropriate guardrails for self-service tooling, including validation, versioning, review, staged promotion, and rollback capabilities. Instrument AI infrastructure end to end, monitoring latency, throughput, failures, costs, and output-quality signals. Develop meaningful alerting and incident-management practices that identify real system degradation and drive lasting improvements. Own secrets management, access controls, and environment isolation in a regulated environment with sensitive data and financial implications. Diagnose production issues across application, networking, and infrastructure layers and implement durable solutions rather than temporary workarounds. Continuously improve platform reliability, deployment speed, operational efficiency, and developer experience. Requirements: Demonstrated experience deploying and operating containerized services in production under real-world traffic. Hands-on experience with production rollouts, autoscaling, failure isolation, resource limits, monitoring, and rollback. Strong infrastructure-as-code experience, with the ability to build declarative, reproducible environments and modernize manually configured infrastructure. Proven experience building and maintaining CI/CD pipelines used by engineering teams for automated testing, environment promotion, safe releases, and rapid recovery. Strong Python skills, including the ability to read, modify, profile, troubleshoot, and improve application code. Strong operational judgment and ability to debug complex production issues across application, network, and infrastructure boundaries. Demonstrated end-to-end ownership across system design, implementation, deployment, monitoring, troubleshooting, and continuous improvement. Ability to work effectively in fast-moving environments, ship incrementally, and continuously improve systems and processes. Experience operating LLM or ML workloads in production is strongly preferred, particularly where cost, latency, and non-deterministic outputs are important considerations. Cloud deployment experience, preferably with AWS, including the ability to containerize, deploy, scale, and operate production systems. Experience in FinTech, lending, insurance, or another regulated industry is strongly preferred, particularly involving secrets management, access control, audit trails, and PII handling. Experience building internal developer platforms or self-service tooling that has been adopted by non-infrastructure engineering teams is an advantage. Full-stack development capability and the ability to build lightweight internal interfaces or tools when required. Experience designing multi-environment promotion pipelines where configuration and application logic move together with code is a plus. Familiarity with agent orchestration frameworks and the operational requirements of running them reliably. Experience with observability for non-deterministic systems, including tracing, evaluation signals, and quality monitoring alongside standard infrastructure telemetry. Knowledge of inference optimization techniques such as model serving, batching, caching, and cost reduction is an advantage. Exposure to workflow automation platforms and low-code builders is a plus. Benefits: Competitive compensation. Fully remote, remote-first working environment. Flexible working hours. Unlimited PTO. $2,000 per year professional development budget. Home office setup stipend. Opportunity to own infrastructure powering production AI systems. High-impact role at the intersection of AI, software engineering, and infrastructure. Opportunity to build an AI platform and internal tooling from an early stage rather than maintaining a mature inherited platform. Exposure to complex AI workloads, regulated financial environments, and real-world production systems. Direct impact on the reliability, deployment speed, and scalability of AI-powered products. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
Not disclosed
Location
India
Job type
Full-time
Category
MLOps
Experience
Mid
Posted
Today
Job Highlights
- Mid level role
- 100% Remote — open to candidates in India
- Full-time position
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.