Sr Machine Learning Engineer (MLOps) - Remote India
About Dynatron
Dynatron is transforming the automotive service industry with intelligent SaaS solutions that drive measurable results for thousands of dealership service departments. Our analytics, automation, and AI-powered workflows help service leaders improve profitability, increase operational efficiency, and make smarter business decisions.
As Dynatron expands its AI and machine learning capabilities, building models is only part of the challenge. We need the infrastructure, engineering discipline, and operational rigor to deploy those capabilities reliably, monitor them continuously, and scale them confidently.The Opportunity
We’re looking for a Senior Machine Learning Engineer (MLOps) to own the production infrastructure and operational lifecycle behind Dynatron’s growing AI and machine learning capabilities.
This is a senior, hands-on engineering role focused on turning models into reliable production services. You’ll build and own the pipelines, infrastructure, monitoring, governance, and deployment practices that allow our Data Science team to move from experimentation to production safely and repeatedly.
You’ll support both traditional machine learning and modern generative AI/LLM workloads, including retrieval-based and agentic applications. You’ll be responsible not simply for getting models into production, but for ensuring they remain performant, observable, secure, cost-effective, and maintainable once they get there.
You’ll work closely with our U.S.-based Data Science team. Because many of Dynatron’s data pipelines and model workloads run overnight in U.S. time, they align naturally with the India workday. You will have significant ownership of the production environment during this critical operating window.What You’ll Do
Build & Own the ML Production Lifecycle
Design, build, and maintain deployment pipelines that move models reliably from development through validation and into production.
Establish model versioning, lineage, registry, and automated promotion practices.
Define repeatable production-readiness standards and deployment patterns across ML and AI workloads.
Partner with Data Scientists to make model handoffs efficient, consistent, and production-ready.
Drive ML Reliability & Observability
Own production monitoring across model performance, drift, data quality, inference health, latency, and availability.
Establish alerts and operational thresholds that identify degradation before it materially impacts downstream products or customers.
Diagnose production failures, perform root-cause analysis, and implement durable corrective actions.
Build operational practices that improve reliability as Dynatron’s portfolio of production models grows.
Operationalize Generative AI & LLM Workloads
Deploy and support production LLM applications, including retrieval-based and agentic architectures.
Build evaluation frameworks that measure quality, reliability, and performance of generative AI capabilities.
Monitor token consumption, inference costs, and cost per interaction to ensure AI capabilities remain economically sustainable.
Implement appropriate controls around model access, usage, safety, and production behavior.
Build Training & Retraining Infrastructure
Design and operate infrastructure supporting model training, validation, and retraining.
Build automated retraining pipelines triggered by appropriate performance, data, or business conditions.
Ensure training environments and workflows are reproducible, scalable, and observable.
Partner with Data Engineering and Data Science to ensure reliable movement of data throughout the ML lifecycle.
Establish AI/ML Governance
Implement model access controls, auditability, lineage, and governance standards.
Support model risk classification and appropriate controls based on use case and business impact.
Produce documentation and technical evidence required to support security, compliance, and internal governance requirements.
Help establish responsible production practices as Dynatron expands its use of AI.
Own Production Operations
Take meaningful ownership of the operational health of Dynatron’s production ML and AI services.
Respond to incidents, troubleshoot failures, and coordinate resolution across teams when necessary.
Build runbooks and operational procedures that reduce dependence on tribal knowledge.
Identify recurring operational issues and automate them away wherever practical.
What You Bring
MLOps & Production ML Experience
6+ years of experience in software engineering, data engineering, machine learning engineering, or a related technical discipline.
3+ years of hands-on experience deploying and operating AI/ML systems in production.
Demonstrated experience supporting both traditional machine learning and LLM-based workloads in production.
Strong understanding of the complete model lifecycle from development and validation through deployment, monitoring, retraining, and retirement.
Generative AI & LLM Expertise
Production experience with LLM-powered applications and agentic frameworks.
Experience with retrieval architectures, evaluation methodologies, and production monitoring for generative AI.
Understanding of LLM performance, latency, token utilization, and cost-per-interaction management.
Ability to establish practical operational and governance controls around generative AI systems.
Cloud & MLOps Engineering
Deep experience with a major cloud platform and its managed AI/ML services; AWS strongly preferred.
Hands-on experience with model registries, pipeline orchestration, ML CI/CD, automated retraining, and production monitoring.
Strong Python engineering skills.
Experience with containerization and infrastructure-as-code.
Experience designing reliable, repeatable, and automated production environments.
Production Operations
Experience operating production services with meaningful ownership for reliability and availability.
Strong incident response, trou
Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →