Work mode: Hybrid/ Remote ( Quarterly 5 days in office for Remote Employees)
Number of Openings: 1
Job Description
We are looking for a Lead MLOps Engineer with strong experience in building, deploying, and managing scalable machine learning platforms and production-grade ML/AI solutions.
The role will be responsible for operationalizing the end-to-end ML lifecycle, automating ML workflows, establishing robust deployment and monitoring practices, and ensuring security, governance, and reliability across ML platforms.
The ideal candidate will have strong hands-on experience with Python, MLflow or Kubeflow, AWS SageMaker, AWS Bedrock, ECS, Docker, CI/CD, and Infrastructure as Code (IaC).
Key Responsibilities
- Design, build, deploy, and manage end-to-end ML lifecycle pipelines.
- Automate model training, testing, validation, deployment, and monitoring workflows.
- Implement experiment tracking, model versioning, model registry, and lifecycle management solutions.
- Develop and maintain scalable MLOps platforms and ML deployment infrastructure.
- Monitor model performance, data/model drift, system health, and operational metrics in production.
- Implement observability and alerting mechanisms for ML workloads.
- Collaborate closely with Data Engineering, Data Science, DevOps, and application teams to operationalize ML/AI solutions.
- Establish and maintain security, governance, access control, auditability, and compliance processes for ML platforms.
- Implement CI/CD and DevSecOps practices for ML pipelines and infrastructure.
- Optimize ML infrastructure and deployment workflows for performance, scalability, reliability, and cost efficiency.
- Troubleshoot production ML pipelines, deployments, infrastructure, and model-serving issues.
- Contribute to the adoption of modern MLOps, LLMOps, and Generative AI practices.
Required SkillsProgramming & ML Lifecycle
- Strong proficiency in Python.
- 4+ years of hands-on experience in MLOps / ML lifecycle management.
- Experience taking ML/AI solutions from Proof of Concept (PoC) to Production.
- Strong understanding of model training, deployment, versioning, monitoring, and lifecycle management.
MLOps & ML Platforms Hands-on experience with:
- MLflow or Kubeflow
- Amazon SageMaker
- AWS Bedrock
- Amazon ECS (Elastic Container Service)
Cloud & Infrastructure
- Strong hands-on experience with AWS.
- Good knowledge of Docker and Kubernetes.
- Experience with Infrastructure as Code (IaC) tools such as:
- Terraform
- AWS CloudFormation
- or equivalent tools.
DevOps & Security
- Experience with CI/CD pipelines and automation.
- Understanding of DevSecOps practices.
- Knowledge of cloud security, access control, secrets management, and auditability.
- Experience implementing production-grade deployment and operational practices.
Monitoring & Observability
- Understanding of model monitoring and observability.
- Experience monitoring:
- Model performance
- Data/model drift
- Pipeline health
- Infrastructure health
- Production failures
- Knowledge of performance optimization and reliability engineering for ML systems.
Preferred Skills
- Hands-on experience with AWS ML/AI services and cloud-native architectures.
- Knowledge or hands-on experience with Amazon Bedrock AgentCore (AgentCore).
- Exposure to data engineering technologies such as:
- Databricks
- Apache Spark
- Apache Airflow
- Apache Kafka
- Snowflake
- Understanding of Responsible AI, model governance, risk management, and compliance requirements.
- Exposure to Generative AI, LLMOps, and RAG-based solutions.
- Experience with production deployment and monitoring of LLM/GenAI applications.
Key Competencies
- MLOps & ML Lifecycle Management
- Python
- MLflow or Kubeflow
- AWS SageMaker
- AWS Bedrock
- ECS
- Docker
- CI/CD & DevSecOps
- Infrastructure as Code
- ML Monitoring & Observability
- Model Governance & Security
- Generative AI / LLMOps
- Production ML Systems