Senior DevOps Engineer
Senior DevOps / Cloud Platform Engineer - ML & AI Infrastructure
Job Summary
We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services.
The ideal candidate will have hands-on experience with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, GitHub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments. This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost-efficient platforms for ML services across development and production environments.
In this role you will...
Key Responsibilities
AWS & Kubernetes Infrastructure
- Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.
- Manage Amazon EKS clusters, including cluster provisioning, upgrades, scaling, networking, and troubleshooting.
- Work with AWS SageMaker, AWS Bedrock, EKS, ECS, and related AWS services.
- Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB.
- Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security.
- Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.
- Implement best practices for security, reliability, availability, and scalability.
CI/CD & Azure-to-AWS Migration
- Build and maintain CI/CD pipelines for ML and AI services.
- Develop and manage GitHub Actions and Azure DevOps pipelines using YAML.
- Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS.
- Automate build, test, containerization, deployment, and release processes.
- Establish deployment strategies across development, staging, and production environments.
ML Service Deployment
Want more jobs like this?
Get jobs in Hyderabad, India delivered to your inbox every week.

- Deploy and manage ML services across AWS EKS/ECS and SageMaker.
- Build and maintain Docker containers and Kubernetes deployments.
- Manage environment segregation and configuration across Dev, QA, and Production.
- Develop and maintain Kubernetes manifests and Helm charts.
- Troubleshoot ML service deployment, networking, scaling, and runtime issues.
App Runner to EKS Migration
- Lead migration of existing services from AWS App Runner to Amazon EKS.
- Containerize applications and develop Kubernetes manifests/Helm charts.
- Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.
- Ensure minimal service disruption during migration and establish operational best practices on EKS.
Self-Hosted LLM & AI Infrastructure
- Deploy and manage self-hosted Large Language Models and inference services.
- Work with model serving frameworks such as vLLM.
- Design containerized infrastructure for GPU-based model serving and inference.
- Manage model versions, deployments, configurations, and rollback strategies.
- Support migration of ML services from managed APIs/services to self-hosted models.
- Work with engineering teams on API integration and inference infrastructure.
Databricks Administration
- Administer Databricks workspaces, clusters, permissions, and access controls.
- Manage cluster configuration, policies, and resource utilization.
- Support LMI Insights and related ML/AI workloads.
- Troubleshoot Databricks infrastructure and connectivity issues.
- Implement appropriate security and access-control practices.
Elasticsearch Infrastructure
- Design, deploy, and manage Elasticsearch clusters.
- Perform cluster sizing, scaling, configuration, and performance optimization.
- Manage indices, mappings, retention, and data lifecycle requirements.
- Support Kibana configuration, dashboards, and troubleshooting.
- Monitor Elasticsearch health, capacity, and performance.
Monitoring, Reliability & Auto-Scaling
- Implement monitoring and observability for Kubernetes, AWS, and ML services.
- Use Prometheus, Grafana, and AWS CloudWatch for monitoring and alerting.
- Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.
- Establish proactive alerting for infrastructure and application health.
- Perform capacity planning and resource optimization.
- Identify opportunities for AWS infrastructure and compute cost optimization.
Infrastructure as Code & Automation
- Build and maintain infrastructure using Terraform.
- Develop reusable Terraform modules for AWS and Kubernetes infrastructure.
- Manage Kubernetes deployments using Helm charts.
- Automate infrastructure provisioning, configuration, deployments, and operational tasks.
- Maintain infrastructure documentation and deployment standards.
Cross-Team Collaboration
- Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams.
- Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.
- Coordinate infrastructure requirements for new ML models and services.
- Participate in production incident resolution, root-cause analysis, and continuous improvement.
- Establish engineering standards around deployment, monitoring, security, and operational ownership.
You've got what it takes if you have...
Required Skills & Experience
- 5+ years of experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering.
- Strong hands-on experience with AWS.
- Strong experience administering Amazon EKS and Kubernetes in production.
- Hands-on experience with:
- AWS EKS
- AWS SageMaker
- AWS Bedrock
- AWS ECS
- AWS App Runner
- Kubernetes
- Docker
- NGINX / AWS ALB Ingress
- Cloudflare / Cloudflare Tunnels
- Strong experience with Terraform and Helm.
- Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.
- Strong YAML scripting and Git experience.
- Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub.
- Experience deploying and operating ML/AI services.
- Experience with self-hosted LLM/model serving, preferably vLLM.
- Experience with GPU-based workloads is highly desirable.
- Experience with Databricks administration.
- Experience managing Elasticsearch and Kibana.
- Experience with Prometheus, Grafana, and CloudWatch.
- Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery.
- Strong understanding of cloud networking fundamentals.
- Experience with production troubleshooting, monitoring, capacity planning, and cost optimization.
- Strong understanding of security, IAM, secrets management, and access control.
Preferred / Nice-to-Have Skills- Experience supporting Generative AI / LLM platforms.
- Experience with GPU infrastructure and NVIDIA/CUDA environments.
- Experience with model lifecycle and model version management.
- Experience migrating workloads between managed cloud services and Kubernetes.
- Experience with AWS networking such as VPC, load balancers, security groups, and Route 53.
- Experience with API gateways and microservice architectures.
- Experience with Python or shell scripting for infrastructure automation.
- Experience working with Data Engineering and ML teams in a production environment.
What You'll Own- AWS ML/AI infrastructure
- EKS cluster administration and upgrades
- ML service deployment and production operations
- CI/CD automation
- App Runner → EKS migration
- Self-hosted LLM infrastructure and vLLM
- Databricks platform administration
- Elasticsearch infrastructure
- Monitoring and auto-scaling
- Terraform and Helm-based infrastructure automation
- Cloud cost, reliability, and performance optimization
Ideal Candidate
The ideal candidate is a hands-on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD → AWS infrastructure → EKS deployment → monitoring → scaling → production support.
They should be comfortable working across both traditional DevOps infrastructure and modern AI/ML infrastructure, and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform.
#LI-Onsite
Perks and Benefits
Health and Wellness
- Health Insurance
- Health Reimbursement Account
- Dental Insurance
- Vision Insurance
- Life Insurance
- Short-Term Disability
- Long-Term Disability
- FSA
- HSA
- HSA With Employer Contribution
- Pet Insurance
- Mental Health Benefits
Parental Benefits
- Birth Parent or Maternity Leave
- Non-Birth Parent or Paternity Leave
- Fertility Benefits
- Family Support Resources
- Adoption Leave
Work Flexibility
- Flexible Work Hours
- Remote Work Opportunities
- Hybrid Work Opportunities
Office Life and Perks
- Casual Dress
- Snacks
- Company Outings
- On-Site Cafeteria
- Holiday Events
Vacation and Time Off
- Paid Vacation
- Unlimited Paid Time Off
- Paid Holidays
- Personal/Sick Days
- Leave of Absence
- Summer Fridays
Financial and Retirement
- 401(K) With Company Matching
- Stock Purchase Program
- Performance Bonus
- Relocation Assistance
- Financial Counseling
- Profit Sharing
Professional Development
- Tuition Reimbursement
- Promote From Within
- Work Visa Sponsorship
- Leadership Training Program
- Internship Program
- Shadowing Opportunities
- Access to Online Courses
Diversity and Inclusion
- Employee Resource Groups (ERG)
- Unconscious Bias Training
- Diversity, Equity, and Inclusion Program