Systems Development Engineer (SRE/DevOps)
Job Responsibilities:
• Develop and maintain infrastructure and configuration as code using CloudFormation, Terraform, Ansible, and related automation tools.
• Administer and optimize AWS environments, including core services, networking, security, and architecture for availability, performance, and cost.
• Manage and support Kubernetes clusters and containerized workloads, including configuration, scaling, and upgrades.
• Design, implement, and evolve end‑to‑end monitoring and observability frameworks using tools such as Open Telemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or similar platforms.
• Create and maintain dashboards, logs, traces, SLIs/SLOs, and automated alerting systems to ensure reliability and rapid detection of anomalies.
• Embed observability, CI/CD best practices, and operational readiness into all stages of the software development lifecycle in partnership with engineering teams.
• Lead or participate in incident response, troubleshooting, and root cause analysis for production incidents, using observability data to drive fast resolution.
• Automate operational tasks, runbooks, and incident remediation workflows to reduce toil and improve service reliability.
• Contribute to risk mitigation, backup, and disaster recovery strategies, including periodic testing and continuous improvement.
• Participate in shared after hours support and project work as needed.
Want more jobs like this?
Get Software Engineering jobs in Hyderabad, India delivered to your inbox every week.

Job Qualification:
• 2-4 years of experience designing, implementing, and maintaining CI/CD pipelines (e.g., Harness, GitHub Actions, ArgoCD or similar tools).
• Hands‑on experience with automation tools and Infrastructure as Code / Configuration as Code (CloudFormation, Terraform, Ansible).
• Strong understanding of Infrastructure as Code and Configuration as Code principles and patterns.
• Solid grasp of the software development lifecycle and modern SRE/DevOps practices.
• AWS administration and architecture experience, including networking, security, IAM, and core services.
• Experience operating Kubernetes clusters (EKS or other distributions) and containerized workloads.
• Deep experience with monitoring and observability tools such as OpenTelemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or equivalent, including metrics, logs, and traces.
• Ability to define and track SLIs/SLOs and use them to guide reliability improvements.
• Proficiency in Linux administration, including system configuration, troubleshooting, and performance tuning.
• Programming/scripting skills in at least one language such as Python, Go, or Rust for automation, tooling, and observability integrations.
• Solid understanding of networking, load balancing, and performance tuning.
• Experience troubleshooting complex distributed systems, supporting incident response, and driving root cause analysis.
• Familiarity with risk mitigation, backup, and disaster recovery concepts.
Preferred
• Experience building unified observability platforms or standardized dashboards for multiple services/teams.
• Experience with GitOps workflows and tools for declarative infrastructure and application delivery.
• Background in incident command and post‑mortem frameworks.
• Experience integrating observability and reliability practices into microservices and/or serverless architectures.
• Experience integrating testing, security and compliance checks into CI/CD pipelines.
Perks and Benefits
Health and Wellness
Parental Benefits
Work Flexibility
Office Life and Perks
Vacation and Time Off
Financial and Retirement
Professional Development
Diversity and Inclusion