Site Reliability Engineer

AgileEngine

**About the Role** We are looking for a **DevOps / Site Reliability Engineer** to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major\-incident calls, own remediation follow\-through, and build the playbooks that guide response. **What you will do** * Scale and maintain the ability to drive operational stability across multi\-cloud environments (Azure, AWS, GCP). * Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations. * Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment. * Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads. * Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time\-critical decisions, and coordinating cross\-functional responders under pressure. * Own the post\-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups. * Draft and send clear, accurate, audience\-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle. * Develop, maintain, and socialize divisional / group\-level incident\-management playbooks, runbooks, and escalation procedures that standardize response and reduce time\-to\-resolution. **Must haves** * **5\+ years of experience** . * In\-depth architectural expertise in **multi\-cloud defense** , federated IAM, and **zero\-trust principles** . * Strong practical experience with **Kubernetes** , **Terraform** , **CI/CD orchestration** , and **Python/Go** scripting. * Senior\-level, hands\-on **incident\-command experience** driving major/critical incident calls to resolution in a **24x7 production environment** . * Proven track record of **remediation follow\-up** — coordinating with teams and holding owners accountable until issues are fully closed. * Demonstrated skill drafting and issuing **incident notification communications** to both technical and executive audiences. * Direct experience authoring **divisional/group incident\-management playbooks** and escalation procedures. * Fully autonomous. * Drives the architecture of **complex automated runbooks** and mentors Middle\-level SREs. * Extensive experience deploying and tuning APIs from modern **CNAPP/CSPM platforms** , ideally **Wiz** . * Prior experience building platforms subject to strict financial compliance standards ( **PCI\-DSS** , **SOC2** ). * Upper\-intermediate English level. **Nice to haves** * **PagerDuty** — hands\-on experience with on\-call scheduling, alert routing, and incident orchestration. * **ServiceNow** — familiarity with incident, problem, and change management workflows and reporting. **Perks and Benefits** * **Professional growth** Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps * **Competitive compensation** We match your ever\-growing skills, talent, and contributions with competitive USD\-based compensation and budgets for education, fitness, and team activities * **A selection of exciting projects** Join projects with modern solutions development and top\-tier clients that include Fortune 500 enterprises and leading product brands * **Flextime** Tailor your schedule for an optimal work\-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.