DevOps / Site Reliability Engineer ID70127

AgileEngine

AgileEngine is an Inc. 5000 company that creates award\-winning software for Fortune 500 brands and trailblazing startups across 17\+ industries. We rank among the leaders in areas like application development and AI/ML, and our people\-first culture has earned us multiple Best Place to Work awards. **WHY JOIN US** If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you! **ABOUT THE ROLE** We are looking for a **DevOps / Site Reliability Engineer** to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major\-incident calls, own remediation follow\-through, and build the playbooks that guide response. **WHAT YOU WILL DO** \- Scale and maintain the ability to drive operational stability across multi\-cloud environments (Azure, AWS, GCP). \- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations. \- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment. \- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads. \- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time\-critical decisions, and coordinating cross\-functional responders under pressure. \- Own the post\-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups. \- Draft and send clear, accurate, audience\-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle. \- Develop, maintain, and socialize divisional / group\-level incident\-management playbooks, runbooks, and escalation procedures that standardize response and reduce time\-to\-resolution. **MUST HAVES** \- **5\+ years of experience** . \- In\-depth architectural expertise in **multi\-cloud defense, federated IAM, and zero\-trust principles** . \- Strong practical experience with **Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting** . \- Senior\-level, hands\-on **incident\-command experience** driving major/critical incident calls to resolution in a **24x7 production environment** . \- Proven track record of **remediation follow\-up** — coordinating with teams and holding owners accountable until issues are fully closed. \- Demonstrated skill drafting and issuing **incident notification communications** to both technical and executive audiences. \- Direct experience authoring **divisional/group incident\-management playbooks** and escalation procedures. \- Fully autonomous. \- Drives the architecture of **complex automated runbooks** and mentors Middle\-level SREs. \- Extensive experience deploying and tuning APIs from modern **CNAPP/CSPM platforms, ideally Wiz** . \- Prior experience building platforms subject to strict financial compliance standards ( **PCI\-DSS, SOC2** ). \- Upper\-intermediate English level. **NICE TO HAVES** \- PagerDuty — hands\-on experience with on\-call scheduling, alert routing, and incident orchestration. \- ServiceNow — familiarity with incident, problem, and change management workflows and reporting. **PERKS AND BENEFITS** \- **Growth without limits** : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget \- **Competitive compensation** : get recognition that reflects your skills and impact, with regular performance and compensation reviews \- **Flexibility** : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm \- **Meaningful, modern projects** : build impactful products using modern technologies alongside global teams and leading brands \- **Collaborative culture** : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized \- **Well\-being \& support** : access local well\-being programs and people\-focused support tailored to your location