DevOps SRE
Anitech Solutions · São Paulo, Brazil
**Role** You will be working as part of a dedicated DevOps Team, Voltron. Voltron is responsible for the reliability and operation of our client's on\-prem services. As an SRE on Voltron, you keep those systems available, observable, and recoverable for the students and staff who depend on them every day. The infrastructure is self\-hosted, so reliability here is hands\-on: you own the systems end to end rather than handing failure modes off to a cloud provider. This is an operations and reliability role. You will run production, reduce the manual work required to keep it healthy, and make failures rarer, shorter, and easier to recover from. Our client is a public cyber charter school in the USA, using infrastructure they own and operate. We build and maintain our own software, and the work is bound by one constraint above all: student learning cannot be disrupted. **Responsibilities** ● Keep client's on\-prem, self\-hosted systems available and healthy for students and staff. ● Define, measure, and defend SLOs, and maintain observability across services — metrics, logs, and traces. ● Run incident response: detection, on\-call, mitigation, and postmortems that lead to durable fixes. ● Automate operational work — provisioning, deployment, recovery — to cut toil and manual steps. ● Operate core platform components: Kubernetes, databases, networking, and secrets. ● Improve disaster recovery toward reliable, repeatable datacenter restoration. ● Partner with development teams on production readiness, capacity, and change management. **Requirements** ● Proficient communication in English as we are an English speaking company with a multinational workforce and the client is US based. Solid experience operating production systems as an **SRE, platform, or operations engineer** , with a reliability\-first mindset. ● Strong **Linux fundamentals** and the ability to troubleshoot under pressure. ● **Kubernetes in production** — deployments, networking, and debugging real failures. ● **Observability** experience — metrics, logs, and tracing — and defining and using **SLOs** (Grafana/Prometheus\-style stacks). ● **Automation and scripting** (e.g., Python, Bash, Go) and infrastructure\-as\-code (Helm, GitOps). ● **Incident response and on\-call** experience, including writing and acting on postmortems. ● Comfort with **API and systems integration** and CI/CD (GitLab) **Preferred / bonus** ● **On\-prem / self\-hosted / bare\-metal or virtualization** experience, as opposed to fully cloud\-managed environments.