Platform Observability Engineer
StoneX Group · São Paulo, SP, BR
Overview: **This role can be located in Sao Paulo (Brazil) or Bogota (Colombia)** **Please submit CV in English** **Connecting clients to markets – and talent to opportunity.** With 4,300 employees and over 400,000 retail and institutional clients from more than 80 offices spread across five continents, we’re a Fortune\-100, Nasdaq\-listed provider, connecting clients to the global markets – focusing on innovation, human connection, and providing world\-class products and services to all types of investors. At StoneX, we offer you the opportunity to be part of an institutional\-grade financial services network that connects companies, organizations, and investors to the global market’s ecosystem. As a team member, you'll benefit from our unique blend of digital platforms, comprehensive clearing and execution services, personalized high\-touch support, and deep industry expertise. Elevate your career with us and make a significant impact in the world of global finance. **Corporate:** Engage in a deep variety of business\-critical activities that keep our company running efficiently. From strategic marketing and financial management to human resources and operational oversight, you’ll have the opportunity to optimize processes and implement game\-changing policies. Responsibilities: **Job Purpose:**As a Platform Observability Engineer, you will design, build, and operate the platforms, tooling, and practices that provide end to end visibility into our systems. In the near term your primary focus will be owning and evolving our OpenSearch based logging and search platform running on Kubernetes. Over time you will work more broadly across our observability stack centered on Datadog or a similar platform, integrating complementary open source and cloud native technologies to meet strategic business goals. This role requires expertise in metrics, logs, traces, alerting, SLOs, and instrumentation, strong skills in infrastructure as code and automation, and close collaboration with security, development, product, and operations teams. You will be a hands\-on contributor who improves reliability, accelerates troubleshooting, and enables data driven engineering through first class observability. **Primary duties will include:*** Design, implement, and operate scalable observability platforms and services, with a primary focus on our OpenSearch based logging and search platform on Kubernetes, and Datadog or a similar platform * Operate and improve OpenSearch clusters in production, including scaling, upgrades, index lifecycle management, and troubleshooting performance and reliability issues * Adopt instrumentation standards using OpenTelemetry where appropriate, drive telemetry by default across services and platforms * Design, build, and manage observability data pipelines that extract, transform, and route telemetry from diverse sources using tools like Cribl, Vector, Fluent Bit, or OpenTelemetry Collector * Implement data filtering, enrichment, redaction, and normalization to ensure telemetry quality, compliance, and cost efficiency * Optimize observability data lifecycle management, including tiered storage, retention policies, and archive strategies * Build actionable alerting, dashboards, runbooks, and analytics that reduce noise and improve MTTR * Implement infrastructure as code for observability resources using Terraform, including Datadog or similar providers and reusable modules * Integrate observability into CI and CD workflows, enable self\-service patterns for developers and SREs * Ensure platform stability, performance, security, and cost efficiency through proactive monitoring, tuning, retention management, and incident response * Partner with security, compliance, and architecture to enforce governance, access controls, and data policies * Create and maintain customer facing and service level documentation, and deliver training for the wider IT organization * Participate in an on\-call rotation and contribute to incident response and post incident reviews Qualifications: **To land this role you will need**:* A track record of being self\-driven and solving complex problems with minimal oversight * 5\+ years of experience in observability, SRE, platform, or infrastructure roles * Hands on experience with Datadog or a similar platform, including metrics, logs, APM and tracing, RUM, synthetics, dashboards, alerting, and service catalogs * Experience with open source observability tools such as Prometheus, Grafana, and InfluxDB, and strong hands on experience operating Elasticsearch or OpenSearch clusters in production, ideally on Kubernetes * Experience with index design, query tuning, and index lifecycle policies in Elasticsearch or OpenSearch, including troubleshooting performance and reliability issues * Practical knowledge of OpenTelemetry concepts and instrumentation patterns * Experience with ETL or data pipeline tools (for example Cribl Stream, Fluent Bit, Logstash, Kafka, or OpenTelemetry Collec