AI/ML Engineer

Katalyst Data Management · Barra da Tijuca, RJ, BR

*Join the dynamic and collaborative team at Katalyst Data Management (KDM)! KDM is seeking an AI/ML Engineer to support the operation, optimization, and reliability of the infrastructure powering our AI innovation platform. This hands\-on role will help maintain GPU environments, monitor large\-scale data ingestion and document processing pipelines, support DevOps and platform operations, and contribute to scalable AI/ML solutions that deliver real value to global energy clients. The ideal candidate brings practical experience in systems administration, DevOps, MLOps, or software engineering; strong problem\-solving skills; and the curiosity and adaptability to learn new tools quickly in a fast\-moving technical environment.* * **Position Located in Houston, TX USA; Calgary, AB Canada; or Rio de Janeiro, Brazil** * **8:00 a.m. – 5:00 p.m. Monday to Friday** * **Full\-Time position** * **Hybrid Schedule Availability** **The Company** Katalyst Data Management (KDM) is a global leader in subsurface data management solutions for the energy industry. For over 30 years, we have helped oil and gas companies, national governments, and energy organizations maximize the value of their data through secure storage, quality management, digital delivery, and online data marketing services. Our industry\-leading iGlass™ platform provides customers with reliable, secure, 24/7 access to critical subsurface information, supported by robust system redundancy and data protection controls. Through innovation, technical expertise, and exceptional customer service, KDM continues to deliver trusted solutions to clients around the world. **Key Responsibilities and Accountabilities** Katalyst Data Management is seeking a Senior AI/ML Engineer to serve as a technical lead for segments of our AI innovation platform. This is a hands\-on, high\-impact role where you will own architecture and day\-to\-day development for AI projects that move from prototype to production. You will work directly with the Director of Innovation and a small, high\-velocity team to build, ship, and iterate AI capabilities that deliver real value to leading oil and gas clients worldwide. This is not a research\-only role; it requires sound architectural judgment, practical execution, and the ability to deliver production\-ready solutions in a rapidly evolving technology landscape. The position may be based in Rio de Janeiro, Brazil; Houston, Texas; or Calgary, Alberta. We value adaptability and intellectual curiosity over mastery of any single technology stack, and we are especially interested in candidates who have repeatedly learned new tools quickly, made thoughtful technical decisions under uncertainty, and shipped AI/ML solutions that work in production. **Key Responsibilities:** **GPU Infrastructure \& Server Management** * Monitor, maintain, and optimize machine learning infrastructure, including test environments and production\-grade NVIDIA GPU systems (e.g., the Houston DGX environment). * Manage GPU workload allocation and balancing across services such as embedding generation, OCR processing, and inference pipelines. * Establish and maintain system observability, including health metrics, alerting frameworks, and uptime targets for AI platforms. * Identify and implement improvements to system efficiency, resource utilization, and pipeline performance. * Diagnose and resolve hardware and system\-level issues, including thermal constraints, NVIDIA Xid errors, driver compatibility, and power optimization. **Data Ingestion \& Pipeline Operations** * Operate and monitor large\-scale ingestion pipelines processing high\-volume datasets (500GB\+), supporting thousands of documents in production environments. * Manage preprocessing workflows for domain\-specific data formats, including LAS, DLIS, SEG\-Y, NAV, PDF, and TIFF. * Monitor, troubleshoot, and optimize OCR workflows, with a focus on resolving bottlenecks in scanned legacy document processing. * Develop and maintain dashboards, logging, and alerting mechanisms to ensure pipeline visibility and reliability. * Perform data validation and quality assurance checks on ingested content; escalate anomalies and inconsistencies as required. **DevOps \& Platform Support** * Support CI/CD pipelines, automated testing frameworks, and release processes for the Katapult platform. * Assist in the administration and optimization of Elasticsearch/OpenSearch clusters, including index management and performance tuning. * Develop and maintain infrastructure documentation, including runbooks, troubleshooting guides, and operational playbooks. * Support and enhance containerized deployment environments using Docker and configuration management tools such as Ansible. **Team Collaboration \& Enablement** * Partner with Senior AI/ML Engineers to support RAG pipeline development, testing, and operationalization. * Assist with benchmarking and performance evaluation of embedding models, search configurations, and retrieval