About this role
Location: Remote, only for candidates in LATAM
Job Overview
Ryz Labs is looking for a Principal Machine Learning Engineer to join the core engineering team of one of our key international clients. In this role, you will play a pivotal part in shaping our client’s next-generation AI platforms, sitting at the intersection of production systems, applied AI, and engineering excellence.
You will bridge distributed systems architecture with hands-on MLOps and LLM engineering. We are seeking an exceptional technical leader who can design scalable multi-tenant architectures, build robust AI agent governance frameworks, and serve as a trusted technical authority collaborating directly with product managers and key client stakeholders.
Key Responsibilities
• Production ML & Optimization: Deploy and manage AI models at scale, monitoring performance, hallucination rates, drift, latency, and infrastructure costs.
• Architecture & Delivery: Design distributed, event-driven microservices using Python, Go, or TypeScript while building IaC and CI/CD pipelines to ship your own services.
• AI Security & Governance: Implement agent permission structures, human-in-the-loop workflows, data isolation boundaries, and prompt injection defenses.
• Product Collaboration: Partner with Product Managers from inception to translate business requirements into scalable architectures and present trade-offs to executives.
What You Bring
• 12–15+ years in software engineering, with a clear evolution from Backend/Distributed Systems Architecture into applied Production ML Engineering.
• Production ML Expertise: Deep experience with MLOps, model evaluation rubrics, advanced RAG, vector search (embeddings, HNSW, hybrid search), and fine-tuning. (We are looking for engineers building real systems, not just consuming LLM APIs).
• Software Architecture: Strong mastery of distributed systems, microservices, and asynchronous event-driven patterns in Python, Go, or TypeScript.
• Practical DevOps & Cloud: Hands-on command of Docker, cloud infrastructure (AWS/GCP/Azure), and automated CI/CD pipelines.
• AI Governance & Security: Practical knowledge of LLM safety, threat modeling, data boundary enforcement, and agent security.
• Fluent English & Communication: Ability to articulate complex technical trade-offs (e.g., RAG vs. Fine-tuning, latency vs. accuracy) clearly to client executives and non-technical stakeholders.
Nice-to-Have
• Prior experience with multi-agent orchestration frameworks (e.g., LlamaIndex, Semantic Kernel, AutoGen, CrewAI, MCP).
• Experience in fast-paced consulting, advisory, or high-growth tech platforms.
.webp)