Senior AI/ML Engineer - R01571019

BrillioBengaluru, KarnatakaOn-siteFull-timeMid level, 2–5 yearsListed 53 minutes ago

Apply now

About this role

Senior AI/ML Engineer

Job requirements

AI Production Engineer - Agentic Systems Role Overview We are looking for an AI Production Engineer to turn agentic AI designs into secure, reliable, and scalable production solutions. Working closely with Enterprise Architecture, Security, this role will build and deploy agents, orchestration workflows, integrations, and the supporting production control. This is a hands-on engineering role focused on operationalizing AI, not AI research or model training. The successful candidate should be comfortable working with different agentic design patterns, technology platforms, and models. Key Responsibilities
- Build and deploy production-grade AI agents and multi-step orchestration workflows based on approved architecture.
- Integrate agents with enterprise APIs, knowledge sources, business applications, databases, and automation platforms.
- Implement tool calling, workflow state, memory, human approvals, exception handling, and recovery mechanisms.
- Develop reusable services and APIs that allow agents to interact safely with enterprise systems.
- Apply responsible-AI and security controls, including authentication, authorization, data protection, audit logging, and human oversight.
- Implement automated testing for prompts, tools, integrations, workflows, security controls, and end-to-end agent behaviour.
- Establish monitoring for quality, latency, cost, tool failures, model behaviour, and production incidents.
- Build CI/CD pipelines and support-controlled releases, rollback, versioning, and environment management.
- Troubleshoot production issues and continuously improve agent reliability and performance.
- Partner with Architecture to translate approved patterns and standards into deployable solutions. Required Experience
- 4–5 years of experience in software, cloud, integration, automation, or AI engineering.
- At least 1–2 years of hands-on experience deploying LLM or agentic applications into production.
- Strong development experience in Python and working knowledge of APIs, event-driven integrations, and databases.
- Experience with at least one agent framework or platform, Agents SDK, LangGraph, LangSmith, Semantic Kernel, AWS Bedrock, or a comparable solution.
- Experience implementing retrieval, tool calling, workflow orchestration, structured outputs, and human-in-the-loop processes.
- Practical experience with Git, automated testing, CI/CD, containers, and cloud deployment.
- Experience with production monitoring, logging, alerting, incident investigation, and performance optimization.
- Understanding of enterprise security, identity, secrets management, access controls, and protection of sensitive data.
- Ability to work across architecture, security, platform, and business teams. Preferred Experience
- Experience with Azure or AWS and infrastructure-as-code tools.
- Familiarity with Kubernetes, serverless services, API gateways, message queues, or workflow platforms.
- Experience evaluating agent quality, task completion, groundedness, tool selection, safety, latency, and cost.
- Understanding of tracing and observability across prompts, models, tools, APIs, and workflow steps.
- Experience integrating AI solutions with platforms such as SharePoint, Salesforce, ServiceNow, Jira, or enterprise data services. What Success Looks Like
- Agentic solutions move from approved design to production through a repeatable and governed process.
- Deployments are secure, observable, testable, and recoverable.
- Agent decisions, tool calls, data access, failures, and human approvals are traceable.
- Solutions meet agreed expectations for reliability, quality, latency, cost, and responsible-AI controls.