About this role
Joining Amex Tech means discovering and shaping your contribution to something big. Here, you can work alongside talented tech teams and build a unique career with the Powerful Backing of American Express. With a range of opportunities to work with the latest technologies, and a commitment to back the broader engineering community through open source, our mission is to power your success. Because Amex Tech is powered by our technology, our culture, and our colleagues.
The Technology organization enables and accelerates the company’s growth strategies, delivering global capabilities and services in support of Amex’s customers and colleagues, while maintaining 24/7 servicing and availability to ensure an uninterrupted, high-quality customer experience. Technology provides the foundation for everything we do in the company while driving differentiation through building and leveraging innovative technology and data insights.
At American Express, AI is reshaping the future of commerce and redefining the experiences our commercial customers and card members expect. Within Amex Technology, we are building platforms, products, and governance that enable agentic AI systems to operate responsibly and at scale across the enterprise.
Our focus is on agentic AI development: designing intelligent, adaptive systems that can plan, reason, and act across complex workflows with appropriate levels of autonomy. These systems power autonomous workflows, decision support, and customer-facing experiences—while meeting the high standards for security, explainability, reliability, and compliance required in financial services.
We partner closely with product, design, and business teams to deliver agentic capabilities that reduce operational friction, improve decision-making, and transform how customers interact, transact, and grow.
The Enterprise Data & AI Team creates the platforms, products, and governance that power enterprise AI. Our team enables rapid experimentation, seamless deployment, and responsible operation of AI systems at scale — delivering trusted, compliant, and high-performance AI solutions that drive meaningful business outcomes, partnering with stakeholders to support use cases from inception to implementation.
American Express is building an enterprise-grade Agentic AI Platform designed to enable teams to safely build, deploy, operate, discover, govern, and continuously improve AI agents at scale.
We are looking for a Senior AI Engineer to help design and build the foundational platform capabilities that power agentic AI across the enterprise. This is a hands-on engineering role at the intersection of AI systems, distributed platforms, cloud infrastructure, developer tooling, and enterprise governance .
This role is ideal for an engineer who wants to go beyond building individual AI applications and instead build the platform and infrastructure that enables hundreds of teams and AI agents to operate safely and reliably at enterprise scale .
What You'll Build
You will contribute to the architecture and implementation of capabilities across the Agentic AI Platform, including:
Agent Runtime & Execution
Design and build scalable agent runtime infrastructure using Kubernetes and cloud-native technologies.
Develop runtime abstractions that allow agents to execute across internally managed sandboxed environments and third-party AI platforms.
Build capabilities for agent orchestration, tool execution, state and context management, memory, asynchronous workloads, event-driven execution, and multi-agent workflows.
Enterprise Agent Control Plane
Build a unified control plane for managing agents across heterogeneous execution environments and AI providers.
Develop APIs and services for agent registration, configuration, deployment, versioning, lifecycle management, policy enforcement, and runtime management.
Create abstractions that reduce provider-specific complexity while preserving access to differentiated capabilities across AI ecosystems.
Agent Governance & Security
Engineer platform-level controls for enterprise AI governance, including agent identity, authentication and authorization, tool permissions, policy enforcement, data boundaries, auditability, and lifecycle controls.
Build mechanisms for governing models, prompts, tools, MCP servers, knowledge sources, agent-to-agent interactions, and external integrations.
Partner with security, risk, privacy, and governance teams to translate enterprise requirements into scalable technical controls.
Agent Discovery & Ecosystem
Build registries and catalogs that allow developers and AI systems to discover reusable agents, tools, skills, prompts, models, knowledge sources, and other platform capabilities.
Develop metadata, ownership, versioning, dependency, certification, and discovery mechanisms that enable a healthy enterprise agent ecosystem.
Agent Observability
Build end-to-end telemetry for agent execution, including traces, events, model interactions, tool calls, latency, token consumption, cost, failures, policy decisions, and quality signals.
Develop capabilities that make complex and multi-agent workflows explainable and debuggable.
Enable platform and application teams to understand agent behavior across multiple models, runtimes, tools, and external systems.
Evaluation & Continuous Learning
Develop evaluation frameworks for measuring agent quality, reliability, safety, and task performance.
Build automated evaluation pipelines incorporating offline evaluations, production signals, human feedback, regression testing, and experimentation.
Help create continuous learning loops that turn production telemetry and feedback into actionable improvements to agents, prompts, tools, models, and platform capabilities.
Developer Experience
Build SDKs, APIs, CLIs, templates, local development environments, and self-service workflows that make agent development simple and productive.
Create opinionated paved roads that allow developers to move quickly while automatically incorporating enterprise security, governance, observability, and operational standards.
AI CI/CD & Platform Engineering
Build CI/CD capabilities specifically designed for AI agents, including automated evaluation, policy validation, security checks, artifact/version management, deployment, promotion, rollback, and release controls.
Design GitOps and infrastructure-as-code patterns for deploying and managing agent workloads.
Help establish engineering standards for taking agents from experimentation to production safely and repeatedly.
Minimum Qualifications
- Strong software engineering experience building production systems using languages such as Python, Java, Go, or TypeScript .
- Strong understanding of Generative AI, LLMs, agent architectures, tool/function calling, retrieval, context management, and agent orchestration .
- Experience building production applications or platforms using major model providers or AI platforms.
- Experience with Kubernetes, containers, microservices, distributed systems, and cloud-native architecture .
- Experience designing production APIs and event-driven or asynchronous systems.
- Strong understanding of modern cloud infrastructure and infrastructure-as-code practices.
- Experience with CI/CD, automated testing, production observability, and software delivery practices.
- Strong understanding of security fundamentals including identity, authentication, authorization, secrets, and least-privilege access.
- Ability to navigate ambiguous technical problems and turn emerging technologies into reliable production systems.
- Strong communication skills and ability to collaborate across engineering, architecture, product, security, and governance organizations.
Preferred Qualifications
Experience in several of the following areas is highly desirable:
- Agent frameworks and orchestration technologies such as LangGraph, Semantic Kernel, Google ADK, OpenAI Agents SDK, or similar frameworks .
- Model Context Protocol (MCP) and emerging agent interoperability protocols like A2A.
- Multi-agent systems and agent-to-agent communication patterns.
- Kubernetes operators, controllers, service meshes, or sophisticated Kubernetes platform engineering.
- Agent/model gateways and intelligent model routing.
- Vector databases, retrieval systems, embeddings, and enterprise knowledge architectures.
- LLM and agent evaluation frameworks, LLM-as-judge techniques, human feedback systems, experimentation, and quality measurement.
- AI observability, distributed tracing, OpenTelemetry , Langfuse and production monitoring.
- Experience with Evals
- Policy engines and policy-as-code.
- AI security, prompt injection defenses, tool security, data-loss prevention, and AI-specific threat modeling.
- Platform engineering, internal developer platforms, developer portals, and enterprise service catalogs.
- Large-scale distributed systems operating in highly regulated environments.
What Success Looks Like
You will help make it possible for developers across American Express to move from an agent idea to a governed production deployment through a consistent, self-service platform.
Successful platform capabilities will make the secure and reliable path the easiest path: developers should be able to build once and operate agents across multiple AI ecosystems while receiving governance, identity, observability, evaluation, deployment, and operational capabilities by default.
You will play a key role in establishing the engineering foundations for an enterprise ecosystem in which AI agents can be built, discovered, trusted, governed, observed, evaluated, deployed, and continuously improved at scale .
Depending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions.