About this role
Senior AI Workflow & Systems Engineer
TubeScience: Los Angeles (preferred) — $110,000–$160,000 base (Los Angeles) · $70,000–$120,000 base (remote)
TubeScience is Meta's largest creative partner and AppLovin's #1 creative partner, producing 8,000+ original ads every month from a 100,000 sq. ft. Los Angeles studio, backed by a library of 1.6 million+ performance ads and $3B in annual managed ad spend.
Information Systems builds the internal software that runs the company. We combine AI, engineering and automation to solve complex operational problems at scale, and our systems are used every day across the business.
You'll work directly with the people who use what you build, ship quickly, and own your systems long after launch. You'll be expected to use the latest frontier and open-weight models in every phase of the job, from prototyping to building, testing and deployment.
The role
We're looking for an engineer who evolved from systems engineering into applied AI: someone who enjoys designing reliable production systems, integrating modern AI capabilities, and owning them in production.
This is an internal forward-deployed engineering role. Rather than building products for external customers, you'll work directly with internal stakeholders to find operational bottlenecks, architect AI-powered solutions, deploy them quickly, and keep improving them based on real business needs. You'll report to our VP of Information Systems.
Success means systems that don't just work at launch but keep working reliably after it. You'll own the full lifecycle: architecture, deployment, monitoring, debugging, incident response and continuous improvement.
This is not an AI research or model-training role. We apply state-of-the-art models to enterprise problems through software engineering. If your experience is mainly low-code automation such as Zapier, Make or n8n, or mostly prototypes and prompt engineering, this role is probably not the right fit.
We weigh directly relevant experience heavily: production systems you have built and operated, and AI you have put into them.
Responsibilities
- Design and build production AI applications that automate complex enterprise workflows
- Architect agent-based systems that coordinate LLMs, APIs, internal services, databases and business logic, with reliable orchestration across tools and enterprise platforms
- Deploy AI systems with observability, monitoring, rollback strategies and operational safeguards
- Investigate production issues, analyze logs, debug failures and restore reliability when incidents occur
- Partner with Product, Operations, Creative, Engineering and Business teams to find high-impact automation opportunities
- Prototype, validate, deploy and iterate quickly based on production performance and business outcomes, and keep improving existing systems for reliability, speed and impact
Minimum qualifications
- You have 3+ years of professional software or systems engineering experience, building and operating production software used by real users or internal business teams.
- You came to applied AI from systems engineering, for example from platform, backend, enterprise systems or internal developer platform work, or from infrastructure work with significant software development, rather than starting your career in AI.
- You have strong Python engineering experience.
- You have integrated modern LLMs into production systems, using tools such as the OpenAI or Anthropic APIs, LangGraph, MCP or similar.
- You have designed systems that coordinate multiple APIs, databases, services and enterprise applications.
- You understand distributed systems and production operations: debugging, logging, monitoring, and improving systems after launch.
- You think in complete systems, with the architectural judgment to design end-to-end solutions.
- You work independently and are comfortable in a fast-paced startup environment.
Preferred qualifications
- You have built production systems at a large technology company.
- You have built multi-agent systems or used orchestration frameworks such as LangGraph, MCP or Temporal.
- You have worked with event-driven architectures.
- You have used Docker and Kubernetes, and AWS, GCP or Azure.
- You have built CI/CD pipelines and used observability platforms such as Datadog, Grafana or OpenTelemetry.
- You have built internal developer platforms or enterprise integrations.
The problems to solve
- Workflows that cross many systems. Business processes that span APIs, databases, internal services and enterprise platforms, coordinated by agents that hold up in production.
- AI that keeps working after launch. Monitoring, rollback and safeguards that catch failures before the business does.
- Incidents with a real cost. Production issues in systems people use every day, found and fixed fast.
- Speed without fragility. From prototype to production quickly, then hardening what works.
- Real operational problems. Bottlenecks found with the teams who live with them, not guessed from a distance.
How the hiring works
Three conversations and a short piece of practical work. Recruiter screen, then the interview with a VP of Information Systems, then a practical assignment with a walkthrough session.
Roughly 17 business days end to end if we both move quickly. You will get a decision either way at every stage.