About this role
Most of our users never really touch our app. They text us. Restaurant owners and creators handle almost everything through conversation, and on our side that's a set of AI agents doing the booking, the follow-ups, the scheduling, and the problem solving.
You own those agents. Not the prompts alone, the whole system underneath them: how they understand what someone actually wants, how they call tools without breaking, how they hold context across a conversation that's been going for 6 months, and how they take real actions in the real world without us having to check their work.
Any text that comes in, your systems handle it.
**What you'll own**
**The agents.** Full ownership of both of our agents end to end. How they're built, what they can do, and what happens when they get something wrong.
**The harness.** The layer coordinating LLMs, tools, memory, and async workflows. This is the actual engineering problem here and it's most of your week.
**Tooling** that doesn't fall over. Schema validation, retries, permissions, error handling. An agent that calls a tool correctly 95% of the time is not good enough when it's booking real visits at real restaurants.
**Evals**. Frameworks that measure task completion, accuracy, safety, latency, and failure modes. If we change a prompt on Tuesday, we should know by Tuesday whether it made things worse.
**Observability.** Traces, logs, and failure analysis across agent workflows. When something goes wrong in a conversation, you should be able to see exactly where.
Turning vague into dependable. A restaurant owner texts something ambiguous at 11pm. Getting from that to a reliable agent behavior is the hard part, and it's the part we care most about.
**What we're looking for**
Required
* 2+ years building software, with real experience building agent systems or harnesses. Not just calling an API in a side project
* Strong conversational agent experience. This is the thing we weigh most heavily. Our product is a conversation, and someone who has only built single-turn or task-runner agents will struggle here
* Strong database and system design
* You've shipped agents that real people used and dealt with the fallout when they broke
* Comfortable with ambiguity. There's no established playbook for most of this
* Based in Orange County / willing to relocate and able to work onsite
* This isn't a 9-to-5
Nice to have
* iMessage agent experience
* You've built eval frameworks, not just run them
* Observability and tracing work on LLM systems
* Restaurant industry or creator economy experience
