About this role
ABOUT BERNARD
Half of all appliance repair visits fail on the first trip: the wrong part, the wrong diagnosis, another week of waiting, and another truck roll. Bernard fixes this at the point of first contact.
Our AI answers the call, runs the diagnostic, predicts the right parts against live inventory, and sends the technician out with a game plan. We do not just book the job. We solve it before the truck rolls.
We are live with enterprise customers, growing fast, and building the operating system for appliance repair from our office in New York City. Every hire touches the product, the customer, and the trajectory of the company.
THE ROLE
You will lead with Voice while helping build Bernard’s broader agent platform and common brain. Our agents operate across calls, chat, diagnostics, scheduling, and customer systems. They need shared context, instructions, tools, memory, policies, and evaluation infrastructure so that every channel reasons and acts consistently.
Voice is the most demanding real-time surface for that platform. Every turn requires tight coordination across telephony, speech recognition, models, tools, and speech generation. You will fight for low latency without sacrificing instruction following, natural voice quality, or correct action-taking, while turning what Voice teaches us into reusable capabilities for agents across every channel.
This is not a role where you will be pinned into telephony. You will own the Voice domain end to end, including the contact-center infrastructure that routes and escalates calls reliably, and contribute deeply to the architecture and product direction of Bernard’s unified agent platform.
WHAT YOU’LL OWN
• Build and operate production voice agents for inbound and outbound customer conversations, while developing shared agent-platform capabilities that also power chat and other workflows.
• Improve instruction following, conversational control, interruption and barge-in handling, turn detection, pacing, pronunciation, and voice quality.
• Build automated evaluations, simulations, and regression tests for task completion, policy adherence, latency, audio quality, and human handoff.
• Design and operate Twilio infrastructure, including Programmable Voice, Flex, Studio Flows, routing, queues, transfers, recordings, webhooks, and failover.
• Build the common brain across channels: shared context, instructions, tools, memory, policies, identity, and action-taking for Voice, Chat, Diagnostics, and future agent surfaces.
• Integrate agents with scheduling, diagnostics, CRM, and operational systems while handling retries, idempotency, partial failure, and long-running state.
• Build observability for live conversations: traces, transcripts, audio events, latency budgets, failure classification, and alerts.
• Reduce end-to-end and turn-level latency across telephony, speech-to-text, model inference, tool calls, and text-to-speech while preserving response quality.
• Partner with product and customer teams to review real conversations, identify cross-channel failure patterns, and ship rapid improvements.
• Help shape the architecture and roadmap for Bernard’s unified agent platform, using Voice as the proving ground for capabilities that should generalize across the product.
YOU SHOULD APPLY IF
• You have 2+ years of professional software engineering experience.
• You have shipped production backend, real-time, telephony, or conversational systems.
• You are strong in TypeScript or Python and can reason clearly about asynchronous state, streaming, distributed failure, and observability.
• You understand the tradeoffs among latency, instruction following, reliability, and natural conversational behavior.
• You are excited to start from Voice but think in platform primitives, and you want to build agent systems that work across channels rather than remain confined to one surface.
• You are comfortable debugging across application code, model behavior, audio pipelines, carrier behavior, and customer configuration.
• You move quickly, own outcomes, and learn directly from customers and production calls.
PREFERRED
• Experience with Twilio Programmable Voice, Flex, Studio, TaskRouter, SIP, or contact-center infrastructure.
• Experience with streaming speech-to-text, text-to-speech, VAD, turn detection, interruption handling, or real-time model APIs.
• Experience building LLM agents, shared tool or memory systems, evaluation infrastructure, call simulations, or human-in-the-loop escalation.
• Familiarity with voice AI agent platforms, carrier networks, WebSockets, or event-driven systems.
HOW WE WORK
We are a small team. Everyone owns their domain end to end. There is no middle management, no committees, and no approval chain. If something needs to happen, you make it happen.
We work in person in New York City. We move fast, give direct feedback, and hold each other to a high bar. The pace is startup pace, if that energizes you, you will love it here.