About this role
ABOUT BERNARD
Half of all appliance repair visits fail on the first trip: the wrong part, the wrong diagnosis, another week of waiting, and another truck roll. Bernard fixes this at the point of first contact.
Our AI answers the call, runs the diagnostic, predicts the right parts against live inventory, and sends the technician out with a game plan. We do not just book the job. We solve it before the truck rolls.
We are live with enterprise customers, growing fast, and building the operating system for appliance repair from our office in New York City. Every hire touches the product, the customer, and the trajectory of the company.
THE ROLE
You will build the AI and machine learning platform behind Bernard’s products: the infrastructure for search, retrieval, training, evaluation, inference, and model deployment, plus the models that turn operational evidence into reliable decisions. Your scope spans both Diagnostics and Bernard’s unified brain for customer support and experience, powering voice agents, chat agents, and future agent surfaces with shared context, memory, knowledge, tools, policies, and learning loops.
This is a broad, hands-on systems role. You will work from model creation through production infrastructure, own reliability across the platform, and improve how engineers build and ship software in an AI code generation world. The goal is not isolated experiments. It is a fast, dependable platform that lets the entire team create, evaluate, deploy, and improve intelligent products.
WHAT YOU’LL OWN
• Build and operate machine learning infrastructure for search, retrieval, indexing, training, evaluation, inference compute, model serving, and observability.
• Create and improve models for diagnostics, failure-mode prediction, parts prediction, ranking, extraction, customer-support automation, and other product capabilities.
• Build the common intelligence layer across voice and chat agents, including shared context, memory, retrieval, knowledge, tool use, policies, identity, and cross-channel continuity.
• Design data and feedback pipelines across service history, manuals, equipment metadata, conversations, images, technician actions, and repair outcomes.
• Own platform reliability across distributed services, data pipelines, search systems, and online inference, including monitoring, latency, capacity, failure recovery, and incident prevention.
• Build evaluation systems that connect model and retrieval quality to real customer and repair outcomes across diagnostics, voice, chat, and customer-support workflows.
• Turn conversations, agent actions, customer history, and human feedback into learning loops that continuously improve the unified brain across channels.
• Improve developer experience for an AI-native engineering team through better environments, testing, CI/CD, observability, reusable platform primitives, and safe workflows for AI-generated code.
• Partner with product engineers and domain experts to turn new models and platform capabilities into fast, explainable, production-ready experiences.
• Make pragmatic architecture decisions across build versus buy, compute, storage, model providers, and open-source infrastructure.
YOU SHOULD APPLY IF
• You have 5+ years of professional software or machine learning engineering experience and have shipped production systems.
• You are strong in Python and Typescript and comfortable working across models, backend systems, data infrastructure, and cloud compute.
• You understand modern search and ML systems, including retrieval, indexing, evaluation, serving, and the failure modes of deployed models.
• You can reason clearly about distributed systems, performance, reliability, observability, and cost.
• You care about engineering leverage and have strong opinions about how AI-assisted development should improve speed without lowering quality.
• You can move between research, infrastructure, and product delivery based on what creates the most value.
• You are excited to build intelligence that spans diagnostics and customer experience, rather than being confined to a single model, modality, or product surface.
• You care about measurable customer outcomes more than benchmark wins.
PREFERRED
• Experience with ranking, recommendations, vector or graph search, multimodal models, or real-time inference.
• Experience building internal ML platforms, model gateways, evaluation systems, data platforms, or developer tooling.
• Experience with GPUs, inference optimization, orchestration, CI/CD, infrastructure as code, or production incident response.
• Experience using AI code generation tools deeply and building guardrails or workflows around them.
HOW WE WORK
We are a small team. Everyone owns their domain end to end. There is no middle management, no committees, and no approval chain. If something needs to happen, you make it happen.
We work in person in New York City. We move fast, give direct feedback, and hold each other to a high bar. The pace is startup pace. If that energizes you, you will love it here.