About this role
About the Role
This is a foundational engineering role at an early-stage AI consumer hardware and software startup, where you will own the transcription pipeline end-to-end. You will work hands-on with product and general management leadership to build, tune, and ship a cloud-based ASR system with a narrowly scoped on-device component. Your work directly shapes how well the core product experience feels to real users.
What You'll Do
- Build and iterate on the cloud-based ASR pipeline, from audio capture through post-processing, running in production at scale.
- Own ASR quality and reliability end-to-end, shipping measurable improvements across latency, small-word accuracy, and voice-print reliability.
- Work across data preparation, model training and fine-tuning, evaluation, and deployment to translate product feedback into shipped pipeline changes.
- Collaborate with a Partner Product Engineer on shared backend and pipeline surfaces.
- Coordinate across time zones with R&D, hardware, and supply-chain teams based in China.
- Operate with minimal specification, turning informal asks into concrete, shipped improvements.
What We're Looking For
- 3 or more years building and tuning transcription and ASR pipelines end-to-end in production, primarily in cloud-based settings.
- Demonstrated ownership of production ASR systems across the full lifecycle: data preparation, model training and fine-tuning, evaluation, and deployment.
- Experience building and optimizing latency-sensitive or streaming audio and ASR pipelines.
- Track record of making latency, accuracy, and reliability tradeoffs based on real user feedback.
- Experience debugging and tuning transcription quality issues in production environments.
- Comfort shipping in early-stage or founding engineering environments with small teams and limited specification.
- On-device or embedded ML experience using frameworks such as Core ML or TensorFlow Lite.
- Prior experience with wearable, hardware, or robotics device products.
- Background at AI-native consumer applications focused on transcription or audio.
- Experience building agent or LLM-based product features including tool use, memory, or retrieval systems.
- Ability to work hybrid three days per week in the San Francisco Bay Area.
- Ability to collaborate asynchronously with international teams across time zones.
Compensation & Benefits
Salary range: $150,000 to $200,000 USD annually. Visa sponsorship is not available for this role.
Location
Hybrid, three days per week on-site in the San Francisco Bay Area, California, United States.