About this role
Google Beam is developing advanced software and hardware technology to enable people who are separated by distance to experience a conversation as if they were together. In this role, you will have the opportunity to have a big impact inventing the future of communication with Google products. We are a distributed team with members in Seattle, Mountain View, San Francisco, Los Angeles, New York, and points in between.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google (https://www.google.com/about/careers/applications/benefits/).
Minimum qualifications:
- Bachelor’s degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience working with and evaluating audio or speech processing ML models.
- 5 years of experience with speech enhancement, audio quality assessment, audio signal processing, and sound separation.
- 5 years of experience working in C++ and Python to develop, integrate, and maintain code.
- 5 years of experience analyzing problems, designing solutions, and implementing them with code.
Preferred qualifications:
- Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
- Extensive experience in Python programming, including deep learning frameworks (JAX, Keras, etc.).
- Experience with real-time audio processing and optimization techniques.
- Experience with understanding issues and either managing them or routing them appropriately.
- Excellent communication and teamwork skills, with the ability to grow in a multidisciplinary environment.
- Define the roadmap and set goals and priorities for the Audio ML effort. Coordinate workstreams with supporting team members. Work closely with research scientists, audio engineers, and software developers to integrate your models into the Starline system and launch groundbreaking audio features.
- Train, evaluate, and fine-tune models for speech enhancement, sound separation, classification, and remixing. Develop and maintain the audio infrastructure in C++, ensuring scalability, reliability, and efficiency.
- Optimize model inference for real-time, low-latency performance on both CPU and GPU platforms.
- Design and implement data processing pipelines to prepare and augment training data, leveraging techniques like TFRoomSim and automated labeling.
- Polish models and infrastructure to deliver optimal processing. Participate in On-Duty rotations to address alerts and concerns across the Audio stack, including performing diagnosis and triage of incoming bugs and issues.