About this role
About Liquid AI
Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there.
The Opportunity
This is a rare chance to own applied post-training work end-to-end for vision workloads, adapting Liquid Foundation Models for customers who need multimodal capabilities that run on-device under real constraints.
You will act as the technical bridge between customer requirements and model delivery for vision tasks. You will lead engagements from scoping through evaluation, with full ownership over how vision models are adapted and shipped. Between engagements, you will build reusable applied workflows and tooling that accelerate future delivery.
If you care about multimodal data quality, vision-language evaluation, and making VLMs actually work in production for real customers, this is the role.
What We’re Looking For
We need someone who:
- Takes ownership: Owns customer post-training projects end-to-end for vision workloads, from requirements through delivery and evaluation.
- Thinks multimodally: Can reason across image-text data pipelines, visual grounding, multimodal alignment, and evaluation as a connected system.
- Is pragmatic: Optimizes for model quality and customer outcomes over publications or theory.
- Communicates clearly: Can translate between customer needs and internal technical teams, and push back when needed.
The Work
- Act as the technical owner for enterprise customer post-training engagements involving vision and multimodal workloads
- Translate customer requirements into concrete post-training specifications for VLM and image/video tasks
- Design and execute data generation, filtering, and quality assessment processes for image, video, and multimodal corpora
- Fine-tune and align vision-language models using SFT, preference optimization, and RL-style methods
- Design task-specific evaluations for vision model performance and interpret results
- Build reusable applied tooling and workflows that accelerate future customer engagements
Desired Experience
Must-have:
- Hands-on experience with data generation and evaluation for LLM or VLM post-training
- Experience training or fine-tuning models using SFT, preference alignment, and/or RL
- Strong intuition for data quality and evaluation design in multimodal contexts
- Experience with vision-language models, multimodal training, or image/video data pipelines
- Proficiency with open-source ML ecosystem (Hugging Face, PyTorch) and modern model architectures
Nice-to-have:
- Experience with visual grounding, image-text alignment, or multimodal evaluation frameworks
- Experience delivering applied ML work to external customers with measurable outcomes
- Background in computer vision or visual representation learning
What Success Looks Like (Year One)
- Independently owns and delivers enterprise post-training projects for vision workloads with minimal oversight
- Is trusted by customers as the technical owner for multimodal engagements, demonstrating strong judgment and delivery quality
- Has built reusable applied workflows or tooling that accelerate future customer engagements
What We Offer
- Real ML work: You will fine-tune vision-language models, generate multimodal data, and ship solutions to enterprise customers, not configure API calls.
- Compensation: Competitive base salary with equity in a unicorn-stage company
- Health: We pay 100% of medical, dental, and vision premiums for employees and dependents
- Financial: 401(k) matching up to 4% of base pay
- Time Off: Unlimited PTO plus company-wide Refill Days throughout the year