About this role
About Verita AI
Verita sources white collar experts and leverages expert-powered data infrastructure to support frontier AI labs and enterprises. Our Foundry platform runs the contractor lifecycle from applications and interviews through onboarding, contracts, time tracking, and payouts. The Studio platform leverages that talent for secure generation, curation, enrichment, and evaluation pipelines across task routing, quality, protected media, cost, and client delivery. This role sits within our core engineering team, working at the intersection of research and product to build the RL infrastructure that powers our sourcing, data, and model training pipelines.
The Role
We are seeking an “AI-native” Machine Learning Engineer to join our core engineering team. This role sits at the intersection of research and engineering, requiring the ability to build robust RL environments and iterate rapidly on new product automations. You will be expected to collaborate closely with researchers and product builders to translate high-level hypotheses into concrete, verifiable implementations designed to scale with minimal friction and the highest quality standards.
Core Technical Requirements (Must-Haves)
- RL Environment Proficiency. Demonstrated experience building, managing, and troubleshooting reinforcement learning environments end-to-end.
- Containerized RL. Hands-on experience with containerized RL environments and framework integration (specifically the Harbor framework).
- Verifiable Rewards (RLVR). Practical experience building and implementing verifiers for RL models to ensure consistent, verifiable reward signals.
- Optimization Methodologies. Proficiency in modern reward functions and optimization (PPO, DPO, GRPO).
- QC Pipelines. Expertise in designing and maintaining quality control (QC) pipelines for RL and synthetic data.
- Collaborative Mindset. Proven ability to work in a pair-programming or highly collaborative team setting, specifically with researchers, operations stakeholders, and product owners to cross-verify ideas and iteratively improve model performance.
Distinctions & Specializations (Nice-to-Haves)
We recognize that candidates may lean toward different domains; we are looking for depth in at least one of the following areas:
- “Deterministic” Domain Specialization
– Experience building scalable RL environments for deterministic domains (e.g. math, physics, biology) with advanced verifier architectures.
– Advanced mathematical reasoning applied to system verification.
– Strong background in “programmer-hardcore” coding practices.
– Familiarity with verifiable proofs (e.g. Lean, Rocq, Dafne)
- Product/Qualitative Specialization
– Expertise in RL for non-deterministic domains (e.g. design, UX, writing, linguistics).
– Experience conceptualizing and implementing architecture-agnostic RL environments designed to be natively multimodal.
– Experience with preference capture techniques (triplets, pairwise comparison, etc).
– Familiarity replicating/cloning tools for trace capture (e.g., Adobe, Figma, web-browser automation) to help train AI agents.
– Experience standing up non-deterministic RLHF (Reinforcement Learning from Human Feedback) workflows that don’t collapse qualitative signals.
Candidate Profile
- Impact > Tenure. We care about what you’ve done; not how long you’ve been doing it. AI-native builders whose excellence has led to rapid career progression are ideal co-builders.
- Seasoned Problem Solving. Being able to explore and learn quickly is essential to success here, but we’re not looking to train you on-the-job. Priority is given to those who have already “solved” some of the puzzles we’re facing, e.g. standing up or integrating open-source comparables of closed-source software for RLE, building and improving AI-powered tools like AI interviewers and dynamic administrative assistants, and originating agent-as-a-judge ensembles.
- Collaborative Thinking. You thrive in environments where you bounce ideas off team members, cross-verify logic, and contribute to both research- and product-heavy initiatives.
Compensation & How to Apply
- Base salary. $200,000 to $250,000 base salary (+ generous equity)