Data Scientist– AI Infra & Evaluation Foundations

monday.comTel Aviv, Tel AvivHybridFull-timeJunior, 1–2 yearsListed 1 day ago

Apply now

About this role

About monday.com:

monday.com is the AI work platform powering the most ambitious teams. 250,000+ customers across departments use us to bring people, workflows, and AI agents together on one flexible platform where AI doesn't just assist, it executes. We move fast, build things that matter, and foster an ownership-driven culture where you're empowered to shape how organizations work and outpace their competition.

About the team
 
The AI Infra group builds the foundations, tools, and platforms that every team at monday relies on to ship intelligent, agentic features. We own the core infrastructure—including the AI Gateway and our centralized Evals framework—ensuring every AI feature deployed to production is secure, resilient, cost-effective, and above all, trustworthy.
 
Our focus is on the frontier of agentic AI: dissecting complex agent trajectories, building robust evaluation frameworks, and turning subjective notions of "good AI" into rigorous, actionable metrics. As a Data Scientist on this team, you will bridge the gap between AI research and production infrastructure. You'll partner closely with engineering and product teams across monday to design the judges, metrics, and error analysis workflows that allow us to ship cutting-edge AI agents with speed and confidence.
     
This position is based at our Tel Aviv office (Headquarters).

About the role:

As a Data Scientist in AI Infra, your goal goes far beyond simply building an evaluation framework—you will own the organizational impact of how monday evaluates and trusts AI. You will define how teams measure quality, influence engineering decisions across R&D, and turn fuzzy notions of "good AI" into numbers product teams rely on to ship with confidence.

- Own the evaluation methodology: Design metrics, pipelines, and methodology that teams across monday trust and adopt as their source of truth.
- Transform the AI agent lifecycle: Standardize how AI agents are built, regression-tested, and maintained across the org, embedding continuous evaluation into everyday engineering workflows and post-deployment monitoring.
- Drive organizational impact & enablement: Partner with AI feature teams across monday to translate domain expectations into meaningful datasets, test suites, and continuous evaluation pipelines—leveling up engineers and product managers along the way.
- Build hands-on tools: Prototype and stand up eval pipelines end-to-end, bridging the gap between ambiguous, high-level product requirements into clear, quantifiable evaluation standards that become central to how features are greenlit for production
- Anticipate future failure modes: Stay ahead of evolving agent architectures by proactively designing next-generation evaluation strategies.

Requirements:

- Agentic Systems & Architecture:

3+ years of experience as a Data Scientist in non-academic settings working with complex production running AI systems. Familiarity with current agentic frameworks like LangGraph , LangChain , and SoTA SDKs.
- Deep, practical understanding of how agents operate - models, context, capabilities and harnesses. Deep experience with agentic evaluation methodologies

- Execution, Code & Trace-First Mindset:

Production-grade coding skills with a track record of building, prototyping, and shipping end-to-end data or eval pipelines,
- A "trace-first" diagnostic mindset—comfortable diving into raw agent execution logs, inspecting failure modes, and constructing qualitative error taxonomies.

- Product & Organizational Impact:

Strong product intuition and exceptional communication skills to translate complex evaluation data into clear, actionable guidelines.
- Proven ability to partner closely with software engineers and product teams, taking ownership of driving adoption and raising the quality bar across the organization.

Preferred Qualifications

- Direct experience designing evaluation strategies for complex agentic systems in production
- Prior experience working within centralized platform/infra teams that support multiple product verticals.
- Experience contributing to modern microservice architectures, GitHub workflows, and automated production CI/CD pipelines
- Familiarity with modern agent and eval tooling and observability stacks (e.g., LangSmith, Langfuse, or custom internal platforms).
- Knowledge of TypeScript or experience working within modern platform architectures.

#LI-DNI