Senior Evaluation Expert – Large Language Models & AI Agents

PatsnapShanghai, ShanghaiOn-siteFull-timeSenior, 5–8 yearsListed 4 hours ago

Apply now

About this role

About PatSnap

PatSnap is a global enterprise SaaS company with extensive professional data, knowledge assets, and real-world use cases across patents, scientific literature, chemistry, materials science, and life sciences.

We are integrating large language models, RAG, and AI agents with professional data to support complex workflows including technology research, intellectual property analysis, scientific intelligence, and R&D decision-making.

We are looking for a Senior Evaluation Expert to build the evaluation system for our LLM- and AI agent-powered products and lead the team in providing reliable quality signals for model improvement, product iteration, and production deployment.

Responsibilities

- Own the evaluation framework for PatSnap’s LLM- and AI agent-powered products, covering foundation model capabilities, retrieval and RAG, tool use, agent planning and execution, and end-to-end task performance.
- Develop a deep understanding of professional workflows across patents, scientific research, chemistry, materials science, and life sciences. Work with product, AI, and domain teams to translate complex business requirements into measurable and reproducible evaluation criteria.
- Build and continuously improve benchmark datasets, real-world user task sets, challenging and adversarial cases, and regression test suites, together with standards for annotation, quality assurance, and version management.
- Combine human evaluation, rule-based methods, LLM-as-a-Judge, online experiments, and user feedback to assess accuracy, completeness, professional quality, traceability, instruction following, hallucination, safety, and task completion.
- Establish process-level evaluation and diagnostic methods for AI agents. Analyze task understanding, planning, retrieval, tool use, and response generation to identify root causes and drive improvements across models, data, retrieval, and product workflows.
- Integrate evaluation into development, model training, release, and production monitoring workflows, including automated evaluation, regression testing, and release quality gates.
- Lead the evaluation team by setting objectives, allocating responsibilities, developing team capabilities, and coordinating product, AI, engineering, data, and domain teams to resolve critical quality issues.

##

Requirements

Qualifications

- Bachelor’s degree or above in Computer Science, Artificial Intelligence, Data Science, Mathematics, Statistics, Information Retrieval, or a related discipline.
- At least five years of experience in algorithm evaluation, data science, search and recommendation, NLP, or AI product quality. Experience evaluating LLMs, RAG systems, or AI agents is preferred.
- Demonstrated ability to independently design an evaluation system, including task definition, metric design, dataset construction, experiment design, result analysis, and root-cause diagnosis.
- Solid understanding of common LLM evaluation methodologies and the appropriate use, potential biases, and reliability limitations of human evaluation, automated metrics, LLM-as-a-Judge, and A/B testing.
- Strong understanding of RAG, search, or information retrieval pipelines, including retrieval, ranking, context construction, citation grounding, and response quality. Proficiency in Python, SQL, or similar analytical tools is expected.
- People management and complex project leadership experience, with the ability to set objectives, coordinate resources, develop team members, and maintain delivery quality. Exceptional individual contributors without formal management experience may also be considered if they demonstrate strong project leadership and mentoring experience.
- Strong structured thinking, ownership, and cross-functional influence, with the ability to remain hands-on and turn evaluation findings into concrete product and technical improvements.

Preferred Qualifications

- Experience in patents, intellectual property, scientific research, chemistry, materials science, or life sciences;
- Experience with enterprise SaaS, professional search, knowledge graphs, or scientific databases;
- Experience evaluating agent workflows, tool use, multi-turn tasks, or complex reasoning;
- Experience building an evaluation platform, evaluation team, or AI quality system from the ground up.