About this role
Accountabilities
- Review and assess the quality, accuracy, and realism of programming tasks designed to evaluate AI agents.
- Evaluate coding environments and technical evaluation harnesses for correctness, robustness, and suitability for AI testing.
- Examine provided code structures and walk through the underlying logic, identifying potential flaws, inconsistencies, or technical limitations.
- Assess whether coding challenges accurately reflect realistic software engineering scenarios and industry practices.
- Evaluate the difficulty and complexity of programming tasks to determine whether they provide meaningful tests of engineering capabilities.
- Review the verifiability and technical soundness of evaluation criteria and harnesses.
- Provide clear, detailed feedback on potential improvements to task design, evaluation methodology, and technical implementation.
- Discuss technical architecture, testing approaches, and software engineering practices during the research session.
- Share professional perspectives on what makes coding challenges robust, realistic, and technically meaningful.
Requirements
- Professional experience as a software engineer, software developer, or closely related technical professional.
- Hands-on experience building, reviewing, testing, or evaluating realistic programming tasks.
- Experience with code review, software testing, automated testing, or technical evaluation frameworks.
- Familiarity with evaluation harnesses or similar environments used to verify programming solutions.
- Experience in one or more relevant areas such as full-stack development, backend engineering, test automation, systems architecture, or related software disciplines.
- Strong understanding of software engineering principles, technical architecture, code quality, and testing methodologies.
- Ability to identify technical flaws and explain their implications clearly and logically.
- Strong analytical and critical-thinking skills, with the ability to assess technical challenges objectively.
- Comfortable discussing complex technical concepts, coding practices, evaluation methodologies, and engineering standards.
- Ability to provide clear, constructive feedback based on practical professional experience.
- Comfortable participating in a remote, structured research interview and sharing detailed technical observations.
Benefits
- Compensation: $75 per hour.
- Paid participation in a remote technical research interview.
- Flexible remote participation from within the United States.
- Opportunity to apply your professional software engineering expertise to AI evaluation research.
- Opportunity to influence how AI agents are tested against realistic software engineering standards.
- Exposure to emerging approaches for benchmarking and evaluating AI coding capabilities.
- A focused engagement that allows experienced engineers to contribute specialized technical feedback without a long-term employment commitment.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1