Software Engineers: Paid Interview on AI Evaluation Tasks
Code Reviewsoftware testingtechnical evaluation frameworksfull stack developmentBackend Engineeringtest automationsystem architecture
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for Software Engineers: Paid Interview on AI Evaluation Tasks based in the United States. This is a remote, paid research opportunity for software engineers with hands-on experience evaluating realistic programming tasks and technical systems. You will review coding challenges and the evaluation environments used to assess the performance of AI agents. Your expertise will help determine whether these tasks are technically accurate, appropriately challenging, verifiable, and representative of real-world engineering standards. You will examine evaluation harnesses, walk through their logic, and identify potential technical or structural issues. Your feedback will contribute to improving how AI systems are tested and benchmarked against practical software engineering expectations. The session is designed for experienced technical professionals who can clearly explain their reasoning and assess code quality objectively. This is a flexible opportunity to apply your engineering expertise to the development of more rigorous and realistic AI evaluations. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for Software Engineers: Paid Interview on AI Evaluation Tasks based in the United States. This is a remote, paid research opportunity for software engineers with hands-on experience evaluating realistic programming tasks and technical systems. You will review coding challenges and the evaluation environments used to assess the performance of AI agents. Your expertise will help determine whether these tasks are technically accurate, appropriately challenging, verifiable, and representative of real-world engineering standards. You will examine evaluation harnesses, walk through their logic, and identify potential technical or structural issues. Your feedback will contribute to improving how AI systems are tested and benchmarked against practical software engineering expectations. The session is designed for experienced technical professionals who can clearly explain their reasoning and assess code quality objectively. This is a flexible opportunity to apply your engineering expertise to the development of more rigorous and realistic AI evaluations. Accountabilities Review and assess the quality, accuracy, and realism of programming tasks designed to evaluate AI agents. Evaluate coding environments and technical evaluation harnesses for correctness, robustness, and suitability for AI testing. Examine provided code structures and walk through the underlying logic, identifying potential flaws, inconsistencies, or technical limitations. Assess whether coding challenges accurately reflect realistic software engineering scenarios and industry practices. Evaluate the difficulty and complexity of programming tasks to determine whether they provide meaningful tests of engineering capabilities. Review the verifiability and technical soundness of evaluation criteria and harnesses. Provide clear, detailed feedback on potential improvements to task design, evaluation methodology, and technical implementation. Discuss technical architecture, testing approaches, and software engineering practices during the research session. Share professional perspectives on what makes coding challenges robust, realistic, and technically meaningful. Requirements Professional experience as a software engineer, software developer, or closely related technical professional. Hands-on experience building, reviewing, testing, or evaluating realistic programming tasks. Experience with code review, software testing, automated testing, or technical evaluation frameworks. Familiarity with evaluation harnesses or similar environments used to verify programming solutions. Experience in one or more relevant areas such as full-stack development, backend engineering, test automation, systems architecture, or related software disciplines. Strong understanding of software engineering principles, technical architecture, code quality, and testing methodologies. Ability to identify technical flaws and explain their implications clearly and logically. Strong analytical and critical-thinking skills, with the ability to assess technical challenges objectively. Comfortable discussing complex technical concepts, coding practices, evaluation methodologies, and engineering standards. Ability to provide clear, constructive feedback based on practical professional experience. Comfortable participating in a remote, structured research interview and sharing detailed technical observations. Benefits Compensation: $75 per hour. Paid participation in a remote technical research interview. Flexible remote participation from within the United States. Opportunity to apply your professional software engineering expertise to AI evaluation research. Opportunity to influence how AI agents are tested against realistic software engineering standards. Exposure to emerging approaches for benchmarking and evaluating AI coding capabilities. A focused engagement that allows experienced engineers to contribute specialized technical feedback without a long-term employment commitment. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
Not disclosed
Location
United States
Job type
Contract
Category
AI Engineering
Experience
Mid
Posted
Today
Job Highlights
- Mid level role
- 100% Remote — open to candidates in United States
- Contract position
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.