AYN

Remote | Senior Software Engineer - AI Evaluation / Coding Agents — $100–$150/hour

24-MAG · New York, New York, United States · remote

We are sharing a specialised freelance opportunity for experienced software engineers to evaluate and improve advanced AI coding systems through rigorous code review, repository-based testing, rubric development, and engineering-quality assessment.

Selected professionals will work with coding agents across substantial real-world codebases, review generated implementations, identify technical failure modes, and translate expert engineering judgement into structured evaluation signals. The focus is not primarily on building production applications, but on determining whether AI-generated software is correct, robust, maintainable, and aligned with professional engineering standards.

Key Responsibilities

AI-Generated Code Evaluation

- Review code produced by AI coding agents

- Assess implementations for correctness, robustness, and maintainability

- Determine whether agents selected appropriate technical approaches

- Identify subtle implementation errors and engineering weaknesses

- Explain clearly why generated solutions succeed or fail

Rubric & Preference Evaluation

- Design and refine technical evaluation rubrics

- Define criteria for assessing coding-agent performance

- Review and label preference and evaluation data

- Apply consistent quality standards across repeated assessments

- Capture nuanced differences between competing implementations

Repository & Engineering Analysis

- Work within substantial real-world software repositories

- Analyse code changes in their broader architectural context

- Evaluate repository-level behaviour rather than isolated snippets

- Review implementation quality using senior engineering judgement

- Identify recurring failure patterns across coding tasks

Evaluation Infrastructure & Workflows

- Build and maintain pipelines supporting data generation and evaluation

- Improve infrastructure for collecting and reviewing model outputs

- Support scalable evaluation and feedback workflows

- Contribute to reliable processes for repeated technical assessment

- Help translate qualitative engineering judgement into structured systems

Research & Technical Collaboration

- Collaborate with research and engineering teams

- Summarise evaluation findings in clear written reports

- Provide actionable recommendations based on observed model behaviour

- Help improve evaluation methodologies and coding-agent benchmarks

- Communicate complex technical issues in concise, structured language

Ideal Profile

- 5+ years of hands-on software engineering experience

- Strong proficiency in Python, TypeScript, JavaScript, Go, Java, or another major production language

- Experience working in substantial real-world codebases

- Strong code-review and technical-analysis skills

- Ability to assess implementation correctness and maintainability

- Strong understanding of modern software-engineering practices

- Experience with GitHub-based development workflows

- Familiarity with CI/CD and production engineering processes

- Comfortable analysing unfamiliar repositories and code changes

- Ability to identify subtle technical issues others may overlook

- Strong written communication and structured technical reasoning

- Experience using modern LLMs or AI-assisted coding tools

- Familiarity with open-source development is valuable

- LLM evaluation or coding-agent experience is advantageous

- Experience with RLHF, preference data, rubric design, or post-training is useful but not required

Engagement Details

- Independent contractor engagement

- Remote — North America only

- Compensation: $100–$150/hour

- Candidates should provide a specific hourly rate expectation

- Flexible commitment of approximately 20–40 hours per week

- At least 6 hours of Pacific Time overlap per day is required

- Expected project duration is approximately 3 months

- Start is as soon as possible

- Evaluation process includes an approximately 25-minute AI interview

- A practical code and AI-evaluation exercise of approximately 30 minutes follows

- Final stage includes an approximately 20-minute hiring manager interview

- The practical exercise focuses on reviewing AI-generated code rather than competitive programming or algorithm puzzles

- Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, repositories, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy:

Apply on the employer’s site