AYN

Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

24-MAG · New York, New York, United States · remote

We are sharing a specialised part-time consulting opportunity for experienced software engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research.

Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help research teams understand how advanced models perform throughout realistic software-development workflows.

Key Responsibilities

Code Curation & Solution Development

- Curate high-quality code examples for model training and benchmarking

- Develop precise solutions to software-engineering tasks

- Correct and improve code across multiple programming languages

- Work with Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go

- Maintain strong standards for correctness and maintainability

AI-Generated Code Evaluation

- Evaluate AI-generated code for technical correctness

- Assess solutions for efficiency, scalability, and reliability

- Identify implementation weaknesses and recurring error patterns

- Review code quality against professional engineering standards

- Provide structured rationales supporting evaluation decisions

Verification & Automated Assessment

- Build agents that assess code quality

- Design mechanisms for automatically verifying software solutions

- Identify recurring model-generated coding errors

- Develop reliable checks for engineering tasks

- Support reproducible evaluation across repeated assignments

Software Engineering Lifecycle Evaluation

- Evaluate model capabilities across the software-development lifecycle

- Assess reasoning around prototyping and architecture design

- Review API design and production implementation decisions

- Evaluate launch, experimentation, monitoring, and maintenance scenarios

- Identify areas where models struggle with real-world engineering workflows

Research & Benchmark Collaboration

- Collaborate with research and cross-functional technical teams

- Contribute to datasets used for training and benchmarking

- Help define engineering evaluation strategies

- Compare model performance against professional engineering expectations

- Support iterative improvements to coding-focused evaluation systems

Ideal Profile

- 3+ years of professional software-engineering experience

- Strong full-stack development capabilities

- Experience building scalable, production-grade software

- Strong understanding of software architecture and system design

- Deep knowledge of development, debugging, and code-quality assessment

- Experience reviewing and improving complex software implementations

- Proficiency in one or more of Python, JavaScript, Java, C++, Rust, or related languages

- ReactJS, C, or Go experience may also be relevant to project assignments

- Strong understanding of API design and production implementation

- Familiarity with software monitoring and operational maintenance

- Ability to reason across the complete software-engineering lifecycle

- Strong analytical and problem-solving capabilities

- Excellent written and verbal communication skills

- Ability to provide clear, structured evaluation rationales

- Comfortable collaborating remotely with research and technical teams

Engagement Details

- Part-time independent contractor engagement

- Fully remote

- Candidates must be based in the United States, Canada, or eligible Western European (WEU) countries

- Source examples of WEU locations include Austria, Belgium, France, and Germany

- Minimum commitment: 10 hours per week

- Flexible workload of up to 40 hours per week

- Initial project duration is approximately 1 month

- Extension may be available depending on performance and project fit

- No medical or paid-leave benefits are included under the contractor arrangement

- Application process takes approximately 15–30 minutes

- Completion of an AI video interview is required

- Compensation is not specified in the source materials

- Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy:

Apply on the employer’s site