AI Evaluation Engineer (Remote)

Denmark
Posted 6 days ago
Data Science

About the role

Job summary

This role involves evaluating AI coding agents by creating realistic developer environments and tasks to assess their performance on real-world developer tasks. The position is project-based and fully remote.

Qualifications

  • Minimum of 5 years of experience in software development.
  • Proficiency in English at B2 level or higher.

Responsibilities

  • Develop realistic developer environments, including codebases and infrastructure, to simulate a believable development history.
  • Design tasks based on intermediate states of these environments, defining what constitutes a 'solved' task.
  • Write tests to verify the solutions provided by AI agents, ensuring a balance between strictness and leniency.
  • Iterate on tasks and tests based on quality assurance feedback, refining them to ensure fair evaluations.

Skills

Compensation

  • Strong knowledge of Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis.
  • Experience in writing functional and integration tests.
  • Up to $50 per hour, based on experience and task completion pace, with an estimated 20 hours per task.
Full Access

Ready to apply for this role?

Full Access gives you the company name, full job description, and a direct link to apply. On the 1- and 3-month plans, CV Tailor rewrites your CV for this exact role.

Share this job