AI Evaluation Engineer (Remote)

Denmark
Posted 7 hours, 2 minutes ago
Engineering

About the role

Job summary

This role involves evaluating AI coding agents by creating realistic developer environments and tasks to assess their performance on real-world developer tasks. The position is project-based and fully remote.

Qualifications

  • Over 5 years of experience in software development.
  • Proficient in Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis.
  • Experience in writing functional and integration tests.
  • English proficiency at B2 level or higher.

Responsibilities

  • Develop realistic developer environments that simulate a virtual company with a codebase and context.
  • Design tasks that challenge AI agents, defining what constitutes a 'solved' task.
  • Create tests to verify the solutions provided by AI agents, ensuring a balance between acceptance of valid approaches and rejection of incorrect ones.
  • Refine tasks and tests based on quality assurance feedback, analyzing agent solutions and iterating for fairness and robustness.

Skills

Compensation

  • Strong understanding of AI model limitations and capabilities in coding tasks.
  • Ability to craft complex evaluation criteria and tasks that effectively challenge AI agents.
  • Up to $50 per hour, depending on experience and task completion pace. Tasks are estimated to take around 20 hours each, with flexible scheduling.
Full Access

Ready to apply for this role?

Full Access gives you the company name, full job description, and a direct link to apply. On the 1- and 3-month plans, CV Tailor rewrites your CV for this exact role.

Share this job