Freelance Agent Evaluation Engineer (Remote)

Denmark
Posted 1 day ago
Data Science

About the role

Job summary

This role involves evaluating AI coding agents by creating realistic developer environments and designing tasks that challenge AI models. The position is project-based and fully remote.

Qualifications

  • Over 5 years of experience in software development
  • Proficiency in English at B2 level or higher

Responsibilities

  • Construct realistic developer environments, including codebases and contextual elements
  • Design tasks from intermediate states of these environments, defining success criteria for AI agents
  • Develop tests to verify the solutions provided by AI agents, ensuring a balance between strictness and leniency
  • Refine tasks and tests based on quality assurance feedback, analyzing agent solutions and improving evaluation methods

Skills

  • Strong knowledge of Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis
  • Experience in writing functional and integration tests

Education

  • Relevant degree or equivalent experience in software development is preferred.

Tools

  • Familiarity with AI evaluation frameworks and testing methodologies.
Full Access

Ready to apply for this role?

Full Access gives you the company name, full job description, and a direct link to apply. On the 1- and 3-month plans, CV Tailor rewrites your CV for this exact role.

Share this job