Task Development Engineer

1 hour, 31 minutes ago
Contract
Mid Level
Software Development
METR

METR

METR, or Model Evaluation and Threat Research, is a nonprofit research institute located in Berkeley, California. Founded in August 2022, METR focuses on evaluating advanced AI models to identify capabilities that may pose significant risks to society. The organization conducts pre-deployment empirical evaluations of AI systems, assessing dangerous capabilities such as autonomous replication and cybersecurity threats. METR's mission is to develop scientific methods for assessing the risks associated with AI systems' autonomous capabilities. It provides services to leading AI companies, including OpenAI, Anthropic, and Google DeepMind, helping them understand AI capabilities and risks before deploying new models. The organization also contributes to the development of standardized evaluation methodologies and publishes research to enhance public understanding of AI risks. With a dedicated team, METR aims to promote safe AI development and informed decision-making.

nonprofit organization management
51-200
Founded 2022
$71M raised

Description

  • Develop novel, difficult, well-scoped tasks that remain challenging as AI model time horizons increase.
  • Perform quality assurance on existing tasks to verify solvability, clarity, and appropriate information constraints.
  • Baseline tasks within areas of expertise when useful.
  • Score task completions produced by AI systems or human baseliners.
  • Identify inefficient or low-quality workflows and improve task-development infrastructure and processes.
  • Contribute to evaluations whose results inform policymakers, frontier AI labs, national security stakeholders, and other decision-makers.

Requirements

  • Several years of software engineering experience with complex projects and codebases.
  • Experience building difficult AI evaluations, ideally agent-based evaluations.
  • Familiarity with evaluations such as RE-Bench, HCAST, SWE-bench Verified, Cybench, or GPQA.
  • High attention to detail, including the ability to identify ambiguity, misspecifications, and small errors.
  • Experience with the Inspect evaluation framework preferred.
  • Prior experience with METR’s Hawk infrastructure preferred.
  • Familiarity with the methodology behind METR’s Time Horizons work preferred.
  • Ability to overlap with Pacific Coast Time for at least 1 hour daily, ideally 4 hours.

Benefits

  • Remote, worldwide contract or freelance work.
  • Flexible schedule of 20–40 hours per week.
  • Compensation of $150–$300 per hour, with the top of the range reserved for exceptional candidates.
  • Contributors working more than 80 hours may receive acknowledgment in the final research output, if desired.
  • Collaborative, mission-driven research culture focused on truth-seeking, integrity, and high-quality science.
  • Opportunity to contribute to research informing major decisions about frontier AI risks and progress.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Mathematician with Python Proficiency - AI Trainer

kake 1-10 Internet Software & Services

A company is building a talent pool of Mathematics professionals with Python proficiency to support project-based AI initiatives focused on evaluating and improving frontier AI models.

Linux Python
1 day, 2 hours ago

AI Trainer - Advanced Mandarin Fluency - France

Prolific 51-250 Professional Services

Prolific is hiring remote AI Trainer - Mandarin participants to join live video conversations that help train AI models using real human speech and interaction.

3 days, 1 hour ago

AI Training - Executive Assistant - Memphis, US

Prolific 51-250 Professional Services

Prolific is seeking experienced Executive Assistants to join its expert network and help train and evaluate AI models by documenting real administrative workflows.

3 days, 1 hour ago

AI Trainer - Vietnamese - France

Prolific 51-250 Professional Services

Prolific is hiring fluent Vietnamese speakers for a remote, task-based AI training role centered on live conversation data collection to help improve how AI models understand natural human speech.

Machine Learning
3 days, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers