Machine Learning Engineer - Model Evaluation & Experimentation

1 week, 5 days ago
Contract
Junior
Data Science and Analytics
Weekday

Weekday

Weekday helps companies hire engineers who are vouched by other software engineers, enabling passive income for engineers. They offer services like drafting outreach messages, shortlisting candidates, and conducting reference checks. Backed by Y Combin...

Construction & Engineering
11-50
Founded 2020

Description

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
  • Implement machine learning solutions in Python and produce reference implementations that demonstrate correct methodology and expected outcomes.
  • Execute experiments, run training pipelines, and analyze model behavior and results.
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
  • Collaborate with AI researchers and subject matter experts to improve benchmark quality, technical rigor, and evaluation consistency.
  • Document experimental methodologies and technical findings clearly.

Requirements

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
  • Practical experience conducting machine learning experiments, including setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Familiarity with reinforcement learning concepts, including reward functions, policy optimization, and training behavior, is preferred.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and ability to solve complex, open-ended technical problems independently.
  • Ability to commit approximately 35 hours per week on a consistent basis.
  • Experience developing or evaluating large language models, foundation models, or generative AI systems is preferred.
  • Background in reinforcement learning, deep learning, distributed training, or model optimization is preferred.
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation is preferred.
  • Experience contributing to research publications, open-source machine learning projects, or advanced AI systems is preferred.
  • Strong written communication skills for documenting experimental methodologies and technical findings.

Benefits

  • Compensation of $60-$90 per hour.
  • Fully remote engagement.
  • Flexible working hours.
  • Approximately 35 hours per week.
  • Weekly payments based on approved work completed.
  • Potential for project extension depending on requirements and performance.
  • Reasonable accommodations available throughout the application and engagement process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Social Media Content Evaluator (Korean)

RWS Group 5K-10K Internet Software & Services

RWS is hiring a freelance, part-time media analyst in South Korea to review short-form videos and generate labeled language data for AI training.

Machine Learning
44 minutes ago

PointClickCare - (US) Senior Clinical AI Nurse SME - 6 months contract

PointClickCare 1K-5K Health Care Providers & Services

PointClickCare is hiring a part-time clinician to support AI-enabled clinical data labeling, evaluation, and consensus-building for healthcare workflows and model quality.

LLM Machine Learning NLP
44 minutes ago

Geography Content Reviewer (AI Evaluation)

Gramian Consultancy Group Professional Services

Gramian Consultancy is hiring a remote Geography Content Reviewer to quality-check geography-focused data and tasks used to evaluate large language models.

44 minutes ago

SWE Fellow - Human Frontier Collective (UK)

Scale AI 251-1K Diversified Consumer Services

The Human Frontier Collective Fellowship at Scale is a fully remote contract role supporting AI research by designing, evaluating, and interpreting advanced generative AI systems with interdisciplinary experts and partner labs.

C++ Generative AI Java JavaScript Machine Learning Python Rust Swift
23 hours, 59 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers