Machine Learning Engineer - Model Evaluation & Experimentation

5 days, 4 hours ago
Contract
Junior
Data Science and Analytics
Weekday

Weekday

Weekday helps companies hire engineers who are vouched by other software engineers, enabling passive income for engineers. They offer services like drafting outreach messages, shortlisting candidates, and conducting reference checks. Backed by Y Combin...

Construction & Engineering
11-50
Founded 2020

Description

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
  • Implement machine learning solutions in Python and produce reference implementations that demonstrate correct methodology and expected outcomes.
  • Execute experiments, run training pipelines, and analyze model behavior and results.
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
  • Collaborate with AI researchers and subject matter experts to improve benchmark quality, technical rigor, and evaluation consistency.
  • Document experimental methodologies and technical findings clearly.

Requirements

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
  • Practical experience conducting machine learning experiments, including setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Familiarity with reinforcement learning concepts, including reward functions, policy optimization, and training behavior, is preferred.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and ability to solve complex, open-ended technical problems independently.
  • Ability to commit approximately 35 hours per week on a consistent basis.
  • Experience developing or evaluating large language models, foundation models, or generative AI systems is preferred.
  • Background in reinforcement learning, deep learning, distributed training, or model optimization is preferred.
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation is preferred.
  • Experience contributing to research publications, open-source machine learning projects, or advanced AI systems is preferred.
  • Strong written communication skills for documenting experimental methodologies and technical findings.

Benefits

  • Compensation of $60-$90 per hour.
  • Fully remote engagement.
  • Flexible working hours.
  • Approximately 35 hours per week.
  • Weekly payments based on approved work completed.
  • Potential for project extension depending on requirements and performance.
  • Reasonable accommodations available throughout the application and engagement process.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Software Engineer ML - Contractor position

Janea Systems 11-50 Internet Software & Services

Janea Systems (USA) is seeking a Senior ML Engineer contractor to support remote, client-facing consulting engagements that take enterprise ML and LLM systems from prototype to production.

Apache Airflow AWS Azure CI/CD Dagster Datadog GCP Kubeflow Kubernetes LLM MLOps NLP Prefect Python PyTorch TensorFlow
3 hours, 25 minutes ago

Remote Work From Home: Audio Evaluation Project

CrowdGen by Appen Internet Software & Services

CrowdGen is hiring a remote independent contractor for an audio evaluation project to compare paired recordings and select the better one for AI speech improvement.

3 hours, 40 minutes ago

Maps Personalization Relevance Rater - Gujarati(India)

Welo Global Professional Services

Freelance raters are needed in India to evaluate personalized search and location recommendations for a weekly project using Google Maps history and defined guidelines.

3 hours, 40 minutes ago

Remote Work From Home: Audio Evaluation Project

CrowdGen by Appen Internet Software & Services

CrowdGen is hiring a remote, project-based Independent Contractor to compare pairs of audio recordings for AI training and select the better recording according to evaluation guidelines.

3 hours, 40 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers