AI Evaluation Specialist

1 week, 3 days ago
Contract
Junior
Artificial Intelligence and Machine Learning
Weekday

Weekday

Weekday helps companies hire engineers who are vouched by other software engineers, enabling passive income for engineers. They offer services like drafting outreach messages, shortlisting candidates, and conducting reference checks. Backed by Y Combin...

Construction & Engineering
11-50
Founded 2020

Description

  • Review AI-generated content for accuracy, logical reasoning, completeness, and clarity.
  • Identify factual errors, reasoning gaps, inconsistencies, and unsupported conclusions in AI outputs.
  • Evaluate responses using structured assessment frameworks and detailed quality guidelines.
  • Write clear, concise, evidence-based rationales for evaluation decisions.
  • Highlight both strengths and areas for improvement in AI-generated responses.
  • Apply consistent judgment across a wide range of evaluation tasks.
  • Follow project instructions and standardized assessment criteria to ensure objective, reproducible evaluations.
  • Complete assignments independently while maintaining high quality standards.

Requirements

  • Bachelor's degree from a globally recognized university, with top-ranked institutions preferred.
  • Excellent analytical thinking and problem-solving abilities.
  • Strong written communication skills with the ability to explain complex reasoning clearly and precisely.
  • Exceptional critical reading skills, including the ability to identify nuanced arguments, implicit meaning, logical inconsistencies, missing context, and weak or unsupported reasoning.
  • Strong attention to detail and ability to consistently apply structured evaluation guidelines.
  • Ability to work independently and manage assigned tasks efficiently.
  • Native-level English fluency.
  • Experience in content evaluation, research, quality assurance, editing, or analytical review (preferred).
  • Familiarity with artificial intelligence, large language models, or AI evaluation methodologies (preferred).
  • Experience working with structured annotation or assessment frameworks (preferred).

Benefits

  • Compensation of $70 per hour.
  • Fully remote work with flexible working hours.
  • Work completed on your own schedule.
  • Weekly payments processed through supported payment platforms.
  • Independent contractor engagement.
  • Opportunity to contribute to the development of next-generation AI technologies.
  • Reasonable accommodations are available upon request.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Hydrus -Voice Contributor — English (US)

Welo Global Professional Services

Welo Data is hiring a freelance native-level English (US) speaker to record conversational customer-service audio for a remote AI voice project.

1 day, 13 hours ago

AI / ML Consultant

Crosslake 251-1K Capital Markets

Crosslake is hiring a US-based remote consultant to help technology and private equity clients evaluate, improve, and implement AI and software-related initiatives across the investment lifecycle.

Agile Computer Vision Generative AI LLM Machine Learning MLOps NLP Statistics
1 day, 13 hours ago

Hydrus -Voice Contributor — Spanish (Spain)

Welo Global Professional Services

Welo Data is hiring a freelance native-level Spanish (Spain) speaker to record conversational customer-service voice data for a remote conversational AI project.

1 day, 13 hours ago

Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Mindrift.ai: Be the “I” in AI Internet Software & Services

Mindrift is seeking an experienced software development specialist to create realistic coding-agent evaluation tasks and tests for project-based AI opportunities.

Docker FastAPI JavaScript Kafka PostgreSQL Python React Redis TypeScript
2 days, 13 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers