Empty pale yellow rectangle with rounded corners and a thin purple border.
BLOG
Case Study

How to Hire HLE Experts: Skills, Screening, and AI Evaluation

September 4, 2026
VectorVector

Table of Content

Industry
Stage
Country

TL;DR: Hiring HLE experts requires matching PhDs and domain experts to the specific subject and evaluation task. They need to validate expert-level questions, verify reference answers, identify errors in model responses, and apply evaluation criteria consistently. Companies should verify domain expertise, use HLE-style work samples, and calibrate evaluators before scaling benchmark work.

Humanity’s Last Exam (HLE) tests advanced AI models on highly specialized academic questions. Evaluating that work requires more than following a general annotation guide. HLE experts need enough subject-matter depth to determine whether a question is valid, its expected answer is defensible, and a model’s response is technically sound.

Hiring for that level of evaluation requires both domain expertise and consistent expert judgment. This guide explains what HLE professionals do, the skills and qualifications they need, who can evaluate advanced AI models, and how companies can screen and hire HLE experts for benchmark development and evaluation.

What Do HLE Experts Actually Do?

HLE experts are researchers, senior practitioners, and domain experts who work as subject-matter specialists in advanced AI evaluation. They develop, validate, and evaluate expert-level questions and model responses for benchmarks such as Humanity’s Last Exam.

HLE experts can support several parts of benchmark development and evaluation:

  • Expert-level question development: Write closed-ended questions that challenge advanced models while remaining precise and answerable.
  • Reference answer verification: Independently confirm that the expected answer is accurate and defensible.
  • Question validation: Identify ambiguity, flawed assumptions, or other issues that could make a benchmark question unreliable.
  • Rubric-based grading: Apply defined evaluation criteria consistently to model responses.
  • Model response evaluation: Assess whether an answer is correct and its reasoning is sound.
  • Expert review and adjudication: Resolve disagreements between evaluators and review edge cases.
  • Benchmark quality control: Catch issues that could affect the reliability of the evaluation.

HLE expert describes a category of specialized AI evaluation talent, not a single discipline or standardized job title. Advanced physics questions, for example, require different expertise from questions in law or biology.

General AI annotators typically work from predefined instructions. [AI training roles](https://www.athyna.com/blog/what-is-an-ai-trainer?) can involve more judgment around model outputs, errors, and rubrics. HLE professionals need an additional level of domain expertise to independently validate the question, reference answer, and model response.

What Is Humanity’s Last Exam and How Is It Evaluated?

Humanity’s Last Exam (HLE) is an AI benchmark that measures how advanced models perform on expert-level academic questions across more than 100 subjects. It contains 2,500 closed-ended questions spanning mathematics, natural sciences, humanities, social sciences, and other specialized fields.

The peer-reviewed Humanity’s Last Exam research published in Nature reports that nearly 1,000 subject-matter experts from more than 500 institutions across 50 countries contributed to the benchmark. HLE emerged as older tests became less useful for comparing advanced models, with leading systems already exceeding 90% accuracy on Massive Multitask Language Understanding (MMLU), a benchmark covering dozens of academic subjects.

HLE questions must have precise, verifiable answers and pass through expert review before inclusion. Models are then evaluated against the accepted reference answers, giving researchers a consistent way to measure performance on expert-level material.

What Skills Do HLE Professionals Need?

HLE professionals need deep subject-matter expertise, error detection, consistent rubric-based judgment, and clear technical communication. These skills allow HLE benchmark evaluators to independently assess advanced questions and model responses while applying the same evaluation standards across their work.

1. Deep Domain Expertise

HLE professionals need expertise that matches the subject they evaluate. They should be able to work through an advanced question independently, verify the expected answer, and judge the model response without treating a supplied answer key as automatically correct.

2. Error and Ambiguity Detection

HLE benchmark evaluators need to catch subtle technical or factual errors, incorrect assumptions, incomplete reasoning, ambiguous questions, and reference answers that require further review.

3. Consistent Rubric-Based Evaluation

HLE professionals need to apply defined evaluation criteria consistently across comparable responses and document why an answer passes or fails.

4. Clear Technical Communication

Evaluators need to explain their decisions in a way another subject-matter expert can review. Clear rationales are especially important when reviewers disagree, and a response requires further adjudication.

Who Can Evaluate Advanced AI Models?

Advanced AI models should be evaluated by subject-matter experts whose knowledge matches the domain and difficulty of the benchmark. For Humanity’s Last Exam, HLE experts can include PhDs and domain experts across mathematics, physics, chemistry, biology, medicine, computer science, law, and the humanities.

The expertise required depends on what the model is being asked to solve:

Benchmark domain Relevant HLE expertise
Mathematics Advanced mathematical reasoning, proofs, and problem validation
Physics Theoretical or applied expertise in the relevant physics subfield
Chemistry Advanced chemical principles and specialized problem solving
Biology & medicine Life sciences, biomedical, or medical expertise relevant to the question
Computer science Algorithms, theory, and advanced technical reasoning
Law & humanities Specialized disciplinary research and analysis

Previous AI evaluation experience can help, but it does not establish subject-matter expertise on its own. HLE professionals need enough knowledge of the underlying field to independently verify questions and expected answers, then identify technical or reasoning errors in model responses.

A benchmark project can span several disciplines. In that case, companies may need different specialists for different parts of the evaluation rather than relying on one general evaluator across every subject.

How Do Companies Hire Humanity’s Last Exam Experts?

Companies hire Humanity’s Last Exam experts by matching specialists to the exact benchmark domain and evaluation task, verifying their subject-matter depth, screening them with HLE-style work, and calibrating expert judgments before increasing evaluation volume. This helps ensure the people evaluating advanced models have both the domain knowledge and consistency the benchmark requires.

1. Match Expertise to the Domain and Task

Start with the subject the model is being evaluated on. Broad searches for “AI experts” or even “science, technology, engineering, and mathematics (STEM) experts” may miss the depth required for a specialized question set.

Match candidates across two dimensions:

  • Domain: Match expertise as closely as possible to the subject being tested. An organic chemistry benchmark, for example, calls for more specific knowledge than a general chemistry background.
  • Task: Determine whether the expert will develop questions, verify reference answers, grade model responses, develop rubrics, validate benchmark items, or adjudicate disagreements.

These tasks draw on overlapping skills, but the level and type of expertise required can differ.

2. Verify Subject-Matter Depth

Look for evidence that candidates can work independently at the required level. Relevant signals include academic, research, or professional credentials, specialization in the subject, research or publication experience where relevant, and strong written communication.

For HLE work, credentials are a useful starting point rather than the final screen. Candidates still need to demonstrate that their expertise translates into the specific evaluation task.

The same principle applies to adjacent domain-expert AI roles. Athyna’s guide to hiring a STEM AI trainer explains how subject expertise, evaluation ability, and role requirements come together in broader AI training work.

3. Screen Candidates With HLE-Style Work

Use a work sample that reflects the judgments the candidate will make on the project. Give them an expert-level question, a proposed reference answer, and a model response, then ask them to:

  1. Determine whether the question is valid and sufficiently precise.
  2. Independently solve or verify the expected answer.
  3. Evaluate the model response.
  4. Identify ambiguity, flawed assumptions, or subtle domain errors.
  5. Explain the reasoning behind their judgment.

This tests whether subject-matter expertise translates into the question validation, answer verification, and model assessment required from HLE benchmark evaluators.

4. Calibrate HLE Benchmark Evaluators Before Scaling

Have multiple qualified experts review the same controlled sample. Compare where their judgments differ, refine the rubric, clarify edge cases, and establish how disputed evaluations will be adjudicated.

Calibration helps teams determine whether experts are applying the evaluation criteria consistently before those judgments are repeated across a larger dataset.

For advanced AI benchmarks, evaluation quality should be established before volume increases. Validate the expertise, criteria, and review process first, then scale the benchmark work.

When Latin America Is a Strong Sourcing Market for AI Evaluation Experts

Latin America can be a strong sourcing market when AI evaluation projects require PhDs and  ⁠domain experts across specialized fields. For HLE-style work, this can help teams access expertise across mathematics, physics, chemistry, biology, computer science, law, and the humanities as benchmark needs evolve.

Latin America is also well suited to projects where expert needs change throughout the evaluation process. Teams can bring in domain specialists, add evaluation capacity as needed, and collaborate with US research teams during overlapping working hours. That overlap is particularly useful for rubric calibration, ambiguous question review, and adjudication.

Match HLE Experts to Your Evaluation Needs With Athyna Intelligence

Hiring HLE experts starts with matching the right subject-matter expertise to the work being evaluated. Athyna Intelligence connects AI labs and research teams with PhDs and domain experts across Latin America, matched to the subject and evaluation task their project requires. Candidates are screened for domain credentials and English fluency, with qualified specialists matched within 48 to 72 hours.

Experts work within the team’s existing evaluation process and can support question development, reference answer verification, rubric-based grading, model evaluation, and domain-specific benchmark design. For private evaluation projects, they can also create and grade original HLE-style questions in the specific subjects a team needs to test.

As evaluation needs change, teams can adjust the expertise, hours, and tasks required for the project. Find HLE experts matched to your benchmark domain and evaluation needs.

Role
Typical US Salary
With Athyna
Athyna Content Team

Frequently asked questions

What is an HLE expert?

An HLE expert is a subject-matter specialist who helps evaluate advanced AI models on expert-level questions. They may develop questions, verify reference answers, assess model responses, and identify ambiguity or technical errors.

What skills do HLE experts need?

HLE experts need deep knowledge in their specific field, strong error detection, consistent evaluation judgment, and clear written communication. General AI annotation experience alone is not enough for advanced benchmark work.

Who can evaluate Humanity's Last Exam?

Humanity's Last Exam should be evaluated by specialists whose expertise matches the benchmark domain. This can include PhDs, researchers, and senior practitioners in fields such as mathematics, physics, biology, computer science, law, and the humanities.

How do companies screen HLE experts?

The strongest screening method is an HLE-style work sample. Candidates should validate an expert-level question, verify the reference answer, assess a model response, identify errors or ambiguity, and explain their reasoning.

Why is evaluator calibration important for HLE projects?

Calibration helps ensure multiple experts apply the same evaluation criteria consistently. Before scaling, teams should compare expert judgments on a shared sample, clarify edge cases, and establish a process for resolving disagreements.

More articles like this

Talk to us

Let's match you with the right talent

Fill this form and we’ll get in touch with you 🚀
Please enter a valid business email
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Download logo as SVG
Download logo as PNG
Downloaded!