


TL;DR: Hiring HLE experts requires matching PhDs and domain experts to the specific subject and evaluation task. They need to validate expert-level questions, verify reference answers, identify errors in model responses, and apply evaluation criteria consistently. Companies should verify domain expertise, use HLE-style work samples, and calibrate evaluators before scaling benchmark work.
Humanity’s Last Exam (HLE) tests advanced AI models on highly specialized academic questions. Evaluating that work requires more than following a general annotation guide. HLE experts need enough subject-matter depth to determine whether a question is valid, its expected answer is defensible, and a model’s response is technically sound.
Hiring for that level of evaluation requires both domain expertise and consistent expert judgment. This guide explains what HLE professionals do, the skills and qualifications they need, who can evaluate advanced AI models, and how companies can screen and hire HLE experts for benchmark development and evaluation.
HLE experts are researchers, senior practitioners, and domain experts who work as subject-matter specialists in advanced AI evaluation. They develop, validate, and evaluate expert-level questions and model responses for benchmarks such as Humanity’s Last Exam.
HLE experts can support several parts of benchmark development and evaluation:
HLE expert describes a category of specialized AI evaluation talent, not a single discipline or standardized job title. Advanced physics questions, for example, require different expertise from questions in law or biology.
General AI annotators typically work from predefined instructions. [AI training roles](https://www.athyna.com/blog/what-is-an-ai-trainer?) can involve more judgment around model outputs, errors, and rubrics. HLE professionals need an additional level of domain expertise to independently validate the question, reference answer, and model response.
Humanity’s Last Exam (HLE) is an AI benchmark that measures how advanced models perform on expert-level academic questions across more than 100 subjects. It contains 2,500 closed-ended questions spanning mathematics, natural sciences, humanities, social sciences, and other specialized fields.
The peer-reviewed Humanity’s Last Exam research published in Nature reports that nearly 1,000 subject-matter experts from more than 500 institutions across 50 countries contributed to the benchmark. HLE emerged as older tests became less useful for comparing advanced models, with leading systems already exceeding 90% accuracy on Massive Multitask Language Understanding (MMLU), a benchmark covering dozens of academic subjects.
HLE questions must have precise, verifiable answers and pass through expert review before inclusion. Models are then evaluated against the accepted reference answers, giving researchers a consistent way to measure performance on expert-level material.
HLE professionals need deep subject-matter expertise, error detection, consistent rubric-based judgment, and clear technical communication. These skills allow HLE benchmark evaluators to independently assess advanced questions and model responses while applying the same evaluation standards across their work.
HLE professionals need expertise that matches the subject they evaluate. They should be able to work through an advanced question independently, verify the expected answer, and judge the model response without treating a supplied answer key as automatically correct.
HLE benchmark evaluators need to catch subtle technical or factual errors, incorrect assumptions, incomplete reasoning, ambiguous questions, and reference answers that require further review.
HLE professionals need to apply defined evaluation criteria consistently across comparable responses and document why an answer passes or fails.
Evaluators need to explain their decisions in a way another subject-matter expert can review. Clear rationales are especially important when reviewers disagree, and a response requires further adjudication.
Advanced AI models should be evaluated by subject-matter experts whose knowledge matches the domain and difficulty of the benchmark. For Humanity’s Last Exam, HLE experts can include PhDs and domain experts across mathematics, physics, chemistry, biology, medicine, computer science, law, and the humanities.
The expertise required depends on what the model is being asked to solve:
Previous AI evaluation experience can help, but it does not establish subject-matter expertise on its own. HLE professionals need enough knowledge of the underlying field to independently verify questions and expected answers, then identify technical or reasoning errors in model responses.
A benchmark project can span several disciplines. In that case, companies may need different specialists for different parts of the evaluation rather than relying on one general evaluator across every subject.
Companies hire Humanity’s Last Exam experts by matching specialists to the exact benchmark domain and evaluation task, verifying their subject-matter depth, screening them with HLE-style work, and calibrating expert judgments before increasing evaluation volume. This helps ensure the people evaluating advanced models have both the domain knowledge and consistency the benchmark requires.
Start with the subject the model is being evaluated on. Broad searches for “AI experts” or even “science, technology, engineering, and mathematics (STEM) experts” may miss the depth required for a specialized question set.
Match candidates across two dimensions:
These tasks draw on overlapping skills, but the level and type of expertise required can differ.
Look for evidence that candidates can work independently at the required level. Relevant signals include academic, research, or professional credentials, specialization in the subject, research or publication experience where relevant, and strong written communication.
For HLE work, credentials are a useful starting point rather than the final screen. Candidates still need to demonstrate that their expertise translates into the specific evaluation task.
The same principle applies to adjacent domain-expert AI roles. Athyna’s guide to hiring a STEM AI trainer explains how subject expertise, evaluation ability, and role requirements come together in broader AI training work.
Use a work sample that reflects the judgments the candidate will make on the project. Give them an expert-level question, a proposed reference answer, and a model response, then ask them to:
This tests whether subject-matter expertise translates into the question validation, answer verification, and model assessment required from HLE benchmark evaluators.
Have multiple qualified experts review the same controlled sample. Compare where their judgments differ, refine the rubric, clarify edge cases, and establish how disputed evaluations will be adjudicated.
Calibration helps teams determine whether experts are applying the evaluation criteria consistently before those judgments are repeated across a larger dataset.
For advanced AI benchmarks, evaluation quality should be established before volume increases. Validate the expertise, criteria, and review process first, then scale the benchmark work.
Latin America can be a strong sourcing market when AI evaluation projects require PhDs and domain experts across specialized fields. For HLE-style work, this can help teams access expertise across mathematics, physics, chemistry, biology, computer science, law, and the humanities as benchmark needs evolve.
Latin America is also well suited to projects where expert needs change throughout the evaluation process. Teams can bring in domain specialists, add evaluation capacity as needed, and collaborate with US research teams during overlapping working hours. That overlap is particularly useful for rubric calibration, ambiguous question review, and adjudication.
Hiring HLE experts starts with matching the right subject-matter expertise to the work being evaluated. Athyna Intelligence connects AI labs and research teams with PhDs and domain experts across Latin America, matched to the subject and evaluation task their project requires. Candidates are screened for domain credentials and English fluency, with qualified specialists matched within 48 to 72 hours.
Experts work within the team’s existing evaluation process and can support question development, reference answer verification, rubric-based grading, model evaluation, and domain-specific benchmark design. For private evaluation projects, they can also create and grade original HLE-style questions in the specific subjects a team needs to test.
As evaluation needs change, teams can adjust the expertise, hours, and tasks required for the project. Find HLE experts matched to your benchmark domain and evaluation needs.
An HLE expert is a subject-matter specialist who helps evaluate advanced AI models on expert-level questions. They may develop questions, verify reference answers, assess model responses, and identify ambiguity or technical errors.
HLE experts need deep knowledge in their specific field, strong error detection, consistent evaluation judgment, and clear written communication. General AI annotation experience alone is not enough for advanced benchmark work.
Humanity's Last Exam should be evaluated by specialists whose expertise matches the benchmark domain. This can include PhDs, researchers, and senior practitioners in fields such as mathematics, physics, biology, computer science, law, and the humanities.
The strongest screening method is an HLE-style work sample. Candidates should validate an expert-level question, verify the reference answer, assess a model response, identify errors or ambiguity, and explain their reasoning.
Calibration helps ensure multiple experts apply the same evaluation criteria consistently. Before scaling, teams should compare expert judgments on a shared sample, clarify edge cases, and establish a process for resolving disagreements.
