
A frontier AI lab evaluating model reasoning across physics, chemistry, and other STEM domains ran into a difficult problem: standard problem sets were no longer testing what the lab needed to measure.
Frontier models have effectively ingested the public internet, textbooks included. Hand one a lightly adapted version of a known problem, swap a constant, change a compound, and the model can still recognize the shape of the question. The result looks like a new task and functions like an old one.
For meaningful frontier model evaluation, the lab needed:
This wasn’t work for generalist annotators. The lab needed PhD-level researchers who could judge whether the science, sourcing, and solution held up within their own domains.
Athyna Intelligence matched the lab with PhD researchers across math, physics, chemistry, and biology, each with active domain expertise in the areas the lab needed covered. Each accepted task took more than a full working day of research, authoring, computation, and verification.
The researchers developed expert-authored reasoning tasks from literature review through verified gold solution, giving the lab original STEM training data built specifically for frontier model evaluation.
The researchers built expert-authored reasoning tasks from literature review through verified gold solution, giving the lab original STEM training data designed specifically for frontier model evaluation.
Athyna’s researchers built each task from their own literature review rather than adapting a published problem set. There was no prior version to pattern-match against because none existed.
That gave the lab genuinely original evaluation material instead of familiar problems with different inputs.
Athyna’s researchers traced every constant, variable, governing equation, and reaction pathway back to a peer-reviewed publication, complete with a DOI and author attribution. Recent literature was preferred where newer research had updated the underlying science.
That level of sourcing mattered because a scientifically plausible value isn’t necessarily the correct one for a particular application. Using an outdated input or applying it under the wrong conditions could change the gold answer entirely.
Each researcher solved their own task by hand, with no model assistance, and the gold solution was then independently verified.
Tasks were authored in LaTeX and rendered to a code-based submission format, so the work required technical writing fluency alongside domain depth. If an answer couldn’t be derived and checked without a model in the loop, the task wasn’t ready.
Every task had to meet two standards before acceptance:
Athyna’s researchers had to establish both from scratch. Each task then went through multiple rounds of review, with revision cycles on sourcing and specification, before acceptance.
That combination gave the lab what it needed from adversarial STEM benchmarks: original problems with verifiable answers that still tested the limits of current frontier models.


Athyna Intelligence connects AI labs with PhD-level experts who build original, expert-authored reasoning tasks across physics, chemistry, and other STEM domains, with independently verified answers.