Evals and post-training data from experts who actually do the work
Verified specialists in finance, tech, healthcare, and legal, custom-sourced per task and identity-verified, with fraud below 1%. We recruit and manage the domain experts who evaluate frontier models, generate post-training data, and audit annotation quality.
Why NewtonX
The human data market is shifting from bulk annotation toward expert-driven evaluation and post-training. NewtonX is built for that, because we already do the hardest part: finding, verifying, and managing niche B2B professionals at scale. Our Knowledge Graph recruitment engine and multi-step verification process source experts who commit to sustained, high-volume work. Not one-off survey responses, but hundreds of hours of judgment-driven evaluation over weeks and months.
Quality evals
Structured human evaluation programs: preference ratings, rubric scoring, red teaming, and benchmark validation, scored by experts whose day jobs match your domain.
Post-training data delivery
Preference pairs, expert demonstrations, and rubric-scored datasets for SFT, RLHF, DPO, RLVR, and agentic RL, delivered on a recurring schedule with QA summaries.
Quality-as-a-Service
An independent audit of the data you already buy or produce: gold sets, calibration, inter-annotator agreement monitoring, and vendor benchmarking.
Overview
We translate your model goals, whether reasoning, hallucination, safety, or agent capability, into evaluation frameworks, rubrics, and dataset specifications.
Our Knowledge Graph finds specialist physicians, engineers, lawyers, and financial professionals, then runs multi-step identity and expertise checks before anyone touches your data.
We handle scheduling, task routing, cohort performance, replacements, and scale-ups, so your team stays on model development.
Calibration sets, inter-annotator agreement monitoring, outlier rejection, and clear escalation paths, recalibrated as your tasks and models evolve.
Start with a
scoped pilot
Bring us the domain your current data can't reach. We scope it as a fixed pilot with clear acceptance criteria, recruit and verify the experts, and return the data with a QA summary.
Fixed scope and acceptance criteria
Experts verified before work starts
QA summary with every delivery


The numbers behind the work
1.1B+
verified professionals in our network
<1%
fraud rate, the lowest in the industry
>0.81
inter-annotator agreement on every batch
96%
of clients start a second project with us