NewCUA Dataset – long-horizon tasks for agent training

Expert Data for
Frontier Research

Frontier models have read everything ever written. We turn what experts know beyond that corpus into what the next model learns.

NVIDIAAWSGoogle CloudIBMServiceNowDatabricksSnowflakeGumGumTwelve LabsFireworks AIKörberGetYourGuideTaranisFloREM People
Benchmark item Mathematics

Prove or disprove: every sufficiently large even integer is the sum of a prime and a semiprime under the stated constraint.

Written by Number theorist
Result SOTA fails

Item bank · written past every training cutoff

Why SuperAnnotate

Why frontier labs work with us

Expert depth, one platform for every vendor, and operations that own delivery – the three pieces every research data engagement needs.

Experts with real depth

Practicing specialists – physicians, proof-writers, litigators – whose knowledge lives outside any training set. They author the benchmark items that locate where your model breaks, and carry the authority to grade the answers.

BenchmarksAuthored by expertsGraded by experts

One platform, every vendor

Bring every data vendor you work with onto the #1-rated platform. Consensus scoring, inter-annotator agreement, gold sets, item-level audit trails – the same QA depth on every pipeline, whoever runs it.

Open to all vendorsConsensus & IAAItem-level audit

Fully managed operations

Dedicated PMs and AI Ops who own quality and speed on every delivery – scoping, staffing, and item-level tracking, so your researchers stay on research.

PMs & AI OpsScopingOn-time delivery
What we run for labs

Across the full post-training loop

Expert signal across training, agent environments, evaluation, and benchmark design.

  1. Train RLHF & SFT
  2. Agents RL Environments & Agents
  3. Evaluate Evals & Red Teaming
  4. Benchmark Benchmark Design
The Platform

Infrastructure built for research pipelines

Every program above runs on the same platform – interfaces, review layers, and orchestration that reshape around whatever the research calls for.

  • Multimodal, low-code interfaces

    Interfaces fully customized to any task and combination of data types – text, image, audio, or video – in seconds, so complex datasets never wait on tooling.

  • Model-in-the-loop

    Human annotators augmented with AI in one workflow – model pre-labels and flags, experts judge and correct. Faster, more consistent dataset creation.

  • Advanced workflows

    Code execution, database queries, guardrails, and routing between multiple experts – orchestration for pipelines that plain annotation tools can’t express.

From the labs

What researchers say

SuperAnnotate's deep understanding of our research requirements, combined with their ability to recruit highly qualified human reviewers and deliver thoughtful, framework-driven solutions, has super-powered our ability to efficiently build high-quality benchmarks.

Databricks DatabricksResearch Engineer

Combining our domain experts with outsourced teams inside one platform significantly improved our iteration speed. Spotting errors early saved us considerable downstream effort.

ServiceNow ServiceNowApplied Scientist

They stood up a bench of PhD-level reviewers for our agent evaluations in days, not months. Expert throughput simply stopped being our bottleneck.

Frontier LabResearch Scientist
Catalog

Off-the-shelf data for agents, robots, and the models after that.

A growing licensed catalog: computer-use trajectories, RL environments, robotics data. Samples available first.

Browse the catalog

Push past what
today's models can do

Bring us the problems your models still fail – we'll bring the experts, the platform, and the operations.