Skip to content

What we build for AI teams

Proprietary data, custom benchmarks, and human-expert evaluations for any company building serious AI — from frontier labs to enterprise ML teams.

LLM Evals
Evaluation

LLM Evals

Rubric-graded evals for code, math, reasoning, instruction following, and safety — authored by domain experts.

SWE Benchmarks
SWE

SWE Benchmarks

SWE-bench-style agentic coding evals across real repos in Python, TypeScript, Go, Rust, Java, and C++.

Cybersecurity
Security

Cybersecurity

Red-team prompts, exploit traces, CTF transcripts, and SOC triage data authored by practicing security experts.

Code Generation
Code

Code Generation

Multi-language code with executable verification, difficulty banding, and senior-engineer reference solutions.

Math & Reasoning
Math

Math & Reasoning

Olympiad-grade problems and Lean / Coq formal proofs with step-level annotation and verification.

Reasoning Chains
Reasoning

Reasoning Chains

Step-level correctness labels and first-error localization — ready-made training data for process reward models.

Preference Data
RLHF

Preference Data

RLHF / DPO / KTO preference pairs with explicit rubrics and inter-rater reliability reporting.

Agentic Traces
Agents

Agentic Traces

Browser, terminal, IDE, and tool-use traces from experts completing real multi-step work — every step labeled.

Red-Team & Safety
Safety

Red-Team & Safety

Adversarial prompts and jailbreak taxonomies refreshed continuously as new attack classes emerge.

Custom Benchmarks
Benchmarks

Custom Benchmarks

Bespoke leaderboards co-designed with your research team — spec, items, grader, and hosting end to end.

Expert Review
Annotation

Expert Review

PhD-grade and senior-engineer review at scale through Hatch — our vetted contractor network.

Continuous Evals
Pipelines

Continuous Evals

Recurring drops aligned to your release cadence with delta reports across checkpoints.

Backed by 5,000+ vetted experts across multiple domains. Every engagement scoped under an MSA + SOW with named experts, explicit rubrics, and per-item provenance.

From the teams we work with

How research teams describe working with us

Customer identities are abstracted at their request. Named references available under NDA.

NobleStark stood up a 12k-item SWE-agentic eval for us in five weeks. Their senior engineers wrote every reference solution, and the rubric documentation made our internal calibration trivial.
Sarah K.
Head of Evaluations, Frontier AI Lab
We needed Lean 4 proof annotations from people who actually do formal mathematics. Hatch's expert pool delivered — IMO and Putnam medalists who could write and verify the chains, not just label them.
Marcus J.
Research Scientist, Foundation Model Team
The red-team coverage NobleStark builds gets refreshed faster than our model release cadence. New jailbreak classes show up in the eval drop before we see them in the wild.
Emily R.
Safety Eval Lead, Top-5 AI Lab
NobleStark stood up a 12k-item SWE-agentic eval for us in five weeks. Their senior engineers wrote every reference solution, and the rubric documentation made our internal calibration trivial.
Sarah K.
Head of Evaluations, Frontier AI Lab
We needed Lean 4 proof annotations from people who actually do formal mathematics. Hatch's expert pool delivered — IMO and Putnam medalists who could write and verify the chains, not just label them.
Marcus J.
Research Scientist, Foundation Model Team
The red-team coverage NobleStark builds gets refreshed faster than our model release cadence. New jailbreak classes show up in the eval drop before we see them in the wild.
Emily R.
Safety Eval Lead, Top-5 AI Lab
Engagement

Two ways to work with us

Every engagement runs under a signed MSA and a scoped Statement of Work with named experts, explicit rubrics, per-item provenance, and an acceptance protocol your research team controls. Pricing is engagement-specific — talk to us for a quote.

Standard SOW

Most teams start here

Project-scoped engagement. You define the rubric and acceptance criteria — we deliver the dataset, eval, or benchmark.

Typical engagement4–10 weeks · fixed fee
Fixed-scope deliverable: dataset, eval suite, or benchmark
Domain-expert authors matched to your vertical
Pilot batch (typically 5%) before full production
Per-item provenance: author, rubric version, review trail
Inter-rater reliability metrics included in delivery
Acceptance protocol controlled by your research team
Typical turnaround: 4–10 weeks depending on scope

Embedded Expert Pod

Long-running

A dedicated team of vetted experts on retainer, integrated with your research org over a multi-quarter horizon.

Typical engagementMulti-quarter retainer
Dedicated 5–25 expert pod with a named program lead
Embeds into your Slack / Linear / Phabricator workflow
Continuous eval refresh aligned to your release cycle
Calibration tasks, qualifier rounds, ongoing QA
Quarterly capability reviews and rubric evolution
Single contracting entity for compliance, payroll, IP
Priority access to scarce verticals: cybersecurity, formal math, low-resource languages
MSA · DPA · Subprocessors · NDA per project

Tell us about your data needs

Tell us what you're scoping. A solutions engineer replies within one business day.

NDAs available on request