What we build for AI teams
Proprietary data, custom benchmarks, and human-expert evaluations for any company building serious AI — from frontier labs to enterprise ML teams.

LLM Evals
Rubric-graded evals for code, math, reasoning, instruction following, and safety — authored by domain experts.

SWE Benchmarks
SWE-bench-style agentic coding evals across real repos in Python, TypeScript, Go, Rust, Java, and C++.

Cybersecurity
Red-team prompts, exploit traces, CTF transcripts, and SOC triage data authored by practicing security experts.

Code Generation
Multi-language code with executable verification, difficulty banding, and senior-engineer reference solutions.

Math & Reasoning
Olympiad-grade problems and Lean / Coq formal proofs with step-level annotation and verification.

Reasoning Chains
Step-level correctness labels and first-error localization — ready-made training data for process reward models.

Preference Data
RLHF / DPO / KTO preference pairs with explicit rubrics and inter-rater reliability reporting.

Agentic Traces
Browser, terminal, IDE, and tool-use traces from experts completing real multi-step work — every step labeled.

Red-Team & Safety
Adversarial prompts and jailbreak taxonomies refreshed continuously as new attack classes emerge.

Custom Benchmarks
Bespoke leaderboards co-designed with your research team — spec, items, grader, and hosting end to end.

Expert Review
PhD-grade and senior-engineer review at scale through Hatch — our vetted contractor network.

Continuous Evals
Recurring drops aligned to your release cadence with delta reports across checkpoints.
Backed by 5,000+ vetted experts across multiple domains. Every engagement scoped under an MSA + SOW with named experts, explicit rubrics, and per-item provenance.
How research teams describe working with us
Customer identities are abstracted at their request. Named references available under NDA.
Two ways to work with us
Every engagement runs under a signed MSA and a scoped Statement of Work with named experts, explicit rubrics, per-item provenance, and an acceptance protocol your research team controls. Pricing is engagement-specific — talk to us for a quote.
Standard SOW
Project-scoped engagement. You define the rubric and acceptance criteria — we deliver the dataset, eval, or benchmark.
Embedded Expert Pod
A dedicated team of vetted experts on retainer, integrated with your research org over a multi-quarter horizon.
Tell us about your data needs
Tell us what you're scoping. A solutions engineer replies within one business day.