Skip to content
Back to home

Research & notes

How we think about data quality, benchmark design, evaluation methodology, and the changing economics of expert-grade annotation. Written by the NobleStark research and engineering team.

Benchmarks2026-05-018 min read

What we look for in a SWE-Agentic benchmark

How we design repo-level agentic coding evaluations that resist contamination, measure real generalization, and survive the next six months of model progress.

Coming soon
Methodology2026-04-156 min read

Inter-rater reliability is the only metric that matters

Three years of expert-annotation data, one consistent finding: if your reviewers don't agree, your model doesn't learn. How we operationalize Cohen's κ in production pipelines.

Coming soon
Hatch2026-04-025 min read

Why we pay our cybersecurity experts what we pay them

A 5-figure-per-month differential between annotation-farm rates and what we pay practicing red-teamers — and why the data quality difference is even larger.

Coming soon
Benchmarks2026-03-187 min read

Held-out integrity in the age of public benchmarks

Public leaderboards are a contamination magnet. Our contamination-prevention protocol — what we hold back, when we refresh, and how we audit submissions.

Coming soon
RLHF2026-03-046 min read

Process reward models need process-level data

Step-level annotation is more work than outcome-level annotation. It's also where the next 20 points of reasoning-eval performance live.

Coming soon
Methodology2026-02-207 min read

Continuous evaluation as a release primitive

How leading labs treat evaluation drops as a fundamental piece of their release cycle — and how we structure delivery to match.

Coming soon

Get the next post in your inbox

Low-frequency updates from the research and engineering team. No marketing email.

Subscribe via email