Benchmark · people data
People Search API Benchmark
When you ask a data API for “{role} in {country}”, how many of the returned people actually hold that role, at that seniority, in that country — and how deep and fast is the result? We asked each provider the same 10 frozen role+geo queries and blind-judged every person they returned.
Updated 2026-07-30 · judge openai/gpt-4.1-mini · 10 queries · K=25 · query set v2026-07-30.1· every raw request & judge call is published for verification.
Leaderboard
| Provider | Precision (of returned) | Precision@10 | Depth (avg/25) | Coverage | Latency |
|---|---|---|---|---|---|
| #1Crustdata | 79.1% | 77% | 25 | 10/10 | 7,305ms |
| People Data Labspartial run | 100% | 100% | 5 | 2/10 | 1,715ms |
| ContactOutpartial run | 16.7% | 8.3% | 1 | 6/10 | 474ms |
“Partial run” = the provider’s test account hit a quota or credit limit before finishing all 10 queries; its numbers cover only the queries that ran (see coverage) and are not ranked against complete runs.
Precision of returned results
Share of returned people whose current title genuinely matches the requested function + seniority in the requested country (blind LLM judge).
Verdict
- Crustdata is the only provider to complete all 10 queries this run — 79.1% precision at full depth (25/25 average results) with 100% title, 100% company and 100% location fill.
- People Data Labs showed excellent precision (100%) on the 2 queries it completed before hitting its plan limit — a re-run on a paid tier will complete the picture.
- ContactOut showed mixed precision (16.7%) on the 6 queries it completed before hitting its plan limit — a re-run on a paid tier will complete the picture.
Quota-limited this run
Apollo, Coresignal, RocketReach returned no results because the test account hit a free-tier / subscription limit. They are excluded from the leaderboard until re-run on a paid tier — not scored as zero.
Methodology
A frozen, versioned set of 10 queries each fixes a role (with accepted title variants), a seniority bar, and a country. Each provider is asked for up to 25 matching people via its native people-search API, with the query expressed as faithfully as each API allows.
Every returned person is judged blind — the judge (openai/gpt-4.1-mini) sees only current title, company and location, never the provider — on two checks: the current title matches the requested function AND seniority, and the location is in the requested country. Duplicate people across providers are judged once. Field fill-rates (title/company/location present) are deterministic.
Every HTTP request/response and every judge call is cached under runs/ so the leaderboard can be independently re-graded. Judge cost for this run: $0.0223. Providers that hit quota mid-run are labeled partial; providers that could not run at all are excluded, never scored as zero.
Per-query results
| Query | Crustdata | People Data Labs | ContactOut |
|---|---|---|---|
| vp-sales-us | 50% · 25 | — | — |
| talent-lead-uk | 70% · 25 | 100% · 25 | 50% · 5 |
| cfo-germany | 80% · 25 | — | — |
| marketing-vp-us | 90% · 24 | 100% · 25 | — |
| devops-india | 100% · 25 | — | — |
| founder-ceo-france | 80% · 25 | — | 0% · 1 |
| product-manager-canada | 70% · 25 | — | 0% · 1 |
| revops-director-us | 80% · 25 | — | 0% · 1 |
| data-engineer-uk | 90% · 25 | — | 0% · 1 |
| cto-australia | 60% · 25 | — | 0% · 1 |
cell = Precision@10 · results returned
Part of the PeopleDataBenchmarks open benchmark suite · dataset released CC BY 4.0.