peopledatabenchmarks

Benchmark · people data

People Search API Benchmark

When you ask a data API for “{role} in {country}”, how many of the returned people actually hold that role, at that seniority, in that country — and how deep and fast is the result? We asked each provider the same 10 frozen role+geo queries and blind-judged every person they returned.

Updated 2026-07-30 · judge openai/gpt-4.1-mini · 10 queries · K=25 · query set v2026-07-30.1· every raw request & judge call is published for verification.

Leaderboard

ProviderPrecision (of returned)Precision@10Depth (avg/25)CoverageLatency
#1Crustdata79.1%77%2510/107,305ms
People Data Labspartial run100%100%52/101,715ms
ContactOutpartial run16.7%8.3%16/10474ms

“Partial run” = the provider’s test account hit a quota or credit limit before finishing all 10 queries; its numbers cover only the queries that ran (see coverage) and are not ranked against complete runs.

Precision of returned results

Share of returned people whose current title genuinely matches the requested function + seniority in the requested country (blind LLM judge).

Crustdata
79.1%
People Data Labs
100%
ContactOut
16.7%

Verdict

Quota-limited this run

Apollo, Coresignal, RocketReach returned no results because the test account hit a free-tier / subscription limit. They are excluded from the leaderboard until re-run on a paid tier — not scored as zero.

Methodology

A frozen, versioned set of 10 queries each fixes a role (with accepted title variants), a seniority bar, and a country. Each provider is asked for up to 25 matching people via its native people-search API, with the query expressed as faithfully as each API allows.

Every returned person is judged blind — the judge (openai/gpt-4.1-mini) sees only current title, company and location, never the provider — on two checks: the current title matches the requested function AND seniority, and the location is in the requested country. Duplicate people across providers are judged once. Field fill-rates (title/company/location present) are deterministic.

Every HTTP request/response and every judge call is cached under runs/ so the leaderboard can be independently re-graded. Judge cost for this run: $0.0223. Providers that hit quota mid-run are labeled partial; providers that could not run at all are excluded, never scored as zero.

Per-query results

QueryCrustdataPeople Data LabsContactOut
vp-sales-us50% · 25
talent-lead-uk70% · 25100% · 2550% · 5
cfo-germany80% · 25
marketing-vp-us90% · 24100% · 25
devops-india100% · 25
founder-ceo-france80% · 250% · 1
product-manager-canada70% · 250% · 1
revops-director-us80% · 250% · 1
data-engineer-uk90% · 250% · 1
cto-australia60% · 250% · 1

cell = Precision@10 · results returned

Part of the PeopleDataBenchmarks open benchmark suite · dataset released CC BY 4.0.