Methodology
How this data is built
Two content layers, one rule for both: don't publish a page unless there's something real to say.
Two layers
Editorial content — company profiles, software analysis, salary bands, city guides, the glossary, and long-form guides and analysis — is researched and written directly. It carries an "updated" date on every page and is revised as facts change.
The knowledge graph — repositories, packages, research papers, researchers, universities, technologies, AI models and datasets, countries, cities, places and job postings — is ingested on a schedule from public external APIs, then run through entity resolution and a data-density score before anything is published.
Entity resolution
The same real-world company, person or paper often shows up in more than one source. Records are matched in order of confidence: an exact external identifier match first (the same GitHub id, DOI or ISO country code seen again), then a domain match for the same entity type, then a case-insensitive name match as a last resort. A second source describing something already in the graph merges into the existing record rather than creating a duplicate page.
What decides whether a page is indexed
Having a record in the database isn't a reason to publish a page for it. Every knowledge-graph entity is scored on how much real content backs it — summary quality and length, how many relationships it has to other entities, a popularity signal, whether a verifiable source URL exists, and whether more than one independent source corroborates it. That score sorts every entity into one of three states:
- Index — real content: published, submitted in our sitemap, and open to search engines.
- Draft — close to the bar but not there yet. Kept and re-scored on every sync, but not published until it earns more data.
- Insufficient data — kept in the database only. Not published.
The same check decides whether a page exists at all, whether it appears in our sitemap, on-site search and "related" links — so none of those can ever point at something that didn't meet the bar. In practice most imported records never qualify: of more than 23,000 imported cities, only a few hundred have a page.
How job postings are kept honest
Job listings carry an extra rule beyond the density score: any posting past its listed expiration date is removed from the site entirely on the next build, regardless of how complete its data otherwise is. A live-looking page for a closed role is worse than no page.
Corrections
Public data is imperfect and goes stale. If something's wrong, tell us which page and what's incorrect — corrections to sourced data are handled ahead of everything else.
Sources
30 external data sources currently active
| Source | Covers | Attribution |
|---|---|---|
| GitHub | repository, organization, person | Data from the GitHub REST API. Not affiliated with or endorsed by GitHub. |
| npm | package | Package metadata from the npm registry. |
| PyPI | package | Package metadata from the Python Package Index (PyPI). |
| crates.io | package | Package metadata from crates.io, the Rust community's crate registry. |
| RubyGems | package | Package metadata from RubyGems.org. |
| Wikidata | company, person, product, organization | Data from Wikidata, available under CC0. |
| arXiv | paper | Thank you to arXiv for use of its open access interoperability. |
| Hugging Face | ai_model, dataset, organization | Model and dataset metadata from the Hugging Face Hub API. |
| OpenAlex | paper, author, institution, journal, topic | Data from OpenAlex (openalex.org), available under CC0. |
| Crossref | paper, author, journal | Publication metadata from Crossref. |
| OpenStreetMap | place, city | Geodata © OpenStreetMap contributors, ODbL. |
| World Bank | country | Indicators from the World Bank Open Data API. |
| SEC EDGAR | company | Filing data from the U.S. SEC EDGAR system. |
| Stack Exchange | topic | Data from the Stack Exchange API, CC BY-SA. |
| Wikimedia / Wikipedia | topic, person, company | Summaries from Wikipedia, CC BY-SA 4.0. |
| GeoNames | city | GeoNames.org, CC BY 4.0. |
| Greenhouse Job Board | job_posting | Job listings via each employer's public Greenhouse job board. |
| Lever Postings | job_posting | Job listings via each employer's public Lever job board. |
| Remotive | job_posting | Remote job listings from Remotive (remotive.com). |
| Arbeitnow | job_posting | Job listings from the Arbeitnow Job Board API. |
| USAJobs | job_posting | Job listings from USAJobs, the U.S. federal government's official job site. |
| Hipolabs University Domains | institution | University directory data from the Hipolabs University Domains API. |
| Adzuna | job_posting | Job listings from the Adzuna Jobs API, aggregated across employer sites and boards. |
| Packagist | package | Package metadata from Packagist, the PHP/Composer package repository. |
| RemoteOK | job_posting | Remote job listings from RemoteOK (remoteok.com). |
| Hacker News | topic | Story data via the Algolia Hacker News Search API. Not affiliated with Y Combinator. |
| Docker Hub | software | Container image metadata from Docker Hub. |
| OSV.dev (vulnerability enrichment) | package | Vulnerability data from the Open Source Vulnerabilities (OSV) database. |
| Y Combinator (yc-oss) | company | Company directory data from Y Combinator, via the yc-oss community mirror. |
| Product Hunt | launch | Launch data from the Product Hunt API. |