Skip to content
The Internet Compass

Methodology

How this data is built

Two content layers, one rule for both: don't publish a page unless there's something real to say.

Two layers

Editorial content — company profiles, software analysis, salary bands, city guides, the glossary, and long-form guides and analysis — is researched and written directly. It carries an "updated" date on every page and is revised as facts change.

The knowledge graph — repositories, packages, research papers, researchers, universities, technologies, AI models and datasets, countries, cities, places and job postings — is ingested on a schedule from public external APIs, then run through entity resolution and a data-density score before anything is published.

Entity resolution

The same real-world company, person or paper often shows up in more than one source. Records are matched in order of confidence: an exact external identifier match first (the same GitHub id, DOI or ISO country code seen again), then a domain match for the same entity type, then a case-insensitive name match as a last resort. A second source describing something already in the graph merges into the existing record rather than creating a duplicate page.

What decides whether a page is indexed

Having a record in the database isn't a reason to publish a page for it. Every knowledge-graph entity is scored on how much real content backs it — summary quality and length, how many relationships it has to other entities, a popularity signal, whether a verifiable source URL exists, and whether more than one independent source corroborates it. That score sorts every entity into one of three states:

  • Index — real content: published, submitted in our sitemap, and open to search engines.
  • Draft — close to the bar but not there yet. Kept and re-scored on every sync, but not published until it earns more data.
  • Insufficient data — kept in the database only. Not published.

The same check decides whether a page exists at all, whether it appears in our sitemap, on-site search and "related" links — so none of those can ever point at something that didn't meet the bar. In practice most imported records never qualify: of more than 23,000 imported cities, only a few hundred have a page.

How job postings are kept honest

Job listings carry an extra rule beyond the density score: any posting past its listed expiration date is removed from the site entirely on the next build, regardless of how complete its data otherwise is. A live-looking page for a closed role is worse than no page.

Corrections

Public data is imperfect and goes stale. If something's wrong, tell us which page and what's incorrect — corrections to sourced data are handled ahead of everything else.

Sources

30 external data sources currently active

SourceCoversAttribution
GitHubrepository, organization, personData from the GitHub REST API. Not affiliated with or endorsed by GitHub.
npmpackagePackage metadata from the npm registry.
PyPIpackagePackage metadata from the Python Package Index (PyPI).
crates.iopackagePackage metadata from crates.io, the Rust community's crate registry.
RubyGemspackagePackage metadata from RubyGems.org.
Wikidatacompany, person, product, organizationData from Wikidata, available under CC0.
arXivpaperThank you to arXiv for use of its open access interoperability.
Hugging Faceai_model, dataset, organizationModel and dataset metadata from the Hugging Face Hub API.
OpenAlexpaper, author, institution, journal, topicData from OpenAlex (openalex.org), available under CC0.
Crossrefpaper, author, journalPublication metadata from Crossref.
OpenStreetMapplace, cityGeodata © OpenStreetMap contributors, ODbL.
World BankcountryIndicators from the World Bank Open Data API.
SEC EDGARcompanyFiling data from the U.S. SEC EDGAR system.
Stack ExchangetopicData from the Stack Exchange API, CC BY-SA.
Wikimedia / Wikipediatopic, person, companySummaries from Wikipedia, CC BY-SA 4.0.
GeoNamescityGeoNames.org, CC BY 4.0.
Greenhouse Job Boardjob_postingJob listings via each employer's public Greenhouse job board.
Lever Postingsjob_postingJob listings via each employer's public Lever job board.
Remotivejob_postingRemote job listings from Remotive (remotive.com).
Arbeitnowjob_postingJob listings from the Arbeitnow Job Board API.
USAJobsjob_postingJob listings from USAJobs, the U.S. federal government's official job site.
Hipolabs University DomainsinstitutionUniversity directory data from the Hipolabs University Domains API.
Adzunajob_postingJob listings from the Adzuna Jobs API, aggregated across employer sites and boards.
PackagistpackagePackage metadata from Packagist, the PHP/Composer package repository.
RemoteOKjob_postingRemote job listings from RemoteOK (remoteok.com).
Hacker NewstopicStory data via the Algolia Hacker News Search API. Not affiliated with Y Combinator.
Docker HubsoftwareContainer image metadata from Docker Hub.
OSV.dev (vulnerability enrichment)packageVulnerability data from the Open Source Vulnerabilities (OSV) database.
Y Combinator (yc-oss)companyCompany directory data from Y Combinator, via the yc-oss community mirror.
Product HuntlaunchLaunch data from the Product Hunt API.