Where the rows come from
Every row traces back to one of three source kinds. Brands' own store locators, extracted directly rather than taken from a third-party directory. Government registries, such as insured bank branches, certified hospitals, schools, and airports, each carrying its own identifier. Open geospatial data, from Foursquare Open Source Places and Overture Maps, used to fill the category and geographic coverage the first two do not reach.
If a location is not in a brand's own locator, a government registry, or an open dataset, it is not in the dataset. There is no device location data in it, no carrier or app panel, and no location-history modelling.
Every row carries a provenance tier
Each business location carries one of four tiers, shown wherever the row appears.
Brand-extractedBrand-extracted rows come from a brand's own store locator.
Gov-registryGov-registry rows come from a government source and carry that source's own identifier.
Foursquare-derivedFoursquare-derived rows fill branded, unbranded, and sub-location coverage.
OvertureOverture rows are the open-data supplement, used only where the first three tiers do not reach.
The tier tells you where a row came from and how it was verified, not just that it is present. Ground Truth US carries the first two tiers only. The Full US master carries all four.
Fusion and dedup
The same physical location often appears in more than one source: a brand's own locator and an open dataset, or two overlapping registries. Records are matched on name, address, and geographic proximity, then conflated into a single canonical business location when they describe the same place.
The surviving record's fields are chosen by source priority. A brand's own locator outranks an open-data supplement for that brand's own locations. Every canonical location carries a stable, content-derived id and, where one exists, the original source's identifier, so a row can always be traced back to what produced it.
Confidence is scored, not hidden
Field-level confidence follows source agreement. A name, category, or address corroborated across independent sources scores high. A single-source field, or one with conflicting values across sources, scores lower.
Confidence drives internal quality sampling and category assignment. It does not drop or hide rows. The per-field fill rates published on each brand page are the customer-facing view of the same signal.
Quality gates
Every monthly release runs a fixed set of automated tests and release checks before it publishes. They cover the schema contract, per-tier row-count sanity, address and category field validity, id-format compliance, and duplicate detection. A release that fails a gate does not ship, and the store always serves the last release that passed all of them.
Release cadence and change events
Most brand locators are refreshed weekly, so a single brand can be more current than the dataset's own publish cycle. The full dataset is reassembled, conflated, and published once a month. That monthly release is what every product page, price, and sample reads from, so the whole dataset stays point-in-time consistent even though sources refresh on their own schedules.
Change between releases is reported as distinct events: openings, closures, relocations, renames, corrections, and source-only changes. Net location change reconciles to the event totals. Every published metric carries a reproducible query and a release identifier.
What ships with a download
Every download ships with its own NOTICE file, and that file reflects what is actually in the archive. Downloads that include Foursquare-derived or Overture rows carry the upstream open-data license texts alongside it. Government-registry sources are public records, and where an agency publishes its own attribution requirement, that credit travels with the rows.
Ground Truth US carries no Foursquare or Overture attribution, because neither source is in it. The NOTICE is fixed per release, not per purchase.
Known limitations
Brand coverage grows every month but is not exhaustive. Smaller and regional chains are still being onboarded. Independent, non-chain businesses are covered only where a registry or open dataset includes them, so the dataset is strongest on chains and registered institutions.
There is a detection lag between a real-world opening or closing and its appearance in the change feed, bounded by that brand's refresh cadence. Foursquare-derived and Overture rows carry lighter per-field verification than the brand-extracted and gov-registry tiers. That is what the tier system is for: it tells you which guarantee applies to which row, instead of asserting one confidence level over the whole file.
What is not in it
The dataset describes places, not people. It carries no mobile-device identifiers, no movement traces, and no individual-level records, and it is built without location signals. That is true of the dataset and of every product built from it.
Mogean's Exposure Measurement services take a separate input of hashed mobile location signals. The dataset gives that work its context and is never built from it. the Exposure Measurement section describes that boundary in full.