Proof

Methodology

Mogean sells four things: a dataset of United States places, geospatial analytics computed against it, out-of-home Exposure Measurement, and a natural-language query layer over the catalog. This page is the record of how each one is built, what it is built from, and where its limits are.

How to read this page

Each section answers the same three questions in the same order: what the inputs are, what is done to them, and what the result does not cover. The four capabilities use different inputs and carry different guarantees, so read the section for the capability you are buying. Where a statement is narrower than it looks, the section says so rather than leaving it to be assumed.

  • 01 Inputs
  • 02 What is done
  • 03 What it does not cover

Places

The POI dataset

Where the rows come from

Every row traces back to one of three source kinds. Brands' own store locators, extracted directly rather than taken from a third-party directory. Government registries, such as insured bank branches, certified hospitals, schools, and airports, each carrying its own identifier. Open geospatial data, from Foursquare Open Source Places and Overture Maps, used to fill the category and geographic coverage the first two do not reach.

If a location is not in a brand's own locator, a government registry, or an open dataset, it is not in the dataset. There is no device location data in it, no carrier or app panel, and no location-history modelling.

Every row carries a provenance tier

Each business location carries one of four tiers, shown wherever the row appears.

Brand-extracted

Brand-extracted rows come from a brand's own store locator.

Gov-registry

Gov-registry rows come from a government source and carry that source's own identifier.

Foursquare-derived

Foursquare-derived rows fill branded, unbranded, and sub-location coverage.

Overture

Overture rows are the open-data supplement, used only where the first three tiers do not reach.

The tier tells you where a row came from and how it was verified, not just that it is present. Ground Truth US carries the first two tiers only. The Full US master carries all four.

Fusion and dedup

The same physical location often appears in more than one source: a brand's own locator and an open dataset, or two overlapping registries. Records are matched on name, address, and geographic proximity, then conflated into a single canonical business location when they describe the same place.

The surviving record's fields are chosen by source priority. A brand's own locator outranks an open-data supplement for that brand's own locations. Every canonical location carries a stable, content-derived id and, where one exists, the original source's identifier, so a row can always be traced back to what produced it.

Confidence is scored, not hidden

Field-level confidence follows source agreement. A name, category, or address corroborated across independent sources scores high. A single-source field, or one with conflicting values across sources, scores lower.

Confidence drives internal quality sampling and category assignment. It does not drop or hide rows. The per-field fill rates published on each brand page are the customer-facing view of the same signal.

Quality gates

Every monthly release runs a fixed set of automated tests and release checks before it publishes. They cover the schema contract, per-tier row-count sanity, address and category field validity, id-format compliance, and duplicate detection. A release that fails a gate does not ship, and the store always serves the last release that passed all of them.

Release cadence and change events

Most brand locators are refreshed weekly, so a single brand can be more current than the dataset's own publish cycle. The full dataset is reassembled, conflated, and published once a month. That monthly release is what every product page, price, and sample reads from, so the whole dataset stays point-in-time consistent even though sources refresh on their own schedules.

Change between releases is reported as distinct events: openings, closures, relocations, renames, corrections, and source-only changes. Net location change reconciles to the event totals. Every published metric carries a reproducible query and a release identifier.

What ships with a download

Every download ships with its own NOTICE file, and that file reflects what is actually in the archive. Downloads that include Foursquare-derived or Overture rows carry the upstream open-data license texts alongside it. Government-registry sources are public records, and where an agency publishes its own attribution requirement, that credit travels with the rows.

Ground Truth US carries no Foursquare or Overture attribution, because neither source is in it. The NOTICE is fixed per release, not per purchase.

Known limitations

Brand coverage grows every month but is not exhaustive. Smaller and regional chains are still being onboarded. Independent, non-chain businesses are covered only where a registry or open dataset includes them, so the dataset is strongest on chains and registered institutions.

There is a detection lag between a real-world opening or closing and its appearance in the change feed, bounded by that brand's refresh cadence. Foursquare-derived and Overture rows carry lighter per-field verification than the brand-extracted and gov-registry tiers. That is what the tier system is for: it tells you which guarantee applies to which row, instead of asserting one confidence level over the whole file.

What is not in it

The dataset describes places, not people. It carries no mobile-device identifiers, no movement traces, and no individual-level records, and it is built without location signals. That is true of the dataset and of every product built from it.

Mogean's Exposure Measurement services take a separate input of hashed mobile location signals. The dataset gives that work its context and is never built from it. the Exposure Measurement section describes that boundary in full.

Analysis

Geospatial analytics

Every figure is computed against a named release

Analytics runs against a named monthly release of the places dataset, not a live scrape and not a blend of vintages. The release identifier travels with the result. The same question against the same release returns the same answer.

There is a reproducible query behind every figure

Each figure in a report is the output of a query we can hand you, run against the release named beside it. If a number changes between reads, the reason is a new release or a changed question, and the report says which.

What we analyze

Trade areas and catchments, for a location, a brand, or a category.

Co-location and competitive density: which brands cluster, which avoid each other, and where the gaps between them fall.

White space and site scoring across markets, using place attributes and geospatial context.

Change analysis: openings, closures, relocations, and renames by month, reconciled to release totals.

Out-of-home exposure measurement is the flagship service family inside this pillar, and the Exposure Measurement section describes it.

What analytics is built on

Place-level analytics use place data only: the location, category, brand, operating status, and provenance of places, plus public geographic reference data. Exposure Measurement is the exception, and its input and boundary are set out in the Exposure Measurement section. Where an engagement uses any additional source, that source and its licensing are named in the statement of work.

How it is delivered

As a written report with the query behind every figure. As a recurring feed into your own tools. Or as an evaluation cut scoped to your footprint.

Custom analytics engagements are contact-led, and there is no published price for analytics work. Self-serve dataset purchases are priced in the store; analytics work is scoped with you first.

Out-of-home

Exposure Measurement

What is measured

Out-of-home advertising runs in the physical world, so we measure it there. Mogean matches advertising screen interactions against mobile device observations to identify the devices that could see a campaign. It then follows those devices to the advertiser's locations.

Devices are matched to screens on proximity, trajectory, velocity, and timing. Devices are matched to destinations on the same signals plus dwell time, which separates visitors from passers-by.

Exposure is resolved per screen

Exposure is resolved per screen, per creative, per venue, and per timestamp. Every device carries that context through to delivery, so a result can be read down to the individual screen rather than only at market level.

The same pipeline runs one screen and one destination, or tens of thousands of screens and thousands of venues. The method does not change with the size of the campaign.

Lift is measured against a comparable unexposed group

Mogean assembles the exposed audience, then assembles a statistically comparable unexposed control group, then measures visitation on both sides. The difference between the two visitation rates is the incremental lift.

Results break out by screen, market, screen type, creative, and destination. Measurement is cumulative, so a mid-campaign read and a final read sit on the same basis.

Results are measured on the full exposed population, not modelled up from a small panel of devices.

When a client brings an exposed device list from another source, Mogean maps the destinations, matches the visits, builds the control group, and measures lift using the identical method.

Two inputs, two jobs

Places

The POI dataset is built from brand locators, government registries, and open geospatial data. It describes places, not people, and carries no device identifiers or movement traces. That is the data sold in the store, and it is built without location signals.

Exposure

Measuring an out-of-home campaign takes a second input: hashed mobile location signals, handled under separate agreements, with our opt-out honoured at mogean.com/opt-out. Inside the measurement services, the POI dataset is what gives those signals their context: which place a visit landed at, and what kind of place it is.

How the signals are handled

Mogean receives these signals under separate agreements covering their source, their permitted uses, and the opt-outs that must be honoured. Mogean does not collect them from this website.

Inside the processing pipeline the signals are hashed. Foot traffic reports are delivered as aggregate results, not as device lists. Where an exposed audience is delivered, it goes to the client who commissioned it, under a separate agreement governing what that client may do with it.

Signals and identifiers are kept for as long as we need them to produce and support the measurement a client commissioned, and are then removed from the working data.

Opting out

Mogean honours opt-outs of mobile advertising identifiers in these services. To opt out, email your mobile advertising identifier toinfo@mogean.com orprivacy@mogean.com. Email is sufficient; no form is required.

When a mobile advertising identifier is received, it is added to Mogean's deletion and suppression list and handled through the daily processing pipeline, after which it is excluded from the Exposure Measurement services. The opt-out page carries the same route.

Query

AI and natural-language query

The model proposes; the catalog decides

The natural-language layer turns a typed question into a cut of a named release. The language model proposes a reading of the question. The catalog decides what that reading actually resolves to. A verified catalog fact always beats a model guess.

Terms resolve against the live catalog

Brand, category, leaf-category, and state mentions are resolved against the live catalog dictionary rather than guessed. A term the catalog does not carry does not silently become a near-miss.

Ambiguity is surfaced, not resolved silently

When a term matches more than one thing, the possible matches are shown for you to choose. The layer does not pick one on your behalf and carry on.

Counts and prices come from the release, not the model

Every result is a deterministic cut of a named release. Its row count, its sample, and its price come from the catalog and the release, never from the model. The same question against the same release returns the same cut.

Where the layer runs

In the store explorer, for building and buying custom cuts. Embedded in your own product under an embedded license. And inside analytics engagements, where the same layer drives ad-hoc questions against your results.

Ask for the query

Ask for the query

If a figure on this site matters to a decision you are making, ask us for the query behind it and the release it ran against. We will send both.

Examples, sample data, and pilots are available for the measurement services.