Semantic model
This page defines what Extralt data means. It intentionally does not duplicate physical ClickHouse schemas or public response types. Explore AI receives the current versioned field catalog at runtime, while the generated OpenAPI document owns exact public request and response schemas.
Three connected layers
| Layer | Purpose | Main records |
|---|---|---|
| Extract evidence | Preserve what was collected from a page or imported catalog, with URL and observation lineage | Capture |
| Enrich output | Normalize one successful Capture into a durable page-grain record | Item |
| Analytical model | Connect exact configurations, Store Listings, and observed commerce evidence across pages | Product, Variant, Listing, Offer, Review, Store, Change |
Captures and Items remain useful independently. Explore normally reads the connected analytical model published before an Enrichment becomes terminal.
Records and entity grains
| Record or entity | One record represents | Use it for |
|---|---|---|
| Capture | One successfully extracted or imported page record | Original page content, source URL, observation time, extracted SKUs and Offers, export, and Enrich input |
| Item | One successfully enriched Capture | Normalized English content, taxonomy, attributes, signals, options, embedded normalized SKU data, and Capture lineage |
| Product | One current canonical product family in an organization dataset | Family-level title, brand, taxonomy, attributes, signals, option axes, and media |
| Variant | One current exact product configuration under a Product | Cross-store identity for the same precise option or identifier combination |
| Listing | One Store SKU identity in one country | Source URL, Store, exact source identity, Variant resolution, and first/last/removal evidence |
| Offer | One seller, price, currency, condition, availability, and stock observation for a Listing | Current price and availability, marketplace evidence, and history |
| Review | One review-count and normalized-score observation for a Listing | Current review evidence and observed aggregate history |
| Store | One hostname and country pair in an organization dataset | Market context for Listings, Offers, and Reviews |
| Change | One observed semantic change to one entity aspect | Before/after evidence for additions, changes, removals, and reappearances |
An Item is not a canonical Product. It is the normalized evidence used to publish Products, Variants, Listings, Offers, Reviews, Stores, and Changes.
Relationships
Capture -> Item
Product -> Variant -> Listing -> Offer observations
|-> Review observations
Store -------------------+- A Product has one or more exact Variants.
- A Variant can resolve equivalent Listings observed in several Stores.
- A Listing belongs to one Store market and one current Variant resolution.
- Offer and Review observations belong to a Listing and retain observation lineage.
- Product, Variant, brand, and category context for an Offer or Review is reached through its Listing.
These relationships are supported by Explore AI, but they are not a promise of unrestricted access to physical tables. See Access surfaces for the supported interfaces.
Observation records and current state
Capture, Item, Offer, Review, and Change rows preserve observations. Product, Variant, Listing, and Store rows represent current dimensions in Explore.
Purpose-built views and the Query model apply these current-state rules:
- active Listings are latest Listing states without explicit removal;
- current Offers are all Offers from each active Listing's latest observed Extract run, selected before Offer filters;
- current Reviews are the latest known review aggregate for each active Listing;
- observed assortment is the distinct set of Variants connected to active Listings in the selected Store and country;
- new means first observed in this organization dataset;
- removed requires explicit Listing removal evidence;
- availability preserves available, unavailable, unknown, absent, and removed as different states.
See Query semantics for the complete comparison and interpretation rules.
Identity boundaries
Product and Variant ids are Extralt canonical identities inside one organization's final dataset. Source product ids, handles, SKUs, GTINs, and MPNs are evidence used during normalization and matching; none is guaranteed to be a universal primary key.
Matching quality depends on the identifiers and product evidence available in the collected pages. The model supports exact cross-store Variant comparison, not substitutes, aesthetic similarity, or exhaustive market identity.
A Store is a storefront in one country. It is not a Robot, seller, legal company, retailer group, traffic estimate, or sales account. A Robot is operational extraction tooling and has a separate lifecycle.
Data boundaries
The model contains observed catalogs and product details, identifiers, options, prices, compare-at prices when exposed, availability, categorical stock, sellers, review aggregates, media, timestamps, and page URLs.
It does not contain sales, transactions, revenue, turnover, traffic, demand, margin, market share, review text, sentiment, or unobserved promotion codes. Coverage is limited to the Stores, countries, URLs, and observation times the organization selected.
For the available dashboard, API, export, Query, and Agent paths, continue with Access surfaces.