ExtraltExtralt

Semantic model

This page defines what Extralt data means. It intentionally does not duplicate physical ClickHouse schemas or public response types. Explore AI receives the current versioned field catalog at runtime, while the generated OpenAPI document owns exact public request and response schemas.

Three connected layers

LayerPurposeMain records
Extract evidencePreserve what was collected from a page or imported catalog, with URL and observation lineageCapture
Enrich outputNormalize one successful Capture into a durable page-grain recordItem
Analytical modelConnect exact configurations, Store Listings, and observed commerce evidence across pagesProduct, Variant, Listing, Offer, Review, Store, Change

Captures and Items remain useful independently. Explore normally reads the connected analytical model published before an Enrichment becomes terminal.

Records and entity grains

Record or entityOne record representsUse it for
CaptureOne successfully extracted or imported page recordOriginal page content, source URL, observation time, extracted SKUs and Offers, export, and Enrich input
ItemOne successfully enriched CaptureNormalized English content, taxonomy, attributes, signals, options, embedded normalized SKU data, and Capture lineage
ProductOne current canonical product family in an organization datasetFamily-level title, brand, taxonomy, attributes, signals, option axes, and media
VariantOne current exact product configuration under a ProductCross-store identity for the same precise option or identifier combination
ListingOne Store SKU identity in one countrySource URL, Store, exact source identity, Variant resolution, and first/last/removal evidence
OfferOne seller, price, currency, condition, availability, and stock observation for a ListingCurrent price and availability, marketplace evidence, and history
ReviewOne review-count and normalized-score observation for a ListingCurrent review evidence and observed aggregate history
StoreOne hostname and country pair in an organization datasetMarket context for Listings, Offers, and Reviews
ChangeOne observed semantic change to one entity aspectBefore/after evidence for additions, changes, removals, and reappearances

An Item is not a canonical Product. It is the normalized evidence used to publish Products, Variants, Listings, Offers, Reviews, Stores, and Changes.

Relationships

Capture -> Item

Product -> Variant -> Listing -> Offer observations
                         |-> Review observations
Store -------------------+
  • A Product has one or more exact Variants.
  • A Variant can resolve equivalent Listings observed in several Stores.
  • A Listing belongs to one Store market and one current Variant resolution.
  • Offer and Review observations belong to a Listing and retain observation lineage.
  • Product, Variant, brand, and category context for an Offer or Review is reached through its Listing.

These relationships are supported by Explore AI, but they are not a promise of unrestricted access to physical tables. See Access surfaces for the supported interfaces.

Observation records and current state

Capture, Item, Offer, Review, and Change rows preserve observations. Product, Variant, Listing, and Store rows represent current dimensions in Explore.

Purpose-built views and the Query model apply these current-state rules:

  • active Listings are latest Listing states without explicit removal;
  • current Offers are all Offers from each active Listing's latest observed Extract run, selected before Offer filters;
  • current Reviews are the latest known review aggregate for each active Listing;
  • observed assortment is the distinct set of Variants connected to active Listings in the selected Store and country;
  • new means first observed in this organization dataset;
  • removed requires explicit Listing removal evidence;
  • availability preserves available, unavailable, unknown, absent, and removed as different states.

See Query semantics for the complete comparison and interpretation rules.

Identity boundaries

Product and Variant ids are Extralt canonical identities inside one organization's final dataset. Source product ids, handles, SKUs, GTINs, and MPNs are evidence used during normalization and matching; none is guaranteed to be a universal primary key.

Matching quality depends on the identifiers and product evidence available in the collected pages. The model supports exact cross-store Variant comparison, not substitutes, aesthetic similarity, or exhaustive market identity.

A Store is a storefront in one country. It is not a Robot, seller, legal company, retailer group, traffic estimate, or sales account. A Robot is operational extraction tooling and has a separate lifecycle.

Data boundaries

The model contains observed catalogs and product details, identifiers, options, prices, compare-at prices when exposed, availability, categorical stock, sellers, review aggregates, media, timestamps, and page URLs.

It does not contain sales, transactions, revenue, turnover, traffic, demand, margin, market share, review text, sentiment, or unobserved promotion codes. Coverage is limited to the Stores, countries, URLs, and observation times the organization selected.

For the available dashboard, API, export, Query, and Agent paths, continue with Access surfaces.