SDKs

The Python and TypeScript SDKs cover every API v1 operation. They paginate, wait for asynchronous work, retry safely, send idempotency keys, and stream files, so application code only describes what it needs.

Install

pip install extralt

Create a client

Both clients read the API key from the EXTRALT_API_KEY environment variable.

from extralt import Extralt

client = Extralt()  # or Extralt(api_key="...")

Python also has AsyncExtralt, with the same methods as coroutines and async iterators.

Methods follow the API

Each operation in the API reference is one method, grouped by area and resource: client.extract.runs.create(...) sends POST /v1/extract/runs. Field names are the API's own, in both languages.

AreaResources
client.extractrobots, runs, schedules, captures, imports
client.enrichenrichments, items
client.exploreoverview, facets, activity, stores, entities, relationships, changes, markets, variants, analyses

client.query has its methods directly: run, export, and model for SQL queries.

robot = client.extract.robots.create(url="https://example-store.com", country="US")
run = client.extract.runs.create(robot_id=robot.id, budget=100)

Lists

list() returns a paginator. Iterating it fetches the next page when needed; first_page() returns one page with its next_cursor.

for capture in client.extract.captures.list(run_id=run.id):
    print(capture.title)

page = client.extract.runs.list(limit=50).first_page()

Jobs

Building a Robot, a Run, an Import, and an Enrichment each take minutes. The call that starts one polls with backoff and returns it finished, within one hour or the timeout you pass.

  • A job that ends as failed raises JobFailedError, which carries the job as resource. A job that was stopped is returned.
  • When the time runs out, WaitTimeoutError carries the job. It keeps running on Extralt, and wait() with its ID resumes waiting.
  • A Run needs a budget: the most Captures to produce, or "unlimited" to crawl until the store is done or credits run out.
from extralt import JobFailedError, WaitTimeoutError

try:
    run = client.extract.runs.create(robot_id=robot.id, budget=100, timeout=900)
except JobFailedError as error:
    print(error.resource.status_reason)
except WaitTimeoutError as error:
    print(f"{error.resource.id} is still running")

To start a job without waiting for it, for example from a web request, pass wait=False and wait for it elsewhere with its ID:

started = client.extract.runs.create(robot_id=robot.id, budget=100, wait=False)
finished = client.extract.runs.wait(started.id)

Errors and retries

API errors raise a typed error per status, such as NotFoundError, RateLimitError, or ConflictError, carrying the code, message, and request ID of the error envelope. Include the request ID when contacting support.

Connection errors, 408, 429, and 5xx responses are retried with exponential backoff and jitter, honoring Retry-After, but only for operations that are safe to repeat. Create and restart requests send an Idempotency-Key that every retry reuses, so a retried request never creates a second resource. Pass your own key to retry the same logical request across processes.

from extralt import NotFoundError

try:
    client.extract.runs.get("missing")
except NotFoundError as error:
    print(error.code, error.request_id)

Files

Capture exports stream to disk without loading the file in memory, and imports stream catalog files from disk.

with client.extract.captures.export(run_id=run.id, format="parquet") as export:
    export.write_to("captures.parquet")

imported = client.extract.imports.create("catalog.csv", country="US")

Explore

Explore time windows take RFC 3339 timestamps. Python accepts timezone-aware datetime values and spells the from parameter from_, a Python keyword. Prices, price differences, and percentages are Decimal in Python. In TypeScript they are decimal strings, which keep their exact value: convert one with Number() before comparing or adding. Analyses are paginators like lists: the first page holds the summary, coverage, and freshness of the whole scope, and iterating yields every evidence row.

import datetime as dt

week_ago = dt.datetime.now(dt.UTC) - dt.timedelta(days=7)
movements = client.explore.analyses.price_movements(
    from_=week_ago, country="US", currency="USD"
)
print(movements.first_page().summary)
for row in movements:
    print(row.title, row.old_price, row.new_price)

Configuration

from extralt import Extralt

client = Extralt(
    timeout=60.0,  # seconds per request
    max_retries=2,
)

To go through a proxy or trust custom certificates, pass your own HTTP client: an httpx.Client as http_client in Python, your own fetch in TypeScript.

Responses are read without strict validation: fields the API adds within v1 stay available instead of breaking older SDK versions. Set EXTRALT_LOG=info to log each job's progress and every retry, or debug to add every response. In Python, the extralt logger takes your logging configuration; in TypeScript, the logLevel and logger options do. API keys are never logged.