SDKs
The Python and TypeScript SDKs cover every API v1 operation. They paginate, wait for asynchronous work, retry safely, send idempotency keys, and stream files, so application code only describes what it needs.
Install
pip install extraltCreate a client
Both clients read the API key from the EXTRALT_API_KEY environment variable.
from extralt import Extralt
client = Extralt() # or Extralt(api_key="...")Python also has AsyncExtralt, with the same methods as coroutines and async
iterators.
Methods follow the API
Each operation in the API reference is one method, grouped by area
and resource: client.extract.runs.create(...) sends POST /v1/extract/runs.
Field names are the API's own, in both languages.
| Area | Resources |
|---|---|
client.extract | robots, runs, schedules, captures, imports |
client.enrich | enrichments, items |
client.explore | overview, facets, activity, stores, entities, relationships, changes, markets, variants, analyses |
client.query has its methods directly: run, export, and model for
SQL queries.
robot = client.extract.robots.create(url="https://example-store.com", country="US")
run = client.extract.runs.create(robot_id=robot.id, budget=100)Lists
list() returns a paginator. Iterating it fetches the next page when needed;
first_page() returns one page with its next_cursor.
for capture in client.extract.captures.list(run_id=run.id):
print(capture.title)
page = client.extract.runs.list(limit=50).first_page()Jobs
Building a Robot, a Run, an Import, and an Enrichment each take minutes. The
call that starts one polls with backoff and returns it finished, within one
hour or the timeout you pass.
- A job that ends as
failedraisesJobFailedError, which carries the job asresource. A job that was stopped is returned. - When the time runs out,
WaitTimeoutErrorcarries the job. It keeps running on Extralt, andwait()with its ID resumes waiting. - A Run needs a
budget: the most Captures to produce, or"unlimited"to crawl until the store is done or credits run out.
from extralt import JobFailedError, WaitTimeoutError
try:
run = client.extract.runs.create(robot_id=robot.id, budget=100, timeout=900)
except JobFailedError as error:
print(error.resource.status_reason)
except WaitTimeoutError as error:
print(f"{error.resource.id} is still running")To start a job without waiting for it, for example from a web request, pass
wait=False and wait for it elsewhere with its ID:
started = client.extract.runs.create(robot_id=robot.id, budget=100, wait=False)
finished = client.extract.runs.wait(started.id)Errors and retries
API errors raise a typed error per status, such as NotFoundError,
RateLimitError, or ConflictError, carrying the code, message, and
request ID of the error envelope. Include the request
ID when contacting support.
Connection errors, 408, 429, and 5xx responses are retried with
exponential backoff and jitter, honoring Retry-After, but only for operations
that are safe to repeat. Create and restart requests send an Idempotency-Key
that every retry reuses, so a retried request never creates a second resource.
Pass your own key to retry the same logical request across processes.
from extralt import NotFoundError
try:
client.extract.runs.get("missing")
except NotFoundError as error:
print(error.code, error.request_id)Files
Capture exports stream to disk without loading the file in memory, and imports stream catalog files from disk.
with client.extract.captures.export(run_id=run.id, format="parquet") as export:
export.write_to("captures.parquet")
imported = client.extract.imports.create("catalog.csv", country="US")Explore
Explore time windows take RFC 3339 timestamps. Python accepts timezone-aware
datetime values and spells the from parameter from_, a Python keyword.
Prices, price differences, and percentages are Decimal in Python. In
TypeScript they are decimal strings, which keep their exact value: convert one
with Number() before comparing or adding.
Analyses are paginators like lists: the first page holds the summary, coverage,
and freshness of the whole scope, and iterating yields every evidence row.
import datetime as dt
week_ago = dt.datetime.now(dt.UTC) - dt.timedelta(days=7)
movements = client.explore.analyses.price_movements(
from_=week_ago, country="US", currency="USD"
)
print(movements.first_page().summary)
for row in movements:
print(row.title, row.old_price, row.new_price)Configuration
from extralt import Extralt
client = Extralt(
timeout=60.0, # seconds per request
max_retries=2,
)To go through a proxy or trust custom certificates, pass your own HTTP client:
an httpx.Client as http_client in Python, your own fetch in TypeScript.
Responses are read without strict validation: fields the API adds within v1
stay available instead of breaking older SDK versions. Set EXTRALT_LOG=info
to log each job's progress and every retry, or debug to add every response. In Python, the
extralt logger takes your logging configuration; in TypeScript, the logLevel
and logger options do. API keys are never logged.
Query semantics
Interpret grain, current state, time, price, availability, assortment, reviews, coverage, and freshness consistently across Explore.
Authentication
Authenticate API requests with org-scoped Extralt API keys, Bearer tokens, rotation guidance, and safe handling for production integrations.