Skip to main content

Overview

Use the Bulk API to request large datasets from Unify without waiting for a long-running HTTP request to finish. The Bulk API is asynchronous: you create a query job, poll the job until it is finished, and then page through the materialized results. Use the Bulk API when you need to export or sync many records. For small operational reads, use the standard API for the resource instead.
The Bulk API is for requesting data from Unify. To create or update records in Unify, see Send records via API.

Supported resources

The Bulk API currently supports these resources: For object_name, use a standard object such as company or person, or the API name of a custom object. Each resource exposes the same set of job-management endpoints under its query-jobs path:
The {base} for each resource is:

Authentication

The Bulk API requires a user-backed API key. Generate an API key in Settings → Developers and include it with each request:
All examples use the main Unify API base URL:

How the Bulk API works

1

Create a query job

Submit a query job for the resource you want to export. For sequences and events the request body is optional and defaults to an empty filter. For object records, a query with a select is required.
The response includes the job_id, current status, and expires_at time:
2

Poll job metadata

Poll the job metadata endpoint until the job reaches a terminal status. Avoid tight polling loops; use a steady interval or exponential backoff. Most jobs finish within a few seconds.
A finished job includes total_rows:
3

Fetch results

After the job status is FINISHED, fetch results by page. Results are immutable for a finished job, so each page is stable for that job.
4

Store a checkpoint

For recurring syncs, store a checkpoint such as the newest updated_at (or, for events, timestamp) value you processed. Use that checkpoint in the next query job so you only request new or changed records.

Create a query job

Create-job endpoints are resource-specific:
The request body depends on the resource:
  • Events and sequences take an optional filter object that narrows the dataset.
  • Object records take a required query object that selects which attributes to return and, optionally, filters and sorts the records.

Filter sequences and events

For sequences and events, omit the body to export everything, or pass a filter to narrow the dataset. For example, scope a request to exclude a specific set of IDs:

Query object records

Object record jobs take a query with a required select that lists the attributes to return. Use where to filter by attribute value, sort_by to order results, and metadata to filter by base record timestamps:
  • select — Required. Each key is an attribute API name on the queried object. Use true to return the attribute directly, or { "select": ... } to expand a single-reference attribute and select attributes on the referenced object. Nested selects may be at most three levels deep. Selected attributes are sparse: a result row only includes keys for selected attributes that have a value, so attributes with no value are omitted rather than returned as null.
  • where — Filter by one or more attribute conditions. { "equals": ... } matches an attribute directly; a nested object filters through a single-reference attribute on the referenced object.
  • sort_by — Sort by a base record field (id, created_at, or updated_at) in ASCENDING or DESCENDING order.
  • metadata — Filter by base record timestamps (created_at / updated_at). Pairing a created_at / updated_at metadata filter with the matching sort_by enables incremental queries that page through every record changed since a prior checkpoint.
The where and select shape mirrors the standard object records API. See Standard objects and attributes for the built-in attributes you can select and filter on.
Timestamps in the object Bulk API are millisecond precision. Both the timestamps returned in results and the values you supply to metadata and where filters are interpreted at millisecond precision.

Filter fields

Available filter fields depend on the resource.

Sequence filters

Sequence enrollment and enrollment-step jobs share a common set of filters: Enrollment jobs also accept status. Enrollment-step jobs additionally accept status, enrollment (filter by enrollment ID), type, and step_number. Range filters such as updated_at accept gt, gte, lt, and lte bounds. gt is mutually exclusive with gte, and lt with lte. Set filters such as id, status, and type accept in and/or not_in arrays.

Event filters

Events are immutable, so the natural cursor is timestamp, which is also the only datetime field you can filter on.

Job statuses

A job can have one of the following statuses: Jobs and results are available until expires_at, which is currently 24 hours after job creation. Download all required result pages before that time.

List jobs

List jobs for a resource by sending a GET to its query-jobs path:
Query parameters: Example request:
Example response:

Cancel a job

Cancel an in-progress job to stop it from being processed:
Only in-progress jobs can be canceled. Canceling a job that has already finished, failed, or been canceled returns 409 Conflict with job_not_cancelable.

Fetch results

Fetch results for a finished job:
Result pagination is page-based. Rows are ordered by a stable sort with a tie-breaker, so pages remain stable for a finished job.

JSON results

JSON is the default response format:
JSON limits: If a JSON page is too large, the request returns 413 with results_page_too_large. Request a smaller page_size or use NDJSON.

NDJSON results

Use NDJSON for larger streamed pages by providing the following header in your request.
Example request:
NDJSON responses return one JSON object per line, streaming one line at a time. Pagination metadata is returned in response headers instead of a JSON envelope:
The maximum NDJSON page_size is 10000.

Result shapes

Each item in data (or each NDJSON line) is a result row whose shape depends on the resource. The examples below are representative — fields may be added before general availability, so parse defensively and ignore unrecognized keys.
Object record rows use the same object / id / attributes envelope as the standard records API. attributes contains exactly the attributes named in your select; expanded single-reference attributes are nested as their own envelope. See Standard objects and attributes for the built-in attributes.
Events are flat, denormalized rows — they are not wrapped in the object / attributes envelope used by object records. type is the canonical event type (page, track, or identify). Page context (domain, path, referrer_domain, referrer_path), the URL query, UTM parameters, and any custom track-event properties are merged into a single properties object; properties is null when none are present.The revealed company and person are embedded as nested object records and are null when the visitor is unresolved. Unlike a top-level object-records select (which returns only the attributes you ask for), an embedded record carries every standard scalar attribute that has a value. As elsewhere in the API, an embedded record’s attributes are sparse: attributes with no value are omitted rather than returned as null. The person’s own company is a { "id": ... } reference rather than an inlined record.
Enrollment rows include status flags and embed the related sequence, mailbox, and person. reply_email_message is null until the person replies.
Step rows embed a summary of their parent enrollment and, for email steps, the sent email_message with open and click counts. email_message is null for non-email steps.

Job states

The results endpoint validates the job state before returning data: Only fetch results after the job status is FINISHED.

Rate limiting

Bulk API rate limits are applied by operation type. Query job creation is limited to roughly 100 jobs per day, so create jobs only when you need a new snapshot and use filters or checkpoints to keep each job focused. Treat other job management requests, such as checking job status, listing jobs, and canceling jobs, as low-frequency control-plane calls. When polling job status, use a steady interval or exponential backoff instead of a tight loop. Status checks are intended for roughly 5 requests per second, and most jobs finish quickly enough that slower polling is usually sufficient. When fetching results, page through data with bounded concurrency. If you receive 429 Too Many Requests, wait before retrying, honor the Retry-After header to avoid additional rate limiting.

Best practices

  • Request only the data you need. Use filters (and, for object records, a focused select) to keep result sets small and exports fast.
  • Prefer incremental syncs. Use checkpoints such as updated_at (or timestamp for events) so recurring jobs only request new or changed records.
  • Design for retries. Store the job_id, job status, and processing checkpoint in your system.
  • Process results idempotently. Use stable record IDs so retrying a page does not create duplicate downstream records.
  • Use NDJSON for large pages. NDJSON streams results and supports larger pages than JSON.
  • Download before expiration. Jobs expire at expires_at; create a new job if you need the data again after expiration.

Build your own connector

Want to move bulk data into your own warehouse or tool? Use our example connectors as a starting point for building integrations on top of the Bulk API.

Bulk API connector examples

Reference implementations for building your own connectors on top of the Bulk API.

What’s next

Data API reference

Review the Data API for object, attribute, and record operations.

Send records via API

Learn how to create and update records in Unify.