# Bulk export

Bulk export delivers the corpus as files. Each day Arche publishes a snapshot:
a set of Parquet files and a manifest listing every file with its size and
SHA-256 checksum. You download the files directly from storage, so loading the
whole corpus into a warehouse takes minutes instead of millions of API calls.

Bulk export is included in the Scale and Enterprise plans. Other plans are
refused with `403 FEATURE_NOT_AVAILABLE` and `details.required_tier: "scale"`.

## What a snapshot holds

A snapshot is read from the database at a single moment, recorded as `read_at`.
It holds what the API would have served at that moment, and nothing EDGAR
accepted after it. Each statement appears once, as the version the API serves;
earlier versions are not exported as separate rows.

Snapshots build daily at 04:30 UTC and are usually published about an hour
later. If a build fails, nothing is published and the previous snapshot remains
the latest. Arche keeps the seven most recent snapshots. An older snapshot's
files are deleted and requests for it return `410`.

Most builds are incremental: a file whose contents did not change is carried
forward with the same checksum and marked `changed: false`. A full rebuild runs
at least once a week.

## Datasets

| Dataset       | One row per                                                      | Partitioned by |
| ------------- | ---------------------------------------------------------------- | -------------- |
| `companies`   | company (`cik`)                                                  | —              |
| `filings`     | filing (`accession_id`)                                          | `filing_year`  |
| `statements`  | served statement version (`statement_version_id`)                | `fiscal_year`  |
| `financials`  | metric on a statement version (`statement_version_id`, `metric`) | `fiscal_year`  |
| `segments`    | segment or geographic disclosure (`metric_dimension_id`)         | `fiscal_year`  |
| `disclosures` | narrative disclosure (`disclosure_id`)                           | `report_year`  |

Files are Parquet, compressed with zstd. Paths follow `<dataset>/<partition column>_<value>/part-NNNN.parquet`, for
example `statements/fiscal_year_2025/part-0000.parquet`. The partition
directories use `_`, not `=`, so read them as plain directories rather than as
Hive partitions. `companies` is not partitioned.

Decimals are `decimal(38, 12)` and timestamps are UTC. For every column, its
type and what it means, ask the API:

```bash
curl -X GET "https://api.arche.fi/v1/bulk/datasets" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "Accept: application/json"
```

Each dataset lists its `key` columns, its `partition_column`, and its `columns`
in file order with `name`, `type`, `nullable` and `description`.

Bulk export does not include modeling, derived-metric, AI or macroeconomic
data.

## Find a snapshot

List published snapshots, newest first:

```bash
curl -X GET "https://api.arche.fi/v1/bulk/snapshots?page=1&page_size=10" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "Accept: application/json"
```

Read one snapshot's manifest by its ID, or use `latest`:

```bash
curl -X GET "https://api.arche.fi/v1/bulk/snapshots/latest" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "Accept: application/json"
```

```json
{
  "data": {
    "snapshot_id": "20260926T043103Z",
    "build_kind": "incremental",
    "format_version": 1,
    "previous_snapshot_id": "20260925T161127Z",
    "published_at": "2026-09-26T05:23:27Z",
    "read_at": "2026-09-26T04:31:04Z",
    "rows": 82159258,
    "bytes": 6460657978,
    "datasets": [
      {
        "dataset": "companies",
        "schema_version": 1,
        "rows": 23065,
        "bytes": 661805,
        "files": 1,
        "changed_files": 0
      }
    ],
    "files": [
      {
        "path": "companies/part-0000.parquet",
        "dataset": "companies",
        "partition": null,
        "rows": 23065,
        "bytes": 661805,
        "sha256": "f049bf4441e1bd692645021c1333f9878faf4a9ee55ade9a76b97791fe988e6d",
        "changed": false
      }
    ]
  }
}
```

A snapshot ID is the time its build started, as `YYYYMMDDTHHMMSSZ`. Pin one ID
for a whole download, so that every file you load comes from the same moment.

## Download a file

Ask for a download URL for one file in the manifest:

```bash
FILE=companies/part-0000.parquet
curl -X GET "https://api.arche.fi/v1/bulk/snapshots/20260926T043103Z/files/$FILE" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "Accept: application/json"
```

```json
{
  "data": {
    "snapshot_id": "20260926T043103Z",
    "file_path": "companies/part-0000.parquet",
    "url": "https://…",
    "expires_at": "2026-09-26T09:10:00Z",
    "bytes": 661805,
    "sha256": "f049bf4441e1bd692645021c1333f9878faf4a9ee55ade9a76b97791fe988e6d"
  }
}
```

Fetch `url` with a plain `GET` and no API key. It expires ten minutes after it
is issued. After downloading, compare the file's SHA-256 with `sha256` and
discard the file if they differ.

The response is JSON carrying the URL, not a redirect. A path that is not in
the snapshot's manifest returns `404`.

## Daily download allowance

Each organization may be issued download URLs for up to 20 GiB of files in any
rolling 24 hours, enough for about three full snapshots. The full size of a file
counts when its URL is issued, not when you download it, because a URL cannot
be withdrawn once issued.

Asking again for the same file while an earlier URL for it is still valid
costs nothing, so retrying a failed download does not spend the allowance
twice.

When the allowance is spent, the request returns `429` with the code
`BULK_DOWNLOAD_BUDGET_EXCEEDED`. `details` reports `issued_bytes`,
`requested_bytes` and `daily_byte_cap`, and a `Retry-After` header says when
enough of the allowance frees up for the file you asked for.

Requests to these routes are also rate limited; see
[Rate limits](https://docs.arche.fi/rate-limits).

## Keeping a copy current

To refresh a local copy, read the latest manifest and download only the files
whose `sha256` differs from the copy you hold. Files marked `changed: false`
are identical to the previous snapshot's file at the same path. After a full
rebuild every file may change, so compare checksums rather than relying on
`changed` alone.

## Errors

- `FEATURE_NOT_AVAILABLE` (403): Your plan does not include bulk export.
- `BULK_SNAPSHOT_NOT_FOUND` (404): No published snapshot has that ID, or nothing has been published yet when
  you ask for `latest`.
- `BULK_SNAPSHOT_EXPIRED` (410): The snapshot is older than the seven that are kept. Use `latest`.
- `BULK_FILE_NOT_FOUND` (404): The path is not in the snapshot's manifest.
- `BULK_DOWNLOAD_BUDGET_EXCEEDED` (429): The daily download allowance is spent. Wait for `Retry-After`.

## Next steps

- [API reference](https://docs.arche.fi/reference): Browse the live OpenAPI contract rendered from the public schema.
- [Errors](https://docs.arche.fi/errors): Interpret machine-readable error responses and troubleshoot failed requests.
- [Request IDs](https://docs.arche.fi/troubleshooting/request-ids): Correlate a failed request across your logs, the API response, and the portal.
