Bulk export
Bulk export delivers the corpus as files. Each day Arche publishes a snapshot: a set of Parquet files and a manifest listing every file with its size and SHA-256 checksum. You download the files directly from storage, so loading the whole corpus into a warehouse takes minutes instead of millions of API calls.
Bulk export is included in the Scale and Enterprise plans. Other plans are
refused with 403 FEATURE_NOT_AVAILABLE and details.required_tier: "scale".
What a snapshot holds
A snapshot is read from the database at a single moment, recorded as read_at.
It holds what the API would have served at that moment, and nothing EDGAR
accepted after it. Each statement appears once, as the version the API serves;
earlier versions are not exported as separate rows.
Snapshots build daily at 04:30 UTC and are usually published about an hour
later. If a build fails, nothing is published and the previous snapshot remains
the latest. Arche keeps the seven most recent snapshots. An older snapshot's
files are deleted and requests for it return 410.
Most builds are incremental: a file whose contents did not change is carried
forward with the same checksum and marked changed: false. A full rebuild runs
at least once a week.
Datasets
| Dataset | One row per | Partitioned by |
|---|---|---|
companies | company (cik) | — |
filings | filing (accession_id) | filing_year |
statements | served statement version (statement_version_id) | fiscal_year |
financials | metric on a statement version (statement_version_id, metric) | fiscal_year |
segments | segment or geographic disclosure (metric_dimension_id) | fiscal_year |
disclosures | narrative disclosure (disclosure_id) | report_year |
Files are Parquet, compressed with zstd. Paths follow <dataset>/<partition column>_<value>/part-NNNN.parquet, for
example statements/fiscal_year_2025/part-0000.parquet. The partition
directories use _, not =, so read them as plain directories rather than as
Hive partitions. companies is not partitioned.
Decimals are decimal(38, 12) and timestamps are UTC. For every column, its
type and what it means, ask the API:
curl -X GET "https://api.arche.fi/v1/bulk/datasets" \
-H "X-Api-Key: $ARCHE_API_KEY" \
-H "Accept: application/json"
Each dataset lists its key columns, its partition_column, and its columns
in file order with name, type, nullable and description.
Bulk export does not include modeling, derived-metric, AI or macroeconomic data.
Find a snapshot
List published snapshots, newest first:
curl -X GET "https://api.arche.fi/v1/bulk/snapshots?page=1&page_size=10" \
-H "X-Api-Key: $ARCHE_API_KEY" \
-H "Accept: application/json"
Read one snapshot's manifest by its ID, or use latest:
curl -X GET "https://api.arche.fi/v1/bulk/snapshots/latest" \
-H "X-Api-Key: $ARCHE_API_KEY" \
-H "Accept: application/json"
{
"data": {
"snapshot_id": "20260926T043103Z",
"build_kind": "incremental",
"format_version": 1,
"previous_snapshot_id": "20260925T161127Z",
"published_at": "2026-09-26T05:23:27Z",
"read_at": "2026-09-26T04:31:04Z",
"rows": 82159258,
"bytes": 6460657978,
"datasets": [
{
"dataset": "companies",
"schema_version": 1,
"rows": 23065,
"bytes": 661805,
"files": 1,
"changed_files": 0
}
],
"files": [
{
"path": "companies/part-0000.parquet",
"dataset": "companies",
"partition": null,
"rows": 23065,
"bytes": 661805,
"sha256": "f049bf4441e1bd692645021c1333f9878faf4a9ee55ade9a76b97791fe988e6d",
"changed": false
}
]
}
}
A snapshot ID is the time its build started, as YYYYMMDDTHHMMSSZ. Pin one ID
for a whole download, so that every file you load comes from the same moment.
Download a file
Ask for a download URL for one file in the manifest:
FILE=companies/part-0000.parquet
curl -X GET "https://api.arche.fi/v1/bulk/snapshots/20260926T043103Z/files/$FILE" \
-H "X-Api-Key: $ARCHE_API_KEY" \
-H "Accept: application/json"
{
"data": {
"snapshot_id": "20260926T043103Z",
"file_path": "companies/part-0000.parquet",
"url": "https://…",
"expires_at": "2026-09-26T09:10:00Z",
"bytes": 661805,
"sha256": "f049bf4441e1bd692645021c1333f9878faf4a9ee55ade9a76b97791fe988e6d"
}
}
Fetch url with a plain GET and no API key. It expires ten minutes after it
is issued. After downloading, compare the file's SHA-256 with sha256 and
discard the file if they differ.
The response is JSON carrying the URL, not a redirect. A path that is not in
the snapshot's manifest returns 404.
Daily download allowance
Each organization may be issued download URLs for up to 20 GiB of files in any rolling 24 hours, enough for about three full snapshots. The full size of a file counts when its URL is issued, not when you download it, because a URL cannot be withdrawn once issued.
Asking again for the same file while an earlier URL for it is still valid costs nothing, so retrying a failed download does not spend the allowance twice.
When the allowance is spent, the request returns 429 with the code
BULK_DOWNLOAD_BUDGET_EXCEEDED. details reports issued_bytes,
requested_bytes and daily_byte_cap, and a Retry-After header says when
enough of the allowance frees up for the file you asked for.
Requests to these routes are also rate limited; see Rate limits.
Keeping a copy current
To refresh a local copy, read the latest manifest and download only the files
whose sha256 differs from the copy you hold. Files marked changed: false
are identical to the previous snapshot's file at the same path. After a full
rebuild every file may change, so compare checksums rather than relying on
changed alone.
Errors
- Name
FEATURE_NOT_AVAILABLE- Type
- 403
- Description
Your plan does not include bulk export.
- Name
BULK_SNAPSHOT_NOT_FOUND- Type
- 404
- Description
No published snapshot has that ID, or nothing has been published yet when you ask for
latest.
- Name
BULK_SNAPSHOT_EXPIRED- Type
- 410
- Description
The snapshot is older than the seven that are kept. Use
latest.
- Name
BULK_FILE_NOT_FOUND- Type
- 404
- Description
The path is not in the snapshot's manifest.
- Name
BULK_DOWNLOAD_BUDGET_EXCEEDED- Type
- 429
- Description
The daily download allowance is spent. Wait for
Retry-After.