Coverage and data quality

Arche reports what it has and how good it is, rather than making you infer coverage from empty responses. A request that finds nothing is answered, not failed — so before treating an empty result as a gap in the market, check whether it is a gap in coverage.

Check coverage first

Two questions have separate endpoints:

  • Name
    What do you have for this company?
    Type
    /v1/coverage/companies/{cik}
    Description

    Company-level coverage and completeness for one issuer.

  • Name
    What do you have for this metric?
    Type
    /v1/coverage/metrics/{metric}
    Description

    Coverage and completeness for one canonical metric across the corpus.

Both are cheap relative to a full extraction job, so it is worth querying them at the start of a backfill rather than discovering the boundary partway through.

Company coverage

Retrieve the coverage summary for a single issuer:

curl -X GET "https://api.arche.fi/v1/coverage/companies/0000320193" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

List coverage across issuers with deterministic ordering — cik ascending, tie-broken on the canonical company row — using the standard paginated envelope:

curl -X GET "https://api.arche.fi/v1/coverage/companies?page=1&page_size=50" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

For a narrower question — the latest filing on record and statement-version counts by type — use the company completeness summary:

curl -X GET "https://api.arche.fi/v1/edgar/companies/0000320193/completeness" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

Reading the coverage response

Two fields mean something narrower than their names suggest, and both are easy to misread as a quality score.

  • Name
    core_metric_completeness
    Type
    decimal 0–1
    Description

    The share of expected periods that hold at least one core metric. It measures whether a period was populated at all, not how many metrics it contains or how complete any statement is. A company with one metric in every period scores 1.0.

  • Name
    annual_reporting_continuous
    Type
    boolean
    Description

    Whether every annual period across the company's own reporting life is present, tolerating a single year still in progress. This is the field to use for "are there gaps in this filer's history", and the one behind the published gap-free figure.

  • Name
    quarterly_periods_expected
    Type
    integer
    Description

    Three per fiscal year, not four. A fourth quarter is not separately filed — it is derived from the 10-K — so a denominator of four would mark every complete filer incomplete. Companies that file no 10-Q at all, such as foreign private issuers on 20-F or 40-F, expect none.

  • Name
    annual_periods_expected
    Type
    integer
    Description

    One per fiscal year across the company's covered span.

The *_populated counters pair with each *_expected field, and the per-statement ratios — income_statement_completeness, balance_sheet_completeness, cash_flow_statement_completeness — report the same period-level measure for one statement type.

Do not rank companies by core_metric_completeness and treat the bottom as poorly covered. A low value usually means the filer reports fewer periods than a full span would imply — a recent IPO, a changed fiscal calendar, a company that files annually only — not that the data is thin. Use annual_reporting_continuous to ask about gaps.

Choosing companies on Free

Every paid plan covers every company Arche serves. The Free plan covers any 25 companies you choose. There is no list to pick from: a company takes one of your 25 slots the first time a request reaches it, and keeps the slot for as long as you keep querying it. A slot frees itself 30 days after the company was last queried.

See which companies you hold and when each slot frees:

curl -X GET "https://api.arche.fi/v1/coverage/selection" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"
  • Name
    mode
    Type
    string
    Description

    selected when coverage is the companies you chose. Paid plans report full.

  • Name
    limit
    Type
    integer | null
    Description

    How many companies you may hold at once. Null when coverage is not chosen.

  • Name
    window_days
    Type
    integer | null
    Description

    How many days a company keeps its slot without being queried.

  • Name
    slots_used
    Type
    integer
    Description

    How many companies you hold right now.

  • Name
    slots_available
    Type
    integer | null
    Description

    How many more companies you can add.

  • Name
    companies
    Type
    array
    Description

    The companies you hold, oldest first. Each carries cik, claimed_at, last_queried_at and expires_at, when the slot frees if nothing queries the company again.

A mistyped CIK takes a slot like any other. Give one back straight away rather than waiting 30 days:

curl -X POST "https://api.arche.fi/v1/coverage/selection/release" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"cik": "320193"}'

Releasing a company you do not hold succeeds and changes nothing.

When all 25 slots are in use, a request for a new company returns 403 with the code COVERAGE_SELECTION_FULL. details lists the companies the request could not reach (refused_ciks), the ones you hold (held_ciks), and the release endpoint. Release a slot, wait for one to expire, or move to a paid plan, which covers every company.

Metric coverage

Coverage for one canonical metric across the corpus:

curl -X GET "https://api.arche.fi/v1/coverage/metrics/REVENUE" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

Representative response:

{
  "data": {
    "metric": "REVENUE",
    "company_count": 7100,
    "years_covered": 18,
    "statement_count": 42525
  }
}

The metric name must be one Arche stores as a canonical statement metric. A name it does not serve is refused with 400, listing the names it accepts:

{
  "error": {
    "code": "VALIDATION_ERROR",
    "http_status": 400,
    "message": "Unknown metric 'GROSS_MARGIN'. Expected one of: ACCOUNTS_PAYABLE, …, TOTAL_ASSETS.",
    "details": {},
    "trace_id": "…"
  }
}

GROSS_MARGIN is refused here because it is a derived ratio, computed on the modeling routes rather than mapped from a filing. A zero from this endpoint now always means the metric genuinely has no coverage, never that the name was wrong. See Metric catalog for which names live where.

Population statistics are scoped to a cohort and a filing form type, so both parameters are required. form_type takes the SEC's own spelling — 10-K or 10-Q, not FORM_10_K, which returns 400 VALIDATION_ERROR:

curl -X GET "https://api.arche.fi/v1/coverage/metrics?cohort=all&form_type=10-K" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

Representative response:

{
  "page": 1,
  "page_size": 50,
  "total": 4,
  "items": [
    {
      "metric_name": "TOTAL_ASSETS",
      "statement_type": "BALANCE_SHEET",
      "form_type": "10-K",
      "fiscal_period": "FY",
      "cohort": "ALL",
      "populated_rate": "1.000000",
      "sample_size": 43,
      "computed_at": "2026-06-07T19:08:44Z"
    }
  ]
}

These rows are precomputed on a sampled cohort, so sample_size is small and computed_at can lag. Use this endpoint for a rough population rate, and the per-metric route above for corpus-wide counts.

ALL is the only cohort with statistics today. A cohort that has never been computed is refused rather than answered with an empty page, and the refusal names the cohorts that do have rows:

{
  "error": {
    "code": "VALIDATION_ERROR",
    "http_status": 400,
    "message": "No coverage statistics exist for cohort 'sp500'. Cohorts with statistics: ALL.",
    "details": {},
    "trace_id": "…"
  }
}

sp500 and dow30 are valid cohort labels that simply have no computed statistics yet — the refusal distinguishes that from a cohort in which no metric is populated, which an empty page could not.

Results are ordered statement_type, then fiscal_period, then metric_name, all ascending.

Data-quality issues

Persisted data-quality issues are queryable directly. Ordering is deterministic: severity descending, then cik, fiscal_year, fiscal_period, rule_code, and issue_id ascending, with nulls last on the fiscal fields.

curl -X GET "https://api.arche.fi/v1/data-quality/issues" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

To see quality signals in the context of the statement they describe, request the normalized statement with its DQ overlay. This identifies a statement by its full identity, not just the company:

curl -X GET "https://api.arche.fi/v1/edgar/companies/0000320193/statements/dq/overlay?statement_type=INCOME_STATEMENT&fiscal_year=2023&fiscal_period=FY&version_sequence=2" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

Reconciliation

Reconciliation checks answer whether a statement is internally consistent — whether the parts sum, and whether balances roll forward.

The ledger holds a row only where a rule actually produced a result, so most statement identities have none and answer with an empty page. Start with the summary to find the identities worth opening, then read the ledger for one of them. Going the other way round mostly returns total: 0, which says nothing was evaluated, not that nothing failed.

PASS, WARN and FAIL counts by category across a multi-year window:

curl -X GET "https://api.arche.fi/v1/edgar/reconciliation/summary?cik=0000789019&statement_type=BALANCE_SHEET&fiscal_year_from=2015&fiscal_year_to=2026" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

Representative response:

{
  "data": [
    {
      "fiscal_year": 2026,
      "fiscal_period": "Q2",
      "version_sequence": 5,
      "rule_category": "ROLLFORWARD",
      "pass_count": 0,
      "warn_count": 2,
      "fail_count": 1
    }
  ]
}

Each row names a full statement identity. Pass that identity — including version_sequence, which is required — to the ledger for the individual rule results:

curl -X GET "https://api.arche.fi/v1/edgar/reconciliation/ledger?cik=0000789019&statement_type=BALANCE_SHEET&fiscal_year=2026&fiscal_period=Q2&version_sequence=5" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

Representative response:

{
  "page": 1,
  "page_size": 200,
  "total": 3,
  "items": [
    {
      "cik": "0000789019",
      "statement_type": "BALANCE_SHEET",
      "fiscal_year": 2026,
      "fiscal_period": "Q2",
      "version_sequence": 5,
      "rule_id": "cash_rollforward_adjacent_periods",
      "rule_category": "ROLLFORWARD",
      "status": "WARNING",
      "severity": "NONE",
      "expected_value": null,
      "actual_value": null,
      "delta": null,
      "notes": {
        "details": { "reason": "INCOMPARABLE_CASH_SCOPES" }
      }
    }
  ]
}

A WARNING with severity: NONE and a reason in notes is the rule declining to compare rather than finding a discrepancy — INCOMPARABLE_CASH_SCOPES here means the two periods define cash differently. Read status together with severity and notes; a warning is not automatically a defect in the filing.

The point-in-time financials response embeds a condensed version of these signals under _meta.reconciliation and _meta.trust_score, so a single as_of query tells you both what was knowable and how well it reconciles.

Accounting identities

A statement read from /v1/edgar/companies/{cik}/statements with include_normalized=true, or from /v1/edgar/companies/{cik}/filings/{accession_id}/statements, carries the accounting identities checked against it under normalized_payload._meta.identity:

{
  "pass_rate": "1",
  "total_rules_count": 6,
  "supported_rules_count": 3,
  "skipped_rules_count": 3,
  "failed_rules_count": 0,
  "checks": [
    {
      "rule_id": "E11_IDENTITY_BALANCE_SHEET",
      "statement_type": "BALANCE_SHEET",
      "fiscal_year": 2026,
      "fiscal_period": "Q3",
      "status": "PASS",
      "severity": "NONE",
      "expected_value": "383266000000",
      "actual_value": "383266000000",
      "delta": "0",
      "tolerance": "38326600.0000"
    }
  ]
}
  • Name
    status
    Type
    string
    Description

    PASS, FAIL, WARNING or SKIPPED. A check is skipped when it applies to another statement type, or when the filing did not report a figure the identity needs. Skipped checks carry null values and are left out of pass_rate.

  • Name
    delta
    Type
    decimal string
    Description

    actual_value minus expected_value.

  • Name
    tolerance
    Type
    decimal string
    Description

    The largest delta the check accepts. It scales with the size of the figures, so read a delta against it rather than against zero.

These checks are stricter than the checks that decide whether a statement is published, so a published statement can report a failing identity. Use the delta and tolerance to judge whether the difference matters to you.

Deterministic replay

Supported queries can be executed through a deterministic endpoint. It records an immutable manifest with the result, and gives you a replay token that returns that exact response later:

curl -X POST "https://api.arche.fi/v1/query/execute" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Content-Type: application/json" \
  -H "Accept: application/json" \
  -d '{
    "query_type": "financial_state_time_series",
    "as_of_date": "2025-03-31",
    "include_result": true,
    "parameters": {
      "cik": "0000320193",
      "fiscal_year_from": 2023,
      "fiscal_year_to": 2024,
      "period_type": "annual",
      "page": 1,
      "page_size": 50
    }
  }'

query_type is required. parameters carries the arguments for that query type, and as_of_date applies the no-lookahead boundary. Four query types are supported:

query_typeParameters
financial_state_time_seriescik; optional fiscal_period, fiscal_year_from, fiscal_year_to, period_type, page, page_size
metric_time_seriescik, metric; optional statement_type, fiscal_period, fiscal_year_from, fiscal_year_to, period_type, as_reported, page, page_size
metric_bundle_time_seriescik, bundle_code; optional fiscal_period, fiscal_year_from, fiscal_year_to, period_type, page, page_size
financials_as_ofcik; optional strategy_id, audit_mode. Requires as_of_date.

For financials_as_of, the manifest's snapshot_hashes holds the hash of the as-of snapshot the answer was read from. It's the same snapshot_hash the as-of financials endpoint returns. The time-series types read no snapshot, so their list is empty; they are pinned by input_hash instead.

The returned replay token can be looked up later:

curl -X GET "https://api.arche.fi/v1/query/replay/REPLAY_TOKEN" \
  -H "X-Api-Key: $ARCHE_API_KEY" \
  -H "X-Request-ID: $(uuidgen)" \
  -H "Accept: application/json"

Replay returns the original response exactly as it was first returned, in result. It is checked against the manifest's result_hash before it is served. Replay also re-runs the query at the manifest's as-of date, and replay_mode reports what that re-run found:

  • Name
    recomputed_verified
    Type
    replay_mode
    Description

    The re-run reproduced the original hash, and matches_original_result is true. This is the reproducibility proof.

  • Name
    recomputed_diverged
    Type
    replay_mode
    Description

    The underlying data has changed since, and matches_original_result is false. result is still the original; recomputed_result carries what the query returns today. This is a finding, not an error: it is the restatement showing up in your own result.

  • Name
    stored_manifest
    Type
    replay_mode
    Description

    The query can no longer be re-run. result still carries the original.

stored_result_available is false only for tokens minted before results were stored. For those, result is present when the re-run reproduces the original hash, which proves it is the same response.

Was this page helpful?