# Bulk Datasets API Reference | SEC API > Complete API reference for the SEC EDGAR bulk datasets. The master index, the per-dataset index, the container download and the whole-dataset ZIP, with every response attribute, the container path format and the status codes. Source: https://sec-api.io/api-reference/bulk-datasets Filing search and retrieval GET`https://api.sec-api.io/bulk/indicies/master/index.json` Download complete SEC EDGAR datasets as compressed monthly archives instead of paging through a search endpoint. Close to 500 datasets cover the filings and exhibits published since 1993, together with the structured data extracted from them, and all of them are refreshed daily between 22:30 and 23:30 Eastern Time. Four endpoints make up the API: the master index lists the catalogue, the dataset index lists the archive files of one dataset, one path returns a single archive file, and one path returns the whole dataset as a single ZIP. The two index endpoints are public. The two download endpoints need an API key with an active subscription. [Read the guide for this API →](https://sec-api.io/datasets) ## Authentication Send the API key either as a header or as a query parameter. The header is preferred; the query parameter exists for cases where a header cannot be set, such as opening a URL directly in a browser. Authorization: required header The API key on its own. Do not prefix it with Bearer or any other word. Example `Authorization: YOUR_API_KEY` token: optional query parameter The API key, appended to the URL. Use this only when a header is not possible. ## Path parameters Supplied as part of the URL. dataset: required string Identifier of the dataset in the URL, taken from datasetIdInUrl in the master index. The UUID held in datasetId works in the same position. Used by /datasets/.json, /datasets// and /datasets/.zip. Example `audit-fees` container: required string Key of one archive file inside a dataset, in the form / followed by the container format of the dataset, so either .jsonl.gz or .zip. Take it from the key attribute of a container, or use the ready-made downloadUrl. A key that the dataset index does not list returns 404. Example `2026/2026-09.jsonl.gz` ## Response A JSON object. Nested attributes are collapsed; expand one to see its fields. ### Master index response GET /bulk/indicies/master/index.json returns the whole catalogue as a bare JSON array, sorted by name, with no envelope and no pagination. The attributes below describe one item. The same array is served at /datasets/master.json and /datasets/indicies/master.json. No API key is needed. Both index endpoints always answer with Content-Encoding: gzip, so a low-level client has to decompress the body itself, for example with curl --compressed. id: string UUID of the dataset. Always holds the same value as datasetId. datasetId: string UUID of the dataset, for example 1f11ba9b-e03a-6950-a464-a23fcc53ee6f. datasetIdInUrl: string Slug of the dataset, used in every dataset path and on the dataset page at sec-api.io/datasets, for example audit-fees. name: string Name of the dataset, for example Form 10-K - Annual Reports - Filing Contents. The array is sorted by this value. description: string What the dataset holds, which filings it covers and what it leaves out. formTypes: array of string EDGAR form types covered by the dataset, including amendments, for example ["10-K", "10-K/A", "10-K405", "10-KSB"]. containerFormat: string Format of every archive file in the dataset, either ZIP for datasets of original EDGAR files or .jsonl.gz for datasets of extracted structured data. fileTypes: array of string Types of file found inside the archives. Common values are HTML, TXT, JSON, PDF, XML and JSONL. updatedAt: string When the dataset last changed, ISO 8601 in UTC. Every dataset is refreshed daily between 22:30 and 23:30 Eastern Time. earliestSampleDate: string First day the dataset covers, YYYY-MM-DD. The oldest datasets start in October 1993. totalRecords: integer or null Number of records across the whole dataset. Null on the eight datasets published as .jsonl.gz, where record counts are not tracked. totalSize: integer Size in bytes of all archive files added together, compressed. Uncompressed data is roughly 10 to 15 times larger. ### Dataset index response GET /datasets/.json returns one dataset with the full list of its archive files, and it needs no API key either. Use it to find the container keys and download URLs, and to compare sizes before downloading again. The same object is served at /datasets/indicies/.json, and at /bulk/indicies//index.json when you pass the UUID rather than the slug. datasetId: string UUID of the dataset. datasetDownloadUrl: string URL that returns the whole dataset as one ZIP archive. Append your API key as token, or send it as a header. name: string Name of the dataset. description: string What the dataset holds. updatedAt: string When the dataset last changed, ISO 8601 in UTC. Matches the newest container. earliestSampleDate: string First day the dataset covers, YYYY-MM-DD. totalRecords: integer or null Number of records across the whole dataset, or null where record counts are not tracked. totalSize: integer Size in bytes of all containers added together, compressed. formTypes: array of string EDGAR form types covered by the dataset. containerFormat: string Either ZIP or .jsonl.gz. fileTypes: array of string Types of file found inside the archives. containers: array of object One entry per archive file, newest first. Each entry covers one calendar month, grouped by the date EDGAR accepted the filing. A dataset that reaches back to 1993 has around 390 entries. downloadUrl: string Full URL of the archive file, for example https://api.sec-api.io/datasets/audit-fees/2026/2026-09.jsonl.gz. Append your API key as token, or send it as a header. key: string Path of the archive file inside the dataset, / plus the container format, for example 2026/2026-09.jsonl.gz. Use it as the local file path to mirror the layout of the dataset. size: integer Size of the archive file in bytes. Compare it with the size of your local copy to decide whether to download the file again. records: optional integer Number of records in the archive file. Present on ZIP datasets, absent on .jsonl.gz datasets. updatedAt: string When this archive file last changed, ISO 8601 in UTC. Months in the past still change when EDGAR publishes a late filing. ### Container download response GET /datasets// returns one monthly archive file, streamed straight through. This needs an API key with an active subscription. Download containers one at a time to keep a local copy in sync: skip a file whose local size already matches size in the dataset index, and fetch the rest. body: binary The archive file itself, streamed as stored, with no JSON envelope around it. A .jsonl.gz container holds one JSON record per line once decompressed. A .zip container holds the original EDGAR files. Content-Type: response header application/gzip for a .jsonl.gz container, application/zip for a .zip container. Content-Disposition: response header attachment, with the file name of the container, for example attachment; filename="2026-09.jsonl.gz". ### Whole dataset download response GET /datasets/.zip packs every container of the dataset into one ZIP archive and streams it. This needs an API key with an active subscription. It is the simplest way to take a first full copy, though it transfers the whole dataset every time, so a repeated sync is cheaper through the container downloads. body: binary One ZIP archive holding every container of the dataset, built and streamed while you download it. It uses the ZIP64 format, so it carries datasets far beyond 4 GB. The largest datasets pass 300 GB compressed, so prefer the per-container downloads when you plan to sync repeatedly. Content-Type: response header application/zip. Content-Disposition: response header attachment, with a file name built from the dataset name, for example attachment; filename="Audit_Fees.zip". index.json: member of the archive Metadata of the dataset, written into the archive at the top level. It holds downloadedAt plus the same attributes as the dataset index endpoint, including the complete containers list. /: members of the archive One member per container, stored at the same path as its key and in the order newest first. The members keep their own compression and are not compressed a second time. ## Status codes | | | | --- | --- | | `200` | Success. An index endpoint returns JSON, a download endpoint streams the archive file. | | `400` | The archive could not be read while it was being streamed. Retry the download. | | `403` | The API key is missing or not valid, or it carries no active subscription. Both download endpoints require an active subscription. The two index endpoints need no key at all. | | `404` | No dataset carries that name or id, or the container key is not listed in the dataset index. | | `429` | Too many requests. The limit is two requests per second, counted per API key, or per IP address when no key is sent. Slow the request rate and retry. | | `500` | Server error. Retry, and report it if it persists. | ## Request example GET https://api.sec-api.io/bulk/indicies/master/index.json ```bash https://api.sec-api.io/bulk/indicies/master/index.json https://api.sec-api.io/datasets/audit-fees.json https://api.sec-api.io/datasets/audit-fees/2026/2026-09.jsonl.gz?token=YOUR_API_KEY https://api.sec-api.io/datasets/audit-fees.zip?token=YOUR_API_KEY ``` ```python from sec_api import Datasets datasets = Datasets(api_key="YOUR_API_KEY") # no API key needed for the two index calls all_datasets = datasets.get_all() details = datasets.get_dataset_details("form-10k-content") # first run: downloads all containers to ./sec-api-datasets/form-10k-content/ datasets.download("form-10k-content") # later runs: only new or updated containers, the rest are skipped datasets.sync("form-10k-content") # custom directory datasets.download("form-10k-content", path="./my-data/form-10k-content") # or the whole dataset as one ZIP file datasets.download("form-10k-content", strategy="zip") ``` ```javascript import { datasetsApi } from "sec-api"; datasetsApi.setApiKey("YOUR_API_KEY"); // no API key needed for the two index calls const allDatasets = await datasetsApi.getAll(); const details = await datasetsApi.getDetails("form-10k-content"); // first run: downloads all containers to ./sec-api-datasets/form-10k-content/ await datasetsApi.download("form-10k-content"); // later runs: only new or updated containers, the rest are skipped await datasetsApi.sync("form-10k-content"); // custom directory await datasetsApi.download({ name: "form-10k-content", path: "./my-data/form-10k-content", }); // or the whole dataset as one ZIP file await datasetsApi.download({ name: "form-10k-content", strategy: "zip" }); ``` ```bash curl --compressed "https://api.sec-api.io/bulk/indicies/master/index.json" curl --compressed "https://api.sec-api.io/datasets/audit-fees.json" curl -H "Authorization: YOUR_API_KEY" \ -o 2026-09.jsonl.gz \ "https://api.sec-api.io/datasets/audit-fees/2026/2026-09.jsonl.gz" curl -H "Authorization: YOUR_API_KEY" \ -o audit-fees.zip \ "https://api.sec-api.io/datasets/audit-fees.zip" ``` ## Response example 200 OK · application/json or application/gzip or application/zip ```json [ { "id": "1f11ba9b-e03a-6950-a464-a23fcc53ee6f", "datasetId": "1f11ba9b-e03a-6950-a464-a23fcc53ee6f", "datasetIdInUrl": "audit-fees", "name": "Audit Fees", "description": "Structured dataset of annual audit fees extracted from SEC filings. Includes details of audited company (ticker, CIK, company name, disclosure date, etc.), audit fees, audit-related fees, tax fees, and all other fees disclosed, including the auditor of the company.", "formTypes": ["DEF 14A"], "containerFormat": ".jsonl.gz", "fileTypes": ["JSONL"], "updatedAt": "2026-09-26T05:00:03.000Z", "earliestSampleDate": "2001-03-01", "totalRecords": null, "totalSize": 10070006 }, { "id": "1f11bb55-d58b-6080-bace-e7a62567f4b9", "datasetId": "1f11bb55-d58b-6080-bace-e7a62567f4b9", "datasetIdInUrl": "form-10k-content", "name": "Form 10-K - Annual Reports - Filing Contents", "description": "HTML and TXT files of all Form 10-K filings published since 1993. Includes all EDGAR 10-K form variations such as 10-K/A, 10-KSB, and 10-KT. Each file represents the original filing document as published on EDGAR, including inline-XBRL where applicable. Images, exhibits and XML/XBRL files are not included.", "formTypes": [ "10-K", "10-K/A", "10-K405", "10-K405/A", "10-KSB", "10-KSB/A", "10-KT", "10-KT/A" ], "containerFormat": "ZIP", "fileTypes": ["TXT", "JSON", "HTML", "PAPER"], "updatedAt": "2026-09-26T02:46:56.259Z", "earliestSampleDate": "1993-10-01", "totalRecords": 305479, "totalSize": 33989121092 } ] ``` ```json { "datasetId": "1f11bb55-d58b-6080-bace-e7a62567f4b9", "datasetDownloadUrl": "https://api.sec-api.io/datasets/form-10k-content.zip", "name": "Form 10-K - Annual Reports - Filing Contents", "description": "HTML and TXT files of all Form 10-K filings published since 1993. Includes all EDGAR 10-K form variations such as 10-K/A, 10-KSB, and 10-KT. Each file represents the original filing document as published on EDGAR, including inline-XBRL where applicable. Images, exhibits and XML/XBRL files are not included.", "updatedAt": "2026-09-26T02:46:56.259Z", "earliestSampleDate": "1993-10-01", "totalRecords": 305479, "totalSize": 33989121092, "formTypes": ["10-K", "10-K/A", "10-K405", "10-K405/A", "10-KSB", "10-KSB/A", "10-KT", "10-KT/A"], "containerFormat": "ZIP", "fileTypes": ["TXT", "JSON", "HTML", "PAPER"], "containers": [ { "downloadUrl": "https://api.sec-api.io/datasets/form-10k-content/2026/2026-09.zip", "key": "2026/2026-09.zip", "size": 21664150, "records": 220, "updatedAt": "2026-09-26T02:46:56.259Z" }, { "downloadUrl": "https://api.sec-api.io/datasets/form-10k-content/2026/2026-08.zip", "key": "2026/2026-08.zip", "size": 25690554, "records": 282, "updatedAt": "2026-09-01T02:49:27.000Z" } ] } ``` ```json GET /datasets/audit-fees/2026/2026-09.jsonl.gz HTTP/1.1 200 OK Content-Type: application/gzip Content-Disposition: attachment; filename="2026-09.jsonl.gz" GET /datasets/form-10k-content/2026/2026-09.zip HTTP/1.1 200 OK Content-Type: application/zip Content-Disposition: attachment; filename="2026-09.zip" ``` ```json HTTP/1.1 200 OK Content-Type: application/zip Content-Disposition: attachment; filename="Audit_Fees.zip" Archive: Audit_Fees.zip index.json 2026/2026-09.jsonl.gz 2026/2026-08.jsonl.gz ... 2001/2001-02.jsonl.gz 2001/2001-01.jsonl.gz ```