> ## Documentation Index
> Fetch the complete documentation index at: https://docs.narrative.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Ingesting Files from a Managed S3 Bucket

> Load files into a dataset by writing them to a Narrative-managed S3 bucket, with no size limit per file and no API call per upload

A [managed S3 bucket](/reference/glossary#managed-s3-bucket) is an Amazon S3 bucket that Narrative creates in its own AWS account and assigns to your company. You write files into it with your own AWS tools, and Narrative ingests them into the dataset that the folder path names. You make no API call for each file, and nothing caps a file below what S3 itself accepts.

Use a managed bucket when your files are larger than the 3 GB that [single-file uploads](/guides/sdk/uploading-data) support, or when another system writes files on a schedule.

## Prerequisites

* An AWS account, and its 12-digit account ID
* The [AWS CLI](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html), or another S3 client, signed in to that account
* An **active** dataset whose schema matches your files. Narrative rejects ingestion into a dataset that is still `pending`
* To work through the API, an [API key](/account-settings/api-keys) with write access on `resources`

## 1. Create the bucket

Each company can have one managed bucket.

<Tabs>
  <Tab title="Platform UI">
    1. Open **Settings** and, under **Infrastructure**, select **Sources**.
    2. Click **New Source**. The only source type is **Managed AWS S3 Bucket**.
    3. Fill in the form:
       * **AWS Account ID**: the 12-digit account that will write to the bucket.
       * **Resource ID**: a short name that becomes part of the bucket name, such as your company name. Use lowercase letters, digits, and dashes.
       * **Access Type**: see [Choose an access type](#choose-an-access-type).
    4. Click **Create**.

    Once your company has a bucket, **New Source** is disabled.
  </Tab>

  <Tab title="API">
    ```bash theme={null}
    curl -X POST https://api.narrative.io/resources/buckets \
      -H "Authorization: Bearer $NIO_API_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{
        "account_id": "123456789012",
        "resource_id": "your-company",
        "access": { "type": "bucket_policy" }
      }'
    ```

    If you leave out `access`, the bucket uses IAM role access. A second request from the same company returns `400`.

    See [Create bucket](/api-reference/resources/create-bucket) for every field.
  </Tab>
</Tabs>

Narrative names the bucket `nio-<resource ID>-<12 random characters>`, for example `nio-your-company-50ac58dabfa1`. The random suffix keeps the name globally unique, and it stops anyone from guessing that a given company has a bucket. Copy the bucket name from the expanded row on the **Sources** page, or read it from [Get buckets](/api-reference/resources/get-buckets).

### Choose an access type

| Access type | How your AWS account reaches the bucket | Choose it when |
| - | - | - |
| **Bucket Policy** (recommended) | The bucket's policy grants access to your AWS account, so any principal in that account that your own IAM policies allow can use it | You want to copy objects across buckets with your own credentials |
| **IAM Role** | Your principals assume a role that Narrative creates and manages, optionally protected by an external ID | You want access to go through one role you can audit. Cross-account bucket copies are not supported |

You can switch later with [Update access type](/api-reference/resources/update-access-type), unless `is_access_mutable` is `false` on the bucket.

## 2. Grant access in your AWS account

The bucket side is configured for you. What remains is to let the person or service that writes the files reach it.

<Tabs>
  <Tab title="Bucket Policy">
    AWS requires a cross-account request to be allowed on both sides. Narrative's side is the bucket policy. Your side is an IAM policy on the user or role that writes the files:

    ```json theme={null}
    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Action": ["s3:ListBucket", "s3:GetBucketLocation"],
          "Resource": "arn:aws:s3:::nio-your-company-50ac58dabfa1"
        },
        {
          "Effect": "Allow",
          "Action": ["s3:PutObject", "s3:GetObject", "s3:DeleteObject"],
          "Resource": "arn:aws:s3:::nio-your-company-50ac58dabfa1/*"
        }
      ]
    }
    ```

    <Warning>
      Every upload must set the `bucket-owner-full-control` ACL. The bucket policy rejects any `PutObject` without it.
    </Warning>
  </Tab>

  <Tab title="IAM Role">
    Expand the bucket's row on the **Sources** page. Copy the **Role ARN**, then copy the policy shown under **Configure your AWS account to allow access**. Attach that policy to the user or role that writes the files. It allows `sts:AssumeRole` on the Narrative role:

    ```json theme={null}
    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Sid": "AssumeNarrativeBucketAccessRole",
          "Effect": "Allow",
          "Action": "sts:AssumeRole",
          "Resource": "arn:aws:iam::704349335716:role/nio-your-company-50ac58dabfa1-access"
        }
      ]
    }
    ```

    Then add a profile for the role to `~/.aws/config`. Include `external_id` only if you set one when you created the bucket:

    ```ini theme={null}
    [profile narrative-bucket]
    role_arn = arn:aws:iam::704349335716:role/nio-your-company-50ac58dabfa1-access
    source_profile = default
    external_id = your-external-id
    ```

    Pass `--profile narrative-bucket` to the AWS CLI commands in the rest of this guide.
  </Tab>
</Tabs>

## 3. Lay out the files for one ingestion

Narrative works out which dataset a file belongs to from where it sits in the bucket. Put each batch of files in its own folder under this prefix:

```text theme={null}
s3://<bucket>/ingestion/datasets/version=1/datasetId=<dataset ID>/<batch folder>/
```

For example, a batch for dataset `12345`:

```text theme={null}
s3://nio-your-company-50ac58dabfa1/ingestion/datasets/version=1/datasetId=12345/2026-09-30T1400/part-0001.csv.gz
s3://nio-your-company-50ac58dabfa1/ingestion/datasets/version=1/datasetId=12345/2026-09-30T1400/part-0002.csv.gz
```

* **The dataset ID comes from the path.** It must be one of your own datasets, and it must be active.
* **Name the batch folder however you like.** A timestamp keeps batches apart and sorts them. Folders nested inside it are included too.
* **Match the dataset's file format.** The dataset's `file_config` sets the format: CSV for `flat`, JSON Lines for `json`, or Parquet. Flat files can be compressed with gzip (`.gz`) or bzip2 (`.bz2`).
* **For CSV, keep the dataset's column order.** Narrative reads CSV columns by position, not by header name, and casts each one to the schema type at that position.
* **Batches of up to 100,000 files.** A folder with more is rejected.

Upload the files:

```bash theme={null}
aws s3 cp ./batch/ \
  s3://nio-your-company-50ac58dabfa1/ingestion/datasets/version=1/datasetId=12345/2026-09-30T1400/ \
  --recursive \
  --acl bucket-owner-full-control
```

The AWS CLI splits large files into multipart uploads on its own, so a single file can go well past 5 GB.

## 4. Commit the batch

Nothing is ingested until a file named `_NIO_COMMIT` appears in the batch folder. When it does, Narrative ingests every file in that folder at once. Write it last, after every data file has finished uploading:

```bash theme={null}
aws s3 cp /dev/null \
  s3://nio-your-company-50ac58dabfa1/ingestion/datasets/version=1/datasetId=12345/2026-09-30T1400/_NIO_COMMIT \
  --acl bucket-owner-full-control
```

The commit file can be empty. It can also carry an idempotency key, so that if your job commits the same batch twice, the data is still written once:

```json theme={null}
{"idempotency_key": "6f1c2a44-3b7e-4d0e-9a51-2f8c7d1e0b93"}
```

The dataset's [write mode](/reference/glossary#write-mode) decides what each commit does. With `append`, the batch adds rows. With `overwrite`, the batch replaces everything already in the dataset.

## 5. Confirm the ingestion

Ingestion runs in the background and usually finishes within minutes. Larger batches take longer.

<Tabs>
  <Tab title="Platform UI">
    Open the dataset and select the **Changelog** tab. Every completed ingestion is listed, newest first, with the records and bytes it added.
  </Tab>

  <Tab title="API">
    ```bash theme={null}
    curl -s https://api.narrative.io/datasets/12345 \
      -H "Authorization: Bearer $NIO_API_TOKEN"
    # → .stats.active_dataset_stored_records
    ```

    The record count rises by the number of rows in the batch.
  </Tab>
</Tabs>

After a batch is processed, Narrative moves its files out of your bucket, whether ingestion succeeded or failed. An empty batch folder means only that the batch was picked up. Use the dataset, not the bucket, to confirm success.

Anything you leave under `ingestion/` that is never committed is deleted after 7 days.

## Troubleshooting

Narrative does not currently report failed bucket ingestions back to you. If the record count does not move, check the batch against this list, then contact Narrative support with the bucket name and folder path.

| Symptom | Likely cause |
| - | - |
| `AccessDenied` on upload | The `bucket-owner-full-control` ACL is missing (bucket policy access), or your IAM policy does not allow the action |
| Files stay in the folder and nothing ingests | There is no `_NIO_COMMIT` in the folder, or the path is not under `ingestion/datasets/version=1/datasetId=<ID>/` |
| Files disappear but no rows are added | The dataset is not active, the dataset belongs to another company, the folder held only the commit file, or the files do not match the schema |
| CSV values land in the wrong columns | The file's column order differs from the schema's |

## Related content

<CardGroup cols={2}>
  <Card title="Uploading Data" icon="upload" href="/guides/sdk/uploading-data">
    Upload a single file through a signed URL
  </Card>

  <Card title="Hashing PII for Upload" icon="hashtag" href="/guides/ingestion/hashing-pii">
    Hash identifiers before they leave your systems
  </Card>

  <Card title="Amazon S3 Connector" icon="aws" href="/reference/connectors/amazon-s3">
    Deliver data from Narrative to your own S3 bucket
  </Card>

  <Card title="API Key Permissions" icon="key" href="/reference/security/permissions">
    What the `resources` permission covers
  </Card>
</CardGroup>
