Skip to main content
A managed S3 bucket is an Amazon S3 bucket that Narrative creates in its own AWS account and assigns to your company. You write files into it with your own AWS tools, and Narrative ingests them into the dataset that the folder path names. You make no API call for each file, and nothing caps a file below what S3 itself accepts. Use a managed bucket when your files are larger than the 3 GB that single-file uploads support, or when another system writes files on a schedule.

Prerequisites

  • An AWS account, and its 12-digit account ID
  • The AWS CLI, or another S3 client, signed in to that account
  • An active dataset whose schema matches your files. Narrative rejects ingestion into a dataset that is still pending
  • To work through the API, an API key with write access on resources

1. Create the bucket

Each company can have one managed bucket.
  1. Open Settings and, under Infrastructure, select Sources.
  2. Click New Source. The only source type is Managed AWS S3 Bucket.
  3. Fill in the form:
    • AWS Account ID: the 12-digit account that will write to the bucket.
    • Resource ID: a short name that becomes part of the bucket name, such as your company name. Use lowercase letters, digits, and dashes.
    • Access Type: see Choose an access type.
  4. Click Create.
Once your company has a bucket, New Source is disabled.
Narrative names the bucket nio-<resource ID>-<12 random characters>, for example nio-your-company-50ac58dabfa1. The random suffix keeps the name globally unique, and it stops anyone from guessing that a given company has a bucket. Copy the bucket name from the expanded row on the Sources page, or read it from Get buckets.

Choose an access type

You can switch later with Update access type, unless is_access_mutable is false on the bucket.

2. Grant access in your AWS account

The bucket side is configured for you. What remains is to let the person or service that writes the files reach it.
AWS requires a cross-account request to be allowed on both sides. Narrative’s side is the bucket policy. Your side is an IAM policy on the user or role that writes the files:
Every upload must set the bucket-owner-full-control ACL. The bucket policy rejects any PutObject without it.

3. Lay out the files for one ingestion

Narrative works out which dataset a file belongs to from where it sits in the bucket. Put each batch of files in its own folder under this prefix:
For example, a batch for dataset 12345:
  • The dataset ID comes from the path. It must be one of your own datasets, and it must be active.
  • Name the batch folder however you like. A timestamp keeps batches apart and sorts them. Folders nested inside it are included too.
  • Match the dataset’s file format. The dataset’s file_config sets the format: CSV for flat, JSON Lines for json, or Parquet. Flat files can be compressed with gzip (.gz) or bzip2 (.bz2).
  • For CSV, keep the dataset’s column order. Narrative reads CSV columns by position, not by header name, and casts each one to the schema type at that position.
  • Batches of up to 100,000 files. A folder with more is rejected.
Upload the files:
The AWS CLI splits large files into multipart uploads on its own, so a single file can go well past 5 GB.

4. Commit the batch

Nothing is ingested until a file named _NIO_COMMIT appears in the batch folder. When it does, Narrative ingests every file in that folder at once. Write it last, after every data file has finished uploading:
The commit file can be empty. It can also carry an idempotency key, so that if your job commits the same batch twice, the data is still written once:
The dataset’s write mode decides what each commit does. With append, the batch adds rows. With overwrite, the batch replaces everything already in the dataset.

5. Confirm the ingestion

Ingestion runs in the background and usually finishes within minutes. Larger batches take longer.
Open the dataset and select the Changelog tab. Every completed ingestion is listed, newest first, with the records and bytes it added.
After a batch is processed, Narrative moves its files out of your bucket, whether ingestion succeeded or failed. An empty batch folder means only that the batch was picked up. Use the dataset, not the bucket, to confirm success. Anything you leave under ingestion/ that is never committed is deleted after 7 days.

Troubleshooting

Narrative does not currently report failed bucket ingestions back to you. If the record count does not move, check the batch against this list, then contact Narrative support with the bucket name and folder path.

Uploading Data

Upload a single file through a signed URL

Hashing PII for Upload

Hash identifiers before they leave your systems

Amazon S3 Connector

Deliver data from Narrative to your own S3 bucket

API Key Permissions

What the resources permission covers