Prerequisites
- An AWS account, and its 12-digit account ID
- The AWS CLI, or another S3 client, signed in to that account
- An active dataset whose schema matches your files. Narrative rejects ingestion into a dataset that is still
pending - To work through the API, an API key with write access on
resources
1. Create the bucket
Each company can have one managed bucket.- Platform UI
- API
- Open Settings and, under Infrastructure, select Sources.
- Click New Source. The only source type is Managed AWS S3 Bucket.
- Fill in the form:
- AWS Account ID: the 12-digit account that will write to the bucket.
- Resource ID: a short name that becomes part of the bucket name, such as your company name. Use lowercase letters, digits, and dashes.
- Access Type: see Choose an access type.
- Click Create.
nio-<resource ID>-<12 random characters>, for example nio-your-company-50ac58dabfa1. The random suffix keeps the name globally unique, and it stops anyone from guessing that a given company has a bucket. Copy the bucket name from the expanded row on the Sources page, or read it from Get buckets.
Choose an access type
You can switch later with Update access type, unless
is_access_mutable is false on the bucket.
2. Grant access in your AWS account
The bucket side is configured for you. What remains is to let the person or service that writes the files reach it.- Bucket Policy
- IAM Role
AWS requires a cross-account request to be allowed on both sides. Narrative’s side is the bucket policy. Your side is an IAM policy on the user or role that writes the files:
3. Lay out the files for one ingestion
Narrative works out which dataset a file belongs to from where it sits in the bucket. Put each batch of files in its own folder under this prefix:12345:
- The dataset ID comes from the path. It must be one of your own datasets, and it must be active.
- Name the batch folder however you like. A timestamp keeps batches apart and sorts them. Folders nested inside it are included too.
- Match the dataset’s file format. The dataset’s
file_configsets the format: CSV forflat, JSON Lines forjson, or Parquet. Flat files can be compressed with gzip (.gz) or bzip2 (.bz2). - For CSV, keep the dataset’s column order. Narrative reads CSV columns by position, not by header name, and casts each one to the schema type at that position.
- Batches of up to 100,000 files. A folder with more is rejected.
4. Commit the batch
Nothing is ingested until a file named_NIO_COMMIT appears in the batch folder. When it does, Narrative ingests every file in that folder at once. Write it last, after every data file has finished uploading:
append, the batch adds rows. With overwrite, the batch replaces everything already in the dataset.
5. Confirm the ingestion
Ingestion runs in the background and usually finishes within minutes. Larger batches take longer.- Platform UI
- API
Open the dataset and select the Changelog tab. Every completed ingestion is listed, newest first, with the records and bytes it added.
ingestion/ that is never committed is deleted after 7 days.
Troubleshooting
Narrative does not currently report failed bucket ingestions back to you. If the record count does not move, check the batch against this list, then contact Narrative support with the bucket name and folder path.Related content
Uploading Data
Upload a single file through a signed URL
Hashing PII for Upload
Hash identifiers before they leave your systems
Amazon S3 Connector
Deliver data from Narrative to your own S3 bucket
API Key Permissions
What the
resources permission covers
