Skip to main content
This guide shows you how to build an unattended integration that loads rows into a Narrative dataset and has Narrative write them as files into your S3 bucket. Once the dataset is connected, every new batch you add is delivered with no further calls. Before you start, complete the prerequisites and set up API access. You need the ID of an enabled profile, from List your profiles.

1. Create the dataset

Any schema works. The S3 Connector writes whatever columns the dataset has, in the file format of the profile.
Record the id in the response.

Required columns

The delivery interface requires no particular columns, and the columns can be plain strings, numbers, and timestamps. A row of the dataset above looks like this:

2. Activate the dataset

The response is 201 with the dataset, now "status": "active". Activation locks the schema, so activate only once the shape is settled.

3. Upload and ingest a file

Each line of the file is one row that matches the schema you declared in step 1. For the columns the S3 Connector accepts, see the Amazon S3 Connector reference and Required columns. You can load a file into a dataset over the API in two ways:
  • Signed-URL upload. Your integration uploads one file of up to 3 GB, then asks Narrative to ingest it.
  • Managed S3 bucket. You write files to an S3 bucket that Narrative manages, and Narrative ingests each batch on its own. Use a managed bucket for files larger than 3 GB, or for files that another system delivers on a schedule.

Choose a file format

The dataset’s file_config.type sets the format of every file you load into it. Parquet and JSON Lines both hold object columns, which nest properties inside one column:
  • Parquet (parquet) is the most compatible format for connector datasets. It stores nested struct columns and their types natively, and Narrative matches columns to the schema by name at every level.
  • JSON Lines (json) holds one JSON object per line. Narrative matches each nested object to the schema by name.
CSV datasets (flat) hold only scalar columns, so they can’t carry object columns.

Upload a file with a signed URL

Request an upload URL, then send the file straight to storage:
The upload URL is valid for 30 minutes and carries its own signature, so send no authorization header with the PUT.
Keep the path from the response. Narrative assigns its own storage path, which does not match the name you requested, and the ingest request needs Narrative’s path rather than yours.
Then ingest the file into the dataset, passing that path as source_file:
Ingestion runs in the background. Watch the record count on the dataset to know when it has finished:
The count moves from zero to your row count, typically within a couple of minutes. Each ingested file adds a new snapshot to the dataset, and every active connection on the dataset delivers that snapshot.

Write files to a managed S3 bucket

A managed bucket is an S3 bucket that Narrative creates for your company. You write each batch of files into its own folder under the dataset’s path in the bucket, then write an empty _NIO_COMMIT file into that batch folder. Narrative ingests every file in the batch folder when the commit file appears, so you make no upload or ingest request. A file can be as large as S3 accepts. See Ingesting Files from a Managed S3 Bucket to create the bucket, grant your AWS account access, and lay out the folders.

4. Confirm the S3 Connector accepts the dataset

Before you create a connection, ask Narrative which connector interfaces the dataset satisfies:
The response checks the dataset’s schema against every interface of the connectors your company has installed, and sorts the results into two lists:
  • accepted lists each interface you can connect the dataset to, by the connector’s app_id and the interface_id.
  • errors lists each interface the schema does not satisfy. Its details hold the reason, such as "required property '<column>' not found".
Add ?tags=<tag> to check only the interfaces that carry that tag. When the interface you want is under errors, the dataset’s schema doesn’t meet what the interface needs, for example a missing column or property. Activation locks the schema, so create a new dataset that fixes what the error names. For the S3 Connector, add ?tags=s3 and look for "app_id": 7 with the delivery interface in accepted:
The S3 Connector’s delivery interface accepts any dataset schema.

5. Connect the dataset to your bucket

Create a connection. It ties the dataset to the profile, and its quick_settings say where in the bucket the files land and what goes with them:
Before Narrative saves the connection, the connector checks that it can write, list, and delete objects under the prefix in your bucket. Record the connection id. You need it to stop the delivery. What the connection does describes every delivery setting, its allowed values, and its default.

6. Find the files in your bucket

Narrative writes each delivery to its own folder under the prefix. By default the folder is named for the dataset snapshot the delivery came from:
The data files take the extension of the profile’s file format. A folder is complete once _SUCCESS or _NIO_COMMIT appears in it, so have downstream jobs wait for one of those marker files before they read the folder. For the delivery that runs when the connection is created, the marker files go directly under the prefix. Each ingested batch is one snapshot of the dataset. To match folders to batches, list the snapshots:

7. Keep delivering

Upload and ingest the next file the same way as in step 3. Each ingested batch becomes a new snapshot, and the connection delivers it to a new folder under the prefix with no further calls.

8. Read and list connections

The first call returns the connection in the shape shown in step 5. The two list calls return connections in that shape inside records.

Stopping delivery

Deleting the connection archives it and stops further deliveries. Files already in your bucket stay there. Delete them in AWS if you no longer need them. To remove the dataset as well:

Getting help

Contact your Narrative relationship manager with your company ID, the dataset ID, the connection ID, and the failing request and response.

Amazon S3 Connector API

Prerequisites, bucket profile setup, and API access

Amazon S3 Connector

Supported file formats and delivery options

Connector Interfaces

Why a dataset connects to an interface rather than a connector

API Keys

Create and rotate keys for programmatic access