1. Create the dataset
Any schema works. The S3 Connector writes whatever columns the dataset has, in the file format of the profile.id in the response.
Required columns
Thedelivery interface requires no particular columns, and the columns can be plain strings, numbers, and timestamps. A row of the dataset above looks like this:
2. Activate the dataset
201 with the dataset, now "status": "active". Activation locks the schema, so activate only once the shape is settled.
3. Upload and ingest a file
Each line of the file is one row that matches the schema you declared in step 1. For the columns the S3 Connector accepts, see the Amazon S3 Connector reference and Required columns. You can load a file into a dataset over the API in two ways:- Signed-URL upload. Your integration uploads one file of up to 3 GB, then asks Narrative to ingest it.
- Managed S3 bucket. You write files to an S3 bucket that Narrative manages, and Narrative ingests each batch on its own. Use a managed bucket for files larger than 3 GB, or for files that another system delivers on a schedule.
Choose a file format
The dataset’sfile_config.type sets the format of every file you load into it. Parquet and JSON Lines both hold object columns, which nest properties inside one column:
- Parquet (
parquet) is the most compatible format for connector datasets. It stores nested struct columns and their types natively, and Narrative matches columns to the schema by name at every level. - JSON Lines (
json) holds one JSON object per line. Narrative matches each nested object to the schema by name.
flat) hold only scalar columns, so they can’t carry object columns.
Upload a file with a signed URL
Request an upload URL, then send the file straight to storage:PUT.
Then ingest the file into the dataset, passing that path as source_file:
Write files to a managed S3 bucket
A managed bucket is an S3 bucket that Narrative creates for your company. You write each batch of files into its own folder under the dataset’s path in the bucket, then write an empty_NIO_COMMIT file into that batch folder. Narrative ingests every file in the batch folder when the commit file appears, so you make no upload or ingest request. A file can be as large as S3 accepts.
See Ingesting Files from a Managed S3 Bucket to create the bucket, grant your AWS account access, and lay out the folders.
4. Confirm the S3 Connector accepts the dataset
Before you create a connection, ask Narrative which connector interfaces the dataset satisfies:acceptedlists each interface you can connect the dataset to, by the connector’sapp_idand theinterface_id.errorslists each interface the schema does not satisfy. Itsdetailshold the reason, such as"required property '<column>' not found".
?tags=<tag> to check only the interfaces that carry that tag.
When the interface you want is under errors, the dataset’s schema doesn’t meet what the interface needs, for example a missing column or property. Activation locks the schema, so create a new dataset that fixes what the error names.
For the S3 Connector, add ?tags=s3 and look for "app_id": 7 with the delivery interface in accepted:
delivery interface accepts any dataset schema.
5. Connect the dataset to your bucket
Create a connection. It ties the dataset to the profile, and itsquick_settings say where in the bucket the files land and what goes with them:
id. You need it to stop the delivery.
What the connection does describes every delivery setting, its allowed values, and its default.
6. Find the files in your bucket
Narrative writes each delivery to its own folder under the prefix. By default the folder is named for the dataset snapshot the delivery came from:_SUCCESS or _NIO_COMMIT appears in it, so have downstream jobs wait for one of those marker files before they read the folder. For the delivery that runs when the connection is created, the marker files go directly under the prefix.
Each ingested batch is one snapshot of the dataset. To match folders to batches, list the snapshots:
7. Keep delivering
Upload and ingest the next file the same way as in step 3. Each ingested batch becomes a new snapshot, and the connection delivers it to a new folder under the prefix with no further calls.8. Read and list connections
records.
Stopping delivery
Getting help
Contact your Narrative relationship manager with your company ID, the dataset ID, the connection ID, and the failing request and response.Related content
Amazon S3 Connector API
Prerequisites, bucket profile setup, and API access
Amazon S3 Connector
Supported file formats and delivery options
Connector Interfaces
Why a dataset connects to an interface rather than a connector
API Keys
Create and rotate keys for programmatic access

