Important

Helios features are now enabled during weekly update windows and are no longer directly tied to SingleStore engine releases. Refer to the release notes to view the latest features available in your Helios cluster.

INFER PIPELINE

Creates a DDL definition for a pipeline and a target table based on input files. Returns a CREATE PIPELINE statement that can be reviewed, edited, and subsequently used to create the required pipeline. Use this command to view the inferred DDL.

Syntax

INFER PIPELINE AS LOAD DATA {input_configuration}
[FORMAT [CSV | JSON | AVRO | PARQUET | ICEBERG]]
[AS JSON]

Remarks

  • The input_configuration specifies configuration for loading files from Apache Kafka, Amazon S3, a local filesystem, Microsoft Azure, HDFS, and Google Cloud Storage. Refer to CREATE PIPELINE for more information on configuration specifications.

  • All options supported by CREATE PIPELINE are supported by INFER PIPELINE.

  • CSV, JSON, Avro, Parquet , and Iceberg formats are supported.

  • The default format is CSV.

  • TEXT and ENUM types use utf8mb4 charset and utf8mb4_bin collation by default.

  • The AS JSON keyword is used to produce pipeline and table definitions in JSON format.

  • For LOAD DATA CONNECTOR sources, INFER PIPELINE generates a CREATE TABLE and matching CREATE PIPELINE statement based on the connector's schema.

  • When the connector is recognized, INFER PIPELINE appends connector-specific fields to the generated CONFIG. This is currently implemented only for KinesisSourceConnector, which adds:

    • "kafka.topic": "s2-kinesis": Kinesis emits into a single topic; this field is not user-relevant but the connector requires it.

    • "output.topic": "false": Disables the topic field in the output document.

  • INFER PIPELINE for FORMAT JSON operates only on first-level fields of the connector's output. To ingest sub-fields (for example, value::data), edit the generated DDL manually or use an output mode to project the payload.

  • INFER PIPELINE for FORMAT AVRO derives column names from Avro record leaf nodes. If two leaves share a name, the generated column name uses the full path joined by "." (for example, key.partitionKey).

  • CREATE INFERRED PIPELINE does not support IF NOT EXISTS on the generated table. If you drop the pipeline and re-run INFER PIPELINE for a pipeline with the same name, the second CREATE TABLE fails with ER_TABLE_EXISTS_ERROR. Drop the table manually, or use CREATE PIPELINE (not INFERRED) when the table already exists.

Note

If the encoding of the source CSV file is not utf8mb4, multi-byte characters in the source file may be replaced with their corresponding single byte counterparts in the inferred table. This results in incorrect header inference for the inferred table.

To change the encoding of the source CSV file to utf8mb4 on a linux machine, run the following commands:

  1. Determine the current encoding of the CSV file.

    file -i input.csv
  2. Convert the file data into utf8mb4 encoded data.

    iconv -f <input-encoding> -t UTF-8 input.csv -o output.csv

Run the INFER PIPELINE query on the output.csv file to get the correct inference.

Example

The following example demonstrates how to use the INFER PIPELINE command to infer the schema of a Avro-formatted file in an AWS S3 bucket.

This example uses data that conforms to the schema of the books table, as shown in the following.

{"namespace": "books.avro",
"type": "record",
"name": "Book",
"fields": [
{"name": "id", "type": "int"},
{"name": "name", "type": "string"},
{"name": "num_pages", "type": "int"},
{"name": "rating", "type": "double"},
{"name": "publish_timestamp", "type": "long",
"logicalType": "timestamp-micros"} ]}

Refer to Generate an Avro File for an example of generating an Avro file that conforms to this schema.

The following example generates a table and pipeline definition by scanning the specified Avro file and inferring the schema from selected rows. The output is displayed in query definition format.

INFER PIPELINE AS LOAD DATA S3
's3://data_folder/books.avro'
CONFIG '{"region":"<region_name>"}'
CREDENTIALS '{
"aws_access_key_id":"<your_access_key_id>",
"aws_secret_access_key":"<your_secret_access_key>",
"aws_session_token":"<your_session_token>"}'
FORMAT AVRO;
"CREATE TABLE `infer_example_table` (
    `id` int(11) NOT NULL,
    `name` longtext CHARACTER SET utf8 COLLATE utf8_general_ci NOT NULL,
    `num_pages` int(11) NOT NULL,
    `rating` double NULL,
    `publish_date` bigint(20) NOT NULL);
CREATE PIPELINE `infer_example_pipeline`
AS LOAD DATA S3 's3://data-folder/books.avro'
CONFIG '{\""region\"":\""us-west-2\""}'
CREDENTIALS '{\n    \""aws_access_key_id\"":\""your_access_key_id\"",
\n    \""aws_secret_access_key\"":\""your_secret_access_key\"",
\n    \""aws_session_token\"":\""your_session_token\""}'
BATCH_INTERVAL 2500
DISABLE OUT_OF_ORDER OPTIMIZATION
DISABLE OFFSETS METADATA GC
INTO TABLE `infer_example_table`
FORMAT AVRO(
    `id` <- `id`,
    `name` <- `name`,
    `num_pages` <- `num_pages`,
    `rating` <- `rating`,
    `publish_date` <- `publish_date`);"

Refer to Schema and Pipeline Inference - Examples for more examples.

The following example uses INFER PIPELINE to generate a table and pipeline for an Amazon Kinesis stream. The generated CONFIG includes the connector-specific fields described in Remarks.

CREATE INFERRED PIPELINE kinesis_orders
AS LOAD DATA CONNECTOR 'KinesisSourceConnector'
CONFIG '{
"kinesis.stream": "orders-stream",
"kinesis.region": "us-east-1",
"tasks.max": "3"
}'
CREDENTIALS '{
"aws.access.key.id": "<ACCESS_KEY>",
"aws.secret.access.key": "<SECRET_KEY>"
}';

The command produces a CREATE TABLE matching the connector's output schema and a CREATE PIPELINE statement whose CONFIG includes the appended fields:

{
"kafka.topic": "s2-kinesis",
"kinesis.stream": "orders-stream",
"kinesis.region": "us-east-1",
"tasks.max": "3",
"output.topic": "false"
}

For more information on FORMAT, output modes, and column mapping, refer to Connector Pipelines.

Last modified:

Was this article helpful?

Verification instructions

Note: You must install cosign to verify the authenticity of the SingleStore file.

Use the following steps to verify the authenticity of singlestoredb-server, singlestoredb-toolbox, singlestoredb-studio, and singlestore-client SingleStore files that have been downloaded.

You may perform the following steps on any computer that can run cosign, such as the main deployment host of the cluster.

  1. (Optional) Run the following command to view the associated signature files.

    curl undefined
  2. Download the signature file from the SingleStore release server.

    • Option 1: Click the Download Signature button next to the SingleStore file.

    • Option 2: Copy and paste the following URL into the address bar of your browser and save the signature file.

    • Option 3: Run the following command to download the signature file.

      curl -O undefined
  3. After the signature file has been downloaded, run the following command to verify the authenticity of the SingleStore file.

    echo -n undefined |
    cosign verify-blob --certificate-oidc-issuer https://oidc.eks.us-east-1.amazonaws.com/id/CCDCDBA1379A5596AB5B2E46DCA385BC \
    --certificate-identity https://kubernetes.io/namespaces/freya-production/serviceaccounts/job-worker \
    --bundle undefined \
    --new-bundle-format -
    Verified OK

Try Out This Notebook to See What’s Possible in SingleStore

Get access to other groundbreaking datasets and engage with our community for expert advice.