Skip to main content

TiDB Data Source (Beta)

A TiDB data source stores the object storage location, credentials, and optional event queue that TiDB Cloud Lake uses to read data staged by your TiDB cluster. It does not connect to the TiDB server itself. Dumpling and TiCDC write exports and change events to an object storage bucket, and TiDB Cloud Lake reads from that bucket.

Use Cases​

  • Manage the staging bucket, credentials, and optional SQS queue centrally for multiple TiDB sync tasks
  • Avoid re-entering the same bucket and authorization settings in every task
  • Use IAM Role instead of static keys so TiDB Cloud Lake obtains short-lived credentials
  • Update the bucket, role, or queue in one place when it is referenced by multiple tasks

Prepare TiDB for Integration​

TiDB Cloud exports full snapshots (Dumpling) and incremental changes (TiCDC) to an object storage bucket. Due to the differences across TiDB Cloud plans, we recommend a specific setup path for each plan to keep configuration minimal and ensure data compatibility.

Create TiDB Data Source​

  1. Navigate to Data > Data Sources, then click Create.

  2. Select TiDB as the service, then fill in the Name of the data source.

  3. Choose a Storage Provider and an Authentication Method, then fill in the connection details. The fields depend on the combination you select. See Amazon S3 or Alibaba Cloud OSS below.

    Use the same storage provider and region as your TiDB Cloud Lake deployment when possible.

  4. Click Test Connectivity to validate the bucket and credentials. If the test succeeds, click OK to save the data source.

Amazon S3​

When the storage provider is Amazon S3, choose one of the following authentication methods.

Role ARN uses the AssumeRole model. TiDB Cloud Lake assumes an IAM Role in your AWS account and obtains temporary credentials, so you never hand static keys to TiDB Cloud Lake.

Before you can save the data source, your IAM Role's trust policy must trust the two platform roles (setup/validation and data loading) with the corresponding External ID as the sts:ExternalId condition. See Authenticate with AWS IAM Role for the full trust policy setup.

FieldRequiredDescription
NameYesA descriptive name for this data source
Storage ProviderYesSelect Amazon S3
Authentication MethodYesSelect Role ARN
Role ARNYesIAM Role ARN in your AWS account that TiDB Cloud Lake is allowed to assume, for example arn:aws:iam::123456789012:role/tidbcloud-lake-tidb
S3 Bucket NameYesBucket where TiCDC / Dumpling stages the data that TiDB Cloud Lake loads
S3 RegionYesAWS Region of the bucket, for example us-east-1
S3 EndpointNoOnly for S3-compatible storage such as MinIO, for example http://localhost:9000. Leave empty for Amazon S3
SQS Queue URLNoOptional SQS standard queue URL for event-driven mode. See Optional SQS Queue

Authentication Method: Access Key / Secret Key​

Use this method when you prefer static credentials, for example for an S3-compatible store that does not support role assumption.

For more details, see Amazon S3 - Credentials.

FieldRequiredDescription
NameYesA descriptive name for this data source
Storage ProviderYesSelect Amazon S3
Authentication MethodYesSelect Access Key / Secret Key
S3 Access KeyYesAccess key ID with access to the staging bucket
S3 Secret KeyYesSecret access key paired with the access key ID
S3 Bucket NameYesBucket where TiCDC / Dumpling stages the data that TiDB Cloud Lake loads
S3 RegionYesAWS Region of the bucket
S3 EndpointNoOnly for S3-compatible storage. Leave empty for Amazon S3
SQS Queue URLNoOptional SQS standard queue URL for event-driven mode

Alibaba Cloud OSS​

When the storage provider is Alibaba Cloud OSS, TiDB Cloud Lake reads the staging bucket from Alibaba Cloud. OSS supports Access Key / Secret Key authentication only.

FieldRequiredDescription
NameYesA descriptive name for this data source
Storage ProviderYesSelect Alibaba Cloud OSS
OSS Access Key IDYesOSS AccessKey ID
OSS AccessKey SecretYesOSS AccessKey secret
OSS BucketYesOSS bucket where TiCDC / Dumpling stages the data
OSS RegionRead-onlyThe Lake's deployment region. It is displayed automatically and cannot be changed

Optional SQS Queue​

The SQS Queue URL field is optional and applies to Amazon S3 only. When you provide a standard SQS queue that receives S3 ObjectCreated events for the staging bucket, TiDB Cloud Lake can discover newly written changefeed / export objects from the queue instead of waiting for the next poll.

  • The queue must be a standard queue. FIFO queues are not supported, because S3 event notifications cannot be delivered to them.
  • The S3 bucket and the SQS queue should be in the same Region.
  • Polling the bucket remains the authoritative discovery path, and SQS is a latency optimization, not a replacement.
  • If you leave this field empty, the task discovers objects by polling.

For queue, bucket notification, and trust policy setup, see Amazon SQS (S3) - IAM Role.

Next Steps​

After creating this data source, you can use it to create a TiDB Integration Task.

Try Databend Cloud for FREE

Multimodal, object-storage-native warehouse for BI, vectors, search, and geo.

Snowflake-compatible SQL with automatic scaling.

Sign up and get $200 in credits.

Try it today