Apache Kafka
This guide walks you through connecting Kafka as a data source in ADOC, configuring observability, and enabling advanced features like concurrency control and freshness policies.
Prerequisites
Ensure the following requirements are met before you connect Kafka as a data source:
- A running Kafka cluster with accessible topics that can publish messages in a supported format such as JSON, Avro, or Confluent Avro.
- An active ADOC Data Plane with access to Kafka brokers
- Depending on your Kafka setup, you may need the following security credentials:
- Username and password (for SASL or Basic Auth)
- Certificate Authority (CA) or Server Certificate (SSL) certificates (keystore and truststore)
- Kerberos principal and keytab
Add Kafka as a Data Source
Follow these steps to set up HDFS in ADOC:
Step 1: Start Setup
- Navigate to the left main navigation menu, click Control Center -> Integrations.
- From the Integrations page, click Add Data Source.
- Select Kafka from the list of data sources.
- On the Data Source Details page:
- Enter a unique name for this data source.
- Optionally, add a brief description to clarify its purpose.
- Ensure the Data Reliability toggle is enabled and select your data plane from the drop-down list.
- Select Next to proceed.
Step 2: Add Connection Details
Common Fields (Displayed for All Protocols)
Field | Description |
Bootstrap Servers | Kafka broker address (e.g., |
Schema Registry Server | URL of the Schema Registry, if applicable. Optional. |
Schema Registry Authentication Type | If you provide a Schema Registry URL, you must also configure its authentication details (if required by your Kafka setup). Authentication Type: Select Basic Auth as the authentication type and provide the associated username and password. Username and Password: Enter the credentials used to access the schema registry. |
Protocol-Specific Fields
To understand each protocol, refer to the Apache Kafka security protocol documentation.
Security Protocol | Field | Description |
Plain Text | – | No additional fields required. |
SSL | Use data plane for connection files | Toggle ON to fetch SSL files from data plane; OFF to upload manually. If Use data plane for connection files is enabled, the Keystore/Truststore file location fields will not appear. These files are fetched automatically from the Data Plane configuration. |
SSL Keystore FileLocation | Upload the Keystore file containing the client certificate. | |
SSL Keystore Password | Password to unlock the Keystore file. | |
SSL Truststore FileLocation | Upload the Truststore file used to validate the server certificate. | |
SSL Truststore Password | Password to unlock the Truststore file. | |
Kafka Additional Properties | Optional key-value pairs for advanced connection configuration. | |
SASL Plain Text | SASL Mechanism | Select the authentication mechanism (e.g., |
Kerberos Principal | Kerberos identity used to authenticate (e.g., | |
Kerberos Service Name | Kafka service name used during Kerberos authentication (e.g., | |
Kerberos KeyTab File Location | Path to the KeyTab file containing credentials. Required for passwordless authentication. | |
Kafka Additional Properties | Optional key-value pairs, including authentication details. | |
SASL_SSL | SASL Mechanism | Select the SASL authentication mechanism. |
SSL Verification Required | Toggle ON to enforce SSL certificate validation. | |
Use data plane for connection files | Toggle ON to fetch SSL files from data plane; OFF to upload manually. | |
SSL Keystore FileLocation | Upload the Keystore file. | |
SSL Keystore Password | Password for the Keystore file. | |
SSL Truststore FileLocation | Upload the Truststore file. | |
SSL Truststore Password | Password for the Truststore file. | |
Kafka Additional Properties | Optional key-value configuration. |
- Select Test Connection. If successful, you’ll see “Connected.” If the test fails, ensure your bootstrap server is reachable, credentials are correct, and that the ADOC Data Plane service (
) is running.ad-analysis-standalone - Select Next to proceed.
Step 3: Setup Observability
Configure how ADOC will monitor your Kafka topic:
- Asset Name – Descriptive name for this Kafka feed.
- Topic Name – Single topic or topic pattern.
- Topic Format – Choose
,JSON, orAvro.Confluent Avro - Schema Registry URL – Required for Avro formats.
- (Optional) Enable Job Concurrency Control and set Max Slots to manage parallel profiling. For more information, see Control Plane Concurrent Connections and Queueing Mechanism
- Click Submit to finalize setup.
Improving Crawl Performance for Large Data Sources
By default, ADOC crawls Kafka topics one at a time. For data sources with a large number of topics, this can make full crawls slow, and a failure on a single topic can stop the entire crawl.
You can configure ADOC to crawl multiple topics together in a single batch, improving crawl speed and preventing a failure on one topic from affecting the rest of the crawl. This is controlled by the CRAWLER_ASSET_BATCH_SIZE setting (default: 5), which you can configure at two levels:
- Per data source: Navigate to Control Center -> Integrations, find your Kafka data source, click the vertical ellipsis (⋮), and select Crawler Resources. In the Edit Resource Configuration dialog, update the value for CRAWLER_ASSET_BATCH_SIZE, then click Save. This applies to this data source only.
- For all data sources on a Data Plane: Navigate to Control Center -> Settings, select the Data Planes tab, find your Data Plane, click the vertical ellipsis (⋮), and select Application Config. Under Dataplane Config -> Crawler, update the value for CRAWLER_ASSET_BATCH_SIZE. This setting also applies to AWS S3 and Google Cloud Pub/Sub crawlers on the same Data Plane.
The updated value takes effect on the next crawl of the data source.
Note
Don't confuse this setting with CRAWLER_ASSET_PUBLISH_BATCH_SIZE, a separate setting that controls how many crawled assets are accumulated before ADOC publishes them to the catalog. CRAWLER_ASSET_BATCH_SIZE controls how many assets are crawled together; CRAWLER_ASSET_PUBLISH_BATCH_SIZE controls how many crawled results are published together. The two settings serve different purposes and can be configured independently.
What’s Next
- Crawl data from the configured Kafka topic.
- Profile the data source to fetch metrics like message count, size, and freshness.
- Apply ADOC’s data reliability rules and policies to validate the quality of your Kafka topic.

Have a suggestion?