Data Sources and Assets Guide
This guide covers catalog operations with : discovering connections, running crawlers, resolving assets, working with metadata and sample data, and running profiling (full, incremental, and selective, with marker configuration).AdocClient
ADOC can register many data source s, including warehouses, lakes, and messaging systems. In code, use type to filter or AssetSourceTypeerpret catalog responses.int
Datasources
Load one datasource by name or numeric ID, optionally with extended properties on the returned object (or pass when loading).properties=True
from acceldata.client.adoc_client import AdocClientfrom acceldata.models.sdk.catalog import AssetSourceTypeclient = AdocClient( url="https://<your-adoc-url>", access_key="<your-access-key>", secret_key="<your-secret-key>",)# By name, optional properties payloadds = client.get_datasource("snowflake_prod")ds = client.get_datasource("snowflake_prod", properties=True)ds = client.get_datasource(5) # same as get_datasource_by_id(5)# All datasourcesall_sources = client.get_all_datasources() # alias: get_datasources() with no filter# Filter by source typesnowflake_only = client.get_datasources(AssetSourceType.SNOWFLAKE)
Crawler Operations
Crawler start and status are available from the datasource object (recommended) or by name on the client. Both support an optional for transient HTTP or transport issues.RetryConfig
# Using a DatasourceResourceds = client.get_datasource("my_datasource")start = ds.start_crawler()status = ds.get_crawler_status()# Or by name on the clientclient.start_crawler("my_datasource")client.get_crawler_status("my_datasource")
Assets and Asset Types
Resolve an Asset
Pass a string UID (for example, ) or a numeric ID (int, or a string of digits, which the SDK treats as an ID, not a UID).Feature_bag_datasource.feature_1
# By numeric idasset = client.get_asset(1)# By uidasset = client.get_asset("Feature_bag_datasource.feature_1")
The result is an .AssetResource
List Asset Type Definitions
types = client.get_all_asset_types() # alias of get_asset_types()
Each entry is a catalog (name, ID, and related fields as returned by the API).AssetType
Metadata, Custom Metadata, and Sample Data
from acceldata.models.sdk.catalog import CustomAssetMetadataasset = client.get_asset("my_datasource.my_table")# Fetched metadata (as stored for the asset)md = asset.get_metadata()# Or: client.get_asset_metadata(asset_id)# Merge custom key/value pairs (string values; merged then POSTed)asset.add_custom_metadata( [ CustomAssetMetadata("owner", "analytics"), CustomAssetMetadata("pii", "true"), ])# Or pass a flat mapping from the client:client.add_custom_metadata(asset_id, {"owner": "analytics"})# Sample data (default): trigger an async job and block by polling until not IN_PROGRESSsample = asset.sample_data(sync=True) # sync=True is the default# Or: client.sample_data(asset_id, sync=True)# Non-blocking mode: trigger an async job and return immediately with request_idinitial = asset.sample_data(sync=False)# Or: initial = client.sample_data(asset_id, sync=False)request_id = initial.request_id if initial is not None else None# With an async `request_id`, poll the resultif request_id: result = client.get_sample_data_result(request_id)
Profiling: Parameters and Modes
Catalog profiling is driven by , which sets the catalog profiling type and an optional AssetProfilingParams (other start-profile fields use catalog defaults inside markerConfig).to_start_profiling_request()
Field | Type | When Required |
| | Always ( |
| | Required for |
(from the same profiling module as ExecutionType) mirrors the same string values as ProfilingType, for execution-oriented typing in application code.ProfilingType
The preferred pattern for starting a run is , which returns a AdocClient.profile_asset with ProfileRequestResource and get_status(). Pass an optional cancel() to retry transient HTTP or transport failures; default retry status codes include HTTP 409 (for example, when a profile is already running).transient_retry=RetryConfig(...)
(default) is fire-and-forget. It triggers profiling and immediately returns the initial request (with id).profile_asset(..., sync=False)uses the returned request ID to keep checking profile status until it reaches a terminal state.profile_asset(..., sync=True)
from acceldata.client.transient_retry import RetryConfigfrom acceldata.models.sdk.catalog import AssetProfilingParams, ProfilingTypeparams = AssetProfilingParams(profiling_type=ProfilingType.FULL)# Async fire-and-forget (default)profile = client.profile_asset( 123, params, sync=False, # transient_retry=RetryConfig(max_attempts=6, initial_interval_seconds=30.0, max_interval_seconds=30.0),)# Sync mode: block until terminal status by polling with request idcompleted_profile = client.profile_asset( 123, params, sync=True,)
returns the same AssetResource.start_profile(profiling_params=..., sync=...) shape, with status and cancel helpers:ProfileRequestResource
from acceldata.models.sdk.catalog import AssetProfilingParams, ProfilingTyperes = asset.start_profile( AssetProfilingParams(profiling_type=ProfilingType.FULL), sync=False, # default: async fire-and-forget)print(res.to_dict())details = res.get_status()res.cancel()
# Latest status for the asset — response type is MiniProfileRequestlatest = client.get_latest_profile_status(asset_id=123)# Or: asset.get_latest_profile_status() on an AssetResource# Specific profile request by (asset_id, req_id) — also exposed as get_profile_statusby_req = client.get_profile_request_details(asset_id=123, req_id=456)# Or: client.get_profile_status(123, 456)
Latest Status and Lookup by Request ID
Cancel by request ID, when only the ID is available and not a resource instance:
client.cancel_profile(profile_req_id)
delegates cancel and status calls to these client and service methods. You can configure transient read retries for crawler and profile calls with ProfileRequestResource wherever the method allows it.RetryConfig
Trigger Full Profiling
This runs the configured full scan for the asset. The snippet below assumes a client from the earlier section.
from acceldata.models.sdk.catalog import AssetProfilingParams, ProfilingTypeasset = client.get_asset("my_datasource.my_table")params = AssetProfilingParams(profiling_type=ProfilingType.FULL)client.profile_asset(asset.id, params)
Trigger Incremental Profiling
Use incremental profiling when the asset's catalog configuration defines an incremental strategy; a marker is required on the request.
- For an ID-column incremental strategy, use
.id_marker_config - For other strategies, use the matching
helper in*_marker_config(for example,acceldata.services.marker_configfordate_time_marker_config, orDateTimeMarkerConfigforfile_marker_config— the same naming pattern asFileMarkerConfig→BoundsDateTimeMarkerConfig).bounds_date_time_marker_config - You can also set type yourself using
ormarker_type_for, or callAPI_TYPE_BY_MARKER_CLASS, which can fix a wrong type on a typed model instance.as_marker_config
also accepts a generated branch instance or a AssetProfilingParams.marker_config wrapper; the client normalizes them when it sends the request.MarkerConfig
The snippet assumes a configured client from the start of this guide. As with full profiling, AdocClient accepts profile_asset to poll until a terminal status, and sync=True for transient HTTP or transport retries (including while polling in transient_retry=RetryConfig(...) mode).sync
from acceldata.client.transient_retry import RetryConfigfrom acceldata.models.sdk.catalog import AssetProfilingParams, ProfilingTypefrom acceldata.services.marker_config import id_marker_configasset = client.get_asset("my_datasource.my_table")params = AssetProfilingParams( profiling_type=ProfilingType.INCREMENTAL, marker_config=id_marker_config( id_column_name="pk_col", initial_offset=0, ),)client.profile_asset( asset.id, params, sync=False, # default: async; use sync=True to block until terminal # transient_retry=RetryConfig(max_attempts=6, initial_interval_seconds=30.0, max_interval_seconds=30.0),)
Trigger Selective Profiling (Bounds-Style Markers)
Selective profiling must include a marker.
- For ID-bounded slices, use
.bounds_id_marker_config - For date- or file-event–bounded slices, use
orbounds_date_time_marker_config.bounds_file_event_marker_config
Each example below assumes a configured client from the start of this guide. The same sync and options described earlier apply to transient_retry for selective runs.profile_asset
ID-Bounded (Monotonic / Range Column)
from acceldata.client.transient_retry import RetryConfigfrom acceldata.models.sdk.catalog import AssetProfilingParams, ProfilingTypefrom acceldata.services.marker_config import bounds_id_marker_configasset = client.get_asset("my_datasource.my_table")params = AssetProfilingParams( profiling_type=ProfilingType.SELECTIVE, marker_config=bounds_id_marker_config( id_column_name="ID", from_id=0, to_id=1000, ),)client.profile_asset( asset.id, params, sync=False, # transient_retry=RetryConfig(max_attempts=6, initial_interval_seconds=30.0, max_interval_seconds=30.0),)
from acceldata.client.transient_retry import RetryConfigfrom acceldata.models.sdk.catalog import AssetProfilingParams, ProfilingTypefrom acceldata.services.marker_config import bounds_date_time_marker_configasset = client.get_asset("my_datasource.my_table")params = AssetProfilingParams( profiling_type=ProfilingType.SELECTIVE, marker_config=bounds_date_time_marker_config( date_column_name="TO_DATE", format="yyyy-MM-dd", time_zone_id="Asia/Calcutta", from_date="2023-07-01 00:00:00.000", to_date="2024-07-14 23:59:59.999", ),)client.profile_asset( asset.id, params, sync=False, # transient_retry=RetryConfig(max_attempts=6, initial_interval_seconds=30.0, max_interval_seconds=30.0),)
from acceldata.client.transient_retry import RetryConfigfrom acceldata.models.sdk.catalog import AssetProfilingParams, ProfilingTypefrom acceldata.services.marker_config import bounds_file_event_marker_configasset = client.get_asset("my_datasource.my_table")params = AssetProfilingParams( profiling_type=ProfilingType.SELECTIVE, marker_config=bounds_file_event_marker_config( date_column_name="ingest_date", time_zone_id="Asia/Calcutta", from_date="2019-04-01 00:00:00.000", to_date="2024-07-16 23:59:59.999", ),)client.profile_asset( asset.id, params, sync=False, # transient_retry=RetryConfig(max_attempts=6, initial_interval_seconds=30.0, max_interval_seconds=30.0),)
DateTime-Bounded (Column, Format, Timezone, Optional Range Strings)
File Event–Based Bounds
For partitioning, file-based markers, Pub/Sub, offsets, and similar cases, use the matching helper in *_marker_config (each pairs with a generated acceldata.services.marker_config class — see *MarkerConfig), or build the model with an explicit API_TYPE_BY_MARKER_CLASS, then pass the instance the same way as in the examples above.type=
What's Next
After you complete this section, explore:
- Policy Guide – Learn how to fetch, execute, and monitor data quality and reconciliation policies against these assets.
- Tags and Labels Guide – Learn how to attach tags and labels to catalog assets.
- Pipelines Guide – Learn how job inputs and outputs reference the assets described in this guide.

Have a suggestion?