source
Description
The source command is a foundational component of DataPrime. It informs the DataPrime engine which datasource you wish to read from.
While you can start your query with this, the source command is optional and will default to logs.
Basic usage
In DataPrime, you read data by specifying a dataset within an optional dataspace:
source <dataspace>/<dataset>
If no dataspace is provided, the query defaults to the default dataspace. This allows for concise syntax when working within the most common data sources.
Common datasets include:
logs– Application and infrastructure logs.
Default dataset. Equivalent to `source default/logs`.
spans– Distributed tracing data from systems like OpenTelemetry.
Equivalent to `source default/spans`.
enrichments/<name>– custom enrichment tables uploaded via the UI or API.
For example: `source default/enrichments/ip_lookup`
You can also query your High (Frequent Search) tier directly through the frequentsearch dataspace:
frequentsearch/logs– High-priority logs kept in the Frequent Search (hot) tier.frequentsearch/spans– High-priority spans kept in the Frequent Search (hot) tier.
Because frequentsearch datasets behave like any other source, you can reference them in `join` and `union` operations alongside default and system data. For more, see the data layer overview.
You can also query system-generated datasets such as:
system/engine.queries– Logs of all DataPrime query executions.system/alerts.history– Historical records of alert events.
> Dataset names may include dots (e.g., engine.queries) but are still treated as flat identifiers, not nested structures.
This structure supports querying across teams, environments, or pipelines, whether you’re debugging logs, analyzing performance, or auditing notifications.
Query several datasets at once
List more than one dataset in a single source command, separated by commas:
source logs, spans, system/engine.queries
The query reads every dataset you list and concatenates the results, exactly as a chain of `union` commands would. Reach for a comma-separated list when all you need is to widen the set of datasets a query reads, and keep union for the cases where each branch runs its own query first.
A comma-separated list works everywhere a source command appears, including inside `union`, `join`, `enrich`, and subqueries:
source logs
| limit 1
| create foo from (source logs, spans | countby dataset())
Three limits apply:
- Every dataset you list has to be in the same region. Querying across regions isn't supported.
- A single query can reference up to 128 sources.
- Listing the same dataset twice reads it once. To read a dataset twice in one query, use `union` explicitly.
Listing several datasets is DataPrime syntax. A Lucene query still runs against the single dataset you select.
Source parameters
Add parameters in parentheses after a dataset name to control which copy of it the query reads:
teamId='<id>'– read the dataset as it exists for another team you have access to.version=<n>– read a specific version of a dataset that has more than one, such as a dataset a background query overwrote.
Parameters combine with a comma-separated list, so you can pull the same dataset from several teams in one query:
source logs(teamId='571001'), logs(teamId='571002'), default/my_summary_dataset(version=3)
Use `dataspace()` and `dataset()` to attribute each record in the result back to the dataset it came from.
Syntax
(source|from) <dataset>[(<param>=<value>, ...)][, <dataset> ...]
Example 1
Query logs from the default dataspace:
Example query
source logs
// or
source default/logs
Example 2
Query spans from the default dataspace.
Example query
source spans
Example 3
Query system generated data such as engine.query logs/
Example query
source system/engine.queries
Example 4
Query the High (Frequent Search) tier directly through the frequentsearch dataspace:
Example query
source frequentsearch/logs
Example 5
Read several datasets in one query by listing them, separated by commas.
Example query
source logs, spans, system/engine.queries
Example 6
Read the same dataset for two teams, and a specific version of a summary dataset, then count the records that came from each source.
Example query
source logs(teamId='571001'), logs(teamId='571002'), default/my_summary_dataset(version=3)
| create origin.dataset from dataset()
| groupby origin.dataset agg count() as count