> Documentation index: https://docs.beetl.io/llms.txt

# Registered datasets

> Files already in your object store, read in place.

Source: https://docs.beetl.io/sources/registered-datasets/
Last updated: 2026-10-05

A registered dataset points at files that already sit in one of your tenant's object stores, such
as a Delta table or a folder of Parquet exports that another tool writes. Beetl reads the files
where they are and copies nothing. It never writes to, moves or deletes them.

Any user can register a dataset. The files must be in an object store the tenant already has,
usually a bucket a TenantAdmin added (see
[Use your own bucket](https://docs.beetl.io/concepts/tenants-and-roles#use-your-own-bucket)).

## Register a dataset

    On Home, right-click the canvas and choose **Register data set**, then pick the object store under
    "Where is the data?". Clicking an object store's node on Home opens the same dialog on that store.

    Browse to the data. Click a file to register that one file, or open a folder and choose **Use this
    directory** to register everything in it. For a Delta table, use the folder that contains
    `_delta_log`. The browser shows "Delta table detected" when you are in it.

    Check the name. Beetl suggests the file name without its extension, or the folder name. You write
    this name in SQL, and it follows the usual
    [naming rules](https://docs.beetl.io/concepts/datasets-and-tiers#names-and-status).

    Choose **Register data set**. Beetl reads the location, detects the format and schema, and adds an
    Initialized dataset to Home, ready to query. If Beetl cannot read the location, no dataset is
    created and the dialog shows the reason.

Object keys that contain `%`, `?`, `#`, spaces or control characters are greyed out in the browser
and cannot be registered. Rename them in your bucket first.

## Supported formats

| Format     | Point at                                 | Beetl reads                                                 |
| ---------- | ---------------------------------------- | ----------------------------------------------------------- |
| Delta Lake | The table folder                         | The latest version of the table, with its partition columns |
| Parquet    | One `.parquet` file, or a folder of them | Every Parquet file at that location                         |
| CSV        | One `.csv` file, or a folder of them     | Comma-separated files with a header row                     |
| JSON       | One `.json` file, or a folder of them    | Newline-delimited JSON, one object per line                 |

Beetl detects the format for you. A folder that contains `_delta_log` is read as Delta. Any other
location is tried as Parquet, then CSV, then JSON, and the first format that opens is used, so keep
one format per folder. A folder includes the files in its subfolders. Extensions must be lowercase,
and a JSON file whose content is one top-level array is refused.

Partition columns are recorded for Delta tables and marked in the schema on Home. CSV, Parquet and
JSON datasets have none.

## Where you can register from

A location cannot overlap anywhere Beetl writes its own data in that store: the folders of datasets
Beetl writes, File Set uploads and saved query results. A path equal to, inside or above one of
them is refused. Each location can be registered once per store, so registering the same path again
fails until the first dataset is deleted.

## Using a registered dataset

A registered dataset is a Bronze dataset you can query, preview, chart, mention in chat and use as
a pipeline source. Beetl never writes to it:

* It cannot be a pipeline destination, and no sync, webhook or File Set can write to it.
* It never triggers a pipeline run. A pipeline that reads it runs when you press Run, when you
  deploy it, or when one of its other sources triggers it.
* In a query result built from it, drill-down to the records behind a cell is not available.

## When the files change

Beetl reads the location every time a query, preview or pipeline run uses the dataset. New files in
a registered folder, and new versions of a Delta table, appear in the next read without any action
from you.

Beetl does not watch the location. A change starts no pipeline run, and the schema on Home stays
the one Beetl read at registration. If the files are moved or deleted, reads of the dataset fail
until they are back.

## Renaming and deleting

Any user can delete a registered dataset, which leaves your files untouched. To change its name,
contact Beetl; see
[Finding, renaming and deleting](https://docs.beetl.io/concepts/datasets-and-tiers#finding-renaming-and-deleting).

## Related

- [Datasets and tiers](https://docs.beetl.io/concepts/datasets-and-tiers)

- [Tenants and roles](https://docs.beetl.io/concepts/tenants-and-roles)

- [External datasets](https://docs.beetl.io/sources/external-datasets)
