Skip to content

CLI

Crawler that crawls metadata of datasets in a Dataverse collection and exports the metadata to JSON and spreadsheet.

Usage

$ [OPTIONS] COMMAND [ARGS]...

Global Options

Option Value / Type Description
-a, --auth Authentication token to access the dataverse repository. The environment variable API_TOKEN will always override this option. [env var: API_TOKEN]
-r, --report / --no-report Output summary report file of the crawl. [default: report]
-c, --collection_alias Name of the collection to crawl [required]
-v, --version The dataset version to crawl. Options are: "draft" - the draft version, if any "latest" - either a draft (if exists) or the latest published version "latest-published" - the latest published version "x.y" - a specific version, where x is the major version number and y is the minor version number "x" - same as "x.0" Note that the version apply only to individual datasets. The publication-status option can be used to filter datasets by their publication status. For example, if you want the get all published datasets but with the draft version, you can use --publication-status Published --version draft. But this option will also return datasets that is published, but does not have a draft version. [default: latest]
-debug, --debug-log Enable debug logging to a file. This will create a log file in the logs directory
--log-level The logging level for console and file output. Options are: TRACE, DEBUG, INFO, SUCCESS, WARNING, ERROR, CRITICAL. Defaults to LOG_LEVEL in .env, else INFO.
-m, --metadata-source The source of the metadata to crawl. This option can be used to exclude harvested datasets (that does not host directly on the installation).
-ps, --publication-status The publication status of the datasets to look for. Common values are "Published", "Draft", "Unpublished", "Deaccessioned". Depends on the installation.
-sl, --semaphore-limit The maximum number of concurrent tasks when crawling datasets. Please adjust this number based on the expected load on the dataverse repository. Might need some trial and error to find the optimal number. [default: 5]
-ts, --timestamp / --no-timestamp Whether to include timestamp in the exported JSON filenames. [default: timestamp]
-p, --permission / --no-permission Whether to include permission metadata in the exported JSON files. This will make additional API calls to fetch the permission metadata for each dataset. [default: no-permission]
--install-completion Install completion for the current shell.
--show-completion Show completion for the current shell, to copy it or customize the installation.
--help Show this message and exit.

Commands

Command Description
search Search for datasets in the collection.
crawl-metadata Crawl dataset metadata.
crawl-permission Crawl dataset permissions.
export-spreadsheet Export the dataset metadata (and...
run-all Run the full crawl process: search, crawl...

Search for datasets in the collection.

$ search [OPTIONS]
Option Value / Type Description
--help Show this message and exit.

crawl-metadata

Crawl dataset metadata.

$ crawl-metadata [OPTIONS]
Option Value / Type Description
--help Show this message and exit.

crawl-permission

Crawl dataset permissions.

$ crawl-permission [OPTIONS]
Option Value / Type Description
--help Show this message and exit.

export-spreadsheet

Export the dataset metadata (and permissions if available) to spreadsheet.

$ export-spreadsheet [OPTIONS]
Option Value / Type Description
--help Show this message and exit.

run-all

Run the full crawl process: search, crawl metadata, crawl permissions, export spreadsheet.

Export the metadata (with permissions if available) to JSON and spreadsheet.

$ run-all [OPTIONS]
Option Value / Type Description
--help Show this message and exit.