Load datasets in Python

load_dataset, DagnamDataset, the framework converters, splits, and the local cache.

Load datasets in Python dagnam.load dataset downloads and caches a dataset, then hands you a DagnamDataset you can read directly or convert to your framework of choice. Load a dataset load dataset dataset id, api url=None, api key=None, cache dir=None, version=None, split=None, verify=False accepts a dataset UUID one of your own datasets or the name of a built-in dataset. dataset.info returns a dict with id , name , format , type , samples , classes , class names and schema . Built-in datasets Pass a built-in dataset's name instead of a UUID to load it without uploading anything, for example MNIST Handwritten Digits or CIFAR-10 . The name is the one the catalog shows, matched without regard to case, so cifar-10 works too; any other name raises DatasetNotFoundError . dagnam.datasets.list system lists the built-in datasets and their names, and dagnam dataset list browses yours and the built-in ones together. Convert to PyTorch, TensorFlow, Flax or polars CSV, TSV, JSON and JSONL datasets work with every converter below; image and audio folder datasets work with PyTorch, TensorFlow and Flax, but not to polars . to pytorch loader , to tensorflow dataset and to flax dataset share the same shape: split="train" , batch size=32 , val ratio=.1 , test ratio=.1 , seed=42 , image size= 224, 224 , and a column roles mapping for tabular data ignored for image and audio formats . Two lower-level helpers work against built-in datasets and CSV, TSV, JSON and JSON Lines files: iter samples split="train", ... yields raw samples one at a time, and to arrays split="train", ... collects them into NumPy arrays for a generated training script. For any other format they raise ValueError . Splits Pass split="train" , "val" or "test" to a converter. When the dataset declares its own split membership on the platform, the converter reads it; otherwise it partitions rows deterministically with val ratio , test ratio and seed , so the same call always returns the same split. The cache Datasets cache under ~/.dagnam/datasets/ / , checkpoints under ~/.dagnam/checkpoints/ , and models under ~/.dagnam/models/ . The cache evicts the least recently used entries once it passes max cache size configured in Install /docs/dag-lib/install configuration , default 10 GiB . Pass cache dir= to load dataset to use a different directory for one call, run dagnam cache clear to empty it from the CLI, or call dagnam.evict lru ... to reclaim space in code.
Open in Dagnam.AI docs