Training overview

Where training runs, the ways to start a job, and what the Training jobs list shows.

Training overview This page covers where a training job runs, the three ways to start one, and the Training jobs list where every job lands. Where training runs A job runs in one of two places: - The Dagnam platform. The platform provisions the compute and streams progress back to the monitor. This needs cloud GPU access. - Your own hardware. You run the training loop yourself and stream its progress to the same monitor with the CLI. See Train on your own hardware /docs/training/own-hardware . Training on Dagnam's cloud GPUs needs cloud GPU access, which is available on request during the beta. Ask through Support /support . Without it, you can still generate code and train on your own hardware. Start from Studio Open a saved, valid project and click Train . The Train Model dialog shows the configuration read from your architecture Epochs, Batch Size, Learning Rate, Optimizer, Loss, Dataset, Device, GPUs, Precision, LR Scheduler , plus Training Platform , which is set to Dagnam platform other platforms show Coming Soon . Start Training checks the design on the server, creates the job and opens its monitor page. If the server refuses the job no cloud GPU access, not enough credits, or a plan limit , the dialog shows the reason with links to Train on your own hardware and Support. Start from the dashboard A project's row menu on the dashboard offers Start training , which opens the same dialog. While the project is a draft, the item is disabled with "Fix architecture errors before training". A project is a draft when its current architecture version is not valid, and also after its latest job was cancelled. Start from the CLI or SDK \\\n --epochs 10 --batch-size 32 --learning-rate 0.001 \\\n --optimizer adam --loss-function cross entropy \\\n --dataset-id --framework pytorch", , label: "Python", language: "python", code: 'job = dagnam.create training job \n project id,\n epochs=10,\n batch size=32,\n learning rate=0.001,\n optimizer="adam",\n loss function="cross entropy",\n training dataset id=dataset id,\n framework="pytorch",\n ', , / training create also accepts --val-dataset-id , --test-dataset-id , --train-split , --val-split , --test-split , --max-duration-seconds and a raw --config JSON object or file; in the SDK, pass advanced fields as config overrides . dagnam training estimate reports the resources a configuration would need before you create the job. The Training jobs list Training Jobs : "Manage and monitor training runs." A search box matches a job ID, project name, framework or status, alongside New Training Job and Train on your own hardware . - Filters : Status Pending, Initializing, Running, Completed, Failed, Cancelled , Framework All Frameworks, PyTorch, TensorFlow, FLAX , Device All Devices, CPU, CUDA GPU , TPU, MPS Apple Silicon , Date Range All Time, Today, Last 7 Days, Last 30 Days, Last 90 Days . - Sort : Created Date, Updated Date, Status. - Columns : Job ID, Project, Actions, Status, Progress, Created, Duration, Compute. - Row menu : View Details , Cancel Job while running or pending , Download Logs the full log as a text file , Publish to Hub once completed , Copy Job ID . - Bulk actions : Cancel Selected any selected job that is running, pending or initializing , Compare 2 or more selected, opens the comparison page , Delete Selected only when every selected job is completed, failed or cancelled; this also removes their checkpoints and cannot be undone . Job statuses Status Meaning ------------ ------------------------------------------------------------------------------------------------------------- Pending Queued. A caption shows a queue position while one is known; it moves up and down and is not a time estimate. Initializing The job is being prepared to run. Running Training is in progress. Completed Training finished successfully. Failed Training stopped with an error. The job carries an error category and remediation hints. Cancelled You or the platform stopped the job before it finished. Limits Training on Dagnam's cloud GPUs and fine-tuning a base model need cloud GPU access, on request during the beta. Training time counts against your plan's usage period, and how many jobs can run at once depends on your plan. A job on the Dagnam platform is also capped, by plan, on model size and on how long one run may last. TensorFlow and FLAX code generation, and multi-GPU training strategies, are not on the Free plan; PyTorch and single-GPU or CPU training are always available. See Billing & usage /pay for your account's live limits, and Support /support to request access.
Open in Dagnam.AI docs