Monitor a job
The training monitor page, its panels, live logs, notifications, and comparing runs.
Monitor a job Every training job has a monitor page at its job ID. This page covers what each panel shows, how the live connection behaves, and how to cancel, restart, or compare jobs. Open the monitor Open a job from the Training jobs list, or straight from its creation dialog. The header shows the current epoch and percent complete, and four actions: - Publish to Hub , once the job has completed. - Download Code , while the job is running or once it has completed. - Back to Studio , when the job belongs to a project. - Auto-refresh , which polls for new checkpoints while the job is active. If the job does not exist or is not yours, the page says so and offers a link back to the Training jobs list. Queue and connection states While a pending job's place in the queue is known, it shows a Queued, position N badge and "Waiting for platform capacity"; the position moves up and down as the queue changes and is not a time estimate. The live connection carries its own status: Connected , Connecting... , Reconnecting... , Disconnected , or Connection Error , with a Reconnect button. A job attached from your own hardware see Train on your own hardware /docs/training/own-hardware shows its own states instead, for whether your machine's metrics are currently arriving. Metrics and charts - Job Information : job ID, project, framework, device, created time, duration, the run's configuration epochs, batch size, learning rate, optimizer , an Export dag.json button, and, if the job failed, the error category, message and remediation hints. - System Metrics : GPU utilization and memory for a GPU device , CPU utilization, RAM usage, the configured learning rate, and throughput in samples per second. - Training Metrics : a chart of the metrics your run reports loss, accuracy, learning rate and any custom metric name , with smoothing options and a CSV export. View it combined or as separate charts per metric. - Progress : overall progress, epoch and step, current epoch progress, time elapsed, and, while the job runs, time remaining and throughput. - Metric Values : the current, best and final value for each reported metric, with a trend indicator. - Checkpoints : covered on its own page, see Checkpoints /docs/training/checkpoints . Live logs Search the log, filter by level All , Debug , Info , Warning , Error , wrap lines, open a fullscreen view, pause or resume auto-scroll, copy the visible lines, or download the full log history as a text file. Notifications While the monitor page is open, you get a toast for Training Completed , Training Failed , Training Cancelled , and Training Started when a pending job starts running . If you allow browser notifications, Completed and Failed also raise one. Cancel and restart - Cancel Training appears for a running, pending or initializing job. Confirming stops it; checkpoints already saved are kept. - Restart Training appears for a failed or cancelled job. It creates a new job with the same configuration and opens it. Compare runs Select two or more jobs on the Training jobs list and click Compare to open Compare Training Runs . It shows a chart of the selected jobs' metrics over time, plus Metrics Comparison , Configuration Diff and Performance Analysis tabs, and exports the comparison as CSV or JSON.
Open in Dagnam.AI docs