Train on your own hardware
Run training on your own GPU and stream its progress to the same monitor.
Train on your own hardware You can run a training loop on your own machine and still see its progress, logs and metrics in the same monitor a platform job uses. This page covers when to use this path, how to set it up, and how to report from your own code. When to use it Use your own hardware when you want to train on a GPU you already have, or when you do not have cloud GPU access yet. Generating code from Studio and running it locally works without any of this; attaching it to a job also gives you the live monitor, notifications and comparison tools the rest of these docs describe. Set up The CLI signs in with a personal API key. Personal API keys need a paid plan; the Free plan does not include them. Ask through Support /support if you need access during the beta. The metrics path is resolved in this order: a path you pass on the command line, then DAGNAM METRICS PATH , then the saved training metrics path . When attach runs a command and none of these is set, it uses ./dagnam metrics.jsonl in the current directory. Attach to a job Use attach to stream metrics to a job you already have: Without a command, attach watches the metrics file from where it currently is and uploads new lines as they are written; stop it with Ctrl-C. With --replay and no command, it uploads the file's existing contents once and exits. With a command after -- , it runs that command, uploads metrics while it runs, and exits with the command's own exit code: If the command you run also calls dagnam.training.init itself, for example generated training code, pass mode="off" to that call so the run is reported once, through attach , rather than twice. Report from your own code Import dagnam.training and call these functions from your training loop. The report functions never raise; they write to your metrics file, and attach uploads what they write. Function What it reports ---------------------------------------------------------------------------------------------------- ------------------------------------- init project id, , framework="pytorch", name=None, mode="auto" Starts a run. report metric epoch, step, metrics A dict of named metric values. report progress epoch, total epochs, step, total steps Progress through the run. report system gpu utilization=None, gpu memory used=None, gpu memory total=None, cpu percent=None System utilization. report log level, message A log line. report error category, technical summary, epoch=None, step=None, traceback=None An error, with an optional traceback. Generated training code already calls these for you. What the monitor shows Once metrics are attached, the job's monitor page behaves like a platform job: the connection status reflects whether your machine's metrics are currently arriving, and the System Metrics, Training Metrics and Live Logs panels fill in as your code reports. See Monitor a job /docs/training/monitor .
Open in Dagnam.AI docs