Call your model

Predict, batch, schema, health, and streaming against a deployment, from the web, CLI, or SDK.

Call your model Once a deployment is running, you call it with its deployment key. This page covers the request and response shapes, batching, schema and health checks, streaming, and how errors and credits work. Deployment keys Send the deployment key as Authorization: Bearer on predict, batch, schema and health requests. A missing or invalid key answers 401. Streaming is the exception; see Streaming streaming . To rotate the key, see Regenerate the key /docs/deployments/manage regenerate-the-key . The CLI and SDK send the key they are configured with, which after dagnam login is your personal API key. For these calls, give them the deployment key instead: set DAGNAM API KEY for a CLI command, or pass api key= in the SDK. Predict POST /api/v1/inference/ /predict with "input": . The shape of input depends on what the model does: The model does input The prediction -------------------- ---------------------------------------------------- ----------------------------------- Chat "messages": "role": "user", "content": "..." An OpenAI-style choices list Text classification "text": "..." "label": "...", "scores": ... Tabular prediction "rows": ... , ... A list of predictions, one per row Image classification "image": " " "label": "...", "scores": ... Text embedding a string, or a list of strings A list of embedding vectors A successful call answers "prediction": , "latency ms": . A request with the wrong shape answers 400 and names the missing or mismatched field, for example input.rows is required . A bare string is not accepted for chat or text classification. /predict \\\n -H "Authorization: Bearer " \\\n -H "Content-Type: application/json" \\\n -d \' "input": "messages": "role": "user", "content": "Hello" \'', , label: "Shell", language: "bash", code: 'DAGNAM API KEY= dagnam inference run \\\n --input \' "messages": "role": "user", "content": "Hello" \'', , label: "Python", language: "python", code: 'result = dagnam.inference \n deployment id,\n "messages": "role": "user", "content": "Hello" ,\n api key=" ",\n ', , / The CLI's --input and the SDK's second argument are the value of input : both wrap it in "input": ... for you, so pass the value alone, as above. Both take a JSON object, so send a text embedding model's string input with batch or the HTTP API. Batch POST /predict/batch takes "inputs": ... and answers "predictions": ... , "latency ms": , with the predictions in the same order as the inputs. CLI: dagnam inference batch --inputs ' ... ' or --inputs-file PATH . SDK: dagnam.inference batch deployment id, inputs, api key=" " . Both take the list of input values and wrap it for you. Schema and health GET /schema returns the model's input and output schema, plus example requests. GET /health reports whether the deployment is currently able to serve. Both take the deployment key. CLI: dagnam inference schema . SDK: dagnam.inference schema deployment id, api key=" " and dagnam.deployment health deployment id, api key=" " . dagnam inference health is different: it uses your personal API key and shows the platform's own health record for the deployment. Streaming Streaming is available where the model supports it. A chat deployment streams; a text classification deployment refuses a stream request with 400. A stream is opened with your own sign in or personal API key, not the deployment key, and only for a deployment you own. to open a session; it returns a session ID and a short-lived token.', "GET /predict/stream/ ?token= to open the stream.", "Read server sent events: token for each piece of output, complete when the response is done, error on failure.", / CLI: dagnam inference stream --input ' ... ' . SDK: dagnam.inference stream deployment id, inputs , which yields each event. Both use your signed in key, and both wrap the value in "input": ... for you. Errors and credits A paused or not yet running deployment answers 503 and names its current status. Each prediction your model produces consumes credits from your account, even when the response is then refused for example, because it is too large or is not valid JSON , and so does each item of a batch served before the batch failed. A request your model never answered is not charged.
Open in Dagnam.AI docs