Skip to main content
Beta: Deploying models into Environments is in beta for private cloud organizations. To join the waitlist, contact support.
A deployment is something you run inside one environment. Deploying a model from the console sets up two things in the environment:
  • The model, which serves an endpoint: a URL you send requests to. It works with the OpenAI API format.
  • A secret, which stores the access key that lets you call the endpoint.
The console lists them together as one deployment.

What it does

  • Dedicated hardware. The model runs on your environment cluster. No other customer shares it.
  • OpenAI-compatible endpoint. OpenAI SDKs and tools work with it.
  • Multiple deployments. You can deploy the same model to several environments, for example dev and prod.

Key words

Profile types:
  • Balanced: a good mix of speed and cost
  • Latency-optimized: faster replies for each request
  • Throughput-optimized: handles more requests at once
  • Custom: a special setup made for a specific need

Before you start

You need an environment with the status Ready.

Deploy a model

1

Start a deployment

Go to AI Services → Models. Click + Deploy Model.You can also start from:
  • A model’s page: click Deploy
  • An environment: click ⋯, then Deploy Model
The Models page with the Deploy Model button.
2

Name the deployment and pick a model

Under Deployment Details:
  • Deployment Name: lowercase letters, numbers, and dashes. Up to 62 characters. Must be unique for this model.
  • Model: search for and select the model.
The Deployment Details step with the Deployment Name and Model fields.
3

Select an environment

Under Select Environment, pick where the model runs. Only Ready environments show in the list.No environment yet? Click + Create Environment. It opens in a new tab.
4

Select a profile

Under Select Profile, pick a profile.Each profile card shows the hardware it needs. The console checks this against your environment’s cluster.
If a card says No matching hardware in this environment, the cluster doesn’t have the right GPUs. Pick another profile, or use an environment on a different cluster.
The Select Profile step with profile cards and hardware check.
5

Review and create

Check the Review your Deployment summary. Click Create Deployment.
6

Copy your access key

A Copy your Access Key window opens. It shows a new access key.
Copy the key now. You can’t see it again. Store it in a password manager or secrets store.
Click Copy Key & Create Deployment. This copies the key and starts the deployment. Nothing is created until you click it.
The Copy your Access Key window with the key and the Copy Key & Create Deployment button.
What happens next: The console opens the model’s Deployments tab. Your deployment shows Deploying. When it shows Running, you can call it.
Large models take longer to download and start. You can leave the page.

Deployment status

The environment’s Deployments tab uses different words for the same status: Pending, In progress, Succeeded, Failed, and Cancelled.

Call your model

  1. Open the model’s page and go to Overview.
  2. Find the Access endpoint card. If you have more than one deployment, pick one from the menu at the top of the card.
  3. Copy the example in cURL, Python, or TypeScript.
You can also copy the Endpoint from the Deployments tab.
The Access endpoint card with code examples.
Example request:
Replace:
  • <your-endpoint> with the Endpoint from the console
  • <your-access-key> with the key you copied
  • <model-name> with the model name shown on the model page

Lost your access key? Roll it

You can’t view an access key again. If you lose it, create a new one.
  1. Open the model’s Deployments tab.
  2. Click Roll Key on the deployment.
  3. Copy the new key. Click Copy Key & Redeploy.
The deployment restarts with the new key. The old key stops working. Update any app that uses it.
Roll Key is not available for profiles that support more than one hardware type. For these, delete the deployment and deploy again.

Delete a deployment

The endpoint stops working straight away. Any app that calls it will fail.
  1. Open the model’s Deployments tab.
  2. Click ⋯ on the deployment. Click Delete Deployment.
  3. Confirm.
You can’t delete a deployment while it is Deploying or Undeploying.