Create Deployment

Add MCP server to your AI tool

Allow AI tools and LLMs to interact with the API documentation portal through MCP.

MCP server URL

https://openapi-v2.exoscale.com/mcp

Standard setup for AI tools providing an mcp.json file

mcp.json
{
  "Exoscale APIv2 MCP server": {
    "url": "https://openapi-v2.exoscale.com/mcp"
  }
}

Close
POST /ai/deployment

Deploy a model on an inference server

application/json

Body Required

  • gpu-count integer(int64) Required

    Number of GPUs (1-8)

    Minimum value is 1.

  • inference-engine-version string

    Inference engine version

    Values are 0.12.0, 0.15.1, 0.16.0, 0.17.0, 0.18.0, 0.18.1, 0.19.0, 0.19.1, 0.20.0, 0.20.1, 0.20.2, 0.21.0, 0.22.0, 0.22.1, 0.23.0, 0.24.0, 0.25.0, 0.25.1, or 0.26.0. Default value is 0.26.0.

  • name string Required

    Deployment name

    Minimum length is 1.

  • gpu-type string Required

    GPU type family (e.g., gpua5000, gpu3080ti)

  • product-name string

    Billing identifier for this deployment. Used by the Router for usage counters and Kafka events.

    Minimum length is 1.

  • replicas integer(int64) Required

    Number of replicas (>=1)

    Minimum value is 1.

  • inference-engine-parameters array[string]

    Optional extra inference engine server CLI args

  • model object Required

    Model reference. Provide either id or name.

    Hide model attributes Show model attributes object
    • name string

      Associated model name

      Minimum length is 1.

    • id string(uuid)

      Associated model ID

Responses

  • 412 application/json

    412 (probably insufficient GPUs)

    Hide response attributes Show response attributes object
    • type string(uri-reference) Required

      An absolute or relative URI reference pointing to human-readable documentation concerning the specific problem type encountered.

    • title string Required

      A brief summary defining the class of failure, optimal for quick user interface groupings.

    • status integer Required

      Minimum value is 100, maximum value is 599.

    • detail string Required

      A highly contextual, readable explanation breaking down explicitly what triggered this error scenario.

  • 403 application/json

    Forbidden

    Hide response attributes Show response attributes object
    • type string(uri-reference) Required

      An absolute or relative URI reference pointing to human-readable documentation concerning the specific problem type encountered.

    • title string Required

      A brief summary defining the class of failure, optimal for quick user interface groupings.

    • status integer Required

      Minimum value is 100, maximum value is 599.

    • detail string Required

      A highly contextual, readable explanation breaking down explicitly what triggered this error scenario.

  • 200 application/json

    OK

    Hide response attributes Show response attributes object
    • id string(uuid)

      Operation ID

    • reason string

      Operation failure reason

      Values are incorrect, unknown, unavailable, forbidden, busy, fault, partial, not-found, interrupted, unsupported, or conflict.

    • reference object

      Related resource reference

      Hide reference attributes Show reference attributes object
      • id string(uuid)

        Reference ID

      • command string

        Command name

    • message string

      Operation message

    • state string

      Operation status

      Values are failure, pending, success, or timeout.

  • 400 application/json

    Bad Request

    Hide response attributes Show response attributes object
    • type string(uri-reference) Required

      An absolute or relative URI reference pointing to human-readable documentation concerning the specific problem type encountered.

    • title string Required

      A brief summary defining the class of failure, optimal for quick user interface groupings.

    • status integer Required

      Minimum value is 100, maximum value is 599.

    • detail string Required

      A highly contextual, readable explanation breaking down explicitly what triggered this error scenario.

POST /ai/deployment
curl \
 --request POST 'https://api-ch-gva-2.exoscale.com/v2/ai/deployment' \
 --header "Content-Type: application/json" \
 --data '{"gpu-count":42,"inference-engine-version":"0.26.0","name":"string","gpu-type":"string","product-name":"string","replicas":42,"inference-engine-parameters":["string"],"model":{"name":"string","id":"string"}}'
Request examples
{
  "gpu-count": 42,
  "inference-engine-version": "0.26.0",
  "name": "string",
  "gpu-type": "string",
  "product-name": "string",
  "replicas": 42,
  "inference-engine-parameters": [
    "string"
  ],
  "model": {
    "name": "string",
    "id": "string"
  }
}
Response examples (412)
{
  "type": "string",
  "title": "string",
  "status": 42,
  "detail": "string"
}
Response examples (403)
{
  "type": "string",
  "title": "string",
  "status": 42,
  "detail": "string"
}
Response examples (200)
{
  "id": "string",
  "reason": "incorrect",
  "reference": {
    "id": "string",
    "link": "string",
    "command": "string"
  },
  "message": "string",
  "state": "failure"
}
Response examples (400)
{
  "type": "string",
  "title": "string",
  "status": 42,
  "detail": "string"
}