API · Migration guide

Deployments, endpoints,throughput.

Auxen deploys models and returns endpoints. How they're served is ours to run. In 2026 we renamed the API to say exactly that: you deploy a model, get a deployment with a stable endpoint, and choose its throughput — Standard, Double, Quad or Max — when you need more simultaneous requests at the same answer quality.

Nothing breaks today.

Every older route, field, status, error code, webhook event and MCP tool name keeps working until April 2, 2027 (2027-04-02). Responses return both the old and new field names until then. Your endpoint URLs and API keys don't change.

How you'll know you're on an old name

Calls to an older route return these headers. Log them in your client and you'll find every call site that needs updating.

Response headers on older routes
Deprecation: true
Sunset: Fri, 02 Apr 2027 00:00:00 GMT
Link: <https://auxen.ai/docs/api/migration>; rel="deprecation"

Status values and error codes follow the route you call. On /v1/deploy and /v1/deployments/*, status and error.code use the new values, with the old ones in legacy_status and error.legacy_code. On the older routes they keep the old values, with the new ones in deployment_status and error.renamed_to. Routes that both generations share — inference, webhooks and billing — keep the old error code plus renamed_to until the sunset date.

Webhooks: POST /v1/webhooks accepts both the older and the new event names until the sunset date. Each subscription receives events under the names it chose, in the event field and the X-Auxen-Event header; subscribe to both and you get one delivery, under the new name. A webhook_url set on the deployment itself keeps the older name and adds the new one in deployment_event.

Old → new, side by side

Routes

BeforeNowNotes
POST /v1/provisionPOST /v1/deploy
GET /v1/instancesGET /v1/deployments
GET /v1/instances/:idGET /v1/deployments/:id
PATCH /v1/instances/:idPATCH /v1/deployments/:id
DELETE /v1/instances/:idDELETE /v1/deployments/:id
POST /v1/instances/:id/scalePOST /v1/deployments/:id/throughputBody { "throughput": "double", "confirm": true } — or 2. confirm acknowledges the rate change
POST /v1/instances/:id/switch-modelPOST /v1/deployments/:id/switch-model
POST /v1/instances/:id/system-promptPOST /v1/deployments/:id/system-prompt
PATCH /v1/instances/:id/spending-capPATCH /v1/deployments/:id/spending-cap
GET /v1/instances/:id/scheduleGET /v1/deployments/:id/schedule
PUT /v1/instances/:id/schedulePUT /v1/deployments/:id/schedule
—GET | POST /v1/deployments/:id/knowledge-baseNew — list, or upload files (multipart field files) / attach document_ids
—GET /v1/deployments/:id/diagnosticsNew
POST /v1/instances/:id/pausePOST /v1/deployments/:id/pause
POST /v1/instances/:id/wakePOST /v1/deployments/:id/resume
POST /v1/{id}/chatPOST /v1/{deployment_id}/chatUnchanged — same URL, same key
POST /v1/{id}/v1/chat/completionsPOST /v1/{deployment_id}/v1/chat/completionsUnchanged — OpenAI-compatible

Response fields

BeforeNowNotes
instance_iddeployment_idSame value — IDs keep their existing format
new_instance_idnew_deployment_idswitch-model
data.instancedata.deploymentSingle-deployment responses
data.instancesdata.deploymentsList responses
capacitythroughput + throughput_multiplier"standard" | "double" | "quad" | "max", and 1 | 2 | 4 | 8
previous_capacityprevious_throughput + previous_throughput_multiplier
estimated_ready_atready_estimate
hourly_raterate_per_hourString with currency, e.g. "$0.20/hour"
previous_rateprevious_rate_per_hour
destroyed_atdeleted_at
on_demand.active_instanceson_demand.active_deploymentsGET /v1/billing/balance
capacity (request body)throughput (request body)Both accepted on requests; throughput wins if both are sent

Status values

BeforeNow
provisioningdeploying
runningready
paused_idle / paused_scheduled / stoppedpaused
destroyeddeleted
warming, switching, suspended, errorunchanged

Error codes

BeforeNow
INSTANCE_NOT_FOUNDDEPLOYMENT_NOT_FOUND
INSTANCE_NOT_RUNNINGDEPLOYMENT_NOT_READY
PROVISIONING_FAILEDDEPLOYMENT_FAILED
CONCURRENT_LIMIT_EXCEEDEDDEPLOYMENT_LIMIT_EXCEEDED
MAX_INSTANCES_REACHEDMAX_DEPLOYMENTS_REACHED

Webhook events

BeforeNowNotes
instance.readydeployment.ready
instance.errordeployment.error
instance.destroyeddeployment.deleted
instance.auto_scaleddeployment.throughput_changed
instance.host_migrateddeployment.relocatedPayload carries only deployment_id and relocated_at
instance.suspendeddeployment.suspended
instance.resumeddeployment.resumed
instance.idledeployment.idle
instance.pauseddeployment.paused
instance.warmingdeployment.warming
instance.schedule_changeddeployment.schedule_changed

MCP tools

BeforeNowNotes
auxen_list_modelsauxen_list_modelsUnchanged
auxen_provision_modelauxen_deploy_model
auxen_get_instance_statusauxen_get_deployment
auxen_list_instancesauxen_list_deployments
auxen_destroy_instanceauxen_delete_deployment
auxen_pause_instanceauxen_pause_deployment
auxen_wake_instanceauxen_resume_deployment
—auxen_set_throughputNew
auxen_get_balanceauxen_get_balanceUnchanged
auxen_get_schedule / auxen_set_scheduleauxen_get_schedule / auxen_set_scheduleUnchanged names; take deployment_id (instance_id still accepted)

Deploying with the new API

The examples use qwen3.5-9b, the Balanced default. Read current IDs from GET /v1/models rather than hard-coding them.

Deploy
curl -X POST https://api.auxen.ai/v1/deploy \
  -H "Authorization: Bearer $AUXEN_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5-9b",
    "throughput": "standard",
    "label": "my-agent-prod"
  }'
202 Accepted
{
  "success": true,
  "data": {
    "deployment_id": "inst_3f9c2a71b0d4",
    "status": "deploying",
    "model": "qwen3.5-9b",
    "throughput": "standard",
    "throughput_multiplier": 1,
    "rate_per_hour": "$0.20/hour",
    "label": "my-agent-prod",
    "ready_estimate": "2026-10-04T15:04:05.000Z",
    "polling_url": "https://api.auxen.ai/v1/deployments/inst_3f9c2a71b0d4"
  },
  "meta": { "requestId": "req_…", "timestamp": "…" }
}

Poll GET /v1/deployments/:id until status is ready, or subscribe to deployment.ready. Need to handle more simultaneous requests later? Change throughput — the endpoint and key stay the same. confirm: true acknowledges that the hourly rate changes:

Change throughput
curl -X POST https://api.auxen.ai/v1/deployments/$DEPLOYMENT_ID/throughput \
  -H "Authorization: Bearer $AUXEN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"throughput": "double", "confirm": true}'

Throughput and model are separate decisions: more throughput serves more requests at the same quality; a more capable model gives better answers. Use /switch-model for the latter.

Legacy model map

Older model IDs keep working in /v1/deploy until 2027-04-02; after that they return MODEL_DEPRECATED with the replacement in the error body. Existing deployments on older models keep running untouched. Switching to the replacement takes about three minutes and issues a new endpoint URL and API key.

Older model → replacement

Older modelReplacementNotes
llama3.2-3bqwen3.5-4bFast
mistral-7bqwen3.5-4bFast
gemma2-2bgemma4-e4bFast
phi3-miniqwen3.5-4bFast
llama3.1-8bqwen3.5-9bBalanced
qwen2.5-14bqwen3.5-9bBalanced
mistral-nemo-12bqwen3.5-9bBalanced
gemma2-9bgemma4-12bBalanced
phi3-mediumqwen3.5-9bBalanced
mistral-small-24bqwen3.6-35b-a3bPowerful
qwen2.5-32bqwen3.6-35b-a3bPowerful
gemma2-27bgemma4-31bPowerful
command-rqwen3.6-35b-a3bPowerful
llama3.1-70bqwen3.5-122b-a10bMaximum
qwen2.5-72bqwen3.5-122b-a10bMaximum
mixtral-8x22bqwen3.5-122b-a10bMaximum

Questions about migrating?

Email [email protected] with your request ID and we'll help.