Deployments, endpoints,throughput.
Auxen deploys models and returns endpoints. How they're served is ours to run. In 2026 we renamed the API to say exactly that: you deploy a model, get a deployment with a stable endpoint, and choose its throughput — Standard, Double, Quad or Max — when you need more simultaneous requests at the same answer quality.
Every older route, field, status, error code, webhook event and MCP tool name keeps working until April 2, 2027 (2027-04-02). Responses return both the old and new field names until then. Your endpoint URLs and API keys don't change.
How you'll know you're on an old name
Calls to an older route return these headers. Log them in your client and you'll find every call site that needs updating.
Deprecation: true Sunset: Fri, 02 Apr 2027 00:00:00 GMT Link: <https://auxen.ai/docs/api/migration>; rel="deprecation"
Status values and error codes follow the route you call. On /v1/deploy and /v1/deployments/*, status and error.code use the new values, with the old ones in legacy_status and error.legacy_code. On the older routes they keep the old values, with the new ones in deployment_status and error.renamed_to. Routes that both generations share — inference, webhooks and billing — keep the old error code plus renamed_to until the sunset date.
Webhooks: POST /v1/webhooks accepts both the older and the new event names until the sunset date. Each subscription receives events under the names it chose, in the event field and the X-Auxen-Event header; subscribe to both and you get one delivery, under the new name. A webhook_url set on the deployment itself keeps the older name and adds the new one in deployment_event.
Old → new, side by side
Routes
| Before | Now | Notes |
|---|---|---|
| POST /v1/provision | POST /v1/deploy | |
| GET /v1/instances | GET /v1/deployments | |
| GET /v1/instances/:id | GET /v1/deployments/:id | |
| PATCH /v1/instances/:id | PATCH /v1/deployments/:id | |
| DELETE /v1/instances/:id | DELETE /v1/deployments/:id | |
| POST /v1/instances/:id/scale | POST /v1/deployments/:id/throughput | Body { "throughput": "double", "confirm": true } — or 2. confirm acknowledges the rate change |
| POST /v1/instances/:id/switch-model | POST /v1/deployments/:id/switch-model | |
| POST /v1/instances/:id/system-prompt | POST /v1/deployments/:id/system-prompt | |
| PATCH /v1/instances/:id/spending-cap | PATCH /v1/deployments/:id/spending-cap | |
| GET /v1/instances/:id/schedule | GET /v1/deployments/:id/schedule | |
| PUT /v1/instances/:id/schedule | PUT /v1/deployments/:id/schedule | |
| — | GET | POST /v1/deployments/:id/knowledge-base | New — list, or upload files (multipart field files) / attach document_ids |
| — | GET /v1/deployments/:id/diagnostics | New |
| POST /v1/instances/:id/pause | POST /v1/deployments/:id/pause | |
| POST /v1/instances/:id/wake | POST /v1/deployments/:id/resume | |
| POST /v1/{id}/chat | POST /v1/{deployment_id}/chat | Unchanged — same URL, same key |
| POST /v1/{id}/v1/chat/completions | POST /v1/{deployment_id}/v1/chat/completions | Unchanged — OpenAI-compatible |
Response fields
| Before | Now | Notes |
|---|---|---|
| instance_id | deployment_id | Same value — IDs keep their existing format |
| new_instance_id | new_deployment_id | switch-model |
| data.instance | data.deployment | Single-deployment responses |
| data.instances | data.deployments | List responses |
| capacity | throughput + throughput_multiplier | "standard" | "double" | "quad" | "max", and 1 | 2 | 4 | 8 |
| previous_capacity | previous_throughput + previous_throughput_multiplier | |
| estimated_ready_at | ready_estimate | |
| hourly_rate | rate_per_hour | String with currency, e.g. "$0.20/hour" |
| previous_rate | previous_rate_per_hour | |
| destroyed_at | deleted_at | |
| on_demand.active_instances | on_demand.active_deployments | GET /v1/billing/balance |
| capacity (request body) | throughput (request body) | Both accepted on requests; throughput wins if both are sent |
Status values
| Before | Now |
|---|---|
| provisioning | deploying |
| running | ready |
| paused_idle / paused_scheduled / stopped | paused |
| destroyed | deleted |
| warming, switching, suspended, error | unchanged |
Error codes
| Before | Now |
|---|---|
| INSTANCE_NOT_FOUND | DEPLOYMENT_NOT_FOUND |
| INSTANCE_NOT_RUNNING | DEPLOYMENT_NOT_READY |
| PROVISIONING_FAILED | DEPLOYMENT_FAILED |
| CONCURRENT_LIMIT_EXCEEDED | DEPLOYMENT_LIMIT_EXCEEDED |
| MAX_INSTANCES_REACHED | MAX_DEPLOYMENTS_REACHED |
Webhook events
| Before | Now | Notes |
|---|---|---|
| instance.ready | deployment.ready | |
| instance.error | deployment.error | |
| instance.destroyed | deployment.deleted | |
| instance.auto_scaled | deployment.throughput_changed | |
| instance.host_migrated | deployment.relocated | Payload carries only deployment_id and relocated_at |
| instance.suspended | deployment.suspended | |
| instance.resumed | deployment.resumed | |
| instance.idle | deployment.idle | |
| instance.paused | deployment.paused | |
| instance.warming | deployment.warming | |
| instance.schedule_changed | deployment.schedule_changed |
MCP tools
| Before | Now | Notes |
|---|---|---|
| auxen_list_models | auxen_list_models | Unchanged |
| auxen_provision_model | auxen_deploy_model | |
| auxen_get_instance_status | auxen_get_deployment | |
| auxen_list_instances | auxen_list_deployments | |
| auxen_destroy_instance | auxen_delete_deployment | |
| auxen_pause_instance | auxen_pause_deployment | |
| auxen_wake_instance | auxen_resume_deployment | |
| — | auxen_set_throughput | New |
| auxen_get_balance | auxen_get_balance | Unchanged |
| auxen_get_schedule / auxen_set_schedule | auxen_get_schedule / auxen_set_schedule | Unchanged names; take deployment_id (instance_id still accepted) |
Deploying with the new API
The examples use qwen3.5-9b, the Balanced default. Read current IDs from GET /v1/models rather than hard-coding them.
curl -X POST https://api.auxen.ai/v1/deploy \
-H "Authorization: Bearer $AUXEN_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.5-9b",
"throughput": "standard",
"label": "my-agent-prod"
}'{
"success": true,
"data": {
"deployment_id": "inst_3f9c2a71b0d4",
"status": "deploying",
"model": "qwen3.5-9b",
"throughput": "standard",
"throughput_multiplier": 1,
"rate_per_hour": "$0.20/hour",
"label": "my-agent-prod",
"ready_estimate": "2026-10-04T15:04:05.000Z",
"polling_url": "https://api.auxen.ai/v1/deployments/inst_3f9c2a71b0d4"
},
"meta": { "requestId": "req_…", "timestamp": "…" }
}Poll GET /v1/deployments/:id until status is ready, or subscribe to deployment.ready. Need to handle more simultaneous requests later? Change throughput — the endpoint and key stay the same. confirm: true acknowledges that the hourly rate changes:
curl -X POST https://api.auxen.ai/v1/deployments/$DEPLOYMENT_ID/throughput \
-H "Authorization: Bearer $AUXEN_KEY" \
-H "Content-Type: application/json" \
-d '{"throughput": "double", "confirm": true}'Throughput and model are separate decisions: more throughput serves more requests at the same quality; a more capable model gives better answers. Use /switch-model for the latter.
Legacy model map
Older model IDs keep working in /v1/deploy until 2027-04-02; after that they return MODEL_DEPRECATED with the replacement in the error body. Existing deployments on older models keep running untouched. Switching to the replacement takes about three minutes and issues a new endpoint URL and API key.
Older model → replacement
| Older model | Replacement | Notes |
|---|---|---|
| llama3.2-3b | qwen3.5-4b | Fast |
| mistral-7b | qwen3.5-4b | Fast |
| gemma2-2b | gemma4-e4b | Fast |
| phi3-mini | qwen3.5-4b | Fast |
| llama3.1-8b | qwen3.5-9b | Balanced |
| qwen2.5-14b | qwen3.5-9b | Balanced |
| mistral-nemo-12b | qwen3.5-9b | Balanced |
| gemma2-9b | gemma4-12b | Balanced |
| phi3-medium | qwen3.5-9b | Balanced |
| mistral-small-24b | qwen3.6-35b-a3b | Powerful |
| qwen2.5-32b | qwen3.6-35b-a3b | Powerful |
| gemma2-27b | gemma4-31b | Powerful |
| command-r | qwen3.6-35b-a3b | Powerful |
| llama3.1-70b | qwen3.5-122b-a10b | Maximum |
| qwen2.5-72b | qwen3.5-122b-a10b | Maximum |
| mixtral-8x22b | qwen3.5-122b-a10b | Maximum |
Questions about migrating?
Email [email protected] with your request ID and we'll help.