For builders and agents

A private model endpoint you pay for by the minute.

Deploy your own private open-weight model. Get a stable, OpenAI-compatible endpoint. Pay per minute while it's active. We run everything underneath — single-tenant, never routed to a shared third-party API, with full lifecycle control over an API and a remote MCP server.

Isolation and control, exposed as an API.

Everything you need to put a model in production — for your agents, your app, or a compliance-constrained team — without sending data through a shared service.

Single-tenant by construction

One private deployment runs one model, for you only. No shared inference service, and no other customer's requests answered beside yours.

No third-party routing

Requests are answered by your deployment and go nowhere else — not to OpenAI, Anthropic, or Google. Its data is erased when you delete the deployment.

OpenAI-compatible endpoint

A stable per-deployment URL that drops into any framework speaking OpenAI — openai-python, Vercel AI SDK, LangChain, LlamaIndex. No rewrites.

OAuth 2.1 + PKCE

The control plane is authenticated with OAuth 2.1 + PKCE; each deployment carries its own scoped API key. Least-privilege access your security team can reason about.

Programmatic lifecycle control

Deploy, pause, schedule, and delete over a REST API or a remote MCP server. Your agents can manage their own endpoints end to end.

Open-weight models, run for you

Pick an open-weight model from the catalog by tier — Fast, Balanced, Powerful, or Maximum. We run everything underneath; you get the endpoint.

A URL and a key. That's the integration.

Point your existing OpenAI client at your deployment's endpoint and go.

curl https://auxen.ai/v1/$USER/$DEPLOYMENT/api/chat \
  -H "Authorization: Bearer auxk_..." \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role":"user","content":"Hello"}]}'

Full lifecycle control,
in three steps

Connect, deploy, pause, and delete — from your agent, your CI, or your own tooling.

Connect

Add Auxen's MCP server to your agent, or use an API key directly. OAuth 2.1 + PKCE works with any MCP-capable client (Claude, Cursor, your own framework).

Standard MCP

Deploy

Pick a model and get a private, single-tenant deployment with a stable OpenAI-compatible endpoint and its own API key. No shared service, no third-party routing.

Ready in minutes

Use, pause, delete

Call /v1/chat/completions on your endpoint, pause the deployment when idle, or delete it when you're done. Its data is erased with it.

Full lifecycle control

Pay per minute while it's active.

A private deployment billed by the minute it's active — from $0.15/hour. Pause it and the meter stops. Priced for privacy and control; at low, sporadic volume a per-token API can be cheaper, and our calculator will tell you so.

Private deployment · per hour, billed per minute
Model tierStandardDoubleQuadMax
Fast
Quickest and cheapest — sorting, pulling out details, short answers
$0.15/hr$0.29/hr$0.56/hr$1.13/hr
Balanced
The everyday workhorse — chat, summaries, agents' day-to-day work
$0.20/hr$0.38/hr$0.75/hr$1.50/hr
Powerful
Deeper reasoning, long documents, coding
$0.65/hr$1.24/hr$2.44/hr$4.88/hr
Maximum
The most capable open models — can lead a cost-sensitive stack
$2.25/hr$4.28/hr$8.44/hr$16.88/hr

Rates are per hour, billed per minute of active time. Throughput levels increase simultaneous request handling at the same model quality.

Warm-up is free. Once your model is ready, each session is billed for at least 30 minutes (35 on Maximum).

Shared option: the Fast tier is also available at $0.05/hr on a shared, multi-tenant deployment — the one exception to single-tenant, clearly labeled when you choose it.

  • A private deployment — your model, your endpoint, your data
  • Full API access + OpenAI-compatible endpoint
  • Knowledge Base, Persona Studio, website widget
  • Pause anytime — the per-minute meter stops after the session minimum
Start with $10$10 minimum deposit · Credits never expire · No contract
Frontier models — contact us
Frontier models — the largest open releases — are available on request.
Contact us→
Volume or bespoke
Running many deployments, need a volume rate, or want something set up specially — your own models, on-premise, or extra compliance? Talk to us.
Contact us→
Cost calculator

Would Auxen cost you less — or more?

Model tier
Throughput
Active per day
Tokens per month
Auxen, private deployment
$48.64 /month
A frontier API at list price
$300 /month

At this volume, a private Auxen deployment costs about $251 less per month — and your data stays private. Make sure your chosen tier can handle the volume; if not, add throughput.

Estimate only. Auxen bills per minute of active time regardless of tokens. Warm-up is free. Once your model is ready, each session is billed for at least 30 minutes (35 on Maximum). The per-token figure assumes 75% input / 25% output at $3/$15 per million tokens. Open-weight models aren't identical to frontier models — for high-volume, well-specified work the quality gap is usually smaller than the cost gap.

More requests, or smarter answers?

Two different ways to grow. Choose the one that matches what you need.

Increase throughput

More requests, same quality

When you're getting good answers but need more of them at once.

  • · Stay on the same model
  • · Move from Standard to Double, Quad, or Max
  • · More throughput handles more simultaneous requests at the same quality.
  • · Double, Quad and Max cost less per request than Standard
Best for: high traffic, growing user base, busy periods
Move to a more capable model

Better quality, smarter answers

When your AI needs to be smarter, not just handle more.

  • · Step up a tier — Fast → Balanced → Powerful → Maximum
  • · Better reasoning and accuracy
  • · Switch models anytime from your dashboard
  • · Higher-tier models answer a little more slowly
Best for: complex reasoning, domain expertise, specialized tasks

Auxen watches your usage and suggests which path makes sense. You're never locked in — change throughput, switch models, or pause anytime.

How Auxen fits a multi-layer AI system
Where a private, open-weight workhorse sits alongside frontier APIs in a real architecture.
Read the architecture→