Private AI, pay-as-you-go

Your own private AI model. Pay only while it's working.

Deploy your own private open-weight model. Get a stable endpoint. Pay per minute while it's active. We run everything underneath — and your data is never sent to a shared third-party API or used to train anyone's model.

The Auxen promise

Where your data goes, in plain English.

A private deployment, just for you

Your model runs in a single-tenant deployment. It is never shared with another customer, and no one else's requests are answered beside yours.

Never used for training

Your prompts, conversations, and uploaded documents are never used to train any model — ours or anyone else's.

Never routed to a third party

Your data is not forwarded to OpenAI, Anthropic, Google, or any outside API. It stays with the deployment that answers you, and is erased when you delete it.

Every Auxen deployment is single-tenant by default. The low-cost Shared option (Fast tier only) is the one opt-in exception, clearly labeled wherever it appears.

Two ways in

Built for your business, open to your builders.

Buy an outcome, or get a private endpoint. Same isolation guarantees underneath either path.

For your business

A private AI assistant for your team

Answers from your own documents, safe for sensitive data, ready to use through a chat panel, a shareable link, or a chat bubble on your website. No code, nothing to manage — and you only pay while it's on.

  • Answers from your own company documents
  • Add it to your website in one line
  • Pay by the minute — pause it and the meter stops
Explore for business→
For builders and agents

A private, stable endpoint for your agents and apps

Deploy an open-weight model and get an OpenAI-compatible endpoint that's yours alone — with full lifecycle control over an API and a remote MCP server. Pay per minute while it's active.

  • OpenAI-compatible REST — drop into any framework
  • OAuth 2.1 + PKCE, per-deployment API keys
  • Deploy, pause, and delete over API or MCP
Explore for builders→

Private AI,
in three steps

No infrastructure to manage. You bring the knowledge; Auxen keeps it private.

Bring your knowledge

Upload the documents your AI should know — policies, product info, past answers. They stay private to your deployment and are never used to train any model.

Your documents

We run it privately

Auxen deploys a private model just for you — no other customer shares it. Nothing is routed to an outside API, and you pay per minute only while it's active.

A private deployment

Put it to work

Use it through a chat panel, a shareable link, a chat bubble on your website, or a private API endpoint — whichever fits how your team works.

Chat · widget · link · API

Everything included.
Nothing leaves your control.

One private model, reachable however your team works — through a screen, a link, your website, or your code.

A private deployment

Your model, your endpoint, your data. One single-tenant deployment answers only you — nothing shared with other customers, and its data is erased when you delete it.

Answers from your documents

Upload your files and Persona Studio grounds every answer in your own knowledge — policies, products, past answers. Never used to train any model.

Website chat widget

One line of code adds a private chat bubble to any site — Squarespace, Wix, Webflow, WordPress, Shopify. Your visitors, your data, your model.

Shareable chat link

Share a private, access-controlled, branded link with your team or your customers. Ready in seconds — no app to install.

OpenAI-compatible endpoint

For technical teams: a stable endpoint that drops into any framework speaking OpenAI — openai-python, Vercel AI SDK, LangChain, LlamaIndex.

Programmatic control

Deploy, pause, and delete your model over a REST API or a remote MCP server, secured with OAuth 2.1 + PKCE. Built for agents that manage their own endpoints.

Simple pay-as-you-go pricing.

One simple rate card for everyone. Pick how capable your model should be and how many requests it handles at once — then pay per minute while it's active. Pause it and the meter stops.

Private deployment · per hour, billed per minute
Model tierStandardDoubleQuadMax
Fast
Quickest and cheapest — sorting, pulling out details, short answers
$0.15/hr$0.29/hr$0.56/hr$1.13/hr
Balanced
The everyday workhorse — chat, summaries, agents' day-to-day work
$0.20/hr$0.38/hr$0.75/hr$1.50/hr
Powerful
Deeper reasoning, long documents, coding
$0.65/hr$1.24/hr$2.44/hr$4.88/hr
Maximum
The most capable open models — can lead a cost-sensitive stack
$2.25/hr$4.28/hr$8.44/hr$16.88/hr

Rates are per hour, billed per minute of active time. Throughput levels increase simultaneous request handling at the same model quality.

Warm-up is free. Once your model is ready, each session is billed for at least 30 minutes (35 on Maximum).

Shared option: the Fast tier is also available at $0.05/hr on a shared, multi-tenant deployment — the one exception to single-tenant, clearly labeled when you choose it.

  • A private deployment — your model, your endpoint, your data
  • Full API access + OpenAI-compatible endpoint
  • Knowledge Base, Persona Studio, website widget
  • Pause anytime — the per-minute meter stops after the session minimum
Start with $10$10 minimum deposit · Credits never expire · No contract
Frontier models — contact us
Frontier models — the largest open releases — are available on request.
Contact us→
Volume or bespoke
Running many deployments, need a volume rate, or want something set up specially — your own models, on-premise, or extra compliance? Talk to us.
Contact us→
Cost calculator

Would Auxen cost you less — or more?

Model tier
Throughput
Active per day
Tokens per month
Auxen, private deployment
$48.64 /month
A frontier API at list price
$300 /month

At this volume, a private Auxen deployment costs about $251 less per month — and your data stays private. Make sure your chosen tier can handle the volume; if not, add throughput.

Estimate only. Auxen bills per minute of active time regardless of tokens. Warm-up is free. Once your model is ready, each session is billed for at least 30 minutes (35 on Maximum). The per-token figure assumes 75% input / 25% output at $3/$15 per million tokens. Open-weight models aren't identical to frontier models — for high-volume, well-specified work the quality gap is usually smaller than the cost gap.

More requests, or smarter answers?

Two different ways to grow. Choose the one that matches what you need.

Increase throughput

More requests, same quality

When you're getting good answers but need more of them at once.

  • · Stay on the same model
  • · Move from Standard to Double, Quad, or Max
  • · More throughput handles more simultaneous requests at the same quality.
  • · Double, Quad and Max cost less per request than Standard
Best for: high traffic, growing user base, busy periods
Move to a more capable model

Better quality, smarter answers

When your AI needs to be smarter, not just handle more.

  • · Step up a tier — Fast → Balanced → Powerful → Maximum
  • · Better reasoning and accuracy
  • · Switch models anytime from your dashboard
  • · Higher-tier models answer a little more slowly
Best for: complex reasoning, domain expertise, specialized tasks

Auxen watches your usage and suggests which path makes sense. You're never locked in — change throughput, switch models, or pause anytime.

How Auxen compares
to the alternatives.

Honest side-by-side comparisons against the platforms you're probably evaluating. Each page says where the competitor wins and where Auxen's single-tenant privacy wins — no spin.