We run the models
We operate the GPUs. You write the application; we run the models.
Per-project, inclusive of GPU.
Run, govern, and budget every enterprise AI application. A budget on every project, one audit trail on every app, one endpoint from experiment to production.
Payloads never logged by default — one disclosed exception.
Your teams' AI applications
Bespin Works is the AI Applications Platform: the layer that runs, governs, and budgets every enterprise AI application.
One endpoint: our GPUs, or your provider keys.
Per-project budgets, per-project access, one audit trail.
Cost is known before a single call runs.
Any open-weight model on our compute, or the big provider APIs on your keys: same client you already run — switch anytime.
We operate the GPUs. You write the application; we run the models.
Per-project, inclusive of GPU.
Route OpenAI, Anthropic, or Google traffic through Bespin Works with your own keys. We govern and audit; we never touch or mark up your provider spend.
Flat per-project fee.
Billed per project, inclusive of GPU, each with its own budget.
| Feature | Bespin Works | Typical Token Gateways |
|---|---|---|
| Billing Model | Per-project, inclusive of GPU Flat per-project fee on Managed Provider |
Token Markup |
| Budget Control | Per-Project Budgets | Post-payment Billing |
| Governance | Project Report | Usage console, no governance report |
| Provider Lock-in | None on Managed Provider | Their platform in the request path |
No per-token markup to model. Bring us one project and we'll scope its budget with you.
See who is driving spend — and set the budget before it runs. A budget on every project; hit the budget and we throttle, not bill.
One audit trail attributes every request to the project and API key that made it. You build. We govern.
Delivered monthly · one per project
The guarantees
Start as an experiment, graduate to production — on the same endpoint. No rebuilding, no re-platforming, no new infrastructure.
# Today
curl https://<your current endpoint>/v1/chat/completions \ -H 'Authorization: Bearer <your API key>' \ -H 'Content-Type: application/json' \ -d '{"model": "<model>", "messages": [...]}'
# With Bespin Works — same call, two fields swapped
curl https://<project endpoint>/v1/chat/completions \ -H 'Authorization: Bearer <Bespin API key>' \ -H 'Content-Type: application/json' \ -d '{"model": "<model>", "messages": [...]}'
# That's the whole switch.
The short answers to the questions we hear most.
An AI Applications Platform is the layer that runs, governs, and budgets enterprise AI applications. SaaS gives you the application runtime; GPU clouds give you the inference substrate; gateways route between models. Bespin Works combines all three, and prices by the project, not the token.
Token-based gateways bill by the token, which makes invoices hard to predict. Bespin Works bills per project, inclusive of GPU, with a budget on every project, so you can model the cost before you deploy.
Every project runs under a per-project plan with its own budget. Managed Inference projects are billed per project, inclusive of GPU; Managed Provider projects bill a flat per-project fee. Hit the budget and we throttle, not bill.
No. With Managed Inference, Bespin Works runs open-weight models on managed compute. You write the application, we run the models. No GPU ownership, no infrastructure to provision.
Managed Inference runs open-weight models on GPU infrastructure Bespin Works operates for you; Managed Provider sends requests from your application, through Bespin Works, to the provider your API keys belong to. You can also point Managed Inference at your own model servers.
Yes. With Managed Provider, you bring your own keys and Bespin Works routes, governs, and audits the traffic without touching, reselling, or marking up your provider spend.
You're not tied to a single model provider. Switch a project between Managed Inference and your own provider keys anytime, on the same endpoint, without re-platforming.
Every project gets its own budget, every request is attributed to a project and audited — including the ones we reject — and every month you get a Project Report. Governance covers cost, access, and routing — model allowlists, project budgets, and per-request traces with latency metrics: one point of policy for all your enterprise AI.
Encryption in transit, least-privilege access control, request and audit logging, and infrastructure-as-code.
Payload content never logged. Logs and telemetry carry envelope metadata — never prompts, completions, or embeddings. One disclosed exception: an opt-in, per-project response cache, off by default and audit-logged.
Request a POC — we scope it on a 30-minute call, then your team gets Enterprise AI Sandbox access. Email [email protected].
A configuration change: point your client at your project endpoint with your Bespin key — no code rewrite. Managed Provider projects keep using your existing provider keys.
One POC, scoped on a 30-minute call. Your approver sees the Project Report before you commit.
Email the founder directly — this opens your mail client with the subject and the three lines we need.