Main Service Panel
Per-project budget Your tools connect unchanged

Run AI. Keep it under control.

Run, govern, and budget every enterprise AI application. A budget on every project, one audit trail on every app, one endpoint from experiment to production.

Payloads never logged by default — one disclosed exception.

[email protected]

Your teams' AI applications

Managed Inference Managed Provider Runtime mode: Managed Inference

One platform. Every AI app.

Bespin Works is the AI Applications Platform: the layer that runs, governs, and budgets every enterprise AI application.

The stack, in one panel

Run

One endpoint: our GPUs, or your provider keys.

Govern

Per-project budgets, per-project access, one audit trail.

Budget

Cost is known before a single call runs.

Your compute or ours

Any open-weight model on our compute, or the big provider APIs on your keys: same client you already run — switch anytime.

Managed Inference

We run the models

We operate the GPUs. You write the application; we run the models.

Per-project, inclusive of GPU.

Managed Provider

Use your existing API keys

Route OpenAI, Anthropic, or Google traffic through Bespin Works with your own keys. We govern and audit; we never touch or mark up your provider spend.

Flat per-project fee.

The mode switchIllustrative mockup — not a product screenshot.

Budgeted, not metered

Billed per project, inclusive of GPU, each with its own budget.

One project, one envelope
Comparison between Bespin Works and Typical Token Gateways
Feature Bespin Works Typical Token Gateways
Billing Model Per-project, inclusive of GPU
Flat per-project fee on Managed Provider
Token Markup
Budget Control Per-Project Budgets Post-payment Billing
Governance Project Report Usage console, no governance report
Provider Lock-in None on Managed Provider Their platform in the request path
The bill, two ways

No per-token markup to model. Bring us one project and we'll scope its budget with you.

Every request. On the record.

See who is driving spend — and set the budget before it runs. A budget on every project; hit the budget and we throttle, not bill.

The console

Three projects, three budgetsIllustrative mockup — not a product screenshot.
support-copilot · Cycle 2026-09 100% of budget
At budget — throttled
Sample data

One audit trail attributes every request to the project and API key that made it. You build. We govern.

support-copilot · Project Report Sample data
  • Spend$5,000 of $5,000 budget
  • Requests served41,208
  • Requests rejected112 throttled at budget
  • Latency p95812 ms

Delivered monthly · one per project

The record

The guarantees

  • Every request accounted for. Attributed to a project and audited, including the ones we reject. Chargeback-ready cost allocation, in the Project Report.
  • Payload content never logged. Logs and telemetry carry envelope metadata — never prompts, completions, or embeddings. One exception, directly below.
  • One disclosed exception: an opt-in, per-project response cache, off by default and audit-logged.
  • Your data estate stays yours. We never touch your training data or your databases.
The ledger

One endpoint. Zero rebuilds.

Start as an experiment, graduate to production — on the same endpoint. No rebuilding, no re-platforming, no new infrastructure.

Sandbox to production, one wire

# Today

curl https://<your current endpoint>/v1/chat/completions \
  -H 'Authorization: Bearer <your API key>' \
  -H 'Content-Type: application/json' \
  -d '{"model": "<model>", "messages": [...]}'

# With Bespin Works — same call, two fields swapped

curl https://<project endpoint>/v1/chat/completions \
  -H 'Authorization: Bearer <Bespin API key>' \
  -H 'Content-Type: application/json' \
  -d '{"model": "<model>", "messages": [...]}'

# That's the whole switch.

Questions, answered

The short answers to the questions we hear most.

What is an AI Applications Platform?

An AI Applications Platform is the layer that runs, governs, and budgets enterprise AI applications. SaaS gives you the application runtime; GPU clouds give you the inference substrate; gateways route between models. Bespin Works combines all three, and prices by the project, not the token.

How is Bespin Works different from a token-based gateway?

Token-based gateways bill by the token, which makes invoices hard to predict. Bespin Works bills per project, inclusive of GPU, with a budget on every project, so you can model the cost before you deploy.

How does billing work?

Every project runs under a per-project plan with its own budget. Managed Inference projects are billed per project, inclusive of GPU; Managed Provider projects bill a flat per-project fee. Hit the budget and we throttle, not bill.

Do I need to own GPUs or build infrastructure?

No. With Managed Inference, Bespin Works runs open-weight models on managed compute. You write the application, we run the models. No GPU ownership, no infrastructure to provision.

Where do the models run?

Managed Inference runs open-weight models on GPU infrastructure Bespin Works operates for you; Managed Provider sends requests from your application, through Bespin Works, to the provider your API keys belong to. You can also point Managed Inference at your own model servers.

Can I use my existing OpenAI, Anthropic, or Google keys?

Yes. With Managed Provider, you bring your own keys and Bespin Works routes, governs, and audits the traffic without touching, reselling, or marking up your provider spend.

What does "no vendor lock-in" mean?

You're not tied to a single model provider. Switch a project between Managed Inference and your own provider keys anytime, on the same endpoint, without re-platforming.

How does governance work, and what does it cover?

Every project gets its own budget, every request is attributed to a project and audited — including the ones we reject — and every month you get a Project Report. Governance covers cost, access, and routing — model allowlists, project budgets, and per-request traces with latency metrics: one point of policy for all your enterprise AI.

What security controls are in place today?

Encryption in transit, least-privilege access control, request and audit logging, and infrastructure-as-code.

Are my prompts and completions logged?

Payload content never logged. Logs and telemetry carry envelope metadata — never prompts, completions, or embeddings. One disclosed exception: an opt-in, per-project response cache, off by default and audit-logged.

How do I get started?

Request a POC — we scope it on a 30-minute call, then your team gets Enterprise AI Sandbox access. Email [email protected].

What does my team change to switch?

A configuration change: point your client at your project endpoint with your Bespin key — no code rewrite. Managed Provider projects keep using your existing provider keys.

Run AI. Keep it under control.

One POC, scoped on a 30-minute call. Your approver sees the Project Report before you commit.

1
POC scoping
2
Enterprise AI Sandbox access
3
Project Report

Email the founder directly — this opens your mail client with the subject and the three lines we need.

[email protected]