Platform

Integrate by changing one URL. Govern it from the console.

Situra sits between your applications and the model providers. It speaks both the OpenAI and the Anthropic dialect, picks an allowed route for every request and records what happened — without keeping what was said.

From your app to the model, in five steps.

An illustrative request from a project with ES residency: Azure’s edge in Spain, the gateway’s checks, route selection and metering. Step through it at your own pace.

Illustrative request · project with ES residency

1Only the base URL and the key change. The rest of your code stays the same.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.situra.ai/v1",
    api_key="situ_live_…",
)
r = client.chat.completions.create(model="gpt-5-mini", …)

The OpenAI and Anthropic APIs. No rewrite.

The official OpenAI and Anthropic SDKs, the frameworks built on them and coding agents work by pointing the base URL at Situra. Errors come back in the caller’s dialect.

MethodPathCompatible with
POST/v1/chat/completionsOpenAI Chat Completions, including streaming, tools and images
POST/v1/embeddingsOpenAI Embeddings
GET/v1/modelsOpenAI List models — only the models the key’s project may use
POST/v1/messagesAnthropic Messages, including streaming, tools and images
POST/v1/messages/count_tokensAnthropic token counting (estimated if the upstream lacks it)

Authenticate with Authorization: Bearer situ_… or, Anthropic style, x-api-key: situ_….

Any catalogue model can be called in either format: the gateway translates requests and responses, streaming included.

Quickstart

Every request goes through the same checks, in the same order.

The data plane never queries the database on the hot path: it works from a snapshot of keys, balances, limits and routes that the control plane pushes over NATS.

The tier never widens

A fallback can only be another route in the same tier or a stricter one. If none is left, the answer is a 503 (no allowed route) or a 502/504 (upstream failure after all fallbacks).

Every served response carries the tier and region of the route that served it (x-situra-residency, x-situra-route-region); every response carries x-situra-attempts. The optional x-situra-residency: es|eu request header can only tighten the project’s tier.

  1. 1

    Key

    Format and checksum are validated, the key is looked up in the snapshot and its HMAC verified with the key’s pepper version.

  2. 2

    Status

    The key has not expired and the organisation is active (or pending verification, with test keys only).

  3. 3

    Limits

    Requests per minute for the key; requests and tokens per minute for the organisation.

  4. 4

    Model

    The model is in the intersection of the key’s and the project’s allowlists.

  5. 5

    Balance and budget

    Positive balance (prepaid) or within the credit limit (invoiced), and no hard-stop budget exhausted.

  6. 6

    Routes

    Active routes the project’s tier allows, zero-retention if the project requires it, by priority and weight.

  7. 7

    Call

    If the route fails in a retryable way (429, 5xx, timeout, connection) before the first byte is streamed, the next allowed route is tried.

  8. 8

    Metering

    Provider-reported tokens, cost at the route’s current price, latency and the route used.

Budgets with alerts and a hard stop.

Scoped budgets
Per organisation, project or key, monthly, in euros.
Alerts
At 50, 80 and 100% by default, configurable.
Hard stop
Optional per budget: once exhausted, requests get a 402.
Prepaid
Each request debits a real-time counter; at zero balance, access stops.
Invoiced accounts
A credit limit per account and a monthly invoice built from the credit ledger.
Auto-recharge
Below a threshold you choose, with the fee itemised.

Per-request observability. Metadata only by default.

By default we record metadata: who, which model, how many tokens, what it cost, how long it took and how it ended. Prompt and response logging is opt-in per project, with 0–365 days of retention, encrypted under your organisation’s data key.

  • request_id
  • model and route
  • project and key
  • residency and region
  • input and output tokens
  • cache tokens
  • cost in EUR
  • latency and TTFT
  • attempts
  • status and error type

Breakdowns by model, project, key, provider, region and residency; latency percentiles; CSV export; comparison with the previous period.

Console usage view Console usage view
Real console with demo data.

A key is shown exactly once.

Keys look like situ_live_<id>_<secret>, with at least 256 bits of entropy in the secret. We store only an HMAC-SHA256 with a pepper held in Azure Key Vault.

  • Project scope, mandatory expiry (1–365 days, 90 by default) and an optional model allowlist.
  • Per-key requests-per-minute limit.
  • Rotation with a grace period (0–168 h) for zero-downtime deploys.
  • Revocation takes effect immediately on every gateway replica.
  • situ_test_ test keys against a deterministic sandbox, at no cost.
  • Planned: GitHub’s secret-scanning partner programme, so leaked keys get revoked.

Where Situra fits.

There are good AI gateways and workspaces. Situra focuses on per-request residency, ENS and the Spanish public sector.

Comparison with other AI gateway and workspace offerings
OfferingWhat it offers (public description)Where Situra differs
Situra this service Managed AI gateway in Azure Spain Central (Madrid), compatible with OpenAI and Anthropic clients; ES ⊂ EU ⊂ Global residency enforced per project on every request. Not certified yet: designed for ENS category Alta; ENS Media + ISO 27001 audit targeted for Q2–Q3 2027. Private beta in Q4 2026.
nexos.ai Gateway, chat and agents, adoption observer; EU-hosted. ENS as a certification target, Spain-hosted, Spanish-first sales and support.
OpenRouter Broad model aggregator behind one API. Per-project data residency, a certification roadmap, and enterprise governance (budgets, audit, four-eyes).
Open WebUI Open-source UI with a gateway layer, operated by each organisation. Managed, multi-tenant service with certification as a target.
Cohere North Enterprise workspace and agents. Model-neutral gateway, ENS as a target, focus on the Spanish public sector.

nexos.ai publicly lists GDPR, SOC 2 Type II, ISO 27001 and ISO 42001; ENS was not listed when we reviewed it in September 2026. (source: nexos.ai)

Based on vendors’ public websites as reviewed in September 2026. If something has changed or is wrong, tell us and we will correct it.

A data plane and a control plane. All in Azure Spain Central.

The data plane is a Go service that embeds Bifrost’s provider adapters (Apache 2.0) at a pinned release; authentication, residency, budgets and metering are our own code.

Situra architecture Clients call Azure Front Door, which passes the request to the data plane. The data plane picks an allowed route and calls the provider. The control plane pushes configuration over NATS. Usage events go to ClickHouse and the credit ledger; payments to Stripe. Everything except the providers and Stripe runs in Azure Spain Central. Azure Spain Central · Madrid Outside the platform config usage Client apps SDKs · agents Claude Code Azure Front Door WAF · TLS Data plane Go · key · residency · budget routes · fallbacks · streaming Console + back office OIDC · roles · four-eyes Control plane tenants · keys · policies Postgres with RLS NATS JetStream config · usage ClickHouse usage analytics Credit ledger Postgres · append-only Upstream routes by project tier ES · Madrid EU · regional Global Stripe payments · VAT
Keeps serving from cached config if the control plane is down.

Private beta · Q4 2026

We are looking for three to five organisations for the beta.

Spanish public administrations and regulated organisations that want to use AI models with data residency under control. We work with each organisation on ENS categorisation, the DPA and the integration.