summaryrefslogtreecommitdiff
path: root/docs/commercial-readiness.md
diff options
context:
space:
mode:
Diffstat (limited to 'docs/commercial-readiness.md')
-rw-r--r--docs/commercial-readiness.md72
1 files changed, 46 insertions, 26 deletions
diff --git a/docs/commercial-readiness.md b/docs/commercial-readiness.md
index 6dc9064..cbf6846 100644
--- a/docs/commercial-readiness.md
+++ b/docs/commercial-readiness.md
@@ -7,7 +7,7 @@ commercial feature.
## What works now
-- OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages proxying, streaming, routing,
+- OpenAI Chat Completions, OpenAI Responses, OpenAI Embeddings, and Anthropic Messages proxying, streaming, routing,
retry, authentication, persistent usage, prepaid billing, quotas, rate limits,
concurrent request limits, RBAC, audit logs, and a PostgreSQL-backed console.
- Stripe-hosted manual top-up and payment-method setup, off-session automatic
@@ -18,20 +18,21 @@ commercial feature.
synchronize them to Stripe Customer with a stable idempotency key. Stripe
failures preserve the local profile for retry; tax IDs stay in Stripe-hosted
Checkout or Customer Portal rather than this database.
-- PostgreSQL is the source of truth. Redis accelerates invalidation and shared
- counters but is not required for startup, control-plane writes, billing, or
- balance correctness.
+- PostgreSQL is the source of truth. Redis accelerates invalidation, shared
+ counters, and short-lived route-health propagation but is not required for
+ startup, control-plane writes, billing, or balance correctness.
- Verified self-service registration, invitation acceptance, password reset,
persistent login throttles, per-device session revocation, encrypted email
outbox delivery, TOTP with recovery codes, and WebAuthn Passkeys.
- Tenant-scoped default/fallback model preferences and RBAC-separated low-balance
- notification thresholds, persisted in PostgreSQL and applied by Quickstart and
- the notification worker.
-- Per-key model restrictions, monthly spend caps, expiration, tags, and last-use
- tracking. Restrictions are enforced by the runtime snapshot and billing
+ and anomalous-spend notification settings, persisted in PostgreSQL and applied
+ by Quickstart and the notification worker. Unconfigured tenants inherit
+ deployment defaults.
+- Per-key model restrictions, daily/monthly spend caps, RPM/TPM, expiration,
+ tags, prefix/suffix-only display, disable/enable, atomic rotation, and last-use tracking. Restrictions are enforced by the runtime snapshot and billing
transaction, not only rendered by the console. The developer console also has
- a page-memory API Playground and displays current-month settled spend, pending
- reservations, request count, remaining cap, and last use for each key.
+ a page-memory API Playground and displays current-day/month settled spend, pending
+ reservations, request count, remaining caps, rate policy, and last use for each key.
- The authenticated model catalog includes a customer-safe detail view, current
price-version cost estimates for input/output/cache tokens, copyable model IDs,
developer filtering, release/price/context sorting, one-click Playground
@@ -39,8 +40,8 @@ commercial feature.
authentication, balance, model, rate-limit, and provider errors.
- The unauthenticated `/admin/models` catalog exposes only globally available
models and supports search, protocol/input/developer filters, release/price/context
- sorting, model details, versioned token prices, aggregate route availability,
- and a pre-registration cost estimate. Tenant/key allowlists, upstream model IDs,
+ sorting, server-rendered per-model canonical URLs with protocol code examples,
+ versioned token prices, and aggregate route availability. Tenant/key allowlists, upstream model IDs,
provider IDs, URLs, and routing weights are excluded by a dedicated public type.
- Quickstart colocates balance state, direct starter-key creation, OpenAI and
Anthropic SDK base URLs, copyable REST endpoints, environment configuration,
@@ -49,13 +50,21 @@ commercial feature.
- Usage events open into a privacy-safe request diagnostic with the complete
request ID, route, retry, protocol, latency, throughput, cache-token, and
settlement fields; copied JSON excludes prompts, responses, and secrets.
-- Ledger-backed model cost ranking and provider performance views share the
+- Ledger-backed model, API-key, and provider performance views share the
Usage filters and report period-over-period charge change, success rate,
- cache hit, P95 latency, and missing-usage exposure without sampling browser data.
-- Runtime routes use a per-instance 100-attempt availability window, response-header
- latency EWMA, and a 3-failure/30-second circuit breaker. The customer model detail
- compares safe provider runtime fields and prevents launching a model while every
- route is cooling down.
+ cache hit, total-latency/TTFT P50 and P95, and missing-usage exposure without
+ sampling browser data. The request list uses stable cursor pagination; tenant
+ results redact provider IDs, names, and upstream model identifiers server-side.
+- Runtime routes use a 100-attempt availability window, response-header latency
+ EWMA, and a 3-failure/30-second circuit breaker. An optional bounded Redis Stream
+ shares recent outcomes and TTFT between instances and replays them after restart;
+ its non-blocking publisher falls back to the local window during Redis failure.
+ The customer model detail identifies shared samples, compares safe provider
+ runtime fields, and prevents launching a model while every route is cooling down.
+- Optional active provider probes make one authenticated, non-inference `/models`
+ request per provider, validate the JSON API response, feed the same circuit breaker,
+ expose probe counters/timestamps in the console and Prometheus, and remain disabled
+ unless `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED=true` is explicitly injected.
- Developers can keep automatic failover or pin a request with
`model-id:provider-slug`. Provider slugs are stable public identifiers returned by
the safe model catalog and selectable in Quickstart and Playground; pinned
@@ -73,7 +82,8 @@ commercial feature.
- Explicit tax treatment after registrations are confirmed still needs legal
and product sign-off. Invoice details, hosted invoice/PDF/receipt links,
payment history, refunds, disputes, reconciliation, and CSV ledger export are
- implemented.
+ implemented. Usage debit rows link to their exact request details, and the
+ lookup remains tenant-scoped even when a caller knows another request ID.
- Operational separation of the public inference listener from the management
listener, HTTPS-only cookies behind a trusted proxy, backup/restore drills,
migration rollback policy, secret rotation, and alerting for usage settlement
@@ -84,13 +94,13 @@ commercial feature.
### P1 for ZenMux-like breadth
-- OpenAI Embeddings, Images, Speech and Transcriptions; Gemini native
- APIs; rerank and other media endpoints. The existing protocol field does not
- make these APIs implemented.
-- Active provider probes, first-token latency, throughput-aware selection,
- cross-instance health aggregation, and customer-visible status history. Runtime
- circuit breaking and request-derived route health are implemented; historical
- success, total latency, cache hit, missing usage, and cost come from the Usage Ledger.
+- OpenAI Images, Speech and Transcriptions; Gemini native APIs; rerank and other
+ media endpoints. Token/image/second metering units exist, but a typed unit does
+ not make those endpoints implemented.
+- Throughput/cost-aware selection and customer-visible status history. Runtime
+ TTFT/availability-aware selection, bounded exploration, circuit breaking,
+ cross-instance short-lived health aggregation, and request-derived route health are implemented; historical
+ success, total latency, TTFT, cache hit, missing usage, and cost come from the Usage Ledger.
- Provider price ranges, non-token search/image/audio pricing units, and richer
deprecation notices. Provider runtime comparison and release sorting are implemented.
- Usage exports, scheduled reports, organization invites, custom roles,
@@ -117,11 +127,21 @@ with mode `0600`.
| Stripe result URLs | `AIGW_STRIPE_SUCCESS_URL`, `AIGW_STRIPE_CANCEL_URL`, `AIGW_STRIPE_PORTAL_RETURN_URL` |
| Console public URL | `AIGW_PUBLIC_URL` |
| Public inference/API URL used by customer examples | `AIGW_INFERENCE_PUBLIC_URL` |
+| Optional authenticated provider probes | `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED` |
+| Optional cross-instance provider health | `AIGW_PROVIDER_SHARED_HISTORY_ENABLED` |
| SMTP endpoint/sender | `AIGW_SMTP_ADDRESS`, `AIGW_SMTP_FROM_ADDRESS` |
| SMTP credentials | `AIGW_SMTP_USERNAME`, `AIGW_SMTP_PASSWORD` |
| WebAuthn RP/origins | `AIGW_WEBAUTHN_RP_ID`, `AIGW_WEBAUTHN_ORIGINS` |
| Static upstream endpoint/key | provider `base_url_env`, `api_key_env` |
+Before enabling Stripe in a deployment, run `go run ./cmd/stripe-preflight` with
+the test restricted key. It performs only authenticated list requests, pins the
+SDK API version, refuses live-mode keys, and returns per-resource permission
+results without printing the key or Stripe response messages. Checkout, Portal,
+automatic top-up, refund, Webhook, and reconciliation write permissions still
+require the full test-mode workflow because the preflight intentionally creates
+no Stripe objects.
+
External service values are never embedded in versioned JSON. When mail is
enabled, startup validates the sender and SMTP endpoint; username/password must
either both be provided or both be empty. Production should use authenticated