1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
|
# Commercial readiness gap analysis
This document compares the current gateway and console with the customer-facing
model catalog at `https://zenmux.ai/models?sort=newest`. It separates working
capabilities from product gaps so an unfinished control is never presented as a
commercial feature.
## What works now
- OpenAI Chat Completions, OpenAI Responses, OpenAI Embeddings, and Anthropic Messages proxying, streaming, routing,
retry, authentication, persistent usage, prepaid billing, quotas, rate limits,
concurrent request limits, RBAC, audit logs, and a PostgreSQL-backed console.
- Stripe-hosted manual top-up and payment-method setup, off-session automatic
top-up, signed/idempotent Webhook crediting, refund/dispute handling, and
reconciliation. The gateway never accepts card details and never credits a
success redirect.
- Tenant billing profiles persist invoice name, email, and postal address and
synchronize them to Stripe Customer with a stable idempotency key. Stripe
failures preserve the local profile for retry; tax IDs stay in Stripe-hosted
Checkout or Customer Portal rather than this database.
- PostgreSQL is the source of truth. Redis accelerates invalidation, shared
counters, and short-lived route-health propagation but is not required for
startup, control-plane writes, billing, or balance correctness.
- Verified self-service registration, invitation acceptance, password reset,
persistent login throttles, per-device session revocation, encrypted email
outbox delivery, TOTP with recovery codes, and WebAuthn Passkeys.
- Tenant-scoped default/fallback model preferences and RBAC-separated low-balance
and anomalous-spend notification settings, persisted in PostgreSQL and applied
by Quickstart and the notification worker. Unconfigured tenants inherit
deployment defaults.
- Per-key model restrictions, daily/monthly spend caps, RPM/TPM, expiration,
tags, prefix/suffix-only display, disable/enable, atomic rotation, and last-use tracking. Restrictions are enforced by the runtime snapshot and billing
transaction, not only rendered by the console. The developer console also has
a page-memory API Playground and displays current-day/month settled spend, pending
reservations, request count, remaining caps, rate policy, and last use for each key.
- The authenticated model catalog includes a customer-safe detail view, current
price-version cost estimates for input/output/cache tokens, copyable model IDs,
developer filtering, release/price/context sorting, one-click Playground
selection, cURL/Python/Node examples, and actionable diagnostics for
authentication, balance, model, rate-limit, and provider errors.
- The unauthenticated `/admin/models` catalog exposes only globally available
models and supports search, protocol/input/developer filters, release/price/context
sorting, server-rendered per-model canonical URLs with protocol code examples,
versioned token prices, and aggregate route availability. Tenant/key allowlists, upstream model IDs,
provider IDs, URLs, and routing weights are excluded by a dedicated public type.
- Quickstart colocates balance state, direct starter-key creation, OpenAI and
Anthropic SDK base URLs, copyable REST endpoints, environment configuration,
code examples, and the live Playground. A newly created key is scoped to the
selected project/model and is kept only in page memory after its one-time reveal.
- Usage events open into a privacy-safe request diagnostic with the complete
request ID, route, retry, protocol, latency, throughput, cache-token, and
settlement fields; copied JSON excludes prompts, responses, and secrets.
- Ledger-backed model, API-key, and provider performance views share the
Usage filters and report period-over-period charge change, success rate,
cache hit, total-latency/TTFT P50 and P95, and missing-usage exposure without
sampling browser data. The request list uses stable cursor pagination; tenant
results redact provider IDs, names, and upstream model identifiers server-side.
- Runtime routes use a 100-attempt availability window, response-header latency
EWMA, and a 3-failure/30-second circuit breaker. An optional bounded Redis Stream
shares recent outcomes and TTFT between instances and replays them after restart;
its non-blocking publisher falls back to the local window during Redis failure.
The customer model detail identifies shared samples, compares safe provider
runtime fields, and prevents launching a model while every route is cooling down.
- Optional active provider probes make one authenticated, non-inference `/models`
request per provider, validate the JSON API response, feed the same circuit breaker,
expose probe counters/timestamps in the console and Prometheus, and remain disabled
unless `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED=true` is explicitly injected.
- Developers can keep automatic failover or pin a request with
`model-id:provider-slug`. Provider slugs are stable public identifiers returned by
the safe model catalog and selectable in Quickstart and Playground; pinned
requests never fail over to another provider, while billing and key allowlists
remain keyed by the canonical base model.
## Customer product gaps
### P0 before a public commercial launch
- Production mail provider DNS authentication (SPF/DKIM/DMARC) and provider-side
bounce/complaint wiring remain deployment tasks; the signed feedback endpoint,
suppression table, low-balance notifications, retries, and dead-letter mail
outbox are implemented. Local development uses Mailpit.
- Explicit tax treatment after registrations are confirmed still needs legal
and product sign-off. Invoice details, hosted invoice/PDF/receipt links,
payment history, refunds, disputes, reconciliation, and CSV ledger export are
implemented. Usage debit rows link to their exact request details, and the
lookup remains tenant-scoped even when a caller knows another request ID.
- Operational separation of the public inference listener from the management
listener, HTTPS-only cookies behind a trusted proxy, backup/restore drills,
migration rollback policy, secret rotation, and alerting for usage settlement
or Webhook backlogs.
- Enterprise identity integrations (OIDC/SAML/SCIM), custom roles, and approval
workflows are not included in the current console; password, invite, session,
TOTP, Passkey, RBAC, and audit flows are implemented.
### P1 for ZenMux-like breadth
- OpenAI Images, Speech and Transcriptions; Gemini native APIs; rerank and other
media endpoints. Token/image/second metering units exist, but a typed unit does
not make those endpoints implemented.
- Throughput/cost-aware selection and customer-visible status history. Runtime
TTFT/availability-aware selection, bounded exploration, circuit breaking,
cross-instance short-lived health aggregation, and request-derived route health are implemented; historical
success, total latency, TTFT, cache hit, missing usage, and cost come from the Usage Ledger.
- Provider price ranges, non-token search/image/audio pricing units, and richer
deprecation notices. Provider runtime comparison and release sorting are implemented.
- Usage exports, scheduled reports, organization invites, custom roles,
OIDC/SAML SSO, SCIM, and support impersonation with approval and full audit
evidence. Model/provider cost attribution is implemented in the console.
## Environment boundary
Versioned JSON configuration stores only environment variable names for external
services. Local values live in `.env.debug`, which is ignored by Git and created
with mode `0600`.
| Integration value | Environment variable |
| --- | --- |
| HTTP listen address | `AIGW_SERVER_ADDRESS` |
| PostgreSQL URLs | `AIGW_DATABASE_URL`, `AIGW_DATABASE_URL_DOCKER` |
| Redis URLs | `AIGW_REDIS_URL`, `AIGW_REDIS_URL_DOCKER` |
| Provider credential encryption | `AIGW_CREDENTIAL_KEY` |
| Bootstrap administrator | `AIGW_ADMIN_TOKEN` |
| Stripe integration switch | `AIGW_STRIPE_ENABLED` |
| Stripe application key | `AIGW_STRIPE_API_KEY` |
| Stripe CLI development key | `AIGW_STRIPE_CLI_API_KEY` |
| Stripe Webhook signing secret | `AIGW_STRIPE_WEBHOOK_SECRET` |
| Stripe result URLs | `AIGW_STRIPE_SUCCESS_URL`, `AIGW_STRIPE_CANCEL_URL`, `AIGW_STRIPE_PORTAL_RETURN_URL` |
| Console public URL | `AIGW_PUBLIC_URL` |
| Public inference/API URL used by customer examples | `AIGW_INFERENCE_PUBLIC_URL` |
| Optional authenticated provider probes | `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED` |
| Optional cross-instance provider health | `AIGW_PROVIDER_SHARED_HISTORY_ENABLED` |
| SMTP endpoint/sender | `AIGW_SMTP_ADDRESS`, `AIGW_SMTP_FROM_ADDRESS` |
| SMTP credentials | `AIGW_SMTP_USERNAME`, `AIGW_SMTP_PASSWORD` |
| WebAuthn RP/origins | `AIGW_WEBAUTHN_RP_ID`, `AIGW_WEBAUTHN_ORIGINS` |
| Static upstream endpoint/key | provider `base_url_env`, `api_key_env` |
Before enabling Stripe in a deployment, run `go run ./cmd/stripe-preflight` with
the test restricted key. It performs only authenticated list requests, pins the
SDK API version, refuses live-mode keys, and returns per-resource permission
results without printing the key or Stripe response messages. Checkout, Portal,
automatic top-up, refund, Webhook, and reconciliation write permissions still
require the full test-mode workflow because the preflight intentionally creates
no Stripe objects.
External service values are never embedded in versioned JSON. When mail is
enabled, startup validates the sender and SMTP endpoint; username/password must
either both be provided or both be empty. Production should use authenticated
`starttls` or implicit `tls`; the checked-in local example uses unauthenticated
Mailpit with `tls_mode: none`.
|