# Commercial readiness gap analysis This document compares the current gateway and console with the customer-facing model catalog at `https://zenmux.ai/models?sort=newest`. It separates working capabilities from product gaps so an unfinished control is never presented as a commercial feature. ## What works now - OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages proxying, streaming, routing, retry, authentication, persistent usage, prepaid billing, quotas, rate limits, concurrent request limits, RBAC, audit logs, and a PostgreSQL-backed console. - Stripe-hosted manual top-up and payment-method setup, off-session automatic top-up, signed/idempotent Webhook crediting, refund/dispute handling, and reconciliation. The gateway never accepts card details and never credits a success redirect. - Tenant billing profiles persist invoice name, email, and postal address and synchronize them to Stripe Customer with a stable idempotency key. Stripe failures preserve the local profile for retry; tax IDs stay in Stripe-hosted Checkout or Customer Portal rather than this database. - PostgreSQL is the source of truth. Redis accelerates invalidation and shared counters but is not required for startup, control-plane writes, billing, or balance correctness. - Verified self-service registration, invitation acceptance, password reset, persistent login throttles, per-device session revocation, encrypted email outbox delivery, TOTP with recovery codes, and WebAuthn Passkeys. - Tenant-scoped default/fallback model preferences and RBAC-separated low-balance notification thresholds, persisted in PostgreSQL and applied by Quickstart and the notification worker. - Per-key model restrictions, monthly spend caps, expiration, tags, and last-use tracking. Restrictions are enforced by the runtime snapshot and billing transaction, not only rendered by the console. The developer console also has a page-memory API Playground and displays current-month settled spend, pending reservations, request count, remaining cap, and last use for each key. - The authenticated model catalog includes a customer-safe detail view, current price-version cost estimates for input/output/cache tokens, copyable model IDs, developer filtering, release/price/context sorting, one-click Playground selection, cURL/Python/Node examples, and actionable diagnostics for authentication, balance, model, rate-limit, and provider errors. - The unauthenticated `/admin/models` catalog exposes only globally available models and supports search, protocol/input/developer filters, release/price/context sorting, model details, versioned token prices, aggregate route availability, and a pre-registration cost estimate. Tenant/key allowlists, upstream model IDs, provider IDs, URLs, and routing weights are excluded by a dedicated public type. - Quickstart colocates balance state, direct starter-key creation, OpenAI and Anthropic SDK base URLs, copyable REST endpoints, environment configuration, code examples, and the live Playground. A newly created key is scoped to the selected project/model and is kept only in page memory after its one-time reveal. - Usage events open into a privacy-safe request diagnostic with the complete request ID, route, retry, protocol, latency, throughput, cache-token, and settlement fields; copied JSON excludes prompts, responses, and secrets. - Ledger-backed model cost ranking and provider performance views share the Usage filters and report period-over-period charge change, success rate, cache hit, P95 latency, and missing-usage exposure without sampling browser data. - Runtime routes use a per-instance 100-attempt availability window, response-header latency EWMA, and a 3-failure/30-second circuit breaker. The customer model detail compares safe provider runtime fields and prevents launching a model while every route is cooling down. - Developers can keep automatic failover or pin a request with `model-id:provider-slug`. Provider slugs are stable public identifiers returned by the safe model catalog and selectable in Quickstart and Playground; pinned requests never fail over to another provider, while billing and key allowlists remain keyed by the canonical base model. ## Customer product gaps ### P0 before a public commercial launch - Production mail provider DNS authentication (SPF/DKIM/DMARC) and provider-side bounce/complaint wiring remain deployment tasks; the signed feedback endpoint, suppression table, low-balance notifications, retries, and dead-letter mail outbox are implemented. Local development uses Mailpit. - Explicit tax treatment after registrations are confirmed still needs legal and product sign-off. Invoice details, hosted invoice/PDF/receipt links, payment history, refunds, disputes, reconciliation, and CSV ledger export are implemented. - Operational separation of the public inference listener from the management listener, HTTPS-only cookies behind a trusted proxy, backup/restore drills, migration rollback policy, secret rotation, and alerting for usage settlement or Webhook backlogs. - Enterprise identity integrations (OIDC/SAML/SCIM), custom roles, and approval workflows are not included in the current console; password, invite, session, TOTP, Passkey, RBAC, and audit flows are implemented. ### P1 for ZenMux-like breadth - OpenAI Embeddings, Images, Speech and Transcriptions; Gemini native APIs; rerank and other media endpoints. The existing protocol field does not make these APIs implemented. - Active provider probes, first-token latency, throughput-aware selection, cross-instance health aggregation, and customer-visible status history. Runtime circuit breaking and request-derived route health are implemented; historical success, total latency, cache hit, missing usage, and cost come from the Usage Ledger. - Provider price ranges, non-token search/image/audio pricing units, and richer deprecation notices. Provider runtime comparison and release sorting are implemented. - Usage exports, scheduled reports, organization invites, custom roles, OIDC/SAML SSO, SCIM, and support impersonation with approval and full audit evidence. Model/provider cost attribution is implemented in the console. ## Environment boundary Versioned JSON configuration stores only environment variable names for external services. Local values live in `.env.debug`, which is ignored by Git and created with mode `0600`. | Integration value | Environment variable | | --- | --- | | HTTP listen address | `AIGW_SERVER_ADDRESS` | | PostgreSQL URLs | `AIGW_DATABASE_URL`, `AIGW_DATABASE_URL_DOCKER` | | Redis URLs | `AIGW_REDIS_URL`, `AIGW_REDIS_URL_DOCKER` | | Provider credential encryption | `AIGW_CREDENTIAL_KEY` | | Bootstrap administrator | `AIGW_ADMIN_TOKEN` | | Stripe integration switch | `AIGW_STRIPE_ENABLED` | | Stripe application key | `AIGW_STRIPE_API_KEY` | | Stripe CLI development key | `AIGW_STRIPE_CLI_API_KEY` | | Stripe Webhook signing secret | `AIGW_STRIPE_WEBHOOK_SECRET` | | Stripe result URLs | `AIGW_STRIPE_SUCCESS_URL`, `AIGW_STRIPE_CANCEL_URL`, `AIGW_STRIPE_PORTAL_RETURN_URL` | | Console public URL | `AIGW_PUBLIC_URL` | | Public inference/API URL used by customer examples | `AIGW_INFERENCE_PUBLIC_URL` | | SMTP endpoint/sender | `AIGW_SMTP_ADDRESS`, `AIGW_SMTP_FROM_ADDRESS` | | SMTP credentials | `AIGW_SMTP_USERNAME`, `AIGW_SMTP_PASSWORD` | | WebAuthn RP/origins | `AIGW_WEBAUTHN_RP_ID`, `AIGW_WEBAUTHN_ORIGINS` | | Static upstream endpoint/key | provider `base_url_env`, `api_key_env` | External service values are never embedded in versioned JSON. When mail is enabled, startup validates the sender and SMTP endpoint; username/password must either both be provided or both be empty. Production should use authenticated `starttls` or implicit `tls`; the checked-in local example uses unauthenticated Mailpit with `tls_mode: none`.