diff options
Diffstat (limited to 'docs')
| -rw-r--r-- | docs/architecture.md | 44 | ||||
| -rw-r--r-- | docs/commercial-readiness.md | 72 | ||||
| -rw-r--r-- | docs/runbook.md | 9 |
3 files changed, 77 insertions, 48 deletions
diff --git a/docs/architecture.md b/docs/architecture.md index f8e617f..a88cb5e 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -22,6 +22,8 @@ flowchart LR W --> L D["PostgreSQL 控制面"] -. "事实数据" .-> S["Redis generation / PubSub"] S -. "变更广播" .-> A + P -. "非阻塞健康事件" .-> H["Redis route-health Stream"] + H -. "跨实例回放/消费" .-> R D -. "启动与轮询快照" .-> A D -. "模型和路由快照" .-> R ``` @@ -29,27 +31,27 @@ flowchart LR ## 热路径 1. 入口生成不可预测的请求 ID,并通过 Bearer 或 `x-api-key` 解析客户身份。 -2. 身份包含 `key_id`、`tenant_id`、`project_id`、scopes、密钥模型白名单、月度上限和过期时间;控制面模式使用 PostgreSQL 摘要快照,代理只依赖 `Authenticator` 接口。 +2. 身份包含 `key_id`、`tenant_id`、`project_id`、scopes、密钥模型白名单、日/月上限、RPM/TPM 和过期时间;控制面模式使用 PostgreSQL 摘要快照,代理只依赖 `Authenticator` 接口。 3. 请求体在配置上限内读取一次,以提取公开模型名并支持故障转移时重放。 -4. 项目策略从 PG 热更新快照读取。Redis Lua 原子占用 RPM、估算 TPM 和并发额度;Redis 不可用时退回本机窗口。 -5. Router 按协议过滤路由并跳过仍处于冷却期的 route,再按优先级分组、在同级内按权重选择首选上游。请求使用 `model:provider-slug` 时,基础模型先经过租户/API Key 白名单校验,然后只保留该公开 slug 的兼容 route,绝不跨供应商回退;全部匹配 route 熔断时直接返回可重试的 503。 -6. 计费开启时,请求进入上游前在锁定钱包的 PostgreSQL 事务中同时检查项目月度额度、API Key 月度额度和未结冻结,再冻结保守估算额度;余额不足返回 402,任一月度额度耗尽返回 429。 -7. Provider Adapter 重写上游模型名和凭证,使用进程级共享 Transport 发送请求。每次上游尝试记录响应头延迟与可重试失败;连续 3 次连接失败或 429/502/503/504 后熔断该模型/供应商 route 30 秒。上游返回成功头后即锁定路由;SSE 逐块 flush。 -8. 请求结束后释放并发 lease,并用 `request_id` 幂等持久化 UsageEvent、更新月度汇总、按实际 token 结算;结构化日志仅是异步副本。 +4. 项目和 API Key 策略从 PG 热更新快照读取。Redis Lua 在一次调用中原子占用项目/Key RPM、估算 TPM 和项目并发额度;任一层拒绝会回滚本次全部计数。Redis 不可用时退回本机窗口。 +5. Router 按协议过滤路由并跳过仍处于冷却期的 route,再按优先级分组。同级 route 冷启动时按配置权重选择;积累 5 个可用性样本或 3 个 TTFT 样本后优先近期成功率和 TTFT 更好的 route,并每 20 次保留一次原权重探索。开启共享历史时,各实例从有界 Redis Stream 回放 TTL 内样本并实时消费;写入走非阻塞队列,Redis 不可用时继续使用本机窗口。请求使用 `model:provider-slug` 时,基础模型先经过租户/API Key 白名单校验,然后只保留该公开 slug 的兼容 route,绝不跨供应商回退;全部匹配 route 熔断时直接返回可重试的 503。 +6. 计费开启时,请求进入上游前在锁定钱包的 PostgreSQL 事务中同时检查项目月度额度、API Key 日/月额度和未结冻结,再冻结保守估算额度;余额不足返回 402,任一额度耗尽返回 429。 +7. Provider Adapter 重写上游模型名和凭证,使用进程级共享 Transport 发送请求。每次上游尝试记录响应头延迟与可重试失败,响应转发再记录最终实际 route 的首个有效输出 TTFT;连续 3 次连接失败或 429/502/503/504 后熔断该模型/供应商 route 30 秒。可选主动探测器按供应商去重调用认证 `/models`,拒绝 HTML/非标准 JSON,并把无客户流量时的可用性写入同一熔断器;它默认关闭且不发推理请求。上游返回成功头后即锁定路由;SSE 逐块 flush。 +8. 请求结束后释放并发 lease,并用 `request_id` 幂等持久化 UsageEvent(含首个有效输出 TTFT)、更新月度汇总、按实际 token 结算;结构化日志仅是异步副本。 ## 已固定的扩展边界 | 边界 | 当前实现 | 下一阶段替换 | | --- | --- | --- | -| 客户身份 | PostgreSQL 快照、内存 SHA-256 索引、密钥过期/模型白名单/月度上限 | IP/CIDR 策略、短期服务身份 | +| 客户身份 | PostgreSQL 快照、内存 SHA-256 索引、一次性明文与前缀/六字符后缀显示、密钥过期/模型白名单/日月上限/RPM/TPM、停用与原子轮换 | IP/CIDR 策略、短期服务身份 | | 控制台权限 | 邮箱验证/邀请/重置、登录限流、设备会话、TOTP/恢复码、Passkey、CSRF、六角色 RBAC、租户 SQL scope、审计日志 | SSO/OIDC、SCIM、组织级 MFA 策略、自定义角色、审批流 | | 权限 | `Principal.Scopes` 中的 `inference`、API Key 模型限制 + 控制台 RBAC | ABAC、IP 与条件策略 | -| 模型目录 | PostgreSQL 快照 + 可选 Redis generation 广播 + PG 轮询兜底;租户安全目录 API、供应商运行状态与 Quickstart | 公开 SEO 目录、状态历史、灰度发布 | -| 路由 | priority + weighted selection + failover + `model:provider-slug` 固定供应商 + 真实流量滑动窗口 + 响应头延迟 EWMA + 熔断 | 主动探测、TTFT/吞吐量、成本/质量策略、跨实例健康聚合 | -| 计量 | PostgreSQL UsageEvent + 月度汇总 + 异步日志副本 | 分区表、持久消息流、供应商账单对账 | +| 模型目录 | PostgreSQL 快照 + 可选 Redis generation 广播 + PG 轮询兜底;租户安全目录 API、服务端渲染公开 canonical 详情页、供应商运行状态与 Quickstart | 状态历史、灰度发布 | +| 路由 | priority + weighted selection + failover + `model:provider-slug` 固定供应商 + 真实流量/可选主动探测滑动窗口 + Redis Stream 跨实例健康聚合/重启回放 + TTFT/可用率自适应择优 + 有界探索 + 熔断 | 吞吐量/成本/质量策略、客户可见状态历史 | +| 计量 | PostgreSQL UsageEvent(总延迟/TTFT)+ 月度汇总 + 游标分页 + 模型/Key/供应商分位数聚合 + 异步日志副本 | 分区表、持久消息流、供应商账单对账 | | 计费 | 版本价格、预付余额、冻结/结算、不可变流水、Stripe 手动/自动充值、退款/争议/对账 | 信用额度、合同价、Metronome 企业合同 | | 限流 | Redis Lua 全局 RPM/估算 TPM/并发,故障时本机降级;PG 月度消费配额 | 滑动窗口、层级策略、边缘 token bucket | -| 协议 | Chat Completions、Responses、Anthropic Messages 同 wire API 透传 | 规范化 IR + OpenAI/Anthropic/Google 双向转换 | +| 协议 | Chat Completions、Responses、Embeddings、Anthropic Messages 同 wire API 透传;计费单位已定义 token/image/second | Images/Audio 端点、规范化 IR + OpenAI/Anthropic/Google 双向转换 | ## 计费数据原则 @@ -57,7 +59,7 @@ flowchart LR Stripe 手动充值使用 Checkout Session:本地先创建 top-up order,Stripe 请求使用 order ID 作为幂等键;Webhook 验证签名后再次核对 event ID、order ID、session ID、金额和币种。自动充值使用 Checkout Setup Session 保存支付方式,低余额 worker 再创建 off-session PaymentIntent;同一订单的同步成功、Webhook 与对账共享幂等账本来源。浏览器成功跳转不具有入账权威性。 -流式请求的 usage 可能只在最后事件出现。当前 observer 会在线解析 Chat Completions、Responses 和 Anthropic SSE 中的 usage,并把 OpenAI details 中属于总 input 子集的缓存 token 拆成互斥计费桶;Anthropic 独立缓存字段不做减法。如果计费上游不返回 usage,请求会进入 `metering_failed` 并保持余额冻结,而不是按零费用结算。每种上游仍需使用供应商账单做日对账。 +流式请求的 usage 可能只在最后事件出现。当前 observer 会在线解析 Chat Completions、Responses、Embeddings 和 Anthropic SSE 中的 usage,并把 OpenAI details 中属于总 input 子集的缓存 token 拆成互斥计费桶;Anthropic 独立缓存字段不做减法。如果计费上游不返回 usage,请求会进入 `metering_failed` 并保持余额冻结,而不是按零费用结算。运营只能通过要求原因且受 `billing.adjust` 保护的接口释放无法恢复的冻结;事务同时写零金额审计流水并把事件标成 `released_unmetered`,不伪造 token 或扣款。每种上游仍需使用供应商账单做日对账。 ## 扩容方式 @@ -91,8 +93,8 @@ Origin,且不接受浏览器 credentials,只向页面暴露 `X-AIGW-Request- 供应商公开 slug、显示名、协议、近期可用率、响应头延迟、样本量和熔断恢复时间。Quickstart 和 Playground 默认使用自动路由,也可以把所选 slug 编入 `model:provider-slug` 来固定供应商。 `GET /v1/models` 和 Anthropic 模型列表同样只公开 slug 与 wire API,不返回内部 UUID、URL、 -凭证、上游模型名或权重。统计窗口是当前实例 -最近 100 次真实尝试;它不冒充主动健康检查、首 token 延迟、持久状态历史或全局 SLA。 +凭证、上游模型名或权重。统计窗口最多保留最近 100 次真实尝试与可选认证 `/models` +主动探测;开启 `AIGW_PROVIDER_SHARED_HISTORY_ENABLED` 后会合并 TTL 内其他实例的 route 结果和 TTFT,详情分别标出共享样本、主动探测次数和最近探测时间。Redis Stream 只保存短期、有界的选路信号,故障时退回本机窗口,因此它不冒充持久状态历史或全局 SLA。 Usage 页不保存 prompt 或响应正文。单请求详情仅把已持久化的身份边界、模型路由、协议、 重试、延迟、token、缓存和结算字段组成可复制诊断 JSON,因此既能支持工单排障,也不会 @@ -100,7 +102,7 @@ Usage 页不保存 prompt 或响应正文。单请求详情仅把已持久化的 `GET /admin/api/usage/analytics` 直接聚合 PostgreSQL Usage Ledger,并复用 Usage 页的租户、 项目、API Key、模型、状态和时间过滤。结果按模型与供应商返回请求量、成功率、token、 -缓存命中、费用、未收金额、缺失 usage 和 P95 延迟,同时用等长前一周期计算费用变化。 +缓存命中、费用、未收金额、缺失 usage、总延迟和 TTFT 的 P50/P95,同时按 API Key 返回同口径归因,并用等长前一周期计算费用变化。 它不在推理热路径执行,也不从浏览器当前加载的有限请求列表推算财务数据。 Quickstart 的 starter key 表单复用正式 `POST /admin/api/keys` 写入链路,为当前租户、所选 @@ -110,16 +112,18 @@ Quickstart 的 starter key 表单复用正式 `POST /admin/api/keys` 写入链è· `GET /admin/api/developer/config`,集中输出两个 SDK Base URL、四个推理/模型端点和不含 真实凭证的环境变量模板。 -密钥表单可设置模型白名单、月度金额上限、到期时间和标签。白名单与密钥摘要在同一个 +密钥表单可设置模型白名单、日/月金额上限、RPM/TPM、到期时间和标签。白名单与密钥摘要在同一个 PostgreSQL 事务内创建,任何未知模型都会让事务整体回滚。`last_used_at` 在 Usage/结算 事务内单调更新;过期时间既用于快照过滤,也在每次鉴权时检查,避免长轮询间隔延迟失效。 -密钥列表还通过 key/time 索引从 PostgreSQL 返回本月已结算费用、待结算冻结和请求数。 +密钥列表还通过 key/time 索引从 PostgreSQL 返回今日/本月已结算费用、待结算冻结和请求数。 +停用/启用会立即切换鉴权快照;轮换在单个事务中复制策略、创建新摘要并撤销旧密钥, +从而不存在两个密钥同时有效的窗口。 -`tenant_preferences` 保存租户默认模型、不同的 fallback 模型和低余额提醒阈值。 +`tenant_preferences` 保存租户默认模型、不同的 fallback 模型、低余额提醒阈值,以及异常消费提醒的启停、相对七日基线倍数和最低金额;异常策略为空时回退到 `admin.mail` 的部署默认值。 `GET /admin/api/developer/preferences` 返回当前租户值;两个独立的写接口分别要求 `developer.preferences.write` 与 `billing.preferences.write`,因此 developer 与 billing 角色 不能越权修改对方的设置。保存默认/fallback 时服务端会重新验证该模型当前对租户可见、 -未退役且至少存在一条启用路由。邮件扫描直接读取 PostgreSQL 中的租户阈值,不依赖 Redis。 +未退役且至少存在一条启用路由。邮件扫描直接读取 PostgreSQL 中的租户提醒策略,不依赖 Redis。 ## 管理面安全 @@ -133,4 +137,4 @@ bootstrap token 只映射为 `platform_admin`,用于首次建号和故障恢å¤ 2. 增加供应商日账单对账和合同价/信用额度。 3. 为不返回 usage 的上游增加可靠 token 计算器,并监控 `uncollected_micros`。 4. 把 UsageEvent 做时间分区和归档,增加定时导出与报告。 -5. 增加主动健康检查、TTFT/吞吐量采样、跨实例状态聚合和按成本/性能路由。 +5. 增加吞吐量/成本感知路由和客户可见的持久状态历史。 diff --git a/docs/commercial-readiness.md b/docs/commercial-readiness.md index 6dc9064..cbf6846 100644 --- a/docs/commercial-readiness.md +++ b/docs/commercial-readiness.md @@ -7,7 +7,7 @@ commercial feature. ## What works now -- OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages proxying, streaming, routing, +- OpenAI Chat Completions, OpenAI Responses, OpenAI Embeddings, and Anthropic Messages proxying, streaming, routing, retry, authentication, persistent usage, prepaid billing, quotas, rate limits, concurrent request limits, RBAC, audit logs, and a PostgreSQL-backed console. - Stripe-hosted manual top-up and payment-method setup, off-session automatic @@ -18,20 +18,21 @@ commercial feature. synchronize them to Stripe Customer with a stable idempotency key. Stripe failures preserve the local profile for retry; tax IDs stay in Stripe-hosted Checkout or Customer Portal rather than this database. -- PostgreSQL is the source of truth. Redis accelerates invalidation and shared - counters but is not required for startup, control-plane writes, billing, or - balance correctness. +- PostgreSQL is the source of truth. Redis accelerates invalidation, shared + counters, and short-lived route-health propagation but is not required for + startup, control-plane writes, billing, or balance correctness. - Verified self-service registration, invitation acceptance, password reset, persistent login throttles, per-device session revocation, encrypted email outbox delivery, TOTP with recovery codes, and WebAuthn Passkeys. - Tenant-scoped default/fallback model preferences and RBAC-separated low-balance - notification thresholds, persisted in PostgreSQL and applied by Quickstart and - the notification worker. -- Per-key model restrictions, monthly spend caps, expiration, tags, and last-use - tracking. Restrictions are enforced by the runtime snapshot and billing + and anomalous-spend notification settings, persisted in PostgreSQL and applied + by Quickstart and the notification worker. Unconfigured tenants inherit + deployment defaults. +- Per-key model restrictions, daily/monthly spend caps, RPM/TPM, expiration, + tags, prefix/suffix-only display, disable/enable, atomic rotation, and last-use tracking. Restrictions are enforced by the runtime snapshot and billing transaction, not only rendered by the console. The developer console also has - a page-memory API Playground and displays current-month settled spend, pending - reservations, request count, remaining cap, and last use for each key. + a page-memory API Playground and displays current-day/month settled spend, pending + reservations, request count, remaining caps, rate policy, and last use for each key. - The authenticated model catalog includes a customer-safe detail view, current price-version cost estimates for input/output/cache tokens, copyable model IDs, developer filtering, release/price/context sorting, one-click Playground @@ -39,8 +40,8 @@ commercial feature. authentication, balance, model, rate-limit, and provider errors. - The unauthenticated `/admin/models` catalog exposes only globally available models and supports search, protocol/input/developer filters, release/price/context - sorting, model details, versioned token prices, aggregate route availability, - and a pre-registration cost estimate. Tenant/key allowlists, upstream model IDs, + sorting, server-rendered per-model canonical URLs with protocol code examples, + versioned token prices, and aggregate route availability. Tenant/key allowlists, upstream model IDs, provider IDs, URLs, and routing weights are excluded by a dedicated public type. - Quickstart colocates balance state, direct starter-key creation, OpenAI and Anthropic SDK base URLs, copyable REST endpoints, environment configuration, @@ -49,13 +50,21 @@ commercial feature. - Usage events open into a privacy-safe request diagnostic with the complete request ID, route, retry, protocol, latency, throughput, cache-token, and settlement fields; copied JSON excludes prompts, responses, and secrets. -- Ledger-backed model cost ranking and provider performance views share the +- Ledger-backed model, API-key, and provider performance views share the Usage filters and report period-over-period charge change, success rate, - cache hit, P95 latency, and missing-usage exposure without sampling browser data. -- Runtime routes use a per-instance 100-attempt availability window, response-header - latency EWMA, and a 3-failure/30-second circuit breaker. The customer model detail - compares safe provider runtime fields and prevents launching a model while every - route is cooling down. + cache hit, total-latency/TTFT P50 and P95, and missing-usage exposure without + sampling browser data. The request list uses stable cursor pagination; tenant + results redact provider IDs, names, and upstream model identifiers server-side. +- Runtime routes use a 100-attempt availability window, response-header latency + EWMA, and a 3-failure/30-second circuit breaker. An optional bounded Redis Stream + shares recent outcomes and TTFT between instances and replays them after restart; + its non-blocking publisher falls back to the local window during Redis failure. + The customer model detail identifies shared samples, compares safe provider + runtime fields, and prevents launching a model while every route is cooling down. +- Optional active provider probes make one authenticated, non-inference `/models` + request per provider, validate the JSON API response, feed the same circuit breaker, + expose probe counters/timestamps in the console and Prometheus, and remain disabled + unless `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED=true` is explicitly injected. - Developers can keep automatic failover or pin a request with `model-id:provider-slug`. Provider slugs are stable public identifiers returned by the safe model catalog and selectable in Quickstart and Playground; pinned @@ -73,7 +82,8 @@ commercial feature. - Explicit tax treatment after registrations are confirmed still needs legal and product sign-off. Invoice details, hosted invoice/PDF/receipt links, payment history, refunds, disputes, reconciliation, and CSV ledger export are - implemented. + implemented. Usage debit rows link to their exact request details, and the + lookup remains tenant-scoped even when a caller knows another request ID. - Operational separation of the public inference listener from the management listener, HTTPS-only cookies behind a trusted proxy, backup/restore drills, migration rollback policy, secret rotation, and alerting for usage settlement @@ -84,13 +94,13 @@ commercial feature. ### P1 for ZenMux-like breadth -- OpenAI Embeddings, Images, Speech and Transcriptions; Gemini native - APIs; rerank and other media endpoints. The existing protocol field does not - make these APIs implemented. -- Active provider probes, first-token latency, throughput-aware selection, - cross-instance health aggregation, and customer-visible status history. Runtime - circuit breaking and request-derived route health are implemented; historical - success, total latency, cache hit, missing usage, and cost come from the Usage Ledger. +- OpenAI Images, Speech and Transcriptions; Gemini native APIs; rerank and other + media endpoints. Token/image/second metering units exist, but a typed unit does + not make those endpoints implemented. +- Throughput/cost-aware selection and customer-visible status history. Runtime + TTFT/availability-aware selection, bounded exploration, circuit breaking, + cross-instance short-lived health aggregation, and request-derived route health are implemented; historical + success, total latency, TTFT, cache hit, missing usage, and cost come from the Usage Ledger. - Provider price ranges, non-token search/image/audio pricing units, and richer deprecation notices. Provider runtime comparison and release sorting are implemented. - Usage exports, scheduled reports, organization invites, custom roles, @@ -117,11 +127,21 @@ with mode `0600`. | Stripe result URLs | `AIGW_STRIPE_SUCCESS_URL`, `AIGW_STRIPE_CANCEL_URL`, `AIGW_STRIPE_PORTAL_RETURN_URL` | | Console public URL | `AIGW_PUBLIC_URL` | | Public inference/API URL used by customer examples | `AIGW_INFERENCE_PUBLIC_URL` | +| Optional authenticated provider probes | `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED` | +| Optional cross-instance provider health | `AIGW_PROVIDER_SHARED_HISTORY_ENABLED` | | SMTP endpoint/sender | `AIGW_SMTP_ADDRESS`, `AIGW_SMTP_FROM_ADDRESS` | | SMTP credentials | `AIGW_SMTP_USERNAME`, `AIGW_SMTP_PASSWORD` | | WebAuthn RP/origins | `AIGW_WEBAUTHN_RP_ID`, `AIGW_WEBAUTHN_ORIGINS` | | Static upstream endpoint/key | provider `base_url_env`, `api_key_env` | +Before enabling Stripe in a deployment, run `go run ./cmd/stripe-preflight` with +the test restricted key. It performs only authenticated list requests, pins the +SDK API version, refuses live-mode keys, and returns per-resource permission +results without printing the key or Stripe response messages. Checkout, Portal, +automatic top-up, refund, Webhook, and reconciliation write permissions still +require the full test-mode workflow because the preflight intentionally creates +no Stripe objects. + External service values are never embedded in versioned JSON. When mail is enabled, startup validates the sender and SMTP endpoint; username/password must either both be provided or both be empty. Production should use authenticated diff --git a/docs/runbook.md b/docs/runbook.md index ddd1a8c..6062e0c 100644 --- a/docs/runbook.md +++ b/docs/runbook.md @@ -44,7 +44,9 @@ deduplicated per recipient and UTC day. ## Database migrations Run `cmd/migrate` before deploying application instances and keep `auto_migrate=false` in -production. Migrations use a PostgreSQL advisory lock and a recorded checksum. Schema rollback +production. Migrations use one stable PostgreSQL advisory lock across all versions and a recorded +checksum. An already-applied matching version exits without replaying DDL, so concurrent process +starts do not contend with normal control-plane queries. Schema rollback is always a reviewed forward migration; restore a database backup only for whole-release rollback after stopping writers. Never edit an already applied migration body. @@ -55,7 +57,10 @@ database credential. Test `scripts/restore-drill.sh` into a disposable isolated least monthly. Record row-count evidence and application smoke tests before deleting the drill. Before each release, run `scripts/load-smoke.sh` against a non-production upstream, then -`scripts/redis-fault-drill.sh` to prove Redis is optional and PostgreSQL polling keeps readiness. +`scripts/redis-fault-drill.sh` to prove Redis is optional, PostgreSQL polling keeps readiness, +and the subscriber reconnects after Redis returns. Set `AIGW_READY_URL`, `AIGW_REDIS_CONTAINER`, +and, when shared provider health is enabled, `AIGW_METRICS_URL`; the metrics assertion verifies +`aigw_provider_health_shared_connected` transitions from 1 to 0 and back to 1. Exercise PostgreSQL failover separately and verify pending settlement jobs resume without duplicate ledger entries. Archive the command output with the release evidence. |
