summaryrefslogtreecommitdiff
path: root/docs
diff options
context:
space:
mode:
authorChia <Chia@93.nz>2026-08-06 15:58:57 +1200
committerChia <Chia@93.nz>2026-08-06 15:58:57 +1200
commit3f702084d20b3c3a3ea916f3110e99b22bda60b3 (patch)
tree517f76c51025ce1ee085ea4898c60f799e5c37ea /docs
parent41e322c53d7b4b796eb377d0df9c29ecd10ba431 (diff)
feat: complete commercial developer workflowspublish-commercial-control-plane
Add tenant-safe usage observability, prepaid billing controls, API key lifecycle management, Embeddings metering, configurable billing alerts, and resilient provider health propagation. Harden Stripe failure handling, migrations, readiness, and the authenticated control-plane UI with end-to-end verification evidence.
Diffstat (limited to 'docs')
-rw-r--r--docs/architecture.md44
-rw-r--r--docs/commercial-readiness.md72
-rw-r--r--docs/runbook.md9
3 files changed, 77 insertions, 48 deletions
diff --git a/docs/architecture.md b/docs/architecture.md
index f8e617f..a88cb5e 100644
--- a/docs/architecture.md
+++ b/docs/architecture.md
@@ -22,6 +22,8 @@ flowchart LR
W --> L
D["PostgreSQL 控制面"] -. "事实数据" .-> S["Redis generation / PubSub"]
S -. "变更广播" .-> A
+ P -. "非阻塞健康事件" .-> H["Redis route-health Stream"]
+ H -. "跨实例回放/消费" .-> R
D -. "启动与轮询快照" .-> A
D -. "模型和路由快照" .-> R
```
@@ -29,27 +31,27 @@ flowchart LR
## 热路径
1. 入口生成不可预测的请求 ID,并通过 Bearer 或 `x-api-key` 解析客户身份。
-2. 身份包含 `key_id`、`tenant_id`、`project_id`、scopes、密钥模型白名单、月度上限和过期时间;控制面模式使用 PostgreSQL 摘要快照,代理只依赖 `Authenticator` 接口。
+2. 身份包含 `key_id`、`tenant_id`、`project_id`、scopes、密钥模型白名单、日/月上限、RPM/TPM 和过期时间;控制面模式使用 PostgreSQL 摘要快照,代理只依赖 `Authenticator` 接口。
3. 请求体在配置上限内读取一次,以提取公开模型名并支持故障转移时重放。
-4. 项目策略从 PG 热更新快照读取。Redis Lua 原子占用 RPM、估算 TPM 和并发额度;Redis 不可用时退回本机窗口。
-5. Router 按协议过滤路由并跳过仍处于冷却期的 route,再按优先级分组、在同级内按权重选择首选上游。请求使用 `model:provider-slug` 时,基础模型先经过租户/API Key 白名单校验,然后只保留该公开 slug 的兼容 route,绝不跨供应商回退;全部匹配 route 熔断时直接返回可重试的 503。
-6. 计费开启时,请求进入上游前在锁定钱包的 PostgreSQL 事务中同时检查项目月度额度、API Key 月度额度和未结冻结,再冻结保守估算额度;余额不足返回 402,任一月度额度耗尽返回 429。
-7. Provider Adapter 重写上游模型名和凭证,使用进程级共享 Transport 发送请求。每次上游尝试记录响应头延迟与可重试失败;连续 3 次连接失败或 429/502/503/504 后熔断该模型/供应商 route 30 秒。上游返回成功头后即锁定路由;SSE 逐块 flush。
-8. 请求结束后释放并发 lease,并用 `request_id` 幂等持久化 UsageEvent、更新月度汇总、按实际 token 结算;结构化日志仅是异步副本。
+4. 项目和 API Key 策略从 PG 热更新快照读取。Redis Lua 在一次调用中原子占用项目/Key RPM、估算 TPM 和项目并发额度;任一层拒绝会回滚本次全部计数。Redis 不可用时退回本机窗口。
+5. Router 按协议过滤路由并跳过仍处于冷却期的 route,再按优先级分组。同级 route 冷启动时按配置权重选择;积累 5 个可用性样本或 3 个 TTFT 样本后优先近期成功率和 TTFT 更好的 route,并每 20 次保留一次原权重探索。开启共享历史时,各实例从有界 Redis Stream 回放 TTL 内样本并实时消费;写入走非阻塞队列,Redis 不可用时继续使用本机窗口。请求使用 `model:provider-slug` 时,基础模型先经过租户/API Key 白名单校验,然后只保留该公开 slug 的兼容 route,绝不跨供应商回退;全部匹配 route 熔断时直接返回可重试的 503。
+6. 计费开启时,请求进入上游前在锁定钱包的 PostgreSQL 事务中同时检查项目月度额度、API Key 日/月额度和未结冻结,再冻结保守估算额度;余额不足返回 402,任一额度耗尽返回 429。
+7. Provider Adapter 重写上游模型名和凭证,使用进程级共享 Transport 发送请求。每次上游尝试记录响应头延迟与可重试失败,响应转发再记录最终实际 route 的首个有效输出 TTFT;连续 3 次连接失败或 429/502/503/504 后熔断该模型/供应商 route 30 秒。可选主动探测器按供应商去重调用认证 `/models`,拒绝 HTML/非标准 JSON,并把无客户流量时的可用性写入同一熔断器;它默认关闭且不发推理请求。上游返回成功头后即锁定路由;SSE 逐块 flush。
+8. 请求结束后释放并发 lease,并用 `request_id` 幂等持久化 UsageEvent(含首个有效输出 TTFT)、更新月度汇总、按实际 token 结算;结构化日志仅是异步副本。
## 已固定的扩展边界
| 边界 | 当前实现 | 下一阶段替换 |
| --- | --- | --- |
-| 客户身份 | PostgreSQL 快照、内存 SHA-256 索引、密钥过期/模型白名单/月度上限 | IP/CIDR 策略、短期服务身份 |
+| 客户身份 | PostgreSQL 快照、内存 SHA-256 索引、一次性明文与前缀/六字符后缀显示、密钥过期/模型白名单/日月上限/RPM/TPM、停用与原子轮换 | IP/CIDR 策略、短期服务身份 |
| 控制台权限 | 邮箱验证/邀请/重置、登录限流、设备会话、TOTP/恢复码、Passkey、CSRF、六角色 RBAC、租户 SQL scope、审计日志 | SSO/OIDC、SCIM、组织级 MFA 策略、自定义角色、审批流 |
| 权限 | `Principal.Scopes` 中的 `inference`、API Key 模型限制 + 控制台 RBAC | ABAC、IP 与条件策略 |
-| 模型目录 | PostgreSQL 快照 + 可选 Redis generation 广播 + PG 轮询兜底;租户安全目录 API、供应商运行状态与 Quickstart | 公开 SEO 目录、状态历史、灰度发布 |
-| 路由 | priority + weighted selection + failover + `model:provider-slug` 固定供应商 + 真实流量滑动窗口 + 响应头延迟 EWMA + 熔断 | 主动探测、TTFT/吞吐量、成本/质量策略、跨实例健康聚合 |
-| 计量 | PostgreSQL UsageEvent + 月度汇总 + 异步日志副本 | 分区表、持久消息流、供应商账单对账 |
+| 模型目录 | PostgreSQL 快照 + 可选 Redis generation 广播 + PG 轮询兜底;租户安全目录 API、服务端渲染公开 canonical 详情页、供应商运行状态与 Quickstart | 状态历史、灰度发布 |
+| 路由 | priority + weighted selection + failover + `model:provider-slug` 固定供应商 + 真实流量/可选主动探测滑动窗口 + Redis Stream 跨实例健康聚合/重启回放 + TTFT/可用率自适应择优 + 有界探索 + 熔断 | 吞吐量/成本/质量策略、客户可见状态历史 |
+| 计量 | PostgreSQL UsageEvent(总延迟/TTFT)+ 月度汇总 + 游标分页 + 模型/Key/供应商分位数聚合 + 异步日志副本 | 分区表、持久消息流、供应商账单对账 |
| 计费 | 版本价格、预付余额、冻结/结算、不可变流水、Stripe 手动/自动充值、退款/争议/对账 | 信用额度、合同价、Metronome 企业合同 |
| 限流 | Redis Lua 全局 RPM/估算 TPM/并发,故障时本机降级;PG 月度消费配额 | 滑动窗口、层级策略、边缘 token bucket |
-| 协议 | Chat Completions、Responses、Anthropic Messages 同 wire API 透传 | 规范化 IR + OpenAI/Anthropic/Google 双向转换 |
+| 协议 | Chat Completions、Responses、Embeddings、Anthropic Messages 同 wire API 透传;计费单位已定义 token/image/second | Images/Audio 端点、规范化 IR + OpenAI/Anthropic/Google 双向转换 |
## 计费数据原则
@@ -57,7 +59,7 @@ flowchart LR
Stripe 手动充值使用 Checkout Session:本地先创建 top-up order,Stripe 请求使用 order ID 作为幂等键;Webhook 验证签名后再次核对 event ID、order ID、session ID、金额和币种。自动充值使用 Checkout Setup Session 保存支付方式,低余额 worker 再创建 off-session PaymentIntent;同一订单的同步成功、Webhook 与对账共享幂等账本来源。浏览器成功跳转不具有入账权威性。
-流式请求的 usage 可能只在最后事件出现。当前 observer 会在线解析 Chat Completions、Responses 和 Anthropic SSE 中的 usage,并把 OpenAI details 中属于总 input 子集的缓存 token 拆成互斥计费桶;Anthropic 独立缓存字段不做减法。如果计费上游不返回 usage,请求会进入 `metering_failed` 并保持余额冻结,而不是按零费用结算。每种上游仍需使用供应商账单做日对账。
+流式请求的 usage 可能只在最后事件出现。当前 observer 会在线解析 Chat Completions、Responses、Embeddings 和 Anthropic SSE 中的 usage,并把 OpenAI details 中属于总 input 子集的缓存 token 拆成互斥计费桶;Anthropic 独立缓存字段不做减法。如果计费上游不返回 usage,请求会进入 `metering_failed` 并保持余额冻结,而不是按零费用结算。运营只能通过要求原因且受 `billing.adjust` 保护的接口释放无法恢复的冻结;事务同时写零金额审计流水并把事件标成 `released_unmetered`,不伪造 token 或扣款。每种上游仍需使用供应商账单做日对账。
## 扩容方式
@@ -91,8 +93,8 @@ Origin,且不接受浏览器 credentials,只向页面暴露 `X-AIGW-Request-
供应商公开 slug、显示名、协议、近期可用率、响应头延迟、样本量和熔断恢复时间。Quickstart
和 Playground 默认使用自动路由,也可以把所选 slug 编入 `model:provider-slug` 来固定供应商。
`GET /v1/models` 和 Anthropic 模型列表同样只公开 slug 与 wire API,不返回内部 UUID、URL、
-凭证、上游模型名或权重。统计窗口是当前实例
-最近 100 次真实尝试;它不冒充主动健康检查、首 token 延迟、持久状态历史或全局 SLA。
+凭证、上游模型名或权重。统计窗口最多保留最近 100 次真实尝试与可选认证 `/models`
+主动探测;开启 `AIGW_PROVIDER_SHARED_HISTORY_ENABLED` 后会合并 TTL 内其他实例的 route 结果和 TTFT,详情分别标出共享样本、主动探测次数和最近探测时间。Redis Stream 只保存短期、有界的选路信号,故障时退回本机窗口,因此它不冒充持久状态历史或全局 SLA。
Usage 页不保存 prompt 或响应正文。单请求详情仅把已持久化的身份边界、模型路由、协议、
重试、延迟、token、缓存和结算字段组成可复制诊断 JSON,因此既能支持工单排障,也不会
@@ -100,7 +102,7 @@ Usage 页不保存 prompt 或响应正文。单请求详情仅把已持久化的
`GET /admin/api/usage/analytics` 直接聚合 PostgreSQL Usage Ledger,并复用 Usage 页的租户、
项目、API Key、模型、状态和时间过滤。结果按模型与供应商返回请求量、成功率、token、
-缓存命中、费用、未收金额、缺失 usage 和 P95 延迟,同时用等长前一周期计算费用变化。
+缓存命中、费用、未收金额、缺失 usage、总延迟和 TTFT 的 P50/P95,同时按 API Key 返回同口径归因,并用等长前一周期计算费用变化。
它不在推理热路径执行,也不从浏览器当前加载的有限请求列表推算财务数据。
Quickstart 的 starter key 表单复用正式 `POST /admin/api/keys` 写入链路,为当前租户、所选
@@ -110,16 +112,18 @@ Quickstart 的 starter key 表单复用正式 `POST /admin/api/keys` 写入链è·
`GET /admin/api/developer/config`,集中输出两个 SDK Base URL、四个推理/模型端点和不含
真实凭证的环境变量模板。
-密钥表单可设置模型白名单、月度金额上限、到期时间和标签。白名单与密钥摘要在同一个
+密钥表单可设置模型白名单、日/月金额上限、RPM/TPM、到期时间和标签。白名单与密钥摘要在同一个
PostgreSQL 事务内创建,任何未知模型都会让事务整体回滚。`last_used_at` 在 Usage/结算
事务内单调更新;过期时间既用于快照过滤,也在每次鉴权时检查,避免长轮询间隔延迟失效。
-密钥列表还通过 key/time 索引从 PostgreSQL 返回本月已结算费用、待结算冻结和请求数。
+密钥列表还通过 key/time 索引从 PostgreSQL 返回今日/本月已结算费用、待结算冻结和请求数。
+停用/启用会立即切换鉴权快照;轮换在单个事务中复制策略、创建新摘要并撤销旧密钥,
+从而不存在两个密钥同时有效的窗口。
-`tenant_preferences` 保存租户默认模型、不同的 fallback 模型和低余额提醒阈值。
+`tenant_preferences` 保存租户默认模型、不同的 fallback 模型、低余额提醒阈值,以及异常消费提醒的启停、相对七日基线倍数和最低金额;异常策略为空时回退到 `admin.mail` 的部署默认值。
`GET /admin/api/developer/preferences` 返回当前租户值;两个独立的写接口分别要求
`developer.preferences.write` 与 `billing.preferences.write`,因此 developer 与 billing 角色
不能越权修改对方的设置。保存默认/fallback 时服务端会重新验证该模型当前对租户可见、
-未退役且至少存在一条启用路由。邮件扫描直接读取 PostgreSQL 中的租户阈值,不依赖 Redis。
+未退役且至少存在一条启用路由。邮件扫描直接读取 PostgreSQL 中的租户提醒策略,不依赖 Redis。
## 管理面安全
@@ -133,4 +137,4 @@ bootstrap token 只映射为 `platform_admin`,用于首次建号和故障恢å¤
2. 增加供应商日账单对账和合同价/信用额度。
3. 为不返回 usage 的上游增加可靠 token 计算器,并监控 `uncollected_micros`。
4. 把 UsageEvent 做时间分区和归档,增加定时导出与报告。
-5. 增加主动健康检查、TTFT/吞吐量采样、跨实例状态聚合和按成本/性能路由。
+5. 增加吞吐量/成本感知路由和客户可见的持久状态历史。
diff --git a/docs/commercial-readiness.md b/docs/commercial-readiness.md
index 6dc9064..cbf6846 100644
--- a/docs/commercial-readiness.md
+++ b/docs/commercial-readiness.md
@@ -7,7 +7,7 @@ commercial feature.
## What works now
-- OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages proxying, streaming, routing,
+- OpenAI Chat Completions, OpenAI Responses, OpenAI Embeddings, and Anthropic Messages proxying, streaming, routing,
retry, authentication, persistent usage, prepaid billing, quotas, rate limits,
concurrent request limits, RBAC, audit logs, and a PostgreSQL-backed console.
- Stripe-hosted manual top-up and payment-method setup, off-session automatic
@@ -18,20 +18,21 @@ commercial feature.
synchronize them to Stripe Customer with a stable idempotency key. Stripe
failures preserve the local profile for retry; tax IDs stay in Stripe-hosted
Checkout or Customer Portal rather than this database.
-- PostgreSQL is the source of truth. Redis accelerates invalidation and shared
- counters but is not required for startup, control-plane writes, billing, or
- balance correctness.
+- PostgreSQL is the source of truth. Redis accelerates invalidation, shared
+ counters, and short-lived route-health propagation but is not required for
+ startup, control-plane writes, billing, or balance correctness.
- Verified self-service registration, invitation acceptance, password reset,
persistent login throttles, per-device session revocation, encrypted email
outbox delivery, TOTP with recovery codes, and WebAuthn Passkeys.
- Tenant-scoped default/fallback model preferences and RBAC-separated low-balance
- notification thresholds, persisted in PostgreSQL and applied by Quickstart and
- the notification worker.
-- Per-key model restrictions, monthly spend caps, expiration, tags, and last-use
- tracking. Restrictions are enforced by the runtime snapshot and billing
+ and anomalous-spend notification settings, persisted in PostgreSQL and applied
+ by Quickstart and the notification worker. Unconfigured tenants inherit
+ deployment defaults.
+- Per-key model restrictions, daily/monthly spend caps, RPM/TPM, expiration,
+ tags, prefix/suffix-only display, disable/enable, atomic rotation, and last-use tracking. Restrictions are enforced by the runtime snapshot and billing
transaction, not only rendered by the console. The developer console also has
- a page-memory API Playground and displays current-month settled spend, pending
- reservations, request count, remaining cap, and last use for each key.
+ a page-memory API Playground and displays current-day/month settled spend, pending
+ reservations, request count, remaining caps, rate policy, and last use for each key.
- The authenticated model catalog includes a customer-safe detail view, current
price-version cost estimates for input/output/cache tokens, copyable model IDs,
developer filtering, release/price/context sorting, one-click Playground
@@ -39,8 +40,8 @@ commercial feature.
authentication, balance, model, rate-limit, and provider errors.
- The unauthenticated `/admin/models` catalog exposes only globally available
models and supports search, protocol/input/developer filters, release/price/context
- sorting, model details, versioned token prices, aggregate route availability,
- and a pre-registration cost estimate. Tenant/key allowlists, upstream model IDs,
+ sorting, server-rendered per-model canonical URLs with protocol code examples,
+ versioned token prices, and aggregate route availability. Tenant/key allowlists, upstream model IDs,
provider IDs, URLs, and routing weights are excluded by a dedicated public type.
- Quickstart colocates balance state, direct starter-key creation, OpenAI and
Anthropic SDK base URLs, copyable REST endpoints, environment configuration,
@@ -49,13 +50,21 @@ commercial feature.
- Usage events open into a privacy-safe request diagnostic with the complete
request ID, route, retry, protocol, latency, throughput, cache-token, and
settlement fields; copied JSON excludes prompts, responses, and secrets.
-- Ledger-backed model cost ranking and provider performance views share the
+- Ledger-backed model, API-key, and provider performance views share the
Usage filters and report period-over-period charge change, success rate,
- cache hit, P95 latency, and missing-usage exposure without sampling browser data.
-- Runtime routes use a per-instance 100-attempt availability window, response-header
- latency EWMA, and a 3-failure/30-second circuit breaker. The customer model detail
- compares safe provider runtime fields and prevents launching a model while every
- route is cooling down.
+ cache hit, total-latency/TTFT P50 and P95, and missing-usage exposure without
+ sampling browser data. The request list uses stable cursor pagination; tenant
+ results redact provider IDs, names, and upstream model identifiers server-side.
+- Runtime routes use a 100-attempt availability window, response-header latency
+ EWMA, and a 3-failure/30-second circuit breaker. An optional bounded Redis Stream
+ shares recent outcomes and TTFT between instances and replays them after restart;
+ its non-blocking publisher falls back to the local window during Redis failure.
+ The customer model detail identifies shared samples, compares safe provider
+ runtime fields, and prevents launching a model while every route is cooling down.
+- Optional active provider probes make one authenticated, non-inference `/models`
+ request per provider, validate the JSON API response, feed the same circuit breaker,
+ expose probe counters/timestamps in the console and Prometheus, and remain disabled
+ unless `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED=true` is explicitly injected.
- Developers can keep automatic failover or pin a request with
`model-id:provider-slug`. Provider slugs are stable public identifiers returned by
the safe model catalog and selectable in Quickstart and Playground; pinned
@@ -73,7 +82,8 @@ commercial feature.
- Explicit tax treatment after registrations are confirmed still needs legal
and product sign-off. Invoice details, hosted invoice/PDF/receipt links,
payment history, refunds, disputes, reconciliation, and CSV ledger export are
- implemented.
+ implemented. Usage debit rows link to their exact request details, and the
+ lookup remains tenant-scoped even when a caller knows another request ID.
- Operational separation of the public inference listener from the management
listener, HTTPS-only cookies behind a trusted proxy, backup/restore drills,
migration rollback policy, secret rotation, and alerting for usage settlement
@@ -84,13 +94,13 @@ commercial feature.
### P1 for ZenMux-like breadth
-- OpenAI Embeddings, Images, Speech and Transcriptions; Gemini native
- APIs; rerank and other media endpoints. The existing protocol field does not
- make these APIs implemented.
-- Active provider probes, first-token latency, throughput-aware selection,
- cross-instance health aggregation, and customer-visible status history. Runtime
- circuit breaking and request-derived route health are implemented; historical
- success, total latency, cache hit, missing usage, and cost come from the Usage Ledger.
+- OpenAI Images, Speech and Transcriptions; Gemini native APIs; rerank and other
+ media endpoints. Token/image/second metering units exist, but a typed unit does
+ not make those endpoints implemented.
+- Throughput/cost-aware selection and customer-visible status history. Runtime
+ TTFT/availability-aware selection, bounded exploration, circuit breaking,
+ cross-instance short-lived health aggregation, and request-derived route health are implemented; historical
+ success, total latency, TTFT, cache hit, missing usage, and cost come from the Usage Ledger.
- Provider price ranges, non-token search/image/audio pricing units, and richer
deprecation notices. Provider runtime comparison and release sorting are implemented.
- Usage exports, scheduled reports, organization invites, custom roles,
@@ -117,11 +127,21 @@ with mode `0600`.
| Stripe result URLs | `AIGW_STRIPE_SUCCESS_URL`, `AIGW_STRIPE_CANCEL_URL`, `AIGW_STRIPE_PORTAL_RETURN_URL` |
| Console public URL | `AIGW_PUBLIC_URL` |
| Public inference/API URL used by customer examples | `AIGW_INFERENCE_PUBLIC_URL` |
+| Optional authenticated provider probes | `AIGW_PROVIDER_ACTIVE_PROBES_ENABLED` |
+| Optional cross-instance provider health | `AIGW_PROVIDER_SHARED_HISTORY_ENABLED` |
| SMTP endpoint/sender | `AIGW_SMTP_ADDRESS`, `AIGW_SMTP_FROM_ADDRESS` |
| SMTP credentials | `AIGW_SMTP_USERNAME`, `AIGW_SMTP_PASSWORD` |
| WebAuthn RP/origins | `AIGW_WEBAUTHN_RP_ID`, `AIGW_WEBAUTHN_ORIGINS` |
| Static upstream endpoint/key | provider `base_url_env`, `api_key_env` |
+Before enabling Stripe in a deployment, run `go run ./cmd/stripe-preflight` with
+the test restricted key. It performs only authenticated list requests, pins the
+SDK API version, refuses live-mode keys, and returns per-resource permission
+results without printing the key or Stripe response messages. Checkout, Portal,
+automatic top-up, refund, Webhook, and reconciliation write permissions still
+require the full test-mode workflow because the preflight intentionally creates
+no Stripe objects.
+
External service values are never embedded in versioned JSON. When mail is
enabled, startup validates the sender and SMTP endpoint; username/password must
either both be provided or both be empty. Production should use authenticated
diff --git a/docs/runbook.md b/docs/runbook.md
index ddd1a8c..6062e0c 100644
--- a/docs/runbook.md
+++ b/docs/runbook.md
@@ -44,7 +44,9 @@ deduplicated per recipient and UTC day.
## Database migrations
Run `cmd/migrate` before deploying application instances and keep `auto_migrate=false` in
-production. Migrations use a PostgreSQL advisory lock and a recorded checksum. Schema rollback
+production. Migrations use one stable PostgreSQL advisory lock across all versions and a recorded
+checksum. An already-applied matching version exits without replaying DDL, so concurrent process
+starts do not contend with normal control-plane queries. Schema rollback
is always a reviewed forward migration; restore a database backup only for whole-release
rollback after stopping writers. Never edit an already applied migration body.
@@ -55,7 +57,10 @@ database credential. Test `scripts/restore-drill.sh` into a disposable isolated
least monthly. Record row-count evidence and application smoke tests before deleting the drill.
Before each release, run `scripts/load-smoke.sh` against a non-production upstream, then
-`scripts/redis-fault-drill.sh` to prove Redis is optional and PostgreSQL polling keeps readiness.
+`scripts/redis-fault-drill.sh` to prove Redis is optional, PostgreSQL polling keeps readiness,
+and the subscriber reconnects after Redis returns. Set `AIGW_READY_URL`, `AIGW_REDIS_CONTAINER`,
+and, when shared provider health is enabled, `AIGW_METRICS_URL`; the metrics assertion verifies
+`aigw_provider_health_shared_connected` transitions from 1 to 0 and back to 1.
Exercise PostgreSQL failover separately and verify pending settlement jobs resume without duplicate
ledger entries. Archive the command output with the release evidence.