fix(agents): ChatGPT-listed models fail on the default Codex runtime …
fix(agents): ChatGPT-listed models fail on the default Codex runtime before any request (#168033) ## What Problem This Solves Fixes: a model that the signed-in ChatGPT account lists (and that the model picker shows as available) fails on the default Codex runtime with `No route-compatible authentication source is configured for openai.` (or `... requires an OpenAI API key profile`) before any request is sent, when OpenClaw has no static ChatGPT route for that model id. ## User Impact User impact: models your ChatGPT account offers, including newly launched ones that OpenClaw does not know yet, now run on the default Codex runtime with your ChatGPT sign-in, matching what the picker shows. Known models, API-key-only users, and the `openclaw` runtime keep their current behavior. Behavior note: a user signed in to ChatGPT **and** holding an OpenAI API key, choosing a model id that OpenClaw does not know and that the catalog lists on both the Platform and ChatGPT rows, now prefers the subscription. This matches the picker and the `openclaw` runtime. ## Why This Change Was Made The picker and the run read different catalogs. The picker reads the published catalog, which includes the account's live ChatGPT rows. A Codex-owned run planned auth from its native placeholder model, whose api/baseUrl describe only the Platform route, so the account-only route was never considered and auth planning failed closed. A native-owned (Codex) foreground run now takes its observed routes from the published catalog captured at Gateway admission, which is the catalog the picker read, and auth planning chooses among those routes. The OpenClaw runtime path is unchanged. The native placeholder model is still not materialized for native-owned runs. The only SDK type change is an optional `observedRoutes` input on `prepareAgentRuntimeAuth`. Known gaps (unchanged from today, not regressions): - Direct CLI runs with no Gateway have no admission-captured catalog and keep today's behavior. - A model that appears after admission fails closed with today's error until the next turn admits a fresh catalog. - Codex isolated/utility completions still plan from the first catalog entry's route; follow-up. Merge danger: low to moderate. The change is confined to native-owned foreground auth planning, which is the path every Codex turn takes. The controls below cover known models, dual-credential users, API-key-only users, and the `openclaw` runtime. ## Evidence **Regression (QA-lab runtime e2e, real Gateway child + fake Codex app-server + loopback ChatGPT/Platform catalog fixture):** `test/e2e/qa-lab/runtime/codex-account-model-route.product-proof.e2e.test.ts`. Profiles are seeded like a fresh sign-in (no explicit auth order). Each case refreshes `models.list`, asserts the model is available, patches the session model, sends a chat turn, and waits for `ok`. It then asserts that the app-server received `turn/start` for that model with the expected login, and that the route errors never appear in Gateway logs. | Case | `origin/main` | This PR | | --- | --- | --- | | ChatGPT-only: account-only id, Platform-only manifest id listed by the account, `gpt-6-sol` (default runtime) | ❌ `No route-compatible authentication source is configured for openai.` | ✅ Codex `turn/start`, login `chatgptAuthTokens` | | Control: ChatGPT + API key, `gpt-6-sol` | ✅ `chatgptAuthTokens` | ✅ `chatgptAuthTokens` | | Control: API key only, `gpt-6-sol` | ✅ `apiKey` | ✅ `apiKey` | Command: `node scripts/run-vitest.mjs --config test/vitest/vitest.e2e.config.ts test/e2e/qa-lab/runtime/codex-account-model-route.product-proof.e2e.test.ts`. It was rerun green after merging current `main`. The file has 3 tests and about 85 s of test time; wall clock was 100–560 s depending on host load. It runs in the e2e tier, not the unit lane. **Real account (device-code ChatGPT sign-in, isolated profile, loopback Gateway built from this branch):** | Model | Runtime | Harness | Status | | --- | --- | --- | --- | | `gpt-daybreak-blue-latest` | default | codex | ✅ ok (responseModel `gpt-daybreak-blue-latest`) | | `gpt-6-sol` (control) | default | codex | ✅ ok | | `gpt-6-sol` (control) | `openclaw` | openclaw | ✅ ok | | simulated new model (`gpt-5.6-terra` removed from OpenClaw's manifest, name rules and a localhost catalog mirror; the account still lists it, and the picker shows it available) | default | codex | ✅ ok (responseModel `gpt-5.6-terra`) | | `gpt-6-sol` on an API-key-only profile (control) | default | codex | ✅ ok | | *Second fresh sign-in, head `37a8a5d6028`:* | | | | | `gpt-daybreak-blue-latest` (same-profile control) | default | codex | ✅ ok | | `gpt-daybreak-blue-latest` | `openclaw` | openclaw | ✅ ok | | `gpt-6-sol` | `openclaw` | openclaw | ✅ ok | | `gpt-6.1-sol` | `openclaw` | openclaw | ✅ ok | | simulated new model `gpt-5.6-terra` (same throwaway build and catalog mirror; the picker shows it available) | `openclaw` | openclaw | ✅ ok (responseModel `gpt-5.6-terra`) | The earlier `main` baseline on the same kind of real account: `gpt-daybreak-blue-latest` on the default runtime failed with `No route-compatible authentication source is configured for openai.`, and the same simulated new model failed the same way, while the `openclaw` runtime served both. The simulation build was a local, never-pushed throwaway. **Inherited CI failure:** `test/e2e/qa-lab/runtime/quota-reset.e2e.test.ts` fails 8/9 in `checks-node-changed`, with `chat.history` returning `UNAVAILABLE`: "Session access facts are unavailable; retry after session storage is ready." The same 8 tests fail locally on `main` `e39552c1810`, and the same single case passes on both. This PR's preload change is opt-in (`platformCatalog`), and those tests do not set it. **Gates:** `node scripts/check-changed.mjs` passed with `OPENCLAW_CHECK_CHANGED_SKIP_DEADCODE=1`. Without that flag, the only failure is the dead-export scan on files this PR does not touch (`extensions/telegram/src/sequential-key.ts`, `src/auto-reply/reply/model-selection.ts`), inherited from `main`. `pnpm plugin-sdk:surface:check` and `pnpm plugin-sdk:check-exports` passed. `pnpm plugin-sdk:api:diff` against the merged `main` base shows 4 changed declarations (149 reachable exports), all the single optional `observedRoutes?: readonly ProviderModelRouteSource[]` input (`PrepareAgentRuntimeAuthPlanParams` and the route-resolution params that carry it). This is additive, with no removals. The touched unit suites (prepare-auth, route intent, model-setup ownership and selected model, auth-plan, capture, catalog view and decisions, openai-model-routes, credential-scoped-model memo, native-picker failures, model-resolution consistency) passed after the merge. Co-authored-by: Ayaan Zaidi
评论
?
参与讨论