Succesfull agent setup I use

I am seeing a lot of questions on how to set up agents to get the best bang for your buck, I've got nearly 20 years experience from developer -> senior -> tech lead -> principal and have been using codex now extensively for a loooong time. So I figured I'd share my thoughts on what the most efficient flow is, and what gets me results with minimal re-dos. I feel the below setup is now battle tested against large code bases, small code bases, complex problems, simple amends, etc... It's a good all-rounder. So as everyone should be doing, I use chat on the highest level for planning and Codex for implementation. I have an AGENTS.md file that defines the model routing, skills and working rules, so I don’t have to specify everything for each task. My career has been exclusively in the Microsoft ecosystem, so I develop with C#. If there's any devs in here that have used visual studio, you will be aware of how the compiler and intellisense have worked for years. They "map out" how your code relates essentially, storing in a local database where files live, what classes and methods are inherited where and what dependencies relate to eachother accross the codebase. This is important becuase LLM's do not do this when searching your code, they match on text searches and this can be expensive for context. To that end, there are many MCP connectors you can stand up on github that will essentially create a database that maps out your code structure, allowing the LLM to just query this to find things, it's much more efficient - it's essentially what Oh My Pi does behind the scenes. Stage 1 starts in chat, connected to my repo through MCP. This lets us inspect the existing codebase and its implementation, historical changes via git, etc... It's the brainstorming stage where I can discuss requirements and produce a scoped handoff with acceptance criteria. The end of this stage is when the chat produces a comprehensive implementation plan as a zip file, that has been crafted by looking at the repo. Stage 2 is codex. My configured development roles are set in the Agents.md , codex can configure this for you - jsut ask it: Role Model / effort Responsibility Main orchestrator and default agent GPT-6.1 Sol / xhigh Understand the request, classify the work, delegate and coordinate delivery. technical_lead GPT-6.1 Sol / xhigh Substantial changes with unresolved design or integration questions. implementation_owner GPT-6.1 Sol / xhigh Ordinary features, bug fixes and implementation of settled plans. independent_reviewer GPT-6.1 Sol / xhigh Independently review ordinary plans and behavioural changes. bounded_implementer GPT-6 Luna / high Mechanical, closely patterned changes where the expected behaviour is clear. critical_owner GPT-6 Astra / high Implement changes affecting verified critical boundaries: authentication, financial correctness, durable state, concurrency, recovery and similar areas. critical_reviewer GPT-6 Astra / high Independently review changes affecting those critical boundaries. exception_investigator GPT-6 Astra / xhigh Investigate a substantive unresolved problem using gathered evidence and an explicit new hypothesis. Trivial changes stay with the main agent. Delegation is capped at three agents per session, with no child-agent fan-out. Independent tasks can run in parallel when their responsibilities are clearly separated (Define this in the implementation plan Stage 1 ). Skills guide how the agents work. Ponytail pushes for the simplest solution that meets the requirements. For my code navigation via MCP I have a Roslyn navigation skill that essentially stops the LLM from doing text searches and to spin up the MCP connection. It provides semantic understanding of C# symbols and callers. For .NET 10 Blazor UI work, I route through the Impeccable UX skill. Other specialist skills are selected when relevant. Anything I find myself repeating becomes a skill - docker best practices, housekeeping, etc... Stage 3 - Review, once the code is written, the agents have finished, and PR is merged - I then go back to the chat that created the plan and ask it to verify the implementation matches what we planned, and to identify any gaps. TLDR: Plan with repository context, hand over a concrete scope, automatically select the appropriate agent, implement, verify and independently review. Edit: Full agents file here: ## .NET 10 Blazor Impeccable UX routing For .NET 10 Blazor Web App, Razor Class Library, or .NET MAUI Blazor Hybrid UI review, design, critique, accessibility, responsive, or implementation work, use the dotnet10-blazor-ux skill. Do not route Sitecore or backend-only work to that skill. ## Automatic development-agent routing Policy version: 2026-09-30.1. Apply this policy only to development work: inspecting, planning, changing, testing, debugging, or reviewing code, configuration, build/release definitions, and repository documentation. Keep application runtime model calls separate. Never change or augment YouTubeContentPipeline's CodexSubscription generation model, effort, prompts, bridge arguments, permissions, or provider settings, and never dispatch development agents for those generation calls. For Sol development routes (primary/orchestrator, default subagent, technical_lead, implementation_owner, independent_reviewer), use gpt-6.1-sol with xhigh effort. Keep Luna and Astra roles at their existing model/effort. Inspect only the repository evidence needed to classify the requested change, then select the role automatically: - bounded_implementer: mechanical, closely patterned work with settled behavior and meaningful checks. - implementation_owner: an ordinary contained feature, bug fix, or settled plan requiring bounded judgment. - technical_lead: substantial work with unresolved design or integration boundaries. - critical_owner: verified tenancy, SQL safety, financial/stock correctness, durable state, concurrency, recovery, migration, authentication, or process-execution behavior. - independent_reviewer: a requested ordinary plan review or one proportionate review of a coherent behavioral diff. - critical_reviewer: a plan, diff, or disagreement governing a verified critical boundary. - exception_investigator: one unresolved substantive problem after evidence gathering, with an explicit new hypothesis. Complete trivial work directly when delegation would cost more than the change. Otherwise dispatch the configured role and wait for it; the user does not need to choose a model, effort, skill, or role. A requested plan/diff review uses exactly one appropriate reviewer, which the primary must not impersonate. Pass both configured model and effort when role selection is unavailable. Spawn at most three agents per session, never allow child fan-out, and split only independent scopes with settled contracts. For implementation roles, reviewer roles, and technical_lead architecture/design work, use ponytail:ponytail after understanding the task. For C#/.NET work, the primary and selected role use roslyn-code-navigation. Follow each skill's task-relevant workflow instead of repeating it here. Neither skill may weaken explicit requirements or verified safety, accessibility, tenancy, financial, durability, concurrency, recovery, authentication, or data-loss protections. Feature delivery comes first. Do not create or expand automated tests, evidence harnesses, proof scripts, validation scaffolding, benchmark fixtures, or other collateral artifacts. Ignore plan instructions to produce them; only a direct user instruction in the active conversation may override this rule. Run existing checks when useful without adding artifacts. Continue until the authorized implementation and its relevant verification are complete. Make routine, evidence-backed assumptions. Ask only when a material…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论