HALLODE / NOTES
Choosing models and effort for coding agents
How I think about planning, implementation, and review without assigning every tool a fixed job.
I use Claude Code, Codex, and Cursor. I do not run every task through all three. More often, I choose one environment and decide how much model capability the task needs at each stage. A difficult plan deserves more thought than a routine edit; a risky change deserves a careful review.
The useful question is not simply which agent is best? It is what does this task need, and which model and effort setting make sense here?
A way to split model effort in Codex
For Codex, I would start by matching the model and reasoning effort to the job. OpenAI’s model-selection guide lays out the tradeoff between capability, cost, and latency. One possible split is:
- Quick exploration — GPT-6 Luna, low effort: find relevant files and facts.
- Complex planning — GPT-6 Astra, high effort: work through an unclear outcome before implementation.
- Scoped implementation — GPT-6 Sol, medium effort: make the change and run checks.
- Risky review — GPT-6 Astra, xhigh effort: challenge the diff and the evidence with fresh context.
This is a suggested starting point, not a fixed setup I use for every task. A small, clear change may need only one agent. I would spend more reasoning effort when an early mistake could shape the rest of the work, or when a review needs to catch something subtle.
What I have tried in Cursor
I have used Opus to plan in Cursor, then switched to Composer to implement. Opus access in Cursor has been a constraint for me, so I would rather spend it on the part that needs deeper reasoning than leave it selected for every edit.
Cursor’s Plan Mode lets me inspect and revise the approach before building it. Once the task is clear, Composer is a practical choice for implementation in the editor. For a small change, I can skip the separate planning step and start directly with Composer. These are choices within Cursor; they do not imply a handoff to Claude Code or Codex.
When Auto makes sense
I used to switch to Auto after I ran out of access to a model I had selected, and it let me keep working at the time. I also find it useful for a quick first pass when the cost of a wrong turn is small: locating a behavior, exploring an unfamiliar file, or trying a small change I can check immediately.
That past experience is not a promise of extra or unlimited usage today. Cursor’s current usage guide describes separate pools for Cursor models and third-party models. Auto can draw from either pool depending on the model it routes to; once included usage runs out, continuing may require on-demand usage or a plan change. The Spending dashboard shows which pool a request used.
For an ambiguous architectural decision, I would choose Opus deliberately and review the plan. If the plan is settled and I want a predictable implementation setup, I would select Composer. Auto is still an option for execution, but I would not write about that run as though it used one fixed worker model: Cursor Router may route each request differently.
Where Cursor exposes Auto’s optimization modes, Cost favors lower spend, Balance weighs quality, speed, and cost, and Intelligence favors stronger models for harder requests. Availability depends on the Cursor plan and team settings. None of these modes removes the need to inspect the diff and run the relevant checks.
A model split to try in Claude Code
Claude Code’s model configuration offers a useful starting point: opusplan uses Opus in Plan Mode and switches to Sonnet for implementation. That is a feature of Claude Code, not a claim that I use it for every task.
- Quick exploration — Haiku: a fast option for simple questions and finding my way around, when it is available. The current Claude Code docs do not list an effort control for Haiku.
- Complex planning — Opus, high effort: investigate tradeoffs and produce a plan I can challenge before building.
- Scoped implementation — Sonnet, medium effort: work through a clear task and verify the result.
- Risky review — Opus, high effort: inspect the original goal, diff, and evidence. I would consider xhigh only when the review genuinely needs deeper reasoning.
- Long, open-ended work — Fable, high effort: consider it when investigation, implementation, and verification may stretch beyond a single sitting.
Fable is positioned for Claude Code’s hardest and longest-running tasks, so it does not need to be the default planner or worker for routine edits. Availability varies, and on some plans it can use separate usage credits. This is a suggested split, not my fixed Claude Code configuration. Anthropic’s effort guide describes medium as a starting point for day-to-day engineering and high for work with edge cases or important verification. Supported levels and defaults depend on the model version and provider. I can also stay on one model for a small change; the roles do not require separate tools or sessions.
Decide from the work, then check the result
I give planning more effort when the goal is ambiguous, the change crosses systems, or several designs could work. I use a faster setting when the work is bounded and the expected behavior is already clear. I give review enough capability to challenge the implementation, especially when a mistake could affect data, security, or a long-lived API.
Model choice does not replace context. The agent still needs the goal, constraints, relevant code, and a way to verify its result. I inspect the diff and the checks it actually ran. If the output is weak, I first ask whether the task or context was underspecified before assuming a more expensive model will fix it.
My setup can change as models, limits, and projects change. The durable habit is to match the effort to the decision in front of me, then verify the work.
About these choices: The Codex and Claude Code splits are examples to adapt, while the Cursor pairing describes a setup I have used. Neither is a measured ranking of the products.