The useful comparison unit
A coding task needs more than a model response. The application reads files, assembles context, offers tools, runs commands, observes output and decides when to stop or ask a person. That surrounding application is the harness. Changing its tools or permissions can change the result even when the model stays the same.
Codex CLI works in a local repository and supports a reviewable edit-and-run loop. Anthropic describes Claude Code's context gathering, action and verification loop. Pi presents a smaller terminal harness with an extension surface. These descriptions tell us what to inspect; they do not rank the products.
References: OpenAI Docs — Codex CLIAnthropic — How Claude Code worksPi coding agent — README at v0.85.1
Questions to ask of each setup
Record the model, product version, operating system, starting repository state and available tools. Then inspect who can read or write files, whether commands can reach the network, how approval works and whether tool results are retained. A permission prompt and an operating-system sandbox solve different problems.
The same product can behave differently in a local CLI, remote task or managed environment. Pi's minimal defaults and extension path also mean a configured installation may differ substantially from its base README. Compare the environment you will actually use, not a feature list assembled across versions.
References: OpenAI Docs — Agent approvals and securityAnthropic — How Claude Code worksPi coding agent — README at v0.85.1
A small, fair evaluation
Prepare several representative bugs or changes in disposable repositories. Write acceptance checks before running any agent. Give each setup the same task description, starting commit, time limit and allowed access. If you cannot use the same model in every harness, describe the result as a comparison of complete systems.
For each run, capture whether the acceptance checks pass, unrelated files changed, a person intervened, a tool failed and the final diff was understandable. Repeat tasks; one successful demonstration is weak evidence. Measure elapsed time and billed usage under the provider's actual accounting, and keep failures in the report.
What to choose for a team
Choose the setup that fits the team's operating controls and review habits. A strong model result is less useful if the environment cannot safely access the repository or produce a diff that maintainers can inspect. A small harness may be appropriate when the team owns the surrounding controls; a managed surface may fit when access, review and administration are more important.
Before adopting any tool broadly, document where credentials live, which paths it may change, how it handles unknown repository instructions and who approves external writes. Revisit that checklist when an extension, model or execution environment changes.
References: OpenAI Docs — Agent approvals and securityAnthropic — How Claude Code worksPi coding agent — README at v0.85.1