CodexFirst version of the paper bot written and saved to GitCodexLocal dashboard, price index feed, first strategies
The agent wrote code. The owner ran it, judged it and deployed it by hand.
This product was built by AI coding agents. It started on OpenAI's Codex until the monthly limit ran out, then moved to Claude Code. Switching tools mid-project created a natural experiment: one owner, one codebase and one change process, worked by two different agents. These seven chapters tell how it went.
The bot began as chat-assisted code. Codex wrote the first version and the dashboard, then on Oct 1 introduced the change process everything else was built on.
CodexFirst version of the paper bot written and saved to GitCodexLocal dashboard, price index feed, first strategiesThe agent wrote code. The owner ran it, judged it and deployed it by hand.
CR-0001Change records, a verify step and a read-only previewBranchesNamed per agent, so history shows who did whatCR-0015Profit Taking Protocol, the largest Codex featureEach change now left a written record. That record later let a different agent take over without losing context.
The CR process made Codex fast. It also made it hungry. In three days of heavy, screenshot-driven work, Codex used up its whole monthly allowance with four weeks still to go.
About 360M tokens went on Oct 1 alone, the same day computer-use calls peaked. Desktop screenshots are token-heavy.
Waiting a month for the reset would have stopped the project. On Oct 3 the work moved to Claude Code, which picked up from the written records mid-stream.
Same owner, same repository, same CR process and the same kind of work, done by two different agents. That rarely happens by design. Chapter 5 compares them.
A new agent with no memory of the project took over in one session. The Codex-era records made that possible.
HandoffClaude read the handoff notes and release history, then tidied upMemoryLessons saved between sessions so mistakes aren't repeatedCR-0035Private GitHub repository; every branch pushedEach costly mistake became a memory note, so later sessions skipped the same trap.
Over five days the agent's reach grew from the laptop to the code repository, the cloud, the network edge and sign-in. Every step came with tests, a written record and a narrow permission.
CR-0067Automatic checks on every change; untested code can't shipCR-0075–77Component tests, browser tests and a replay of every trading decisionCR-0078One-command release and a project guide for agentsThis moved the trust from the agent to the checks. A wrong change fails tests before it ships.
CloudIts own cloud identity, with deny rules on secrets and deletionsCR-0089–90First CRs prepared in a cloud sessionCR-0098Releases deploy from the repository, not a laptopThe owner did the account-level setup for the move. The agent did everything that could be scripted.
CR-0110Sign-in service configured from the command lineCR-0128Small changes batched, cutting build minutes about 70%CR-0130Four skills, two helper agents and a guard hook, kept in the repoRepeated explanations became files the agent loads when it needs them, and cheaper models took over the routine checks.
CR-0135Move to btcmm.app done through the edge API, bot untouchedSafetyThe safety check paused a production step until the owner said yesDocsThese pages drafted privately, then published on this domainAccess came as one narrow, expiring grant, and no secret was ever pasted into chat.
The work is measured three ways: per CR, per commit and per line of code. A category goes to the agent that wins at least two of its three measures.
Bars are lines changed on main (added plus deleted, without lockfiles, generated files or images). The line is tokens spent per 1,000 of those lines. Codex's first bucket combines its September sessions with the Oct 1 commit that captured that backlog. Oct 7 runs high because the sign-in rollout and the domain and server work used many tokens but changed few lines.
*Metered usage only. The Claude plan fee is not included.
How the windows were chosen. Speed, test investment and per-CR / per-commit density use Codex's CR era (Oct 1–3, 33 CRs, 91 commits, 10,142 lines), because Codex had no CRs before Oct 1. Token and cost ratios per commit and per line use Codex's full usage window (Sep 9 – Oct 3: 108 commits, 12,318 lines, 902.6M tokens). Per CR uses the ≈640M tokens of Oct 1–3, read from Codex's usage chart. Claude Code: Oct 3–7, 102 CRs, 252 commits, 22,120 lines, 1,290.8M tokens.
Codex's chart shows the same pattern. Long sessions mostly re-read their own context at a discounted rate.
Codex leaned on computer use and its Build Dashboard skill. Claude Code worked mostly from the command line.
Daily tokens for both agents, with the average context sent per model call on top. Numbered bubbles mark the changes that cut cost or token use.
Codex's busiest day was also its peak for computer-use calls. Screenshots and desktop actions are token-heavy, so 16 CRs cost about 22M tokens each.
From the first Claude session, 98.75% of tokens were cache reads: the conversation re-sent as cached context, billed at a fraction of fresh input.
98.75% cachedTool paths, encoding traps and the laptop clock issue went into memory notes. Later sessions read a few lines instead of re-discovering each trap with failed commands.
One line per step, the full log in a file, and only the last 25 lines on a failure. Test and build output stays out of the agent's context.
Logs on disk, not in contextChecks run on the laptop first, and records go into the same push, so each CR triggers one CI run instead of several.
−70% CI minutesDiff reviews and layout checks run in Sonnet helper agents with their own small context, not in the main Opus session.
19–47K vs ~474K per callA new session for each block of work keeps the re-sent context small. Each call is cheaper even when the day is busy.
504K → 381K per callThe btcmm.app move, two docs sites and these pages ran in a fresh session with the smallest context per call so far.
207K per call, −59%Why Oct 7 is still the tallest bar: it had the most CRs (30), including the sign-in rollout and a server incident, so tokens per CR rose to 16.1M that day. Per-call context still fell, and that's the number that drives cost. Codex's daily values are read from its usage chart, and its Sep 9–30 usage (about 260M tokens in total) is not drawn.
About twice the CRs per day, fewer tokens and dollars per unit of work, more than twice the test code per CR, and work across five platforms instead of one machine.
Codex's changes were bigger chunks of product code: about twice the source and test lines per CR and per commit, at fewer tokens per line. Its built-in computer use made desktop work quick.
The agents did different phases: Codex did early features, Claude did testing, cloud and security. Lines of code reward verbosity and miss work that produces none. Claude's figures leave out its cloud sessions.
Change records, release notes, a project guide and memory let a new agent or session pick up the work without starting over. The Codex-to-Claude handoff took one clean-up change.
A decision replay, coverage floors and checks before release catch an agent's mistakes mechanically. Approvals are kept for what tests can't see.
Each platform got its own narrow grant with deny rules or an expiry date. Secrets stayed out of chat and out of the code.
Claude Code ran into limits too: the Pro weekly cap, then most of the monthly extra usage. This time the answer was a bigger plan, not a different tool. The project moved to Claude Max, and the records, tests and skills built along the way meant nothing had to be re-learned.