The story · BTC Market Monitor

One owner, two AI agents, 135 changes

This product was built by AI coding agents. It started on OpenAI's Codex until the monthly limit ran out, then moved to Claude Code. Switching tools mid-project created a natural experiment: one owner, one codebase and one change process, worked by two different agents. These seven chapters tell how it went.

Agents
OpenAI Codex → Claude Code
Handoff
2026-10-03 (CR-0034)
Models
Opus (main) · Sonnet (helpers)
Surfaces
Desktop app · cloud sessions
Scope
Code, tests, releases, operations
102of 135 CRs done by Claude Code since Oct 3
253commits co-authored by Claude
4–1head-to-head categories won, Claude Code against Codex
2plan limits hit: Codex's monthly limit, then Claude's weekly one
Chapter 1 · Sep 7 – Oct 2

Starting on Codex

The bot began as chat-assisted code. Codex wrote the first version and the dashboard, then on Oct 1 introduced the change process everything else was built on.

Sep 7–22pre-CR
Chat-assisted build
  • CodexFirst version of the paper bot written and saved to Git
  • CodexLocal dashboard, price index feed, first strategies

The agent wrote code. The owner ran it, judged it and deployed it by hand.

Oct 1–2CR-0001 – 0033
Codex with a process
  • CR-0001Change records, a verify step and a read-only preview
  • BranchesNamed per agent, so history shows who did what
  • CR-0015Profit Taking Protocol, the largest Codex feature

Each change now left a written record. That record later let a different agent take over without losing context.

Chapter 2 · Oct 1–2

Hitting the limit

The CR process made Codex fast. It also made it hungry. In three days of heavy, screenshot-driven work, Codex used up its whole monthly allowance with four weeks still to go.

The wall

Monthly allowance used up

Codex · monthly allowance99.6%
resets at the end of the month

About 360M tokens went on Oct 1 alone, the same day computer-use calls peaked. Desktop screenshots are token-heavy.

The switch

Move, don't wait

Waiting a month for the reset would have stopped the project. On Oct 3 the work moved to Claude Code, which picked up from the written records mid-stream.

The opportunity

A natural A/B test

Same owner, same repository, same CR process and the same kind of work, done by two different agents. That rarely happens by design. Chapter 5 compares them.

Chapter 3 · Oct 3

The handoff

A new agent with no memory of the project took over in one session. The Codex-era records made that possible.

Oct 3CR-0034 – 0040
Handoff to Claude Code
  • HandoffClaude read the handoff notes and release history, then tidied up
  • MemoryLessons saved between sessions so mistakes aren't repeated
  • CR-0035Private GitHub repository; every branch pushed

Each costly mistake became a memory note, so later sessions skipped the same trap.

Chapter 4 · Oct 4–8

Claude Code takes over

Over five days the agent's reach grew from the laptop to the code repository, the cloud, the network edge and sign-in. Every step came with tests, a written record and a narrow permission.

Oct 4–5CR-0041 – 0082
Agent checked by machines
  • CR-0067Automatic checks on every change; untested code can't ship
  • CR-0075–77Component tests, browser tests and a replay of every trading decision
  • CR-0078One-command release and a project guide for agents

This moved the trust from the agent to the checks. A wrong change fails tests before it ships.

Oct 6CR-0083 – 0104
Agent operates the cloud
  • CloudIts own cloud identity, with deny rules on secrets and deletions
  • CR-0089–90First CRs prepared in a cloud session
  • CR-0098Releases deploy from the repository, not a laptop

The owner did the account-level setup for the move. The agent did everything that could be scripted.

Oct 7CR-0105 – 0134
Agent tooling as code
  • CR-0110Sign-in service configured from the command line
  • CR-0128Small changes batched, cutting build minutes about 70%
  • CR-0130Four skills, two helper agents and a guard hook, kept in the repo

Repeated explanations became files the agent loads when it needs them, and cheaper models took over the routine checks.

Oct 8CR-0135
Agent at the edge
  • CR-0135Move to btcmm.app done through the edge API, bot untouched
  • SafetyThe safety check paused a production step until the owner said yes
  • DocsThese pages drafted privately, then published on this domain

Access came as one narrow, expiring grant, and no secret was ever pasted into chat.

How the agent reaches each system

Owner sets goals, approves risky steps, does account-level clicks Claude Code · Opus desktop app, plus cloud sessions Memory: lessons, owner preferences Project guide: the change process Skills: release a change, batch small fixes, health check, safe restart Hook: blocks keys in files, asks before risky commits Safety check on every command, pauses for the owner's OK Reviewer Sonnet · code review Layout checker Sonnet · 3 screen sizes Git tools can run pipelines, not change settings Cloud CLI operate, never delete or read secrets Edge API this domain only, expiring access Sign-in CLI settings only, no secret key Browser + test runner phone, tablet and desktop checks Artifacts private drafts before publishing Code repository tests · automatic deploys Cloud hosting server · monitoring · backups Network edge domain · protection · docs Sign-in service user accounts Dashboard previews and btcmm.app claude.ai drafts, then moved to docs.btcmm.app

Who does what

Agent alone

Build, check, release

  • Code, tests, change records and release notes
  • Local checks, then one push through the pipeline
  • Read-only health checks
  • Code review and layout checks through helper agents
Agent after an OK

Anything that touches production

  • Bot restarts, only between trading windows
  • Production settings and service restarts
  • Deleting anything, publishing or sending
Owner only

Identity, money and accounts

  • Creating accounts and access grants
  • Account-level dashboard steps
  • Email verifications and sign-in tests
  • Plans, billing and sharing
Chapter 5 · Head to head

Codex vs Claude Code

The work is measured three ways: per CR, per commit and per line of code. A category goes to the agent that wins at least two of its three measures.

0 5K 10K 0 60M 120M lines changed tokens / 1K lines handoff 9.7K 2.6K 3.5K 6.5K 2.6K 5.1K 4.4K Codex avg 73.3M Claude avg 58.4M 64M 107M 19M 58M 45M 43M 108M Sep 9 – Oct 1 Oct 2–3 Oct 3 Oct 4 Oct 5 Oct 6 Oct 7 Codex Claude Code

Bars are lines changed on main (added plus deleted, without lockfiles, generated files or images). The line is tokens spent per 1,000 of those lines. Codex's first bucket combines its September sessions with the Oct 1 commit that captured that backlog. Oct 7 runs high because the sign-in rollout and the domain and server work used many tokens but changed few lines.

Claude Codewins 3 of 3

Speed per active day · higher is better

CRs
11.0
20.4
Commits
30.3
50.4
Lines changed
3,381
4,424
Claude Codewins 3 of 3

Token efficiency tokens per unit · lower is better

Per CR
≈19.4M
12.7M
Per commit
8.4M
5.1M
Per 1K lines
73.3M
58.4M
Claude Code*wins 3 of 3

Cost efficiency metered cost per unit, Codex = 100% · lower is better

Per CR
100%
11%
Per commit
100%
15%
Per 1K lines
100%
19%

*Metered usage only. The Claude plan fee is not included.

Claude Codewins 3 of 3

Test investment test lines · higher is better

Per CR
22.7
51.2
Per commit
8.2
20.7
Per 100 source lines
10
76
Codexwins 3 of 3

Product-code density source + test lines

Per CR · higher is better
251
118
Per commit · higher is better
91
48
Tokens per 1K lines · lower is better
87.1M
107.1M

How the windows were chosen. Speed, test investment and per-CR / per-commit density use Codex's CR era (Oct 1–3, 33 CRs, 91 commits, 10,142 lines), because Codex had no CRs before Oct 1. Token and cost ratios per commit and per line use Codex's full usage window (Sep 9 – Oct 3: 108 commits, 12,318 lines, 902.6M tokens). Per CR uses the ≈640M tokens of Oct 1–3, read from Codex's usage chart. Claude Code: Oct 3–7, 102 CRs, 252 commits, 22,120 lines, 1,290.8M tokens.

What the lines were

Codex Claude
  • Source code 75% · 31%
  • Tests 9% · 24%
  • Docs and records 14% · 37%
  • Config and ops 1% · 9%

Where Claude Code's tokens went

  • Cache reads 98.75%
  • Cache writes 1.03%
  • Output 0.22%

Codex's chart shows the same pattern. Long sessions mostly re-read their own context at a discounted rate.

How each agent worked

  • Claude: shell 1,481
  • Claude: file edits 1,241
  • Claude: read and search 576
  • Claude: browser 422
  • Codex: plugin calls 198
  • Codex: skill uses 92

Codex leaned on computer use and its Build Dashboard skill. Claude Code worked mostly from the command line.

Chapter 6 · Oct 1–8

Getting cheaper

Daily tokens for both agents, with the average context sent per model call on top. Numbered bubbles mark the changes that cut cost or token use.

−20%tokens per 1K lines changed: 73.3M with Codex → 58.4M with Claude Code
−59%context per call: 504K in the long Oct 5–7 session → 207K in a fresh one
~10×less context per call on Sonnet helpers (47K) than on the main session (474K)
−70%GitHub Actions minutes per CR after verify-locally, push-once
0 200M 400M 0 300K 600K ≈360M ≈200M ≈80M 67M 377M 119M 220M 482M 26M 311K 522K 418K 525K 471K 209K 1 2 3 4 5 6 7 8 Oct 1 Oct 2 Oct 3 ⇄ Oct 4 Oct 5 Oct 6 Oct 7 Oct 8 16 CRs 14 CRs 3 + 7 CRs 33 CRs 9 CRs 22 CRs 30 CRs 1 CR + docs
  1. 1
    Oct 1 · Codex

    Peak day, ≈360M tokens

    Codex's busiest day was also its peak for computer-use calls. Screenshots and desktop actions are token-heavy, so 16 CRs cost about 22M tokens each.

  2. 2
    Oct 3 · Handoff

    Prompt caching does the heavy lifting

    From the first Claude session, 98.75% of tokens were cache reads: the conversation re-sent as cached context, billed at a fraction of fresh input.

    98.75% cached
  3. 3
    Oct 3–4 · Memory

    Learn once, not every session

    Tool paths, encoding traps and the laptop clock issue went into memory notes. Later sessions read a few lines instead of re-discovering each trap with failed commands.

  4. 4
    Oct 5 · CR-0078

    Quiet release script

    One line per step, the full log in a file, and only the last 25 lines on a failure. Test and build output stays out of the agent's context.

    Logs on disk, not in context
  5. 5
    Oct 7 · CR-0128

    Verify locally, push once

    Checks run on the laptop first, and records go into the same push, so each CR triggers one CI run instead of several.

    −70% CI minutes
  6. 6
    Oct 7 · CR-0130

    Routine checks on Sonnet

    Diff reviews and layout checks run in Sonnet helper agents with their own small context, not in the main Opus session.

    19–47K vs ~474K per call
  7. 7
    Oct 7 · Sessions

    Shorter sessions

    A new session for each block of work keeps the re-sent context small. Each call is cheaper even when the day is busy.

    504K → 381K per call
  8. 8
    Oct 8 · Fresh start

    Domain move and docs

    The btcmm.app move, two docs sites and these pages ran in a fresh session with the smallest context per call so far.

    207K per call, −59%

Why Oct 7 is still the tallest bar: it had the most CRs (30), including the sign-in rollout and a server incident, so tokens per CR rose to 16.1M that day. Per-call context still fell, and that's the number that drives cost. Codex's daily values are read from its usage chart, and its Sep 9–30 usage (about 260M tokens in total) is not drawn.

Chapter 7 · Epilogue

What we learned

Claude Code ahead

Speed, cost and tests

About twice the CRs per day, fewer tokens and dollars per unit of work, more than twice the test code per CR, and work across five platforms instead of one machine.

Codex ahead

Product code per change

Codex's changes were bigger chunks of product code: about twice the source and test lines per CR and per commit, at fewer tokens per line. Its built-in computer use made desktop work quick.

Read with care

Not a lab test

The agents did different phases: Codex did early features, Claude did testing, cloud and security. Lines of code reward verbosity and miss work that produces none. Claude's figures leave out its cloud sessions.

Records first

Write it down once

Change records, release notes, a project guide and memory let a new agent or session pick up the work without starting over. The Codex-to-Claude handoff took one clean-up change.

Checks over trust

Let tests judge the agent

A decision replay, coverage floors and checks before release catch an agent's mistakes mechanically. Approvals are kept for what tests can't see.

Least privilege

Narrow, expiring, revocable

Each platform got its own narrow grant with deny rules or an expiry date. Secrets stayed out of chat and out of the code.

Epilogue · Oct 8

Claude Code ran into limits too: the Pro weekly cap, then most of the monthly extra usage. This time the answer was a bigger plan, not a different tool. The project moved to Claude Max, and the records, tests and skills built along the way meant nothing had to be re-learned.