Chasing an edge, finding a process
What 8,079 paper trades say about predicting 15-minute crypto markets, and what eight days of building with two AI coding agents say about cost, speed and scale.
This project set out to test a skeptical claim: that even well-built, faithfully executed trading strategies cannot reliably predict Kalshi's 15-minute crypto markets. Across 16 paper accounts, made of 8 strategies each run with and without a stop-loss, 8,079 closed trades staked $339,181 in paper money between Sep 22 and Oct 8, 2026. The portfolio ended −$3,119 (−0.92%), statistically indistinguishable from zero (t = −0.93). Two strategies stood out, and only one account cleared a conventional significance bar, on just 36 trades. Exit protocols meant to lock in profit made things worse on net: whale-watch stops, stop-losses and profit-taking together cost $1,257 against simply holding to settlement. When the first AI agent's plan ran out mid-project, the work moved to a second agent and its focus moved from tuning strategies to building a lean, scalable product. That second phase, an MVP in 8 days and 135 recorded changes, produced the clearer result. Process changes, not the choice of model alone, cut context per model call by 59%, tokens per line of code by 20% and CI minutes per change by 70%.
1Question and hypotheses
Kalshi lists a new set of binary markets every 15 minutes: will Bitcoin (and a few other coins) settle above a strike when the window closes? Prices are probabilities, settlement is mechanical, and the bot can watch the same index that settles them. If any short-horizon prediction market should reward a careful model, it is this one. The project tested two hypotheses:
- H1 (no reliable edge). Strategies that are faithfully executed and tuned on recent data still won't beat the market reliably over short horizons.
- H2 (exit protocols add profit). Selling early, on a stop-loss, on large "whale" order flow or at a profit target, locks in more profit per bot than holding to settlement.
All trading is simulated ("paper"). No real orders were placed, and dollar figures in sections 2–4 are paper dollars.
2Method
- Data: every paper trade closed between 2026-09-22 18:30 and 2026-10-08 03:37 UTC, read from the production database (the history before the Oct 6 move to the cloud plus everything since). That is 8,079 trades, 96% of them on Bitcoin (7,769), plus ETH, XRP, NEAR, SOL, ZEC, HYPE and DOGE.
- Strategies: a random coin-flip baseline, a volatility-smile challenger, and six "playbook" strategies: fair value, maker, settlement-minute, panic fade, cross-market and hedged. A seventh, complete-set, was removed in CR-0079.
- A/B design: from Oct 1 each strategy also ran as a twin account with the stop-loss protocol switched on, so every strategy has a control (plain) and a treatment (with stop-loss) over the same markets and period.
- Counterfactuals: for every position sold early, the market's actual outcome, taken from settled trades on the same contract, gives what holding would have paid. Protocol value = actual P&L − hold P&L.
- Significance: per-trade P&L mean, standard error, t-statistic and 95% interval for each account. Accounts share forecasts and markets, so trades are not independent, and the t-values flatter the evidence if anything.
3Results: strategies
Half the strategies made money and half lost it. The coin-flip baseline lost 0.9%: the cost of trading with no information. The standout was Settlement, which buys in the final minute when the outcome is nearly fixed: +23.6% on $4,859 staked, 70% of trades won, and 8 of 13 days positive. But its per-trade results swing widely (standard deviation $102), and over 101 trades its 95% interval runs from −$9 to +$31 per trade. Over a few weeks of data, "profitable" and "lucky" can't yet be told apart.
Verdict on H1: supported, with a caveat. The portfolio as a whole showed no edge (−0.92%, t = −0.93). Two strategies, Settlement and the volatility-smile challenger, earned enough to justify a longer test, but two and a half weeks of data cannot separate skill from variance.
4Results: exit protocols
Whale-watch stop
- Exits
- 413
- Sold eventual losers
- 353 (85%)
- Sold eventual winners
- 60
- Actual P&L
- −$13,745
- If held
- −$12,610
Stop-loss
- Exits
- 155
- Sold eventual losers
- 140 (91%)
- Sold eventual winners
- 14
- Actual P&L
- −$5,718
- If held
- −$5,713
Profit-taking (PTP)
- Exits (all Oct 2)
- 20
- Sold eventual losers
- 0
- Sold eventual winners
- 20 (100%)
- Actual P&L
- +$283
- If held
- +$400
The stop-loss mostly sold positions that were already lost, so it barely changed results (−$5 net). The whale-watch stop reacted to large order flow and was right 85% of the time, but each wrong call gave up a full winning payout. Profit-taking ran for one day. All 20 of its sales were in positions that went on to settle as winners, and that day the portfolio lost $2,012. It was switched off. The pattern points to a threshold set too tight rather than a flawed idea, but it was never recalibrated, because the project changed direction the next day.
Verdict on H2: rejected on this data. Together the protocols cost $1,257 against holding (−$1,135 whale-watch, −$5 stop-loss, −$118 profit-taking). The stop-loss twin now helps 5 strategies of 8, for a net +$70, but that result changed sign within two days, so any value is strategy-specific and unproven.
5The pivot
On Oct 2 the first agent, OpenAI's Codex, used 99.6% of its monthly allowance in three heavy days (about 360M tokens on Oct 1 alone). Rather than wait four weeks for the reset, the project moved to Claude Code on Oct 3. The new agent picked up from the written change records in a single session. From then on the question changed. Strategy tuning stopped: after Oct 3 the only strategy-logic change was removing one strategy. The goal became shipping a lean, secure, scalable product with as much AI help as possible, and measuring what that help cost.
- Dashboard and UX
- Strategy, bot and market data
- Infrastructure, CI and ops
- Identity, security and legal
6Results: two agents, one codebase
The switch created a natural experiment: same owner, same repository and the same change process, worked by two agents. They did different phases of the product, so the comparison is indicative, not controlled. Each ratio uses the period when its unit existed. Codex had no CRs before Oct 1.
| Measure | Codex | Claude Code | Better |
|---|---|---|---|
| CRs per active day | 11.0 | 20.4 | Claude Code |
| Commits per active day | 30.3 | 50.4 | Claude Code |
| Lines changed per active day | 3,381 | 4,424 | Claude Code |
| Tokens per CR | ≈19.4M | 12.7M | Claude Code |
| Tokens per 1K lines, all files | 73.3M | 58.4M | Claude Code |
| Tokens per 1K source + test lines | 87.1M | 107.1M | Codex |
| Source + test lines per CR | 251 | 118 | Codex |
| Test lines per CR | 22.7 | 51.2 | Claude Code |
| Test lines per 100 source lines | 10 | 76 | Claude Code |
| Metered cost per CR, Codex = 100% | 100% | 11%* | Claude Code |
| Platforms operated | 1 | 5 | Claude Code |
Table 1. Head-to-head summary for the 8-day MVP (through CR-0135). Re-running with work through Oct 8 (CR-0137: 104 Claude CRs, 1,374M tokens) moves each figure by less than 10% and changes no winner. See the AI Tooling tab for the full scorecard. *Plan fee not included. Codex windows: CR era Oct 1–3 for per-CR and per-day measures, full usage window Sep 9 – Oct 3 for per-line and per-commit tokens. Claude Code: Oct 3–7.
Claude Code won 4 of 5 categories: speed, token efficiency, cost and test investment. Codex won product-code density, since its changes were larger chunks of feature code at fewer tokens per line. The bigger lesson is in what moved the numbers within a single agent.
7Process changes that paid
| Change | When | Before | After | Gain |
|---|---|---|---|---|
| Fresh sessions per block of work | Oct 7–8 | 504K context/call | 207K | −59% |
| Routine checks on Sonnet helper agents | Oct 7 | ~474K context/call | 19–47K | ~10× less |
| Verify locally, push once per CR | Oct 7 | several CI runs per CR | one, ~8 min | −70% CI minutes |
| Prompt caching from the first session | Oct 3 | – | 98.75% of tokens cached | cheap re-reads |
| Quiet release script (logs to disk) | Oct 5 | full test output in context | one line per step | smaller context |
| Agent switch + test-first process | Oct 3 | 73.3M tokens/1K lines | 58.4M | −20% |
| Tests written with each change | Oct 4–5 | 10 test lines per 100 source | 76 | 7.6× |
Table 2. Efficiency gains traced to specific process changes. Context per call is the conversation re-sent with each model request, the main driver of token cost.
None of these needed a better model. They came from how the work was organized: shorter sessions, cheaper models for routine checks, keeping logs out of the conversation, checking locally before using paid CI, and writing tests that let the agent be trusted with production.
8Lean, then scalable
The MVP runs on one small cloud server behind an edge network, with managed sign-in, health alarms, weekly backups and automatic, tested deploys. That's 135 changes and 501 automated tests, in eight days. Scaling follows measured triggers rather than dates. The next step, isolating the bot and caching shared data at the edge, costs about 2–4× today's. Only a real-money decision would justify the isolation and audit work of the final stage (Build Week tab, section F).
9Limitations
- About two and a half weeks of trades and a single market regime. Trades share forecasts and markets and aren't independent.
- Counterfactuals assume a sold position would have settled like the other trades on the same contract.
- The A/B twins started Oct 1 and placed fewer trades than their plain accounts.
- The agents did different phases of work. Lines of code reward verbosity and miss operations work. Codex's daily tokens are read from its usage chart, and Claude Code's figures exclude cloud sessions.
- Profit-taking ran for one day with one threshold, so this is evidence about that setting, not about profit-taking in general.
10Conclusions
- Prediction: no reliable edge over two and a half weeks. The portfolio was flat, and the most promising strategies, Settlement and the volatility-smile challenger, need months of data before anyone should trust them.
- Protection: exit protocols didn't lock in profit. A right-most-of-the-time stop can still lose money when its wrong calls give up full payouts.
- Process: the clearest, most repeatable gains came from how the AI work was organized. Records made the handoff possible, tests made the agent trustworthy, and session hygiene and helper models made it cheaper.
- Next: keep the bot running unchanged to grow the sample, re-test Settlement and the challenger on 8–12 weeks of data, and recalibrate profit-taking before switching it back on.
ASources
- Paper trade ledger: 8,079 closed trades through 2026-10-08 03:37 UTC, read-only from the production database through the dashboard's own trade-history code (
analysis/wp_stats.py). Run against the Oct 6 snapshot, the same script reproduces the earlier 7,125-trade figures exactly. - Release notes and 135 change records (CR-0001 to CR-0136), and git history on main (388 commits since Sep 1).
- Claude Code session logs on the development laptop (2,758 model calls, Oct 3–8). Codex usage page as of Oct 8, 2026.
- The Build Week and AI Tooling tabs of these docs.