White paper · BTC Market Monitor · October 2026

Chasing an edge, finding a process

What 8,079 paper trades say about predicting 15-minute crypto markets, and what eight days of building with two AI coding agents say about cost, speed and scale.

Abstract

This project set out to test a skeptical claim: that even well-built, faithfully executed trading strategies cannot reliably predict Kalshi's 15-minute crypto markets. Across 16 paper accounts, made of 8 strategies each run with and without a stop-loss, 8,079 closed trades staked $339,181 in paper money between Sep 22 and Oct 8, 2026. The portfolio ended −$3,119 (−0.92%), statistically indistinguishable from zero (t = −0.93). Two strategies stood out, and only one account cleared a conventional significance bar, on just 36 trades. Exit protocols meant to lock in profit made things worse on net: whale-watch stops, stop-losses and profit-taking together cost $1,257 against simply holding to settlement. When the first AI agent's plan ran out mid-project, the work moved to a second agent and its focus moved from tuning strategies to building a lean, scalable product. That second phase, an MVP in 8 days and 135 recorded changes, produced the clearer result. Process changes, not the choice of model alone, cut context per model call by 59%, tokens per line of code by 20% and CI minutes per change by 70%.

−0.92%portfolio return on $339.2K paper staked, t = −0.93
1 of 16accounts with a significant edge (t = 3.85, n = 36)
−$1,257net value of the three exit protocols vs holding
−59%context per model call after process changes

1Question and hypotheses

Kalshi lists a new set of binary markets every 15 minutes: will Bitcoin (and a few other coins) settle above a strike when the window closes? Prices are probabilities, settlement is mechanical, and the bot can watch the same index that settles them. If any short-horizon prediction market should reward a careful model, it is this one. The project tested two hypotheses:

  1. H1 (no reliable edge). Strategies that are faithfully executed and tuned on recent data still won't beat the market reliably over short horizons.
  2. H2 (exit protocols add profit). Selling early, on a stop-loss, on large "whale" order flow or at a profit target, locks in more profit per bot than holding to settlement.

All trading is simulated ("paper"). No real orders were placed, and dollar figures in sections 2–4 are paper dollars.

2Method

3Results: strategies

Figure 1Paper P&L of the eight strategies without the stop-loss, Sep 22 – Oct 8. ROI is P&L over paper stake. Every |t| is below 1.2, so no plain strategy's result is distinguishable from luck.

Half the strategies made money and half lost it. The coin-flip baseline lost 0.9%: the cost of trading with no information. The standout was Settlement, which buys in the final minute when the outcome is nearly fixed: +23.6% on $4,859 staked, 70% of trades won, and 8 of 13 days positive. But its per-trade results swing widely (standard deviation $102), and over 101 trades its 95% interval runs from −$9 to +$31 per trade. Over a few weeks of data, "profitable" and "lucky" can't yet be told apart.

Figure 2Daily portfolio paper P&L across all 16 accounts. Best day +$2,074 (Oct 1), worst −$3,387 (Oct 4). Eight of 17 days were negative (Oct 8 is a partial day). Day-to-day swings dwarf the two-week total.

Verdict on H1: supported, with a caveat. The portfolio as a whole showed no edge (−0.92%, t = −0.93). Two strategies, Settlement and the volatility-smile challenger, earned enough to justify a longer test, but two and a half weeks of data cannot separate skill from variance.

4Results: exit protocols

Figure 3A/B test: P&L of each stop-loss twin minus its plain account over the same period (Oct 1–8). Five strategies were helped and three hurt; net +$70. Two days earlier the same comparison read four and four, net −$284, so the effect is still within noise. The twins took fewer trades than their plain accounts, so this measures the whole protocol, entry filters included.

Whale-watch stop

Exits
413
Sold eventual losers
353 (85%)
Sold eventual winners
60
Actual P&L
−$13,745
If held
−$12,610
−$1,135

Stop-loss

Exits
155
Sold eventual losers
140 (91%)
Sold eventual winners
14
Actual P&L
−$5,718
If held
−$5,713
−$5

Profit-taking (PTP)

Exits (all Oct 2)
20
Sold eventual losers
0
Sold eventual winners
20 (100%)
Actual P&L
+$283
If held
+$400
−$118
Figure 4Counterfactual value of each exit protocol: actual P&L minus the P&L of holding to settlement. The whale-watch stop was usually right, but its 60 wrong calls cost more than its 353 right ones saved. Profit-taking sold only positions that went on to win, forfeiting 29% of their upside.

The stop-loss mostly sold positions that were already lost, so it barely changed results (−$5 net). The whale-watch stop reacted to large order flow and was right 85% of the time, but each wrong call gave up a full winning payout. Profit-taking ran for one day. All 20 of its sales were in positions that went on to settle as winners, and that day the portfolio lost $2,012. It was switched off. The pattern points to a threshold set too tight rather than a flawed idea, but it was never recalibrated, because the project changed direction the next day.

Verdict on H2: rejected on this data. Together the protocols cost $1,257 against holding (−$1,135 whale-watch, −$5 stop-loss, −$118 profit-taking). The stop-loss twin now helps 5 strategies of 8, for a net +$70, but that result changed sign within two days, so any value is strategy-specific and unproven.

5The pivot

On Oct 2 the first agent, OpenAI's Codex, used 99.6% of its monthly allowance in three heavy days (about 360M tokens on Oct 1 alone). Rather than wait four weeks for the reset, the project moved to Claude Code on Oct 3. The new agent picked up from the written change records in a single session. From then on the question changed. Strategy tuning stopped: after Oct 3 the only strategy-logic change was removing one strategy. The goal became shipping a lean, secure, scalable product with as much AI help as possible, and measuring what that help cost.

  • Dashboard and UX
  • Strategy, bot and market data
  • Infrastructure, CI and ops
  • Identity, security and legal
Figure 5What each era's change requests were about, by keyword classification of the release notes (Codex 26 entries, Claude Code 103). Strategy work fell from 42% to 15%, and the remaining 15% was mostly market-data plumbing. Infrastructure rose from 4% to 33%.

6Results: two agents, one codebase

The switch created a natural experiment: same owner, same repository and the same change process, worked by two agents. They did different phases of the product, so the comparison is indicative, not controlled. Each ratio uses the period when its unit existed. Codex had no CRs before Oct 1.

MeasureCodexClaude CodeBetter
CRs per active day11.020.4Claude Code
Commits per active day30.350.4Claude Code
Lines changed per active day3,3814,424Claude Code
Tokens per CR≈19.4M12.7MClaude Code
Tokens per 1K lines, all files73.3M58.4MClaude Code
Tokens per 1K source + test lines87.1M107.1MCodex
Source + test lines per CR251118Codex
Test lines per CR22.751.2Claude Code
Test lines per 100 source lines1076Claude Code
Metered cost per CR, Codex = 100%100%11%*Claude Code
Platforms operated15Claude Code

Table 1. Head-to-head summary for the 8-day MVP (through CR-0135). Re-running with work through Oct 8 (CR-0137: 104 Claude CRs, 1,374M tokens) moves each figure by less than 10% and changes no winner. See the AI Tooling tab for the full scorecard. *Plan fee not included. Codex windows: CR era Oct 1–3 for per-CR and per-day measures, full usage window Sep 9 – Oct 3 for per-line and per-commit tokens. Claude Code: Oct 3–7.

Claude Code won 4 of 5 categories: speed, token efficiency, cost and test investment. Codex won product-code density, since its changes were larger chunks of feature code at fewer tokens per line. The bigger lesson is in what moved the numbers within a single agent.

7Process changes that paid

ChangeWhenBeforeAfterGain
Fresh sessions per block of workOct 7–8504K context/call207K−59%
Routine checks on Sonnet helper agentsOct 7~474K context/call19–47K~10× less
Verify locally, push once per CROct 7several CI runs per CRone, ~8 min−70% CI minutes
Prompt caching from the first sessionOct 3–98.75% of tokens cachedcheap re-reads
Quiet release script (logs to disk)Oct 5full test output in contextone line per stepsmaller context
Agent switch + test-first processOct 373.3M tokens/1K lines58.4M−20%
Tests written with each changeOct 4–510 test lines per 100 source767.6×

Table 2. Efficiency gains traced to specific process changes. Context per call is the conversation re-sent with each model request, the main driver of token cost.

None of these needed a better model. They came from how the work was organized: shorter sessions, cheaper models for routine checks, keeping logs out of the conversation, checking locally before using paid CI, and writing tests that let the agent be trusted with production.

8Lean, then scalable

The MVP runs on one small cloud server behind an edge network, with managed sign-in, health alarms, weekly backups and automatic, tested deploys. That's 135 changes and 501 automated tests, in eight days. Scaling follows measured triggers rather than dates. The next step, isolating the bot and caching shared data at the edge, costs about 2–4× today's. Only a real-money decision would justify the isolation and audit work of the final stage (Build Week tab, section F).

9Limitations

10Conclusions

  1. Prediction: no reliable edge over two and a half weeks. The portfolio was flat, and the most promising strategies, Settlement and the volatility-smile challenger, need months of data before anyone should trust them.
  2. Protection: exit protocols didn't lock in profit. A right-most-of-the-time stop can still lose money when its wrong calls give up full payouts.
  3. Process: the clearest, most repeatable gains came from how the AI work was organized. Records made the handoff possible, tests made the agent trustworthy, and session hygiene and helper models made it cheaper.
  4. Next: keep the bot running unchanged to grow the sample, re-test Settlement and the challenger on 8–12 weeks of data, and recalibrate profit-taking before switching it back on.

ASources

  1. Paper trade ledger: 8,079 closed trades through 2026-10-08 03:37 UTC, read-only from the production database through the dashboard's own trade-history code (analysis/wp_stats.py). Run against the Oct 6 snapshot, the same script reproduces the earlier 7,125-trade figures exactly.
  2. Release notes and 135 change records (CR-0001 to CR-0136), and git history on main (388 commits since Sep 1).
  3. Claude Code session logs on the development laptop (2,758 model calls, Oct 3–8). Codex usage page as of Oct 8, 2026.
  4. The Build Week and AI Tooling tabs of these docs.

Independent research project; not affiliated with or endorsed by Kalshi. Results are from simulated paper trading. Simulated results have inherent limitations and do not represent actual trading; no account will or is likely to achieve similar results.