All projects

kalshi-arb

An automated trading system for prediction markets and sports betting. It ran for five and a half months, 1,278 commits, 24 strategies, and 1,983 real-money trades, and finished at −$2,004.37. I archived it because the system worked and the answer it gave was no.

Two numbers explain almost all of the loss. Strip the exchange fees out and the same trades lose $969. Hold the bet size flat at the $2.75 I started with, changing no decision, and they lose $254 — size scaling accounts for 87% of it.

What it did

  • Kalshi trading — automated orders against Kalshi's event contracts (sports winners, crypto price ranges, economic releases) through the v2 API, with position management, fill tracking, and settlement reconciliation
  • Soft-book +EV alerts — compared eight New York-legal sportsbooks against Pinnacle every 10 minutes and pushed mispricings to a companion app and Telegram for manual placement
  • Published its own track record publicly, scored under whichever rule was live on each pick's date
  • Ran unattended on a single EC2 instance with ~40 systemd timers, a FastAPI ops dashboard, and a React Native (Expo) app on TestFlight

Results

  • 1,983 Kalshi trades, −$2,004.37 net, of which $1,035 was fees
  • No strategy decayed. The first two weeks made $171 and everything after lost $2,175, but the same strategies ran in both periods with no significant change. What changed was bet size: $2.75 to $27 in six weeks, each increase following a winning week
  • Kalshi's taker fee is steepest exactly where underdog strategies live. At 25–35¢ the gross edge was +5.2% against a 4.8–5.3% fee: right about direction, wrong by less than the cost of transacting
  • Soft-book record at retirement: 677 picks, 309–368, +$19.15. Closing-line value of +0.30% with a 95% CI of [−0.45, +1.04] — entry prices sat at the closing line, not ahead of it

The part worth keeping

Most of the code that survived contact with reality is not strategy logic. It is the apparatus that stops you from believing a strategy that isn't there.

  • Era-scored history. A published track record that re-derives itself from source will rewrite the past every time a rule changes, and the rewrite always flatters. One tier change would have moved the record from +2.1% to +13.9% overnight with no bet settling. The fix was a dated tuple of rules, with history scored under whatever was true at the time, and a CI test that fails on a new era without a disclosure
  • One definition of "a bet." Five surfaces (Telegram, app, push worker, paper ledger, public export) once disagreed about what counted. A single function with a required, un-defaulted policy argument turned silent divergence into an import error
  • Closing-line value, and the confound that ruined it. Per-bet standard deviation was ~120%, so confirming a +5% edge on P&L needs about 4,500 bets. CLV resolves far faster, and it produced a beautiful monotone result that reproduced in both halves of the data and beat 2,000 placebo cuts. It was an artifact: entry EV and CLV shared a term. Measured cleanly, the correlation vanished. Reproducibility and placebo survival are not evidence of causation
  • Pre-registration in code. Thresholds for a hypothesis test committed before the data existed, with a test that duplicates the constants so moving one takes two files in a single diff

Why it stopped

The economics stopped closing. Odds data cost $120 a month, monthly handle was about $9,600, so break-even needed a 1.25% edge. The measured edge was +0.30% with the entire confidence interval below break-even. The project's own rule was no size increase until forward CLV confirmed an edge, but the edge only pays for itself at much larger size. You cannot afford the proof without the size, and you cannot justify the size without the proof.

Kalshi trading stopped for a simpler reason: seven distinct edge vectors were tested and closed, and the exchange priced sports better than Pinnacle did.

Tech stack

Python 3.12 with FastAPI for the dashboard, SQLite as the only database (9.9M odds observations in one 7.2 GB file), systemd timers instead of a queue, and a React Native (Expo) companion app. One box, no ORM, no container. The Odds API subscription is cancelled, the EC2 host is gone, and the repository is published as a record.