7 MIN READ · 01 SEPT 2026 · trading

I won 93% of my bets and still lost money.

40 days, 2,695 settled trades, real money. My best strategy won 93.2% of its bets and still finished down.

SB
Steven Battilana Quant · Zurich · ex-ETH

I hoped for break-even

I run a bot that trades up/down binary options on Polymarket: will Bitcoin be higher or lower at the close of a five-minute window. On 22 July I gave myself 40 days to find out whether it makes money, and said I would post the number either way.

Well, I hoped that I would be at least break-even, if not slightly in the green.

Between 31 July and 31 August, 2,695 trades settled on chain across 30 trading days. The book returned -5.71% on the money staked.

That is the wallet, not the model. Every figure here is what the blockchain actually paid, tied to the cent against my balance. It matters, because the two do not agree.

The 93% that lost money

My best strategy won 93.2% of its bets. 426 wins, 31 losses, 457 trades.

That sounds like a machine that cannot lose. Here is what it actually earned. The average win returned 8.4% of the stake. The average loss returned -101.2%, because the stake goes to zero and you have already paid the fee on the way in. Put those two together and the strategy needed to win 92.4% of the time just to break even.

It won 93.2%. It cleared its own bar by eight tenths of a point.

Then the sizing. Staked flat, a dollar every time, that strategy returns +0.95% per bet. At the stakes it actually used it returns -0.86%, because the losing bets carried 1.31 times the stake of the winning ones. Thirty-one losses sized slightly above four hundred and twenty-six wins, and the entire edge is gone.

Across the whole book the win rate was 54.8% and my break-even sat at 58.2%. Better than a coin. Worse than my own costs.

The test came back, and it is not inconclusive any more

I don’t know whether it works or not live, whether I’m just within the natural standard deviation, or whether the strategy is actually broken.

I ran two tests on 340 trades and wrote them up as I lost 122 times on Polymarket so you don’t have to. The first test asked whether I beat pure chance, meaning whether I land on the right side more often than a coin toss would. The answer was a clear yes. The second test asked whether I make money. It scored 1.02 where you need about 2, so it could not tell me either way. I wrote then that roughly 1,300 settled trades would end the argument. I have 2,695, so I ran both again on live money.

Ep3, 340 tradesNow, 2,695 trades
1. Am I better than pure chance?5.2, decisive5.03, decisive
2. Do I actually make money?1.02, inconclusive-2.91, decisive

It resolved, and it resolved in the wrong direction.

That answers the second question I could not before. The book as a whole is not inside natural variation. It is losing money, which I could already tell by looking at my wallet.

The best strategy, on its own

The book is a mix, so I ran the same two tests again on the one strategy that wins 93.2%.

scoreverdict
1. Am I better than pure chance?18.48yes, and not narrowly
2. Do I actually make money?0.68no answer

18.48 is about as far from a coin toss as this data can measure. 0.68 is nothing. The test has to clear 2 in one direction or the other, and this sits in the dead zone between them, so the honest verdict is that it cannot tell.

The numbers underneath: that same +0.95% per bet, and the range the data actually supports runs from -1.78% to +3.67%. Zero sits inside that range. It might make about 1% a bet, it might lose about 2%, and 457 trades are too few to statistically separate the two.

Why so unclear after 457 trades? Because it only lost 31 times. Every win looks the same, about eight cents on the dollar. The losses are the big ones, and 31 is not many. Four more of them and the strategy stops breaking even.

So the strategy I am most certain is not luck is the one I cannot statistically show makes money.

The week I raised the stakes

I was running the strategy for maybe half a day, and the majority of it was just green. By the majority, I mean maybe one bad trade within 20 or 30 trades.

It was a really good run, so I decided to increase the stakes. I believed the strategy was genuinely good and wanted to see how hard we could push it, and I was curious how much liquidity there actually is in this market.

The liquidity question got a clean answer. An order that comes back partly filled means I have eaten everything posted at my price. Most of mine filled in full, so there was more on the book I never touched.

The money question got an uglier one. Over those seven days I staked 73.5% of all the money in the experiment and took 95.5% of the loss. Two hours on one afternoon, eight trades, cost more than the other thirty days put together.

Then the part I did not expect. Since I increased the stakes it started to perform less well, but I don’t think that the stakes made a difference. I think it was just a natural variation of the market conditions. I checked, and the data agrees. The picks were fine. The size was not.

That week I won 82.1% of my bets. My best run of the whole month and I was still negative. The big bets landed on the losers, 1.39 times the size of the winners.

Across the whole month it went the other way. The bigger bets landed slightly more often on winners. That is why a flat dollar on every bet returns -6.00%, worse than the -5.71% I got.

The takeaways

  1. A win rate is not a business. That 93.2% strategy needed 92.4% just to break even, because a win paid 8.4% of the stake and a loss took all of it. Work out that ratio before you celebrate a win rate.
  2. Inconclusive is temporary, so go and end it. The test that said nothing on 340 trades said something clear on 2,695. If a number is not significant yet, the job is more data, not a better adjective.
  3. Size does not fix a negative edge, it spends it faster. Restaked flat, the same bets lose the same percentage. Sizing decides how expensive being wrong is, never whether you are wrong.
  4. Never let a system mark its own homework. My model said flat. The wallet said -5.71%, and 340 of the trades I published had no chain price at all.

What comes next

The machine is still running and I still plan to refine it. What changes is that I know which part to work on now, and it is not the plumbing, the latency, or the rounding bug I spent a week killing. It is the signal.

In October I am running a local LLM workshop, two or three sessions, on a Mac first and then on a VPS. Dates, location and how to get a seat go out to my newsletter first, and seats are limited.

I have also started talking to people in physical commodity trading about which step between contract and settlement still eats their day. If that is you, let’s have a thirty minute chat: book a slot here.

Would your P&L survive being published?

Get the next episode first

Every week: the progress unfiltered, plus what I can't post publicly. The build, the number, green or red.

Posted 01 SEPT 2026 · filed under trading