π§ͺ Backtesting
Backtesting replays a leader's past trades against your copy settings and shows what the simulation says you would have made. It answers the question the trader page cannot: the leader was profitable, but would I have been profitable copying them with my ratio, my filters, and my balance?
Those are different questions. A leader pays no copy fees, has no follower lag, and never has a trade filtered out. You have all three. Backtesting is where that gap becomes a number.
Open it from any trader profile: search a wallet address, or pick a trader from Top Wallets, then click Backtesting in the header. You need to be signed in.
β οΈ Read this first: a backtest is a simulation, not a recording. Polymarket does not publish historical order books. Olympus therefore reconstructs executions from the public historical trade tape and marks positions from historical price data, while reporting when either source is missing. See What The Simulation Does Not Model before you trust a number.
Step-By-Step: Running A Backtest
Step 1 β Open A Trader Profile
Find a wallet through Top Wallets, the Polymarket leaderboard, or by pasting an address into the Olympus search bar.
Step 2 β Click Backtesting
The Backtesting button sits in the profile header, next to AI Analyze and Copy trade.
Step 3 β Choose Your Window
Two ways to bound the replay:
- By date (default) β
7d,30d, or aCustomfrom/to range. Windows are capped at 90 days. - By trade count β turn By date off and set Number of last trades instead (default
100, up to80,000). This replays the leader's most recent N trades, scanning back up to 90 days.
Trade count mode exists for fast leaders. A bot firing thousands of trades a day will not fit a full week inside the data limits, so asking for "the last 500 trades" gets you a complete, honest sample where "the last 7 days" would get you a truncated one. If the leader trades fast enough for this to matter, Olympus estimates their pace when the modal opens and warns you before you run.
Step 4 β Pick A Settings Source
- My settings β the settings you already use for this leader, if you follow them
- Manual β start from defaults and edit freely
- Public preset β a shared preset, when one is available
Tabs only appear when you have more than one option. Edits here are local to the backtest and never touch your live copy settings.
Step 5 β Set Your Sizing
Starting balance seeds from your active wallet's equity, and defaults to $1,000 if that is unavailable. Leave it blank to simulate an unlimited bankroll β useful for judging the strategy in isolation, but it changes what ROI means (see Reading The Results).
Ratio % is what fraction of the leader's trade you copy. The refresh icon next to it recalculates the ratio from your starting balance against the leader's balance, both including open positions β the same calculation the live ratio calculator uses.
Everything else lives in collapsible sections: Entry filters, Market filters, Execution mode, and Risk exits.
Step 6 β Run It
The run is queued and polls until it finishes, with a progress bar tracking rows fetched. Identical requests share one run, and results are cached for about 15 minutes.
Reading The Results
The Headline Numbers
| Metric | What it actually means |
|---|---|
| PnL | Realized PnL plus mark-to-market on still-open positions, net of estimated fees |
| ROI | PnL divided by your bankroll β but see the warning below |
| Win rate | Share of positions with a terminal outcome that ended positive |
| Participation | Leader BUY/SELL fills mirrored out of every eligible leader trade fill |
| Follower orders | Filled follower orders; one accumulated order may represent several leader fills |
| Drawdown | The largest peak-to-trough fall during the window |
β οΈ ROI means two different things. If you set a Starting balance, ROI is measured against that bankroll β return on the money you committed. If you leave it blank, ROI is measured against capital at risk: the peak amount deployed at once. The second number is almost always larger and is not a return on your account. Compare two backtests only when both used the same basis.
The Chart
Your simulated equity curve is plotted against the leader's curve on the same grid. The leader's line is deliberately raw β no fees, no sizing, and no filters. Drag the timeline brush to zoom, use Reset zoom to return to the full window, and click either legend series to hide it; Shift-click isolates that series. The Y-axis follows only the visible series and selected time range.
Evidence Gates
Six conservative checks on whether the result is worth believing at all. They are about sample quality, not profit:
| Gate | Passes when |
|---|---|
| Trades | at least 100 copied trades |
| Markets | at least 20 unique markets |
| Active days | at least 14 distinct trading days |
| Top market | no single market drives more than 40% of realized PnL |
| Unrealized | no more than 50% of PnL is still marked-open |
| Unpriced | no more than 25% of capital at risk sits in unpriced positions |
A big green PnL that passes two of six gates is a small sample with a lucky streak, not an edge. The gates exist specifically so that a thin result cannot present itself as a confident one.
The Verdict
Olympus grades each run as Do not follow, Cautiously follow, or Strong candidate, with a confidence level. The rules are fixed:
- Do not follow β the sample is too thin, or the simulation lost money
- Strong candidate β requires the walk-forward test to hold and every quality gate to pass. There is no shortcut to this grade
- Cautiously follow β everything in between
The verdict is a heuristic over the gates above. It is not advice.
Walk-Forward
The covered window is split into three consecutive sub-periods and your settings are replayed in each. If a strategy only worked in one of the three, you are looking at a streak rather than a method.
π‘ What this is not: nothing is trained on one period and tested on another. The same settings run on all three. It catches results that depend on one hot streak; it does not validate that your parameters generalize to a wallet or period the backtest never saw.
Sub-periods where you copied zero trades are excluded rather than counted as failures.
Execution Sensitivity
The replay still reports the existing spread-sensitivity passes for compatibility. Evidence-priced fills do not receive an invented spread adjustment, so these passes may be identical when historical execution coverage is complete. Treat a flat sensitivity result as βthe historical evidence was reused,β not proof that future slippage cannot matter.
Skipped Trades
Every skipped leader trade is tagged with a reason, and the top reasons are counted with the leader dollars behind them. Each reason links to the setting that caused it, so you can go straight to the knob.
Skips are not automatically bad β filtering out a leader's worst markets is the entire point of filters. What matters is which trades you skipped. Skipping their losers is an edge. Skipping their winners is a leak. Some reasons are not filters at all: ACCUMULATING_BUY means small buys were batched, and LEADER_PRICE_EXTREME means the leader traded under 1Β’ or over 99Β’, which Olympus never copies regardless of settings.
Leader vs You
The result itemizes why your number differs from the leader's: position sizing, fees, execution spread, filtered trades, size caps, bankroll limits, and missed fills.
π‘ The leader baseline here is windowed and coverage-matched, so it will not match the lifetime PnL on their trader page. That is intentional β comparing your 30-day simulation against their all-time profit would be meaningless.
Auto-Tune With AI
Auto-tune with AI reads the finished run β your settings, the metrics, the skip reasons, the verdict, and your cached Wallet Analysis if you have one β and suggests settings changes with a written rationale. The changes are applied and the backtest re-runs automatically.
It can adjust 14 knobs: ratio, max trade size, max market size, the odds band, liquidity and volume floors, max days out, the trade-count caps, min trigger, and the buy-at-minimum / accumulation choice. Every value is clamped to a valid range, so a tuned re-run cannot fail validation.
β οΈ Auto-tune does not guarantee a better result. It is a single model suggestion, and nothing compares the tuned run against the previous one or rejects a worse outcome. Treat it as a starting hypothesis and read the new numbers yourself. If the tuned run is worse, that is a real and useful answer.
Auto-tune is limited to once per hour per user, and the button shows a live countdown when it is on cooldown. A failed attempt does not burn your hour. Auto-tune also recalculates your ratio from balances, which overrides whatever ratio the model suggested.
What The Simulation Does Not Model
This is the most important section on the page. Olympus reports these limits inside the result rather than hiding them, and you should read a backtest with all of them in mind.
Executions are reconstructed from trades, not historical books. Polymarket does not expose historical order-book snapshots. For an immediate order, Olympus searches the public trade tape from the next whole second through the next 15 seconds and supports partial fills. For a resting maker order, the tape must trade at least one tick through the submitted limit; merely touching it is classified as uncertain and unfilled. This cannot prove queue position or exact depth, so the result labels this source as trade-proxy evidence rather than full book coverage.
Missing execution evidence never invents a fill. Limit-order runs fetch complete public trade slices only for markets that the selected settings could copy. If the requested interval contains more candidate markets than one run can support, the replay uses the newest contiguous fully evidenced interval and reports the shorter duration. If no non-empty supported interval remains, the run fails instead of treating missing evidence as an unfilled order. Post-only execution also requires historical book state that Polymarket does not provide, so it is reported as unavailable rather than guessed.
Settlement coverage can shorten a fast leader's window. One run can enrich at most 2,500 replay markets with complete lifecycle and filter metadata. When a leader exceeds that budget, Olympus uses the newest contiguous supported suffix and shows both the requested and effective durations. If lifecycle, metadata, reconciliation, execution, or terminal valuation remains incomplete, diagnostics retain the partial accounting but PnL, ROI, win rate, drawdown, leader comparison, and setting advice are shown as unavailable.
Resting orders use their actual evidence time. Cash or shares are reserved while an order rests. Later trades see those reservations, expirations release them, and every order whose configured expiry falls after the replay remains pending at the window boundary. A configured limit-sell fallback runs only when its expiry is reached inside the replay. BUY and SELL expiration settings therefore affect the replay.
Some settings are accepted but not simulated. These do nothing to the replay today, and the result panel lists them explicitly:
- Stop loss
- Take profit
- Max drawdown, and Include unrealized PnL
- Counter trading
- Min spend
Open positions are marked without lookahead. At each replay timestamp, Olympus uses only the latest historical price at or before that time, carries it forward with a reported age, and uses execution price before the first mark. Historical windows never substitute today's price. Positions with no defensible historical or terminal price remain unpriced and are excluded from PnL, ROI, and win rate; the Unpriced gate exposes how much capital this affects.
Coverage may be shorter than you asked for. Activity pulls and limit-execution evidence are bounded by separate row, market, and time budgets. On a fast leader those bind long before the 90-day cap does. The Coverage panel reports the actual interval, the requested interval, why it was shortened, and the candidate markets covered by complete trade evidence.
Fees are modeled at the standard rate. The replay charges the Olympus fee of 0.75% of notional β tapering toward zero below 3Β’ and above 97Β’ β plus Polymarket's own market fee where the market has one. Subscription and community fee discounts are not modeled, so if you have one, your real fees would be lower than simulated. The leader's baseline curve pays no fees at all.
Limits
- Sign-in required. Backtesting itself is not geo-restricted; the start copying action inside the modal is
- 10 runs per 10 minutes per user. Joining an identical in-flight run or getting a cached result does not count against it
- Auto-tune: once per hour per user
- 90-day cap on date windows; up to 80,000 trades in trade-count mode
- Results are cached for roughly 15 minutes; identical requests share one run
- Interactive rows are capped for browser performance, while summaries use the full ledger. Diagnostic CSV export fetches every retained detail page and aborts instead of producing a partial file if any page expired
Reproducibility
The replay itself is deterministic β no randomness anywhere, so the same request returns the same result. But re-running later can legitimately give a different answer, because:
- the
7d/30dpresets and trade-count mode all end at now, so the window slides - current-ending windows may use live terminal marks, and those marks move
- markets resolve and data coverage improves over time
Pin a Custom window if you want a stable basis for comparing settings.
Good Practices
- Backtest before you follow, not after you lose money. That is the entire point of the feature
- Check the gates before the PnL. A result that fails the sample gates is noise, no matter how green it is
- Change one setting at a time. Otherwise you learn that something helped, but not what
- Check execution and valuation coverage. A profitable result with missing trade evidence or stale marks is not a fully evidenced result
- Check what you skipped. If your filters removed the leader's winners, tighten the thesis, not the filters
- Use a fixed Custom window when comparing runs, so you are comparing settings rather than comparing two different weeks
- Do not read a backtest as a forecast. A good backtest tells you a leader was plausibly copyable over one past window. It does not tell you what next month does
β οΈ Past performance is not a promise of future results. A backtest is a simulation over a window the market has already resolved, using reconstructed executions, historical marks, and estimated fees. It cannot know about a leader who changes strategy, a market regime that shifts, or liquidity that dries up. Use it to eliminate wallets you should not copy β that is what it is genuinely good at β rather than to predict what you will earn.
