GuidesCookbook
Backtest a strategy on historical candles
Historical candles are the input to any backtest. Predictefy labels every candle with where it came from and how good it is, and a backtest that ignores those labels will produce confident numbers from data that cannot support them.
Request
Section titled “Request”Candles key on an outcome, not a market — a binary market has one series per side. Identifiers covers why.
curl -s "$PREDICTEFY_API_URL/api/polymarket/fetchOHLCV?outcomeId=OUTCOME_ID&resolution=1h&start=2026-01-01T00:00:00Z&end=2026-06-30T00:00:00Z&limit=5000" \ -H "Authorization: Bearer pk_live_YOUR_KEY"| Parameter | Detail |
|---|---|
outcomeId | The required series selector. id is a compatibility alias; marketId may narrow the lookup but is not accepted alone |
resolution | 1s 5s 10s 30s 1m 5m 15m 30m 1h 4h 6h 1d |
start / end | ISO timestamp or epoch milliseconds |
limit | 1–5000 |
5m, 15m, 30m, 4h and 6h are aggregated at query time from stored history rather
than being stored natively. That is not a defect, but it does mean those buckets inherit the
quality of whatever they were built from — which the response tells you.
The candle record
Section titled “The candle record”{ "timestamp": 1767225600000, "open": 0.41, "high": 0.44, "low": 0.4, "close": 0.43, "volume": 128400, "source": "official", "sourceType": "true-candle", "quality": "ok", "isTrueCandle": true}timestamp, open, high, low, close, volume, source, sourceType, quality, and
isTrueCandle are always present. The volume value may be null; the key itself is required.
source — where the series came from: official, onchain, write-forward, derived.
sourceType — how the bucket was built:
| Value | Meaning |
|---|---|
true-candle | The venue published this candle |
point-derived | Built from point-in-time prices |
trade-derived | Built from the trades tape |
rest-derived | A coarse REST trade-tape candle |
book-derived | Built from order-book state |
rollup | Aggregated from finer buckets |
quality — ok, partial, suspect, or mixed. mixed means the bucket aggregates
inputs of differing quality, which is the expected value for a query-time aggregation.
isTrueCandle — the single boolean that separates published candles from reconstructed
ones.
Filter before you backtest
Section titled “Filter before you backtest”function usable(candle) { // A backtest that mixes published candles with book-derived reconstructions is // measuring two different things and reporting one number. if (candle.quality === 'suspect') return false; if (candle.sourceType === 'book-derived') return false; return true;}
const candles = body.data.filter(usable);const coverage = candles.length / body.data.length;if (coverage < 0.95) { throw new Error( `only ${(coverage * 100).toFixed(1)}% of buckets are usable — widen the window or drop the venue`, );}Decide the rule up front and record it with the result. “Backtested on 1h candles, excluding
suspect and book-derived buckets, 98.2% coverage” is a claim someone can check. A bare Sharpe
ratio is not.
volume is nullable. A null is “not reported”, not zero — a volume filter that treats null as
zero silently discards every venue that does not publish it.
Replay the tape with /v1/history/replay
Section titled “Replay the tape with /v1/history/replay”A candle is a summary: four prices for a stretch of time whose actual path nobody recorded.
GET /v1/history/replay serves the rows underneath it — top-of-book ticks, printed trades, or
stored candles — one NDJSON page at a time, so a strategy can be replayed against what the market
actually did rather than against a bar’s corners.
curl -s "$PREDICTEFY_API_URL/v1/history/replay?venue=kalshi&outcomeId=OUTCOME_ID&kind=tob&since=2026-01-01T00:00:00Z&until=2026-01-08T00:00:00Z&limit=5000" \ -H "Authorization: Bearer pk_live_YOUR_KEY"| Parameter | Required | Detail |
|---|---|---|
venue | yes | Served venue slug. router is rejected — history is venue-scoped. |
outcomeId | yes | The canonical catalog id or the venue-native id; native ids are resolved before the query runs |
kind | yes | Exactly one of tob, trades, candles |
since / until | yes | ISO timestamp or epoch milliseconds. until minus since must be 31 days or less |
resolution | for candles | 1s 5s 10s 30s 1m 1h 1d |
limit | no | Rows per page, 1–20,000 (default 5,000) |
cursor | no | The next value from the previous page |
marketId | no | Disambiguates a venue-native outcome id shared by several markets (see below) |
Pass the optional marketId query parameter when an outcome id is ambiguous on its venue;
the API returns 400 VALIDATION_ERROR with “outcomeId is ambiguous on this venue; pass marketId”
when it needs this selector. Python’s client.replay_history() and ReplayFeed accept
market_id="ob:56:mkt1" and send it as marketId on every page, including cursor pages.
Omitting market_id leaves the query unchanged; pass a non-empty value or omit it (an empty string is
sent and rejected as a mismatch). A known outcome paired with a different market is refused with
400 VALIDATION_ERROR “outcomeId does not resolve on marketId”.
Requires a server that supports marketId (history replay with the ambiguity 400); an older server
ignores the parameter, and on such a server an ambiguous id silently replays an arbitrary market’s tape.
The five query-time aggregations fetchOHLCV also accepts (5m, 15m, 30m, 4h, 6h) are not
replay resolutions. Replay serves stored rows; aggregate them yourself if you want coarser bars.
The page body
Section titled “The page body”The response is application/x-ndjson — one JSON object per line, ordered by ts ascending, with a
cursor line last. A kind=tob request returns tob rows only; the kinds never mix inside a page.
{"k":"tob","ts":1767225600123,"bid":0.41,"ask":0.43,"bidSize":120,"askSize":80,"mid":0.42,"last":0.42,"mechanism":"ws-lossless"}{"k":"trade","ts":1767225600456,"tradeId":"0xabc:1","price":0.42,"size":25,"side":"buy","source":"venue-rest","quality":"ok"}{"k":"candle","ts":1767225600000,"resolution":"1m","open":0.41,"high":0.44,"low":0.40,"close":0.43,"volume":128400,"source":"official","sourceType":"true-candle","quality":"ok","isTrueCandle":true}{"k":"cursor","next":"eyJ0cyI6Li4u","rows":5000}ts is epoch milliseconds on every row — a candle’s ts is the bucket’s open. The nullable
numerics (bid, ask, bidSize, askSize, mid, last, size, volume) arrive as JSON null
when the source did not report them, and a null means “unknown”, never zero.
Paging
Section titled “Paging”The cursor line closes the page and is the only signal that it ended. Repeat the identical query
with cursor=<next> for the following page, and stop when next is null:
# same query as above, plus the previous page's cursorcurl -s "$PREDICTEFY_API_URL/v1/history/replay?venue=kalshi&outcomeId=OUTCOME_ID&kind=tob&since=2026-01-01T00:00:00Z&until=2026-01-08T00:00:00Z&limit=5000&cursor=eyJ0cyI6Li4u" \ -H "Authorization: Bearer pk_live_YOUR_KEY"rows counts the rows on that page. A body that ends without a cursor line was truncated in
transport: retry the page. Treating a truncated page as the end of the window is how a backtest
quietly loses a day of data and still prints a number.
One kind per request
Section titled “One kind per request”Ticks key on ts, trades on (ts, tradeId), candles on the bucket. Three different sort keys cannot
share one honest cursor, so kind takes exactly one value and each kind pages independently.
Merging them into a single ordered stream is the client’s job — the Python ReplayFeed below does
it, and it orders by when each row became knowable, not by ts.
What a page costs
Section titled “What a page costs”Each page is metered as one history read, the same weight as a fetchOHLCV query — see
Credits & billing. A 31-day window of ticks on a busy outcome is many pages, and
the charge is per page, not per window. Pull once, save the NDJSON, and replay from disk.
The honesty fields survive the replay
Section titled “The honesty fields survive the replay”Candle rows keep the whole fetchOHLCV vocabulary — source, sourceType, isTrueCandle, and
quality as ok, partial, suspect or mixed — so the filter above applies to replayed candles
unchanged. Trade rows carry source and a quality of ok, partial or suspect; a single print
is never mixed, because mixed describes a bucket that aggregated inputs of differing quality.
Top-of-book rows carry mechanism instead: ws-lossless, ws-top20, rest-adaptive,
rest-top-of-book, or synthetic-spot, taken from the outcome’s coverage record (the same field
Screen markets documents). It says how the book was captured, and
it decides what the tape can support: a rest-top-of-book mechanism samples a book rather than
observing every change, so a strategy that reacts to individual quotes is being tested on evidence
that never contained them.
Errors use the standard envelope: 400 VALIDATION_ERROR for an unknown kind, a window over 31 days,
a missing resolution, or a cursor that does not decode; 403 PLAN_REQUIRED when the key’s plan is
below Builder; retryable 503 HISTORY_UNAVAILABLE when the history store cannot be reached.
Run it in Python
Section titled “Run it in Python”The predictefy Python package ships predictefy.backtest — a feed that pages and merges replay
rows, an event-driven engine, and a report. It is imported separately from the client surface, so a
plain API user never loads it. The package is a beta line; see the
Python SDK guide for install and pinning.
from predictefy import Predictefyfrom predictefy.backtest import Backtest, ReplayFeed, Strategy
class BuyTheDip(Strategy): """Take a lot when the ask falls to 40c; give it back at 60c."""
LOT = 25.0
def on_event(self, ev, ctx): # One event at a time is all a strategy ever sees — there is no "next bar" handle. if ev.k != "tob" or ctx.open_orders: return if ctx.position == 0 and ev.ask is not None and 0 < ev.ask <= 0.40: # Fills no earlier than the NEXT event, never against the tick that triggered it. ctx.buy(size=self.LOT, price=ev.ask) elif ctx.position > 0 and ev.bid is not None and ev.bid >= 0.60: ctx.sell(size=ctx.position, price=ev.bid)
client = Predictefy(api_key="pk_live_YOUR_KEY")
feed = ReplayFeed( client, "kalshi", "OUTCOME_ID", "2026-01-01T00:00:00Z", "2026-01-08T00:00:00Z", kinds=("tob", "trades"), # add "candles" only with resolution="1m")
result = Backtest(feed, BuyTheDip(), venue="kalshi", initial_cash=10_000).run()report = result.report()
print(report["final_equity"], report["total_return"], report["max_drawdown"])print(report["fills"], "fills,", report["stale_skips"], "stale skips")for note in report["assumptions"]: print("-", note)venue is not decoration: it selects the fee model charged on every fill. report() returns a plain
JSON-serializable dict, and result.to_csv("fills.csv") writes one row per fill with the rule and
book timestamp that produced it.
A metric that cannot be computed honestly comes back as None, never zero — no closed trades means
there is no hit rate, not a hit rate of nought. sharpe ships with sharpe_caveat attached because
equity is sampled once per replayed event and events are irregularly spaced: compare it between runs
on the same feed, never as an annual figure.
Replay from disk
Section titled “Replay from disk”Paying for the same window twice is a waste; save it once and iterate offline.
client.replay_history() yields the rows of a whole window, following the cursors for you.
import jsonfrom predictefy.backtest import Backtest, NdjsonFileFeed
# Save each kind to its OWN file — one file is one ordered stream.for kind, path in (("tob", "tob.ndjson"), ("trades", "trades.ndjson")): with open(path, "w") as out: for row in client.replay_history( "kalshi", "OUTCOME_ID", kind=kind, since="2026-01-01T00:00:00Z", until="2026-01-08T00:00:00Z", ): out.write(json.dumps(row) + "\n")
# The same merge, no network, no credits.feed = NdjsonFileFeed(["tob.ndjson", "trades.ndjson"])result = Backtest(feed, BuyTheDip(), venue="kalshi").run()Cursor lines are skipped on read, so raw pages saved verbatim from curl work too. Mixing kinds inside
one file is where that breaks down: rather than hand the engine an event from the past, the feed
raises ValueError the moment a file’s own order goes backwards. Save each kind separately.
What the engine assumes
Section titled “What the engine assumes”Every number the report prints rests on the list below, and report()["assumptions"] emits it
alongside the metrics so the two travel together. Matching is Fill Model v1 — the same rules the
paper engine runs, pinned to the same golden vectors, so the two implementations cannot silently
drift.
- Every uncertainty declines to fill. No queue position, no hidden liquidity, no price improvement, no market impact, no latency. Where the honest answer is “we do not know”, the model refuses the fill rather than inventing an edge you could not have had.
- Next-event execution. Feeds must be ordered by availability time; late rows are discarded and
counted in
discarded_rowsbefore matching, callbacks, or sampling. The clock never moves backwards; equal timestamps preserve arrival order. An order placed during event N can fill no earlier than N+1: the engine absorbs, matches, delivers, then samples equity including callback cancellation fees.LookAheadErrorguards execution on the order’s own event and later use of a stashed context. Only exact repeats of every event field count asreplayed_rowsand skip the event entirely. Changed signals on the same timestamp and ladder are delivered and sampled, but supply no fresh lot or fills; liquidity identity preserves participation limits. Identity memory is pruned by time with amortized O(1) work per identity; its horizon does not admit late rows. - An order never fills against the observation that caused it. A strategy placing while it handles book B(t) has acted on that quote; filling against B(t) — at that event, or at any print while it still stands — would be trading into its own cause. The order fills on the next book.
- A candle is knowable at its close. Bars enter the stream at
tsplus their resolution, not at their open, so a strategy is never handed a bar whose high it could not have seen. - Participation cap and lot memory. One order takes at most a quarter of a level’s displayed size per book update (the default cap), and one displayed lot cannot fill two orders at the same book timestamp. A trade print is not a book update: an order that has had its allowance against a book gets another only when the book actually changes, and a row repeating a timestamp and ladder already seen is that book replayed, not a fresh lot. A bar repeating a window and prices already seen is likewise that bar replayed: it supplies its volume once. Both are recognised even when other observations arrived in between, which is what an overlapping archive produces.
- Staleness gating. A book older than the budget (30 seconds by default) never fills; the order
waits, whatever its time in force, because a book that could not be read was never a chance to fill
that the order declined. Those skips are counted in
stale_skipsrather than being silently dropped. - Bars are walked pessimistically. In a candle-only feed, a candle is matched at four
synthetic marks — open, the two extremes, close — and the low is visited before the high by
default, so a long position’s stop is reached before its target.
adaptive_path=Truevisits the extreme nearer the open instead: likelier geometry, weaker guarantee. - A bar fills only orders that predate it. Every fill from a bar is stamped at the bar’s close, the moment the bar first became knowable, and a bar whose window was already open when an order was placed cannot fill that order at all — on a mixed trade/candle feed such a bar covers time the order did not exist for.
- Degraded bars trade close-only. A bar that was reconstructed (
isTrueCandle: false) or quality-flagged (partial,suspect,mixed) gets no intrabar marks at all — its high and low were inferred, and filling against them would be inventing a price that may never have traded. Those bars are counted indegraded_bars. - A bar’s own volume is the only liquidity it evidences, split across the marks visited so one
bar cannot supply its volume four times over. A bar that records no volume evidences no
liquidity at all and fills nothing; those bars are counted in
no_liquidity_barsso the report can say why. - A trade tape carries no depth. Orders working when a print arrives with no book to match
against cannot fill; those events are counted in
no_book_eventsso the report can say why nothing filled instead of implying nothing tried. - Long-only, per outcome. A sell closes a long; it does not open a short. Shorting a prediction market means buying the complement, which is a different outcome with its own feed and its own fee.
- Fees are the venue’s verified taker model, charged per fill. A fill on a venue with no proven
model is charged zero and labelled unverified — never claimed as free. Closed-trade P&L is the
gross move minus the closing fill’s fee; the entry fee is counted once, in
fees_paidand in cash. - Every order reserves its fee. A buy reserves its per-fill charges alongside its notional; a sell nets those out of its own proceeds, so a fully invested position can always be liquidated. Both sides hold back the venue’s order-level charge in cash, as a constant that per-fill fees can never eat, because it is taken when the order stops working whether or not the proceeds covered it. An order that could not pay is rejected at placement, exactly as the paper engine rejects it. A fill may then spend only what its own order still reserves plus the cash no order has claimed, so it can neither take cash below zero nor spend another order’s reservation — on either side, since a fractional sell on a venue that rounds every fill up can pay more in fees than it earns in proceeds.
- Order-level fees settle once. A venue’s per-order minimum (opinion’s $0.25) or fixed
per-transaction charge (myriad’s $0.0085) applies to the ORDER, not to each fill, and is charged
when an order that traded stops working. It lands in
fees_paidand in cash and belongs to no fill row — which is why a thin round trip can report a loss its per-fill fees alone would hide. Orders still working when the feed ends are closed, and charged, before the final equity sample, so the reported return carries every charge the run incurred.
None of this makes a backtest a prediction. It makes the result reproducible and the assumptions legible, which is the most a replay can honestly offer.
What history you can actually read
Section titled “What history you can actually read”Two limits apply, and they are different:
- Venue coverage. Not every venue has history for every resolution. Historical data records what exists, including which venues have sub-minute data and in what id format.
- Plan window. Your plan caps how far back you may read. Requesting beyond it returns
PLAN_REQUIREDrather than a silently truncated series — see Credits & billing.
When meta.provenance is present, venue-native means every returned bucket came from the venue,
predictefy-store means our own store served it, and merged means both. Store-only responses may
omit meta, so use each candle’s required source fields as the universal contract. A backtest that
spans a provenance change is comparing two datasets.
History reads are priced above catalog reads, and a backtest is many of them — one per outcome per window. Pull once and cache locally; re-running a strategy should not re-read the API. Current weights are in Credits & billing.
Related
Section titled “Related”- Historical data — coverage per venue and resolution
- Identifiers — why candles key on
outcomeId - Capability-honest data — reading the honesty fields generally
- Python SDK — the client that
predictefy.backtestruns on