OverviewCore concepts
Cross-venue data
The same real-world question often trades on several venues at once. Predictefy’s matching engine (embedding similarity + LLM validation) groups equivalent markets into clusters, and computes indicative price discrepancies between cluster members.
Clusters
Section titled “Clusters”curl -s "$PREDICTEFY_API_URL/v1/clusters?limit=20" \ -H "Authorization: Bearer pk_live_YOUR_KEY"GET /v1/clusters— a page of cross-venue clusters.GET /v1/clusters/:id— one cluster with its per-venue member markets.
Cluster members carry a similarity score — the raw embedding similarity between
the matched markets. It is deliberately not called “confidence”: it is not a
calibrated probability that the markets are equivalent. Cluster ids are stable — they
do not change when a member market delists.
Filtering clusters
Section titled “Filtering clusters”GET /v1/clusters returns a plain page. When you need to narrow the set,
GET /api/{exchange}/fetchMatchedMarketClusters is the filterable projection of the same store:
minSimilarity— floor on the stored matcher score.- venue allow/deny lists — restrict which venues may appear.
category— taxonomy scope.- relation filtering — identity, subset or superset, asking only what the matcher actually verified.
includeRawMatches— up to 100 pairwise matches per cluster, for auditing why a cluster formed.
Members come back as full UnifiedMarket objects with the same delist-stable clusterId.
Two omissions are deliberate. volume24h is absent, because the cluster store has no aggregate
column and a missing field is honest where a zero would be fabricated. And there is no per-cluster
relations field: the filter can ask what the matcher verified, but the response does not assert a
relation it did not compute.
GET /api/{exchange}/fetchMatchedEventClusters is the event-grain equivalent.
Matching at event grain
Section titled “Matching at event grain”Clusters match markets. GET /api/{exchange}/fetchEventMatches matches events, which is a
different question — “is this the same election?” rather than “is this the same contract?”
It has two modes. Without an eventId it browses, returning cross-venue event pairs from a cluster
page as sourceEvent and event. With an eventId it looks up co-member events for the canonical
"{venue}:{eventId}" anchor.
Scores are similarity, never confidence, and reasoning is null — the event matcher emits no
per-match rationale, so the field is present and empty rather than filled with something invented.
There are no prices on this surface; it is discovery, not quoting.
Gated by READS_ENABLE_EVENT_MATCHES.
Indicative price discrepancies
Section titled “Indicative price discrepancies”curl -s "$PREDICTEFY_API_URL/v1/discrepancies" \ -H "Authorization: Bearer pk_live_YOUR_KEY"Pass ?live=true to recompute each discrepancy from live order-book mids instead
of the latest snapshot prices (the response’s meta.live tells you which you got).
Use expand=markets to add the catalog market object to every low and high leg:
title, venue, status, close time, image URL when stored, imageResolved and imageSource
(both nullable), liquidity, and volume.
The optional volume24hSource key is present only when the market’s volume source is known.
The CSV
expand query parameter supports only markets today; any other value returns 400.
The expansion adds no extra metering weight.
Stored mode keeps its existing limit default of 20 and maximum of 100; this
documents pre-existing behavior rather than adding headroom. live=true now has an explicit
maximum of 10 because every live cluster recomputes against real order books. The tighter
cap bounds that cost; the old shared maximum of 100 was an accidental abuse vector on the live path.
How a gap moved
Section titled “How a gap moved”GET /v1/discrepancies is a snapshot of now. GET /v1/discrepancies/history is the record of how it
got there — change-only snapshots, newest first, so a run of identical readings does not pad the
response.
curl -s "$PREDICTEFY_API_URL/v1/discrepancies/history?clusterId=CLUSTER&from=2026-08-01" \ -H "Authorization: Bearer pk_live_YOUR_KEY"from and to accept ISO-8601 timestamps or epoch milliseconds. to defaults to now, from
defaults to 24 hours before to, and every range is capped at 90 days. Filter to one cluster
with clusterId.
Metered with the history weight rather than the cheaper catalog weight.
Executable assessment
Section titled “Executable assessment”fetchArbitrage (GET /api/router/fetchArbitrage) is router-only and assesses a bounded contract size against live asks:
curl -s "$PREDICTEFY_API_URL/api/router/fetchArbitrage?contracts=100&limit=500&executableOnly=true&venues=polymarket,kalshi&minEdge=0.02" \ -H "Authorization: Bearer pk_live_YOUR_KEY"The endpoint is paged with limit up to 500 plus an opaque cursor for subsequent pages (page.nextCursor). A cursor from the published surface pages that exact surface version (meta.seq) for 60 seconds after a newer version replaces it, whatever snapshotTTL says; after that the page answers 400 Cursor has expired and you restart from page 1. The SDK iterator walks pages back to back, well inside that window.
Each base cluster expands into every ordered cross-venue pair (buyYes on venue A and
buyNo on venue B), with 90 rows as a pathology guard. If that guard truncates a pathological
cluster, every emitted row is marked truncated: true. The engine makes one batched live-book read
per venue for the selected outcome books and reuses those results across the pairs.
Each row’s clusterId is a composite key: ${clusterId}:${venueA}:${venueB}. The base
cluster id can contain :, so recover it by stripping the last two colon-delimited segments.
Do not split at the first colon.
fetchArbitrage rows carry a separate rules-comparison matchScore, matchBand, and
matchDifferences list of { kind, detail, points }, including judge-confidence deductions.
This score is distinct from candidate similarity: 100 means verified rules equivalence, 85 to
99 is minor, 60 to 84 is material, and lower scores or rule vetoes hide the pair. Unjudged pairs
remain unscored unless close dates differ by more than 21 days. Only the existing verified and
live-execution gates can authorize the arbitrage label. Qualification carries the same score
explanation in checks.resolutionEquivalence.matchDeductions, with differenceKinds reserved
for judge kinds. See the match score tables.
The match score is a rules comparison produced by an automated judge from the two venues’ published resolution rules; it is not a probability that the markets settle identically, and trading a pair below 100 is your own risk.
The first three query filters match the WebSocket filters. pairsPerCluster is REST-only:
executableOnly=true— return only rows that passed every executable gate. The default returns qualifying rows and visible rejected candidates, excluding hidden match bands, with per-leg and pair-level reasons such assynthetic_book,insufficient_depth,unverified_fees,market_not_open, or an equivalence conflict.venues(orvenue) — comma-separated venue filter. A row remains when either itsbuyYesorbuyNoleg uses one of those venues; the other leg may use a venue outside the list.minEdge— filter by minimum net edge.pairsPerCluster— optional positive integer limiting rows per base cluster. Invalid values return400; omitting it keeps all generated pairs. On the published surface, the cap is applied afterexecutableOnly,venues, andminEdge, then before cursor pagination.
The response meta includes { asOf, seq, source } where source is 'published' (served from the live published surface, fresh within the publisher’s self-declared cadence (about 30s; 3s floor)) or 'computed-fallback' (bounded on-demand computation, top 10 candidate clusters) when the published surface is unavailable.
Re-check the live result immediately before acting because books, depth, and market status can change after the response.
Stake calculator
Section titled “Stake calculator”The optional stake= parameter on fetchArbitrage remains dark by default behind
READS_ENABLE_ARBITRAGE_STAKE. While disabled, supplying it returns
400 VALIDATION_ERROR with stake is not enabled.
When enabled, stake is a USD budget greater than 0 and at most 1,000,000.
Supply a single decimal number with at most two decimal places; repeated values,
signs and exponents are rejected. With stake, limit must be at most 100.
It adds no credits to the existing request weight.
The calculator finds the largest affordable whole-contract allocation from the row’s
recorded executable price levels. It never exceeds the assessed contracts, changes
the assessed size, or fetches deeper books. Each row gains a stake object with
requested (the USD budget), reason and allocation. A priced allocation has
reason: null and these fields:
| Field | Meaning |
|---|---|
contracts | Whole contracts allocated to each leg, no more than the row’s assessed size. |
legs.buyYes, legs.buyNo | Each leg’s venue, canonicalMarketId, side, optional sports selection, whole contracts, ask cost before fees and verified taker fee, including any per-order minimum. Monetary amounts are USD. |
settlementFee | Worst-case winning-leg settlement commission in USD. |
totalCost | Both legs’ costs, taker fees and worst-case settlement commission in USD. |
payout | Allocated contracts times the locked USD 1 payout. |
profit | payout - totalCost, in USD. |
profitable | true only when profit is positive. Small allocations can lose money because of per-order minimum fees. |
roi | profit / totalCost, floored to four decimal places; null when totalCost is zero. |
unallocated | requested - totalCost, in USD. |
exact | The allocation equals the assessed contract size. It does not mean the budget was spent exactly. |
capped | The requested budget exceeds the assessed total cost, so recorded depth caps the allocation. |
When an allocation cannot be priced, allocation is null and reason is:
| Reason | Meaning |
|---|---|
not_executable | The row did not pass the executable gates. |
fills_unavailable | The recorded price levels or fee evidence needed to price the allocation were unavailable or could not be verified. |
below_one_contract | The budget cannot afford a whole contract on each leg, including fees. |
On the published path, fills_unavailable means the recorded price levels or fee
evidence needed for the projection were unavailable or could not be verified;
the assessment row is still served. If the read of those stored levels fails,
meta.degraded is true and
meta.degradedReason is fills_sidecar_unavailable. meta.stake echoes the requested
budget. Without stake, the row’s stake object and meta.stake are omitted.
Why “indicative” — and never anything stronger
Section titled “Why “indicative” — and never anything stronger”A price gap between two venues is only tradeable if executable asks (not midpoints), order-book depth at those prices, per-venue fees and gas, market open-status, and resolution equivalence (the two markets truly settle on the same terms) all check out — live, at execution time. The discrepancy endpoints do not apply those gates.
They tell you where to look, not what to trade:
- By default, prices compared are each cluster member’s stored Yes price from the catalog
snapshot;
live=trueinstead overlays current order-book mid-prices, and on some venues the book itself is reconstructed (synthetic: true) — indicative either way. - Two markets in a cluster may resolve on subtly different terms.
- Fees, gas, spread, and depth routinely exceed a small headline gap.
Treat the output as a research signal and do your own verification. Existing cross-match lookups cost 5 credits; price-gap queries and cross-venue comparisons cost 10 credits.