Skip to content

OverviewCore concepts

The same real-world question often trades on several venues at once. Predictefy’s matching engine (embedding similarity + LLM validation) groups equivalent markets into clusters, and computes indicative price discrepancies between cluster members.

Terminal window
curl -s "$PREDICTEFY_API_URL/v1/clusters?limit=20" \
-H "Authorization: Bearer pk_live_YOUR_KEY"
  • GET /v1/clusters — a page of cross-venue clusters.
  • GET /v1/clusters/:id — one cluster with its per-venue member markets.

Cluster members carry a similarity score — the raw embedding similarity between the matched markets. It is deliberately not called “confidence”: it is not a calibrated probability that the markets are equivalent. Cluster ids are stable — they do not change when a member market delists.

GET /v1/clusters returns a plain page. When you need to narrow the set, GET /api/{exchange}/fetchMatchedMarketClusters is the filterable projection of the same store:

  • minSimilarity — floor on the stored matcher score.
  • venue allow/deny lists — restrict which venues may appear.
  • category — taxonomy scope.
  • relation filtering — identity, subset or superset, asking only what the matcher actually verified.
  • includeRawMatches — up to 100 pairwise matches per cluster, for auditing why a cluster formed.

Members come back as full UnifiedMarket objects with the same delist-stable clusterId.

Two omissions are deliberate. volume24h is absent, because the cluster store has no aggregate column and a missing field is honest where a zero would be fabricated. And there is no per-cluster relations field: the filter can ask what the matcher verified, but the response does not assert a relation it did not compute.

GET /api/{exchange}/fetchMatchedEventClusters is the event-grain equivalent.

Clusters match markets. GET /api/{exchange}/fetchEventMatches matches events, which is a different question — “is this the same election?” rather than “is this the same contract?”

It has two modes. Without an eventId it browses, returning cross-venue event pairs from a cluster page as sourceEvent and event. With an eventId it looks up co-member events for the canonical "{venue}:{eventId}" anchor.

Scores are similarity, never confidence, and reasoning is null — the event matcher emits no per-match rationale, so the field is present and empty rather than filled with something invented. There are no prices on this surface; it is discovery, not quoting.

Gated by READS_ENABLE_EVENT_MATCHES.

Terminal window
curl -s "$PREDICTEFY_API_URL/v1/discrepancies" \
-H "Authorization: Bearer pk_live_YOUR_KEY"

Pass ?live=true to recompute each discrepancy from live order-book mids instead of the latest snapshot prices (the response’s meta.live tells you which you got).

Use expand=markets to add the catalog market object to every low and high leg: title, venue, status, close time, image URL when stored, imageResolved and imageSource (both nullable), liquidity, and volume. The optional volume24hSource key is present only when the market’s volume source is known. The CSV expand query parameter supports only markets today; any other value returns 400. The expansion adds no extra metering weight.

Stored mode keeps its existing limit default of 20 and maximum of 100; this documents pre-existing behavior rather than adding headroom. live=true now has an explicit maximum of 10 because every live cluster recomputes against real order books. The tighter cap bounds that cost; the old shared maximum of 100 was an accidental abuse vector on the live path.

GET /v1/discrepancies is a snapshot of now. GET /v1/discrepancies/history is the record of how it got there — change-only snapshots, newest first, so a run of identical readings does not pad the response.

Terminal window
curl -s "$PREDICTEFY_API_URL/v1/discrepancies/history?clusterId=CLUSTER&from=2026-08-01" \
-H "Authorization: Bearer pk_live_YOUR_KEY"

from and to accept ISO-8601 timestamps or epoch milliseconds. to defaults to now, from defaults to 24 hours before to, and every range is capped at 90 days. Filter to one cluster with clusterId.

Metered with the history weight rather than the cheaper catalog weight.

fetchArbitrage (GET /api/router/fetchArbitrage) is router-only and assesses a bounded contract size against live asks:

Terminal window
curl -s "$PREDICTEFY_API_URL/api/router/fetchArbitrage?contracts=100&limit=500&executableOnly=true&venues=polymarket,kalshi&minEdge=0.02" \
-H "Authorization: Bearer pk_live_YOUR_KEY"

The endpoint is paged with limit up to 500 plus an opaque cursor for subsequent pages (page.nextCursor). A cursor from the published surface pages that exact surface version (meta.seq) for 60 seconds after a newer version replaces it, whatever snapshotTTL says; after that the page answers 400 Cursor has expired and you restart from page 1. The SDK iterator walks pages back to back, well inside that window.

Each base cluster expands into every ordered cross-venue pair (buyYes on venue A and buyNo on venue B), with 90 rows as a pathology guard. If that guard truncates a pathological cluster, every emitted row is marked truncated: true. The engine makes one batched live-book read per venue for the selected outcome books and reuses those results across the pairs.

Each row’s clusterId is a composite key: ${clusterId}:${venueA}:${venueB}. The base cluster id can contain :, so recover it by stripping the last two colon-delimited segments. Do not split at the first colon.

fetchArbitrage rows carry a separate rules-comparison matchScore, matchBand, and matchDifferences list of { kind, detail, points }, including judge-confidence deductions. This score is distinct from candidate similarity: 100 means verified rules equivalence, 85 to 99 is minor, 60 to 84 is material, and lower scores or rule vetoes hide the pair. Unjudged pairs remain unscored unless close dates differ by more than 21 days. Only the existing verified and live-execution gates can authorize the arbitrage label. Qualification carries the same score explanation in checks.resolutionEquivalence.matchDeductions, with differenceKinds reserved for judge kinds. See the match score tables.

The match score is a rules comparison produced by an automated judge from the two venues’ published resolution rules; it is not a probability that the markets settle identically, and trading a pair below 100 is your own risk.

The first three query filters match the WebSocket filters. pairsPerCluster is REST-only:

  • executableOnly=true — return only rows that passed every executable gate. The default returns qualifying rows and visible rejected candidates, excluding hidden match bands, with per-leg and pair-level reasons such as synthetic_book, insufficient_depth, unverified_fees, market_not_open, or an equivalence conflict.
  • venues (or venue) — comma-separated venue filter. A row remains when either its buyYes or buyNo leg uses one of those venues; the other leg may use a venue outside the list.
  • minEdge — filter by minimum net edge.
  • pairsPerCluster — optional positive integer limiting rows per base cluster. Invalid values return 400; omitting it keeps all generated pairs. On the published surface, the cap is applied after executableOnly, venues, and minEdge, then before cursor pagination.

The response meta includes { asOf, seq, source } where source is 'published' (served from the live published surface, fresh within the publisher’s self-declared cadence (about 30s; 3s floor)) or 'computed-fallback' (bounded on-demand computation, top 10 candidate clusters) when the published surface is unavailable.

Re-check the live result immediately before acting because books, depth, and market status can change after the response.

The optional stake= parameter on fetchArbitrage remains dark by default behind READS_ENABLE_ARBITRAGE_STAKE. While disabled, supplying it returns 400 VALIDATION_ERROR with stake is not enabled.

When enabled, stake is a USD budget greater than 0 and at most 1,000,000. Supply a single decimal number with at most two decimal places; repeated values, signs and exponents are rejected. With stake, limit must be at most 100. It adds no credits to the existing request weight.

The calculator finds the largest affordable whole-contract allocation from the row’s recorded executable price levels. It never exceeds the assessed contracts, changes the assessed size, or fetches deeper books. Each row gains a stake object with requested (the USD budget), reason and allocation. A priced allocation has reason: null and these fields:

FieldMeaning
contractsWhole contracts allocated to each leg, no more than the row’s assessed size.
legs.buyYes, legs.buyNoEach leg’s venue, canonicalMarketId, side, optional sports selection, whole contracts, ask cost before fees and verified taker fee, including any per-order minimum. Monetary amounts are USD.
settlementFeeWorst-case winning-leg settlement commission in USD.
totalCostBoth legs’ costs, taker fees and worst-case settlement commission in USD.
payoutAllocated contracts times the locked USD 1 payout.
profitpayout - totalCost, in USD.
profitabletrue only when profit is positive. Small allocations can lose money because of per-order minimum fees.
roiprofit / totalCost, floored to four decimal places; null when totalCost is zero.
unallocatedrequested - totalCost, in USD.
exactThe allocation equals the assessed contract size. It does not mean the budget was spent exactly.
cappedThe requested budget exceeds the assessed total cost, so recorded depth caps the allocation.

When an allocation cannot be priced, allocation is null and reason is:

ReasonMeaning
not_executableThe row did not pass the executable gates.
fills_unavailableThe recorded price levels or fee evidence needed to price the allocation were unavailable or could not be verified.
below_one_contractThe budget cannot afford a whole contract on each leg, including fees.

On the published path, fills_unavailable means the recorded price levels or fee evidence needed for the projection were unavailable or could not be verified; the assessment row is still served. If the read of those stored levels fails, meta.degraded is true and meta.degradedReason is fills_sidecar_unavailable. meta.stake echoes the requested budget. Without stake, the row’s stake object and meta.stake are omitted.

Why “indicative” — and never anything stronger

Section titled “Why “indicative” — and never anything stronger”

A price gap between two venues is only tradeable if executable asks (not midpoints), order-book depth at those prices, per-venue fees and gas, market open-status, and resolution equivalence (the two markets truly settle on the same terms) all check out — live, at execution time. The discrepancy endpoints do not apply those gates.

They tell you where to look, not what to trade:

  • By default, prices compared are each cluster member’s stored Yes price from the catalog snapshot; live=true instead overlays current order-book mid-prices, and on some venues the book itself is reconstructed (synthetic: true) — indicative either way.
  • Two markets in a cluster may resolve on subtly different terms.
  • Fees, gas, spread, and depth routinely exceed a small headline gap.

Treat the output as a research signal and do your own verification. Existing cross-match lookups cost 5 credits; price-gap queries and cross-venue comparisons cost 10 credits.